GPT‑5.6 Sol davant matemàtiques de doctorat: què resol de debò
Una prova de GPT‑5.6 Sol amb dos problemes de recerca mostra utilitat per redactar i explorar, però també verbositat i verificació incompleta.
El canal Easy Riders posa GPT‑5.6 Sol davant dos problemes de matemàtiques de nivell doctoral. La prova no és un benchmark: són preguntes vinculades a la recerca del creador, amb una mostra minúscula i una verificació parcial. Precisament per això és interessant. En lloc de preguntar si el model «sap matemàtiques», mostra quan estalvia feina, quan produeix farciment convincent i quan l’expert encara ha de tornar als articles i fer el càlcul.
OpenAI presenta Sol com el model de màxima capacitat de la família GPT‑5.6 i ofereix nivells de raonament elevats, inclosa una modalitat ultra amb subagents. El vídeo explora aquests modes dins d’un pla de consum que el títol situa en 200 dòlars. El preu no converteix la resposta en una demostració correcta.
1. Dues preguntes amb dificultat molt diferent
A 01:17, el creador defineix el primer encàrrec. Aporta notes amb proposicions i lemes, descriu el camí algebraic i demana que es formulin i provin dos resultats combinant el material. Ho considera una tasca directa: el model ha d’ordenar i escriure un argument que ja està gairebé esbossat.
El segon problema, a 02:22, és molt més obert. Demana asimptòtica en un límit de doble escala mitjançant tècniques avançades de descens més pronunciat no lineal aplicades a determinants. El mateix autor diu que no domina formalment tota la maquinària. Això dificulta saber si una derivació nova és correcta.
La diferència és clau: redactar conseqüències de premisses aportades no és el mateix que resoldre un problema de recerca obert.
2. El primer resultat és útil, però massa llarg
Després d’uns deu minuts de processament, Sol retorna codi LaTeX, dues proposicions i proves. A 04:29, el creador conclou que la identitat principal està ben explicada i que l’encàrrec bàsic s’ha complert.
La crítica és d’edició: el model redefineix símbols, repeteix justificacions i afegeix una observació sobre extreure una constant d’un determinant que l’autor considera trivial. El text sembla sofisticat, però obliga l’expert a separar el pas necessari de la decoració.
Quan rep una pregunta de seguiment, el model accepta que l’observació es pot eliminar. Aquesta correcció és un bon patró d’ús: demanar una versió concisa, assenyalar el punt dubtós i exigir que el defensi o el retiri.
3. El problema difícil produeix orientació, no una prova
Sol dedica uns vint minuts a la segona consulta. A 08:38, l’autor valora positivament que generalitzi una part del problema i introdueixi una quantitat conjunta que ell mateix pensava explorar. El model sembla haver llegit entre línies i detectat l’estructura que faria falta per obtenir el factor esperat.
Però la resposta inicial no executa l’anàlisi asimptòtica amb el rigor demanat. Proposa un anàleg estructural, ofereix intuïció i presenta fórmules plausibles, sense donar a l’investigador una cadena que pugui adoptar amb confiança. En matemàtiques avançades, «té la forma esperada» no és verificació.
El creador en treu una conclusió contundent a 10:33: per fer servir aquella maquinària, encara ha d’aprendre-la i revisar-la de la manera tradicional. El model no elimina l’expert que coneix les hipòtesis del teorema.
4. Reformular el problema canvia la qualitat
En lloc d’aturar la prova, l’autor redueix l’abast. A 11:21, explica que aïlla la peça essencial —el límit de doble escala d’un determinant—, adjunta les notes de recerca i un article amb un cas particular, i utilitza el mateix model per afinar el nou prompt.
També activa Sol amb ultra, la configuració que pot coordinar subagents. La nova resposta triga més i presenta una estructura asimptòtica més detallada. En una sessió separada de Pro apareixen expressions semblants, un indici de consistència, però no una demostració independent.
El salt de qualitat no es pot atribuir només al mode. Alhora han canviat l’enunciat, el context, els documents i el temps de càlcul. La prova ensenya que preparar bé el problema continua essent una part central de la recerca assistida.
5. Què pot fer realment per a un matemàtic
En aquest experiment, Sol és valuós com a redactor d’una derivació esbossada, lector ràpid de notes, generador d’hipòtesis i interlocutor per buscar generalitzacions. També pot produir LaTeX i oferir un primer mapa bibliogràfic o tècnic que l’investigador després comprova.
Les limitacions apareixen quan cal saber si cada hipòtesi s’aplica, controlar un error asimptòtic o demostrar que un límit és uniforme. Una fórmula coherent pot ser falsa per un signe, una regió del pla complex o una condició no satisfeta. La fluïdesa verbal fa l’error més difícil de detectar, no menys important.
Un flux prudent separa tres nivells: idees per explorar, passos verificats contra les fonts i resultats demostrats. Només els dos darrers haurien d’entrar en un article, i el nivell «demostrat» exigeix revisió línia per línia o verificació formal.
6. Per què aquesta prova no mesura una millora global
El vídeo compara informalment amb experiències anteriors amb GPT‑5.5 i altres models, però no repeteix els mateixos prompts a cegues, no defineix una rúbrica prèvia ni presenta jutges independents. Tampoc publica una solució de referència completa per al problema difícil.
Per tant, no es pot concloure que GPT‑5.6 sigui una millora «de salt» en totes les matemàtiques de doctorat. Sí que s’observa una capacitat útil per reorganitzar material tècnic i proposar una via prometedora quan rep context abundant. El cost s’hauria de comparar amb hores estalviades i errors introduïts, no amb el nombre de pàgines generades.
Conclusions
GPT‑5.6 Sol supera la part més acotada de la prova: converteix notes i indicacions en proposicions i proves aprofitables, encara que verboses. Davant la pregunta de recerca, el primer intent aporta estructura però no la resolució rigorosa esperada. La resposta millora quan l’expert redueix el problema i prepara millor el context.
El resultat no és «la IA ja fa un doctorat» ni «no serveix». És una eina potent per accelerar lectura, escriptura i exploració sota supervisió. En matemàtiques, el coll d’ampolla continua essent distingir una expressió plausible d’un teorema correcte; aquesta responsabilitat no es pot delegar al mateix sistema que ha generat la resposta.
Contrast i context
Fonts consultades
-
01
Easy Riders We Tested $200 GPT-5.6 Sol on PhD Level Math
- 02
-
03
OpenAI Developers GPT‑5.6 Sol model
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
So today we're back because OpenAI have just released GPT 5.6. And if you've been on this channel before, you'll see that we've been basically testing every single model against PhD maths problems. To see whether all of the claims that are made on Twitter are actually legitimate. We've tested all of OpenAI's recent pro models because I pay monthly for that subscription, and we've also tested chords, favour, and we've seen some really strong capabilities, but also some pretty huge limitations where these models will fumble in an incredibly basic way. that said, we have also seen the opposite of this happen where the model will make really big claims and they turn out to be true. And in these cases, I'm the one playing catch-up in order to see whether the model has said is actually true and then it does turn out to be true and it took me a while to actually verify it. So it's a mix bag, really. But this model in particular has seen a lot of hype in terms of its ability to do research level math problems. So today we're going to be exploring some problems that I've been working on and we're going to see whether GPT-5.6 Pro can actually help make progress on them. I've got questions regarding what I've already done, questions regarding like general generalizations of another paper in the context of what I'm doing. So, we're all kinds of capabilities of GPT-5.6 for explaining and exploring things that it's given to actually trying to generalize things that it's given with a benchmark given by a different paper, which serves a similar problem but in a different context. And so much more so that we can really see whether GPT-5.6 is actually a step-function improvement
-
1:13
, obre el vídeo en una pestanya nova
in research level maths, or whether it's just another hyped AI model release. Okay, so we set up two questions which are an increasing difficulty, and we'll then ask some following questions based on how the model responds. So, the first task, which I think is much more straightforward than the second one is I've given it a whole bunch of research now, so you can kind of see here, which basically give some propositions and some lamas and I've essentially explained what I'd like it to do is to combine some results in order to make another result actually there's two results that I'd like it to basically state and prove which basically just combine the things which I've given it. So this task seems definitely doable in my opinion because basically it is quite straightforward how you do this is just a whole bunch of algebra that I don't want to type up. It's not necessarily that complicated of a task. And I then outline the argument of exactly how it gets each one. So one does a coordinate transformation introduces a kernel that's basically conversed a factor that we want and then gets the type of symbol from there. And then the other one does a coordinate transformation gets the kernel and then does the same whole argument but then after you've written it more to relatively which is a thing that's important somehow in this type of problem, you then undo the coordinate transformation and then somehow the symbol seems to lose the end dependence. And as you can see we're on GPT5.6 so and put it on pro. Next up is a much more difficult problem.
-
2:24
, obre el vídeo en una pestanya nova
So in this case, I'm asking you to do a very serious calculation, which I don't actually even fully understand to be honest, which uses some really advanced techniques from integral systems, called non-linear steepest to send, which I personally haven't had formal training in, but it does seem like it's kind of the big guns that can solve pretty much any asymptotics problem if you have a determinant and the whole point of what I'm doing. This is trying to represent the thing that I'm trying to solve as a determinant so that then you can get the asymptotics that determined using this machinery that's been developed. Why I want to see really is if a clanker which is much more clever than me at knowing this kind of stuff, whether it can give any kind of outline of results because there are some limits of the thing that I'm trying to solve, which kind of have been predicted already. So if basically the clanker can use this method of non-linear steepers that say maybe it can try and recover known results and then you would know whether it's done it right because there's already formally that have been predicted or like conjectured and so you really would know if it's followed the right reasoning and whether it's worth actually studying what it's given. So really the task is this.
-
3:19
, obre el vídeo en una pestanya nova
It's to compute the larger asymptotics in what's called the double scaling limit. So that's where you take one variable to infinity. And another variable that kind of is in the whole game of what you're doing. That also depends on this thing that you're taking to infinity. So in some sense it's two things of scaling at once, although they're really still end up depending on the thing that you're taking to infinity. So it is pretty cool because when you do that you get different results. So this question is definitely asking a lot more. And it's going to be more difficult for me to verify whether what it's done is right. by I'm interested in seeing what this model can do because I've really asked previous models of this type of question and the answer hasn't been that useful. Okay so the model has replied on both the questions and interestingly if you look at the thinking time for the two questions so the more difficult problem it seems to have thought for 20 minutes if you look to the simpler problem it thought for half of that time which to be honest I'm still surprised it thought for that long given that this task was really much more simple actually than the other ones. I'm surprised actually that the other tasks which will read and see whether it's done it. I mean just looking at the thinking time to be honest to just to me that it hasn't solved the problem. So in this first question it basically gave back a whole bunch of um late-tech code which is some mathematics surrounding what I'd asked it to do. So I
-
4:25
, obre el vídeo en una pestanya nova
posted that into a little section of a late-tech file that I'm already working on. So this first result is what I kind of expected that it would give. Having read through the notes it did do a pretty good job I would say explaining how you would prove the identity. So it did succeed I would say and doing what I'd asked it to do. It typed up to two propositions and gave some proofs. I do find sometimes that these models would just give so much writing. When you actually want to write something yourself, you do want to distill it into what's important. And I do find that even this model is just kind of saying so much crap. It seems like it's capable of answering the question. I just wish it was able to do it in a more concise manner, which is like a bit easier to read. So that I'm not spending hours reading the same thing. But this is interesting. So it's given a little remark. I wonder what it's saying. So it's got this character equivalent one.
-
5:08
, obre el vídeo en una pestanya nova
You could use the un-shifted symbol. Okay, that's fine. Oh. Ready? Why are you telling me this? Why would you put this as a remark? Why would you put that as a remark? I mean, that would be a joke. If you put this in a remark, you'd come across a more on. Well, you know, actually I'm gonna put a remark in this paper that I've released, that actually I could factor this term out of the determinant and have it pick up an end. That's worth a remark. Why are you telling me this? I don't know. I mean, maybe it's suggesting that it's better to have it factor down.
-
5:41
, obre el vídeo en una pestanya nova
Maybe. It also is kind of annoying how often it defines things. It defines so many symbols. By the end you've got an equation that's got so many symbols and you've got to scroll back through all of the crap. I was about to say it would be cool if it defined everything at the beginning, which it kind of does in this, but I think there was another symbol that you use this one. As Q of I, I think why I expected with this was a slightly better quality of answer. from this question, because I've really done this type of question with GPT5.5. And it's the same thing. You get an answer with some useful equations and a whole bunch of filler. And within this problem, I've been using Fable, I've been using GPT Pro.
-
6:16
, obre el vídeo en una pestanya nova
And in both cases actually, the models would generate so much text. It's much more common for it to extend your kind of file, adding loads of code, than to offer something more efficient. I'm gonna ask why it gave that remark, because it seems like a trivial statement, and I'm not sure why you would state that remark. Okay, so as a follow-on question, I've said, can you provide a more concise version? seem to redefine things a lot avoid using things like Delta Resipter Online Z squared or SQZ equal to the or complex additive statistics. It's not it's complex additive statistics. The complex additive statistics. Please explain why you felt the need to write a whole remark regarding the property of a determinant underfecturing in or out the constant. There's just no way you would have that and unless and this is what I'm saying it's happy to make a case for itself. Okay, so I said perhaps you felt necessary because you think it's better to fact around the deterministic drift either argue your case or remove the remark. So we're going to see what it does him and in doing so we're going to read the answer
-
7:12
, obre el vídeo en una pestanya nova
to this. So just to clarify by the way, as I said before, 20 minutes for this problem, I don't believe that it would have done it. I just don't believe it. Okay? Not if in nine minutes it gave this crap and maybe I'm being too critical. Okay, but this, I don't think it was that complicated for it to do. I'm surprised that took nine minutes. and it did have this really stupid remark. This remark is a big deal, okay? I really want Nege put it in your deez as ableed, put it in your deez so that when you have your vibe you look like a complete moron. To justify a fight if you don't know why it's stupid this remark, look up factoring out constant from determine. This is effectively what it's done and it's written a remark about this. So let's see. Maybe this one is pretty much this, which I think is probably an A level video. Is this a question here? If you take your matrix and then you multiply by number and then you take the determinant. Because this constant when you multiply by the matrix, I think it goes into every row and column. So when you take the determinant and you expand it, you're going to get in of them
-
8:09
, obre el vídeo en una pestanya nova
of this constant. So really it just factors out. So it's the determinant and very multiplied by the constant to power of end. This is what it's saying here. Okay. It looks fancy. You know you've got this cool symbol here. This is stupid. Unless crunchy pt argues it's case. And we'll I'm excited to see whether our I guess it's case to be honest because I've happily be wrong, but we'll see what it says. Yeah, I would therefore remove the remark entirely, okay? And keep the drift inside the symbol. That's what I was all with. Let's go back to this question. So, it actually tries to generalize the problem instead of answering the question I asked to. It actually does a more general problem which, so that was cool. This is cool to see that it starts and introducing this factor of the logarithmic derivative. So this is the more general problem. So, really what I was asking about was just this term here. But in general, people would do this. I kind of was under the impression this was just a belief that I had Which I believe is justified Otherwise I probably wouldn't have believed it. I could have just done it the other way But I wanted to see whether the methods worked for this case truth is I couldn't be bothered right now to do it like this was enough Anyway, so it's cool to see that the client has actually been like you know what you may as well just define the joint moment That's cool because actually another question I was thinking of asking was to be like oh, here's some work
-
9:22
, obre el vídeo en una pestanya nova
which I know if you just do it all for the more general version that should it all go through. Can you do that? So it's called see that it's doing that. Okay so I've got to give it credit here. Take me speaking in the question that I asked it is making a fair point because I asked for this a structural analog of a theorem that shows up in a paper. In order to get that you would definitely need to introduce this. So you would need this because otherwise you're not going to get this term. You can just tell that it's sort of read between the lines because otherwise you wouldn't be getting this factor. Then it talks about a known result. Talk to us proposed structural analog. See, this is just, that's fine. This is obviously true. I don't know why it's putting a box around it though. So it's giving a lot of vibe justification of what's going on. So it's giving you a vibe of what's happening. Let me just show you this epic epic paper. I'm sorry, but this, I was asking for, I wanted to know stuff like this, actually get the asymptotic structure of the parts
-
10:18
, obre el vídeo en una pestanya nova
which are important. This is kind of what I've seen it do before to be honest. It's just kind of the same stuff. And you could argue, okay, I'm trying to one-shot it by asking one question. I think it's giving insight on how you would solve the problem. I just think the thing is, if you're actually trying to solve this problem like this, using this advanced machinery that's kind of outlined here, there's really no other way of going about it. You can't just trust a client to do it. So the things you can't actually trust what it's doing for a more complicated problem. So it makes you wonder how you intending on removing the expert, like say the person that wrote this paper. I'm saying, if I wanted to try and tackle the problem that I'm trying to solve with the machinery in here, the only way I can do that pretty much
-
11:01
, obre el vídeo en una pestanya nova
is the old fashioned way of learning how to do. You can't just be like, oh, gone, track GPT's sender. Okay, so I was a little bit disappointed with the first attempt with GPT Pro, but after leaving it for a bit, and unfortunately seeing even more AI hype to do with this model being released and just how good it is with math. I kind of wondered whether maybe I'd asked the wrong question on maybe it was too difficult. So I had to think about maybe how I could isolate the actual problem I wanted to ask about it. So if we go back to the conversation that I had, why have decided to do instead is to basically simplify the problem and try and isolate the part which I thought was the important part which is evaluating this double-scadler method of a settlement and try and get a new prompt based on that simplified problem which I had the help of GPT Pro to basically help me write the prompt, basically starting from the prompt that I've written, that's what's going into here, as you can see it's something like this, it just includes all of that stuff. I've also given all of the research notes for the thing I'm trying to do and this paper which is a specific case out of the thing that I'm trying
-
11:52
, obre el vídeo en una pestanya nova
to get. So what we've done is we're using GPT5.6 sole on Ultra because apparently it's spawned subagents to try and solve the problem. So I thought maybe it was worth giving that ago. In any case we've also re-ass this question like a simplified question because I can't imagine open I've put this much time in this model and it being basically the same feels a little bit more likely the I've made a mistake maybe than the model is basically the same, but maybe it is basically the same who knows at this point. Maybe we'll see how it answers this question whether we can give any more information than it previously did before. Okay so I had to look over the responses last night and I had to say I was actually much more impressed with the answers that it gave this time round with the improved prompt to what it had given yesterday when looking at this problem. So I had a session in GPT work with this sub-agents thing and it's all for 20 minutes And then we had one in GPT Pro, thought for much longer than the other times for some reason, maybe it's just the prompt was better, which I'm looking at actually because there was some news recently that opening I basically proved some new conjecture. And apparently it was one shot and they gave the prompt to actually, so I'll show you guys that on the screen and they even like publish some sort of archive paper. So from there, it seems like maybe I could improve the answers that I get by comparing my current prompts to what they've done. So the main things that are good about this, and this equation looks pretty good. this one.1. I mean I'm not sure how useful it is really. Just looking at it, it does seem decent. Why? Because basically this piece is quite important and reduces to some other known case, and then this piece is a bit like the generalization. So I basically read through this whole thing
-
13:26
, obre el vídeo en una pestanya nova
yesterday and it seemed pretty promising to be honest. So it talks about this limiting tread home determinant, which is kind of what you would expect, because you are taking a double scale limit of a determinant and if we look at the chaturity work session, it kind of gave the same sort of result which is interesting and overall the answer is just much higher quality than what we had before. It goes into way more detail and I think it's much more rigorous and the formula that it's given are much more interesting than what it had before.