Claude Opus 5 vs GPT‑5.6 Sol: tres proves, un guanyador i molts matisos
Claude Opus 5 guanya tres proves visuals contra GPT‑5.6 Sol, però consumeix molt més temps i recursos. Analitzem el resultat i els límits.
En resum
El creador Dubibubi compara Claude Opus 5 i GPT‑5.6 Sol en tres projectes de programació creativa: una introducció animada, un videojoc de trets en primera persona i un planificador d'interiors en 3D. Tots dos models reben el mateix encàrrec, un nivell de raonament elevat i una sola oportunitat, sense revisions posteriors.
El veredicte subjectiu del vídeo és una victòria d'Opus 5 per 3 a 0 en qualitat visual i acabat inicial. GPT‑5.6 Sol, executat mitjançant Codex, acaba les tres proves molt més de pressa i amb menys tokens i cost API equivalent.
| Prova | Preferència del creador | Diferència principal |
|---|---|---|
| Introducció animada | Claude Opus 5 | Més moviment, ritme i sincronització musical |
| Videojoc FPS | Claude Opus 5, per poc | Tots dos són jugables; GPT és molt més ràpid i econòmic |
| Disseny d'interiors 3D | Claude Opus 5 | Millor presentació i lògica per col·locar objectes |
Aquesta prova és útil com a demostració, però no estableix quin model és universalment millor. Només hi ha una execució per tasca, la puntuació és personal i es comparen agents complets —Claude Code i Codex—, no únicament els pesos dels models.
Quins models es comparen
OpenAI va llançar GPT‑5.6 Sol el 9 de juliol de 2026 com a model principal de la família GPT‑5.6. La documentació oficial el descriu com un model de frontera per a treball professional complex, programació, recerca, ús d'eines i disseny.
Anthropic va presentar Claude Opus 5 el 24 de juliol de 2026. El posiciona com un model per a programació agentiva, treball empresarial i tasques llargues, amb millores particulars en verificació, interfícies, visió i coordinació d'eines.
Les especificacions de context són molt pròximes:
| Model | Context màxim | Sortida màxima | Preu API per milió de tokens |
|---|---|---|---|
| GPT‑5.6 Sol | 1.050.000 | 128.000 | 5 $ entrada / 30 $ sortida |
| Claude Opus 5 | 1.000.000 | 128.000 | 5 $ entrada / 25 $ sortida |
Per tant, Opus 5 no és intrínsecament molt més car per token. En tarifa base, l'entrada costa el mateix i la sortida d'Opus és lleugerament més barata. La diferència observada al vídeo prové del volum total processat, la memòria cau, el nombre d'accions i la durada de cada agent.
Metodologia: mateix encàrrec i una sola oportunitat
El vídeo intenta mantenir constants diversos elements:
- el mateix text d'instruccions per als dos models;
- un nivell de raonament «extra high»;
- una única execució sense demanar correccions;
- publicació del resultat perquè es pugui provar;
- revisió visual parcialment a cegues abans de revelar l'autor;
- registre de temps, tokens i cost equivalent mitjançant Tokenmeter.
Tokenmeter és una eina oberta del mateix creador que llegeix els registres locals de Claude Code i Codex. No modifica els models ni injecta instruccions. Comptabilitza tokens d'entrada, sortida i memòria cau, accions amb eines, temps actiu i un cost estimat segons les tarifes configurades.
El repositori adverteix que el cost és «API equivalent». Si l'usuari treballa amb una subscripció, no representa necessàriament els diners addicionals pagats per aquella sessió. També cal configurar manualment els preus i verificar que cada execució hagi utilitzat el model correcte.
Prova 1: una introducció animada de 20 segons
El primer encàrrec demana una introducció de vint segons per al canal, creada amb HyperFrames. És una prova de tipografia, composició, ritme, sincronització amb música i programació d'animacions fotograma a fotograma.
La primera proposta combina lletres, formes i transicions que entren al ritme de la música fins a formar el nom del canal. La segona utilitza imatges generades i moviments més simples. Durant la revisió a cegues, el creador prefereix clarament la primera i, abans de revelar els models, sospita que la segona és de GPT per la incorporació d'imatges.
En revelar el resultat, la introducció preferida correspon a Opus 5. El panell mostra aproximadament 39 milions de tokens totals davant de 15 milions, al voltant d'una hora i mitja davant de 25 minuts, i uns 28 dòlars de cost equivalent davant d'11,42.
Aquestes quantitats inclouen molt context reutilitzat i iteracions internes de l'agent; no equivalen a 39 milions de paraules noves. La victòria d'Opus es basa en l'acabat visual, però Sol ofereix un resultat funcional amb menys temps i recursos.
Prova 2: un videojoc FPS en un únic HTML
La segona instrucció demana un videojoc complet de trets en primera persona dins d'un sol fitxer HTML. Els dos resultats permeten moure's, disparar, afrontar onades d'enemics, recollir recursos i perdre la partida.
El joc de GPT‑5.6 Sol utilitza colors vius i està disponible al cap d'uns deu minuts. El d'Opus 5 incorpora música, enemics amb comportaments més variats i un cap final. El creador s'ho passa bé amb tots dos i inicialment atorga el punt a Opus per la sensació de producte més complet.
La diferència d'eficiència és considerable. El vídeo mostra uns 2,7 milions de tokens per a Sol i uns 31 milions per a Opus. També afirma que Opus triga aproximadament quatre vegades més i té un cost equivalent pròxim a deu vegades superior en aquesta execució.
El mateix creador matisa després la decisió: per temps i preu, el resultat de GPT és especialment competitiu. Una o dues iteracions addicionals amb Sol podrien costar menys que l'execució única d'Opus, però el vídeo no fa aquesta segona fase i manté la regla d'un sol intent.
Prova 3: un planificador d'interiors en 3D
L'últim encàrrec és Room Craft, una aplicació per dissenyar una habitació en tres dimensions des d'un sol fitxer HTML. Ha de permetre afegir mobles i plantes, moure'ls, canviar colors i alternar estils com industrial, Japandi o modern de mitjan segle.
Les dues versions funcionen. Una mostra diferents vistes, calcula el cost total i permet aplicar conjunts decoratius. L'altra ofereix miniatures més informatives, una presentació visual més polida i una lògica espacial que reconeix, per exemple, que un test petit es pot col·locar damunt d'una taula mentre un objecte gran ha de quedar al terra.
La segona aplicació és la preferida i resulta ser d'Opus 5. GPT‑5.6 Sol acaba la seva versió en uns onze minuts; Opus necessita aproximadament quaranta-tres. El creador considera que l'espera compensa perquè la qualitat inicial és superior.
Al final afirma que les tres execucions d'Opus acumulen uns 91 dòlars de cost API equivalent. No presenta al vídeo una taula final completa amb cada categoria de tokens de les sis sessions, de manera que la dada s'ha d'entendre com el resultat del seu panell i la seva configuració de preus.
Per què el cost observat és tan diferent
Els agents de programació no envien una sola pregunta i reben un únic fitxer. Inspeccionen directoris, creen codi, executen comandes, obren el resultat, detecten problemes i tornen a modificar-lo. Cada pas pot reenviar part del context.
Opus 5, segons la seva pròpia documentació, tendeix a verificar el treball i a delegar més. Anthropic recomana limitar explícitament la delegació en tasques sensibles al cost i provar nivells d'esforç inferiors, perquè «low» i «medium» poden mantenir qualitat amb menys tokens i latència.
En aquestes proves, l'esforç màxim afavoreix l'acabat d'un sol intent, però també permet que l'agent dediqui més recursos a comprovar i refinar. No sabem si Opus mantindria l'avantatge amb esforç mitjà ni si Sol l'igualaria amb una segona iteració curta.
També hi ha diferències de memòria cau. Tokenmeter inclou els tokens llegits de la memòria cau dins del total, encara que el preu unitari sigui inferior. Comparar només la xifra total de tokens pot exagerar la diferència econòmica si no se separen entrada nova, lectura de memòria cau, escriptura i sortida.
Una prova d'agents, no només de models
El vídeo titula la comparació com Opus 5 contra GPT‑5.6 Sol, però els sistemes reals són Claude Code i Codex. Cada entorn decideix com planificar, quines eines utilitzar, quan obrir un navegador, com gestionar permisos i quina informació conservar.
Això significa que el resultat mesura una combinació:
model + instruccions del sistema + agent + eines + configuració + entorn d'execució.
És una comparació rellevant per a qui utilitza exactament aquests productes, però no permet inferir que una crida directa a l'API produiria la mateixa relació de qualitat, temps o cost. Tampoc permet separar quina part de l'avantatge visual prové del model i quina de l'estratègia de verificació de l'agent.
Els límits del 3 a 0
El resultat del vídeo és coherent amb els tres artefactes que el creador inspecciona, però té limitacions:
- Una sola mostra per tasca. Els models generatius poden oferir resultats diferents amb la mateixa instrucció.
- Puntuació subjectiva. No hi ha un jurat independent ni una rúbrica numèrica publicada per a cada funció.
- Ceguesa incompleta. El revisor intenta deduir el model a partir de l'estil i veu miniatures al panell.
- Sense proves automatitzades. No es mesuren errors de JavaScript, fotogrames perduts, rendiment, accessibilitat o compatibilitat mòbil.
- Només tasques visuals. No inclou manteniment d'un repositori gran, correcció d'errors, proves, seguretat o revisió de codi.
- Sense iteració. La regla d'un intent beneficia el model amb millor primera versió, però no mesura la velocitat per arribar a un nivell de qualitat objectiu.
Per això, «Opus guanya 3 a 0» descriu la preferència del creador en aquest vídeo, no una classificació general.
Què diuen les fonts oficials
OpenAI presenta GPT‑5.6 Sol com el seu millor model de programació i destaca l'eficiència per token, l'ús d'eines i el treball agentiu. Anthropic descriu Opus 5 com un salt respecte d'Opus 4.8, amb fortaleses en verificació, interfícies, visió i tasques llargues.
Totes dues empreses publiquen avaluacions on el seu producte apareix molt ben posicionat, però no sempre utilitzen el mateix entorn, esforç, versió, cost o data de tall. El llançament de GPT‑5.6 del 9 de juliol encara comparava principalment amb Opus 4.8 i Fable 5, perquè Opus 5 no es va anunciar fins al 24 de juliol.
Els gràfics dels fabricants són útils per entendre capacitats, però una organització ha de provar els models amb els seus propis repositoris, eines i criteris d'èxit.
Com fer una comparació més robusta
Una segona versió de l'experiment podria:
- executar cada tasca almenys cinc vegades per model;
- fixar l'identificador exacte i la versió de cada agent;
- igualar permisos, eines, ordinador i temps màxim;
- separar tokens nous, memòria cau i sortida;
- aplicar proves funcionals automatitzades;
- utilitzar un jurat cec de diverses persones;
- publicar una rúbrica abans de generar els resultats;
- mesurar el cost necessari per arribar a un llindar comú de qualitat;
- informar de la mediana, la variabilitat i els errors, no només del millor exemple.
Aquest disseny permetria respondre dues preguntes diferents: quin model produeix la millor primera versió i quin arriba a una qualitat acceptable amb menys temps i diners.
Conclusions
- Opus 5 guanya la preferència visual del vídeo. Les tres primeres versions semblen més polides al creador.
- GPT‑5.6 Sol guanya en velocitat i eficiència dins d'aquestes sessions. Acaba molt abans i consumeix menys recursos.
- Les tarifes base no expliquen la diferència. Sol costa 5/30 dòlars i Opus 5, 5/25 per milió de tokens d'entrada/sortida; el comportament dels agents genera la separació.
- Un sol intent premia la qualitat inicial. En un flux iteratiu, el model més ràpid pot recuperar terreny.
- La comparació és de productes complets. Claude Code i Codex aporten planificació, eines i estratègies diferents.
- No hi ha un guanyador universal. Per a prototips visuals d'un sol intent, el vídeo afavoreix Opus; per velocitat, pressupost i iteració, Sol resulta especialment atractiu.
La decisió pràctica no hauria de ser quin model «és millor», sinó quin arriba al resultat necessari amb menys supervisió, temps i cost en el flux de treball concret de cada usuari.
Contrast i context
Fonts consultades
- 01
-
02
Dubibubi / GitHub Tokenmeter: compare coding-agent CLIs
- 03
-
04
OpenAI Developers GPT‑5.6 Sol model specifications and pricing
-
05
Anthropic Introducing Claude Opus 5
-
06
Claude Platform Prompting Claude Opus 5
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
So, Claude Opus 5 launched earlier this week, and what can I say? It's a good model. I compared it to Fable recently, and it won by a landslide. However, there's one model in Anthropic's own benchmarks that Opus 5 couldn't beat, GPT 5.6 Soul. Every benchmark I've seen online ranks these two models extremely closely. So, today we're putting them to the test by giving the exact same prompt to both Opus 5 and GPT 5.6 Soul across three massive builds, a motion graphics video, a fully playable first-person shooting game, and finally, an interior design app. Same rules for both, a single one-shot output, no revisions, deployed live to the internet. This is Claude Opus 5 versus GPT 5.6 Soul. Let's get straight into it. Up here on the left-hand side, we have GPT 5.6 Soul on extra high effort. On the right-hand side, we have Opus 5 on extra high effort as well. Now, I recently hired an animator to create intros for my short form. Check this out. >> [singing and music] >> So, what I want to do is I want to see if I can get Claude and Codex to create something similar for my YouTube channel. And we're going to be doing this using HyperFrames. Because using HyperFrames, you can actually get your AI agents to create motion graphics using code. And what's really cool about HyperFrames is it's completely open source. So, you can literally go and try this today if you have an AI subscription. So, by creating motion graphic intro animations for my YouTube channel, this one build is actually going to test a lot at once. We're testing design taste, typography, animation timing, and also frame-perfect code. Because this can't just be a text slide across the screen, the models have to write the whole animation in code, every [snorts] letter, every shape, every transition, timed to the exact frame. So, this could actually be useful if you're making your own intros, running a faceless YouTube channel, or paying an animator every month like I am. Let's see what these two models cook up. I've got the prompt right over here. Use a HyperFrame skill to create a 20-second animated intro for my YouTube channel called Dooby Dooby. So, I'm just going to go ahead and copy this. And if you want access to all the prompts that I'm using in this video, I'll leave a link to this document in the description. Boom, both models are now running. And of course, you know the drill. We're going to be using this usage dashboard to verify everything. We're going to look at the in tokens, output tokens, the time it takes, how much it's going to cost us. And guess what? This dashboard is now open source. So, if you want to run your own test, then I'll leave a link in the description, or you can just put this into your tab to use it. It's really simple. Just paste it into your AI model and just say, "Hey, I want to use this dashboard. Can you set it up for Claude versus Grok?" Or whatever models you want to use. Now, for the scoring rubric, I have Sam as a wizard. He's representing GPT 3.5 Turbo, and I have Boris Cherny, the boxer, representing Claude Opus 5. Now, last time we did this, you guys seemed to really like me doing the blind testing. So, I'm going to do it again, and I'm not going to know what models generated which outputs until after I've tested both of them. Because it's going to show as a little thumbnail in our dashboard. Anyway, I'm going to let these two models cook. I'm going to go for a walk, and I'll see you in a bit. Yeah, I'm not even going to lie. I'm extremely curious as to who's going to win this bench. Because in my own personal use, I'm constantly switching between Codex and Claude, and I find them both to have strengths and weaknesses. For example, Codex feels like a scientist in a lab. He's very analytical, straight to the point, and just gets [ __ ] done. Its ability to execute is amazing, but it does tend to falter at times. Claude on the other hand is like a really hot girlfriend
-
3:58
, obre el vídeo en una pestanya nova
that won't shut up. It left a really good first impression on me. It one shot at a bug I hadn't been able to get fable to fix after weeks of back and forth. However, it's extremely chatty. I find myself reading essays after Opus completes a tiny task. This makes skills like I have ADHD or the caveman skill increasingly more useful as they teach your models to cut the [ __ ] and give it to you straight. Anyway, the models are probably done now. I'm going to finish this snack and then we'll go back and have a look. Imagine how far back you got to walk to collect that camera man. All right, both models are finished. Let's see how they went. I've got them both lined up here. Let's start with this first one. Woah. Woah. Ooh. That is really cool. Woah. Damn. That is a sick intro. That was really freaking cool. I'm not even going to lie. That was dope. I really like how everything was like to the beat. And look how much animation there is. Like look how they come into frame. Signal. Like this is sick. Look at that. That's such a cool animation right there. Boom. How it explodes. Super cool. And then you can kind of see it's like building up to say the full channel name, which is Dooby Booby. That's actually going to be really hard to beat. That might be one of the better generations that I've seen from an AI agent. Let's have a look at number two. >> [music] >> Cool B. Full power. Uh Yeah, I don't know about that one, man. I can almost guarantee that this is GPT, and I'll tell you why. GPT has image generation capabilities. So, you can see here it's generated an image, and then it's tried to use that image in the video, but it just looks kind of lame. All it does is zoom in a little bit. And then you can see over here it's done that again with this weird texture up here. I don't know, I'm not really vibing this one. Let's have a look at these stats. We got Damn, that's a big difference in token usage. 39 million and 15 million. And if we look at the thumbnail, that looks like the first one right here where it goes just build Yeah, right there. Boom. So, that's Opus, that first one. The second one is GPT 5.6. And that is a big difference, man. 25 minutes, an hour and a half. $28, $11.42. And you know what's crazy is I totally expected that because if you look at the benchmarks, the cost difference between GPT 5.6 Soul Max and Opus Max is literally is quite a big jump. I think it gets even worse when you take in the fact that it's way more token hungry. So, you can see here it's nearly done three times as many tokens compared to GPT, therefore costing nearly three times as much. Now, GPT's wasn't bad, but I still think Opus's was marginally better. Not only was the music actually catchy, but everything was on beat. I was actually really impressed. So, I'm going to have to give this one to Opus, man. I like you, Sam. I'm going to have to give you a little frown in face. Looks like a mustache. Anyway, let's go to the next prompt. Voxo first-person shooter. I'm actually really excited for this one. Create a complete fully playable first-person shooter in single HTML. I'm going to go ahead and send this prompt
-
7:54
, obre el vídeo en una pestanya nova
to both agents. And I'll see you when they're done. All right, both models have just finished the build. And you can see over here, GPT only took 10 minutes, which is actually kind of ridiculous. Let's have a look at the build. All right, Voxo blaster. Which one should we do first? Eeny, meeny. Oh, this one's quite colorful. Maybe we start with this one. WASD to move. All right, let's do it. Whoa. Oh, [ __ ] We got monsters. Oh, oh, damn. Oh, cannon. All right, this is pretty fun. Oh, [ __ ] And then you just got to try and survive as long as possible, hey? Is this like a like a patch? Oh, so that must be like ammo or something. Oh, this is a laser, sniper, whatever you want to call it. Oh my god, I've found the hack. Boom. Boom. For a one-shot game, I'm having a lot of fun, guys. And goes to show, anyone can make a game like this. You got cubed. The arena claimed another. 81 kills, six waves. I mean, that was fun, man. That was actually really fun. All right, let's let's give the second one a go. Voxo blaster. Survive the waves, clear the arena, break everything. Oh, okay. This one's got music. Oh, someone's behind me. Let's go up there. Oh, they can come up here. Oh, no. Oh, [ __ ] He went over, dude. That's kind of scary. Imagine I just get really good at this game, and I'm just like the best at this game in the world cuz no one else plays it but me. Oh, there's a boss. Damn. Oh my god. Okay, I healed a little bit. Oh, there's so many monsters. [ __ ] Oh, I died. Oh, that was fun. That was actually really fun. All right, let's have a look at the stats. Damn, Codex 2.7 million tokens. Okay, so it did that first one. It did that first one cuz it had the the brighter colors. 31 million tokens on that one. Opus 5 is super token hungry. It took four times as long and damn, Codex popped that out for like 1/10 the price. I'm sure you can see a trend coming on here, right? Where Codex is clearly faster and cheaper. It comes down to like what do you care about? Personally for me, the price difference is not that big of an issue if the output is better because time is money, right? And so if Opus is producing a better output that is saving me time, then I'm willing to pay a little extra. I know that people watching this are going to have a different opinion to me and they're going to think, "Ooh, you know, Codex is so much better than Opus because it's cheaper and faster and like you can just iterate and and provide a better output. But like, dude, for me, I'm just like, "Damn, it got the job done, you know?" So that's another point to Opus. >> A few moments later. >> For the time and for the price, I think we can't deny it. GPT did a pretty good job there. Let's get rid of your mustache and we're going to be doing things a little bit differently this time around because for the last build, you're going to be blind as well. You're going to be guessing with me which models generated which output. We have the interior design room planner. Create Room Craft, a complete fully functional 3D interior design room planner in a single HTML file. I'll send these prompts off and I'll see you in a bit. Okay, now remember you're doing this blind with me. So we're not going to know which models did what. We have both Room Crafts done here. Maybe let's start with this first one. Design a room you can walk into. Oh, this is kind of cool. Oh, the rotation is quite smooth. Wow. Oh, I can change the color of the walls. Halo side table. Oh, got to add a bit of that. Oh, and I can just put this anywhere. So, there's cool different types of views I can do here. Floor plan.
-
11:52
, obre el vídeo en una pestanya nova
I think 3D view is the best though. So, we got different modes here. Industrial. Oh, that looks pretty cool. And it tells you the total room cost, which is kind of cool. On the left-hand side we have different We can get another olive tree or a ficus tree or I mean, they all kind of look the same. But industrial is kind of cool as well. Bit of a vibe. Leaning art. It's kind of cool. It's not really leaning. Japandi. It's feeling very Japandi. Anyone feeling Japandi right now? And then we got mid-century. Okay. I like it. They all kind of look the same a little bit, but still pretty cool. Let's have a look at the second one. Room craft. Design a real room in 3D. No manual work. Why Woah. Oh, is that like a skyscraper in the background? That's kind of cool to signal a window. We got a table here I can move around. I like this pulsing circle to show that I've selected it. Kind of feels like SimCity a little bit. Each one has a photo, which is really cool. Compared to this one where you're kind of guessing a little bit. Let's have a look at the plants. The plants look nice. Oh, a little succulent. Little succulent over here. Oh, put it on top of the table. See how it recognized that? I don't know if this first version had that. Let's go ahead. Let's be on the floor. It It has to be on the floor. So, this is pretty pretty cool logic right there. Oh, industrial is pretty clean. Let's have a look at Japandi. Oh, Japandi is a bit of a vibe too. And then we got mid-century. Ooh, mid-century is not for me, but I like it. So, if we were to compare this now to mid-century over here, I'm thinking this one looks way nicer. Can I put tables on top of tables? No. I can't. Okay. Well, what are we thinking, guys? I'm thinking this one's a winner. So, can GPT bring it back? Let's have a look at the stats. Okay, so we can see here that looks like the first one. Damn. Did you get it right? Did you guess that that was going to be GPT or did you think that was going to be Opus? You can see it's done 3.1 million tokens, 11 minutes, man. Look at the trend here, 25 minutes, 10 minutes, 11 minutes. Opus taking 43 minutes, 43 minutes well spent if you ask me. That apple looks way better. Opus total cost, $91. See here, that was all Opus. Again here, one nearly close to 1/8 or 1/7 the price of Opus, but the result is significantly better. So, if we look at the total use, I've only really used 12% of my current 5-hour window and 6% of my weekly, okay? So, are you willing to spend 12% of your total usage to get a significantly better result than GPT. But hey, I mean, GPT I've got 100% of my weekly usage left, so hey, I don't know. What do you care about, man? Are you on a budget? Do you care about iterating a couple more times so you can save a couple dollars, or do you just care that the result is good off the rip? If you're like me, then you probably care about the result being good, so that's Opus 3, man. Was this the result you guys were expecting? Let me know in the comments below. And if you want to see more of this style of video, then feel free to leave a like and subscribe because I make videos like this every single week.