GPT-5.6 Sol o Claude Opus 5? La comparativa que evita proclamar un guanyador
Julian Goldie compara GPT-5.6 Sol i Claude Opus 5 i conclou que els benchmarks no decideixen el duel: l’elecció depèn de la feina i de l’entorn agent.
Dos models nous, però cap marcador compartit
Julian Goldie posa cara a cara GPT-5.6 Sol i Claude Opus 5, dos models de frontera publicats amb només quinze dies de diferència. La seva primera conclusió és també la més important: els gràfics de llançament no permeten declarar un guanyador directe.
OpenAI va publicar els resultats de Sol el 9 de juliol de 2026, quan Opus 5 encara no existia. Per això, les seves taules el comparaven sobretot amb Claude Fable 5 i Opus 4.8. Anthropic va presentar Opus 5 el 24 de juliol amb altres proves, configuracions de cost i entorns d’execució. Posar les dues notes una al costat de l’altra és comparar exàmens diferents.
El vídeo proposa canviar la pregunta. En lloc de demanar quin model és «més intel·ligent», cal preguntar quin encaixa millor amb una tasca, un pressupost i un sistema d’agents concrets.
Què aporta GPT-5.6 Sol
Sol és el nivell principal de la família GPT-5.6. Terra busca un equilibri entre capacitat i cost, mentre que Luna prioritza velocitat i volum. OpenAI orienta Sol al raonament complex, la programació, la recerca, l’ús d’ordinadors, el disseny i els fluxos agent de llarga durada.
El model ofereix diversos nivells d’esforç. max li dona més temps que xhigh per explorar alternatives, comprovar resultats i revisar el pla. ultra va més enllà: coordina quatre agents en paral·lel per defecte i intercanvia un consum superior de tokens per més capacitat i menys temps total en problemes exigents.
En els resultats publicats per OpenAI, Sol obté un índex de 80 a l’Artificial Analysis Coding Agent Index. A Terminal-Bench 2.1 arriba al 88,8%, i la configuració Ultra puja al 91,9%. A BrowseComp marca un 90,4% com a agent únic i un 92,2% amb Ultra; a OSWorld 2.0, orientat a l’ús d’ordinadors, registra un 62,6%.
Aquestes xifres descriuen fortaleses reals, però sempre dins d’un harness, unes eines i una configuració determinats. No demostren automàticament que Sol superi Opus 5 en una aplicació concreta.
Crides programàtiques a eines
Goldie destaca especialment el Programmatic Tool Calling de GPT-5.6. Sol pot escriure petits programes que coordinen eines, processen resultats intermedis i conserven només la informació útil. Això redueix els viatges d’anada i tornada entre el model i cada eina.
En un procés amb moltes consultes, fitxers o passos repetitius, aquesta arquitectura pot reduir latència i consum. El model no necessita rebre de nou tot el resultat brut després de cada operació; pot filtrar-lo i decidir el següent pas dins del mateix flux.
Per al creador del vídeo, aquest comportament encaixa amb projectes autònoms llargs: construcció d’aplicacions, programació precisa, iteracions de disseny i tasques on cal continuar fins que el producte funcioni.
Què aporta Claude Opus 5
Anthropic presenta Opus 5 com el seu model d’ús diari per a programació agent complexa i treball empresarial. S’acosta al nivell de Claude Fable 5 en diverses proves, però costa la meitat: cinc dòlars per milió de tokens d’entrada i vint-i-cinc per milió de sortida. Sol comparteix el preu d’entrada de cinc dòlars, però la seva sortida costa trenta dòlars per milió.
Opus 5 ofereix una finestra de context d’un milió de tokens, fins a 128.000 tokens de sortida i pensament adaptatiu. El nivell d’esforç és high per defecte a l’API i a Claude Code, i es pot ajustar per equilibrar qualitat, temps i consum.
Anthropic afirma que el model lidera Frontier-Bench i GDPval-AA en les seves configuracions publicades. També assenyala que la seva taxa d’èxit a Zapier AutomationBench és aproximadament una vegada i mitja la del següent model al mateix cost per tasca, i que a ARC-AGI-3 triplica la puntuació del competidor més proper.
De nou, aquests resultats provenen d’un altre conjunt d’avaluacions. Serveixen per descriure Opus 5, però no per restar-los directament de les xifres de Sol.
Autorevisió i coordinació d’agents
El tret que més interessa Goldie és que Opus 5 revisa el seu propi treball sense que se li demani. Anthropic documenta que el model és més fort a l’hora de verificar, iterar i corregir errors. En una prova de front-end citada al llançament, va obrir la pàgina en mides d’escriptori i mòbil, va detectar elements fora de pantalla i els va reparar abans de lliurar el resultat.
La guia oficial recomana eliminar instruccions antigues com «afegeix una verificació final» o «utilitza un subagent per comprovar-ho». En Opus 5, aquestes ordres poden duplicar una conducta que ja executa i gastar tokens sense millorar la qualitat.
El model també coordina equips de subagents amb patrons d’escriptor i verificador. Anthropic adverteix, però, que la delegació indiscriminada multiplica el cost i el temps. Cal reservar-la per a línies de treball grans i realment independents.
Goldie tradueix aquestes característiques en un perfil pràctic: Opus 5 per planificar, revisar, analitzar i detectar buits; Sol per executar construccions llargues i empènyer la implementació fins al final.
Una prova visual, però no un duel controlat
El vídeo recorda una comparació anterior en què el canal va fer construir a Sol i Claude Fable 5 diversos projectes: un joc de dracs en món obert, un dungeon crawler, un joc de curses 3D, un simulador de vol, una ciutat de neó i una pàgina de vídeo.
Segons la valoració de Goldie, Sol va destacar en acció, precisió, moviment i acabat gràfic. Claude va oferir més atmosfera, detalls i animació. El resultat li va semblar prou equilibrat per considerar-lo un empat.
Aquesta prova ajuda a il·lustrar diferències de comportament, però no és una comparació directa amb Opus 5 ni una avaluació reproduïble. Hi intervenen preferències visuals, instruccions, eines i l’entorn Agent OS del creador. L’article, per tant, tracta aquestes observacions com experiència d’ús i no com a benchmark independent.
L’entorn pot importar més que el model
Un model no treballa sol. El harness decideix quines eines pot utilitzar, com conserva la memòria, quan crea subagents, què veu després de cada pas i com es recupera d’un error. Un model lleugerament inferior en una prova pot rendir millor si l’entorn li dona el context i les eines adequats.
Goldie recomana assignar una carpeta separada a cada model. Executar Sol i Opus 5 simultàniament sobre els mateixos fitxers pot provocar sobreescriptures o canvis incompatibles. També proposa una memòria compartida que tots dos puguin llegir, però amb carrils de treball independents.
La combinació que defensa és:
- Sol per a implementacions llargues, programació i ús intensiu d’eines;
- Opus 5 per a planificació, revisió, anàlisi i control visual;
- una memòria comuna per evitar que cada model treballi sense el context de l’altre;
- carpetes o branques separades per impedir conflictes.
Aquesta divisió no és universal. En un projecte real s’hauria de validar amb un conjunt propi de tasques, mesurant qualitat, cost, temps, errors i percentatge de feina que necessita revisió humana.
Set consells pràctics del vídeo
- No obliguis Opus 5 a verificar dues vegades. La seva autorevisió ja forma part del comportament normal.
- Comença amb esforç
high. Redueix-lo si la qualitat es manté i reservamaxper als problemes que realment ho necessiten. - No deixis dos agents modificar la mateixa carpeta alhora. Separa els espais de treball.
- Assigna cada model a la feina adequada. L’execució i la revisió no han de recaure necessàriament en el mateix agent.
- Demana concisió de manera explícita. Opus 5 tendeix a generar respostes i documents més llargs que els seus predecessors.
- Utilitza el mode ràpid quan la latència sigui el coll d’ampolla. Anthropic indica que Fast mode funciona aproximadament 2,5 vegades més de pressa, però costa el doble.
- Centralitza la memòria del projecte. Els models han de compartir decisions i requisits sense competir pels mateixos fitxers.
Quin hauries de triar?
Si el projecte exigeix terminal, navegació, moltes eines, construcció sostinguda o agents paral·lels, Sol ofereix una arquitectura especialment atractiva. Si el problema demana una planificació acurada, revisió de codi, inspecció visual o coordinació d’escriptors i verificadors, Opus 5 pot encaixar millor.
També hi ha diferències econòmiques. A l’API, tots dos costen cinc dòlars per milió de tokens d’entrada; Opus 5 cobra vint-i-cinc per la sortida i Sol, trenta. Aquest avantatge nominal pot desaparèixer si un model necessita més tokens, més reintents o una configuració d’esforç superior. El cost real és el de completar correctament la tasca.
Conclusions
- Les puntuacions oficials de Sol i Opus 5 no formen un cara a cara comparable.
- Sol destaca per persistència, programació, eines i configuracions
maxiultra. - Opus 5 posa l’accent en autorevisió, treball empresarial, visió i coordinació de subagents.
- El harness, la memòria i l’aïllament dels espais de treball poden influir tant com el model.
- La millor elecció s’ha de fer amb proves representatives del projecte, no amb una única classificació.
La conclusió del vídeo no és un empat diplomàtic. És una advertència metodològica: abans de proclamar un vencedor, cal comparar els dos models amb la mateixa tasca, el mateix pressupost, les mateixes eines i un criteri d’èxit definit. Fins aleshores, la decisió més útil és especialitzar-los.
Contrast i context
Fonts consultades
- 01
-
02
OpenAI Developers GPT-5.6 Sol Model
-
03
Anthropic Introducing Claude Opus 5
-
04
Claude Platform Docs Models overview
-
05
Claude Platform Docs Prompting Claude Opus 5
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
GPT 5.6 SOL versus Claude Opus 5.2 Frontier models. 15 days apart, everyone already picked a winner, but nobody checked the one thing that actually matters. The two score cards were never measuring the same fight. So, I put GPT 5.6 SOL and Claude Opus 5 side by side inside Agent OS. And what came out was not what I expected. I'm the digital avatar of Julian Goldie, and I help people learn AI tools and actually use them in their day-to-day work. In this one, I'll show you what each model really is, where each one pulls ahead, how I run both of them inside Agent OS at the same time, and the one setting almost everyone gets wrong that quietly makes Opus 5 worse. That last one is near the end, so stick with me. Let's start with GPT 5.6 SOL. SOL is OpenAI's new flagship. It went generally available on the 9th of July. GPT 5.6 comes as a family of three. SOL sits at the top, Terra is the balanced one, Luna is the fast one. SOL is the one built for hard work, long coding jobs, agent workflows, research, computer use, design. It also brings two new settings. Max gives it more time to think and check its own approach. Ultra goes further and runs four agents in parallel by default, so it can split a big job and finish faster. Now, Claude Opus 5, Opus 5 landed on the 24th of July, so about 2 weeks after SOL. Anthropic calls it a step change for the Opus tier. It gets close to the intelligence of their top fable 5 model, and they built it to be the one you reach for every day because it works more efficiently than their other models. Thinking is on by default. It has a 1M token context window, and it has an effort dial you control from low all the way up to max. The headline for me is that Opus 5 checks its own work without being told to. That's the part you feel when you're building something real. Here's how I'd use that for the AI Profit Boardroom. I'd hand Opus 5 the topics AI Profit Boardroom members keep asking about and let it map out a full tutorial series inside Agent OS. Outlines, hooks, lesson order, all in one pass. And that self-checking habit is exactly what a job like that needs because a gap in a content plan is the thing that hurts most when you spot it too late. So, what's the actual difference between these two? Soul is the tenacious one. It keeps pushing. OpenAI put it top of the artificial analysis coding agent index with a score of 80. It hit 88.8% on Terminal Bench, which tests command line work. And with Ultra turned on, that goes up to just under 92%. It's also strong on browsing and computer use, 90.4% on Browse Comp and 62.6% on OS World. Opus 5 is the careful one. Anthropic put it at the top of Frontier Bench and GDP Val for coding and knowledge work. On ARC-AGI-3, which throws brand new puzzles at a model, its score is three times the next best model. On Zapier's Automation Bench, which tests whether a model can finish a full business task start to finish, its pass rate is around 1 and 1/2 times the next best. Now, here's the catch nobody mentions. When OpenAI published their charts, Opus 5 didn't exist yet. So, Soul was measured against Claude Fable 5, not Opus 5. And when Anthropic published theirs, they were measuring against their own models. If you line the two blog posts up side by side and try to read a winner off them, you're comparing two different fights on two different days. That's why I stopped reading charts and started building with both Inside Agent OS. And that's the reason we built the Agent OS walkthroughs the way we did inside the AI Profit Boardroom. You can get the full zip file inside the Agent OS ready to install, and we built out a complete 30-day roadmap here as use cases, so you're not staring at a blank screen wondering where to plug each model in. There are daily video tutorials, the 6-week beginner to expert classroom, and four weekly coaching calls where you can share your screen and get help in real time on your own Agent OS setup. Everything about running Soul and Opus 5 side by side is in there. Link is in the comments and description. Now, let's talk about what actually happens when you build with them. The last time we ran a full head-to-head, we tested Soul
-
3:40
, obre el vídeo en una pestanya nova
against Claude Fable 5 model. We built the same things in both, an open world dragon game, a dungeon crawler, a 3D racer, a Skyrim style world, a flight simulator, a neon city, and a full video page using Remotion. The pattern was clear. Soul was better at action and precision. The movement was smoother, the graphics were sharper, things just worked. Claude was better at atmosphere and detail. The worlds felt more interesting. And on the video page, the Claude build was more animated and more fun to look at. So, the split was never about which model is smarter, it was about what each one is good at. Across the whole set, it was close enough to call a draw. And the thing that actually decided it for me wasn't the model at all, it was the harness. Where you run it, how easy is to move around in, whether you can keep three jobs going without losing track. That's why the setup matters more than the scoreboard. With Opus 5, Anthropic's early access partners reported the same kind of thing on the front end. One team said the animations, games, and 3D work were the best they'd seen from an Opus model. Another said it opened its own pages in a browser at desktop and phone size, spotted a button that had drifted off screen, and fixed it before handing the work back. That's the behavior I care about, not the score. The fact it goes back and looks. Here's another way I'd use that. I'd point Opus 5 at a landing page for the AI Profit Boardroom inside Agent OS, and let it check its own layout at desktop and phone size before it hands anything back to me. That mobile checking behavior is the whole reason I'd give it that job for the AI Profit Boardroom instead of doing the pass myself. There's one more difference worth knowing, and it's under the hood. Soul can now write and run small programs that drive your tools for it. OpenAI calls this programmatic tool calling. Instead of every single tool result getting fired back through the model, Soul handles the middle steps itself and only keeps what actually matters. Fewer round trips, less waiting. On tool-heavy jobs, that's a real difference in speed. Opus 5 went the other direction. It got better at running a team. Anthropic says it coordinates sub-agents well using a writer and verifier pattern, and that it rarely has agents overriding each other's work. If you've ever had two agents fight over the same file inside Agent OS, you already know why that matters. Its vision got sharper, too. It reads charts, documents, and diagrams, and it can rebuild an interface from a picture. It's strongest when you give it tools to zoom in, crop, and check itself as it goes. So, Soul drives the tools, Opus 5 runs the team, and both of those land differently once you have them inside Agent OS. Now, the tips. This is the part that saves you the most time. First, that setting I promised you. On Opus 5, do not tell it to add a verification step. Old prompts from earlier models often say things like, "Verify your work at the end." or "Use a sub-agent to check." Anthropic says Opus 5 already does this on its own, and those old instructions make it over-verify. Strip them out. Second, use the effort dial properly. Start at high. Go down when quality holds to save tokens and time. Go up to max only when the job really deserves it. Third, do not run two models on the same project folder at the same time. You will overwrite your own work. Give each model its own lane inside Agent OS. Fourth, match the model to the job. I let Soul run long autonomous builds in Agent OS, while Opus 5 handles the planning and the review pass. That combination beats either one on its own. Fifth, Opus 5 writes longer than the older models by default, and it talks you through what it's doing more as it works. So, if you want short output, ask for short up front. Don't try to trim it afterwards. Sixth, if you need Opus 5 to move quicker, there's a fast mode that runs it at around two and a half times the normal speed. That's the one to reach for when you're iterating on a build and the waiting is what's slowing you down. And seventh, keep your memory system in one place inside Agent OS, so both models can read from it. That's the
-
7:09
, obre el vídeo en una pestanya nova
piece that turns them into one system instead of two separate chats that know nothing about each other. If you want the full process SOPs and 100 plus AI use cases like this one, join the AI Success Lab. Links in the comments and description. You'll get all the video notes from there, plus access to our community of 85,000 members who are crushing it with AI. You can get the full zip file inside the Agent OS, ready to install, and we built out a complete 30-day roadmap here as use cases. And when you go and try this yourself, here's what's going to happen. You'll get both models running, and then you'll hit the real questions. Which effort level for which job? How to stop them stepping on each other? Where to put the memory system so both models can see it? That's exactly what we work through inside the AI Profit Boardroom. You can get the full zip file inside the Agent OS ready to install, and we built out a complete 30-day roadmap here as use cases along with the prompts we actually use, daily tutorials, and four weekly coaching calls where you can share your screen and get unstuck the same day. Over 4,000 members are in there building this stuff right now, and you can meet them on the member map. Come and join us at aiprofitboardroom.com.