Grok 4.5: preu, rendiment i què confirma SpaceXAI
Grok 4.5 promet codi, agents i documents amb una API de 2/6 dòlars. Contrastem la prova del vídeo amb les especificacions oficials.
Grok 4.5 vol competir en programació, agents i feina d’oficina amb una combinació poc habitual: capacitat de model de frontera, resposta ràpida i una API més barata que moltes alternatives tancades. Caleb Writes Code en comenta els primers gràfics, crea un mapa de centres de dades i transforma el resultat en un informe de Word i una presentació.
La cronologia obliga a separar comentari i confirmació. El vídeo es va publicar el 10 de juliol de 2026, mentre la pàgina oficial de llançament de SpaceXAI està datada el 16 de juliol. Algunes afirmacions del vídeo provenen, doncs, d’informació prèvia o no documentada en els materials oficials actuals.
1. Un nou model, però no totes les xifres estan confirmades
A 00:00, el vídeo afirma que Grok 4.5 es basa en una fundació «V9» d’1,5 bilions de paràmetres i relaciona les dades d’entrenament amb Cursor. SpaceXAI sí que diu que el model es va entrenar «al costat de Cursor» i que va utilitzar desenes de milers de GPU NVIDIA GB300, però la seva presentació pública no confirma ni el recompte de paràmetres ni la nomenclatura V9.
Tampoc cal confondre una col·laboració d’entrenament amb una adquisició empresarial. El vídeo parla de la compra d’Anysphere per SpaceXAI; aquesta operació no apareix en l’anunci oficial consultat. Sense una font corporativa o reguladora, s’ha de tractar com una afirmació no verificada.
La lliçó és especialment pertinent en llançaments d’IA: la mida estimada d’un model pot circular abans que la documentació, però no és una especificació fins que el proveïdor la publica.
2. On se situa als benchmarks de codi
A 01:05, Caleb compara Grok amb models oberts i tancats mitjançant l’índex d’Artificial Analysis. La pàgina oficial aporta resultats més concrets: 83,3% a Terminal-Bench 2.1, 64,7% de resolució a SWE-Bench Pro i 29% a SWE Marathon.
No lidera totes les proves. A les mateixes gràfiques del fabricant, Fable obté 80,4% a SWE-Bench Pro i 84,3% a Terminal-Bench; Grok, en canvi, encapçala SWE Marathon entre els models comparats. Cada prova utilitza un arnès, pressupost i tipus de tasca diferents.
Per això «millor model de codi» és una simplificació. Un equip hauria de repetir incidències del seu repositori, amb proves automatitzades i revisió del canvi, i mesurar taxa d’èxit, temps i cost.
3. El preu és l’argument més clar
A 02:22, el vídeo destaca 2 dòlars per milió de tokens d’entrada i 6 per milió de sortida. Són els preus oficials per a context curt. La documentació actual afegeix entrada en memòria cau a 0,30 dòlars i duplica les tarifes a 4 i 12 dòlars quan el context arriba a 200.000 tokens o més.
La finestra publicada és de 500.000 tokens. El vídeo cita una futura ampliació a un milió basada en una declaració a X, però no convé presentar-la com una capacitat disponible fins que la fitxa del model canviï.
El preu per token tampoc és el cost per feina resolta. Compten el nombre d’intents, la mida del context, els tokens de raonament, les eines i el temps de revisió.
4. Menys tokens com una altra forma de velocitat
SpaceXAI anuncia uns 80 tokens per segon i una mitjana de 15.954 tokens de sortida per tasca a SWE-Bench Pro. Segons la seva comparació, Opus 4.8 en necessita 67.020: 4,2 vegades més. El vídeo interpreta aquesta eficiència com una velocitat que no depèn només del ritme de generació.
És una distinció útil. Un model que escriu més ràpid però dona quatre voltes al mateix problema pot acabar més tard i costar més. Ara bé, la comparació l’ha publicada el mateix fabricant i està lligada a un benchmark concret; cal reproduir-la abans d’extrapolar-la a qualsevol agent.
5. La prova de programació del vídeo
Després del patrocini, a 05:05, Caleb demana a Grok Build una web que visualitzi grans centres de dades dels Estats Units. El model cerca informació, genera el codi i entrega un mapa filtrable en pocs minuts.
La demostració prova capacitat per coordinar cerca, codi i interfície, però no comprova exhaustivament les ubicacions ni mostra una bateria de proves. En productes basats en dades, la part difícil pot ser la procedència, actualització i llicència de la informació, no que el mapa s’obri.
Un resultat visualment convincent és un prototip. Abans de publicar-lo cal validar fonts, accessibilitat, errors de càrrega, dispositius mòbils i contingut inventat.
6. De la web a Word i PowerPoint
A 06:08, el creador obre el projecte a Cursor i demana convertir la web en un resum corporatiu de Word i una presentació. Grok produeix dos fitxers amb estructura neta i estil deliberadament senzill.
SpaceXAI posa precisament Excel, PowerPoint i Word entre els casos de coneixement professional del model. El vídeo confirma que la cadena és possible, no que qualsevol document quedi llest per enviar. Fórmules, cites, identitat visual, dades confidencials i detalls de maquetació encara necessiten control.
Conclusions
Grok 4.5 té una proposta competitiva i verificable: 500.000 tokens de context, preu base de 2/6 dòlars, bona puntuació en tasques d’agents i una integració capaç de produir codi i documents. La seva eficiència de tokens podria ser tan important com la velocitat nominal.
El vídeo també mostra el risc de comentar un model abans de la documentació definitiva. Les afirmacions sobre paràmetres, arquitectura i adquisicions no consten a l’anunci oficial consultat. La decisió pràctica és provar Grok amb tasques pròpies i conservar una frontera clara entre el que mostra una demostració, el que afirma el fabricant i el que ha confirmat una avaluació reproduïble.
Contrast i context
Fonts consultades
-
01
Caleb Writes Code Grok 4.5 explained in 8min..
-
02
SpaceXAI Introducing Grok 4.5
-
03
SpaceXAI Docs Pricing
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Grop 4.5 is the next frontier model released by Space XAI. But unlike previous Grop 4 variants, the new Grop 4.5 model is actually based on a newly pre-trained V9 foundation model that is 1.5 trillion parameters in size, nearly three times bigger than the previous V8 base model. Now the reason why this is worth mentioning is because when we have a newly pre-trained model, it typically requires training the model from scratch with a fresh batch of data. And we know that SpaceX recently acquired Kershurs parent company, Anyspher, back in June, and the amount of high quality coding data
-
0:35
, obre el vídeo en una pestanya nova
that went towards the 9-based model opens up so much room for newer capabilities. And this is exactly what Elon tweeted on X work. Kershurs data was supplemental to the new V9 Foundation models training. Another fact worth noting is the size of the model being 1.5 trillion parameter model, which is comparable to open models like Kimey K2.6, which is a one-chryl-in-primator model
-
0:57
, obre el vídeo en una pestanya nova
and deep-seek V4 Pro, which is 1.6-primator model. And despite that, GROP 4.5 pulls ahead in performance when you compare them against all open models, as you can see here. But just a bit shy of state-of-the-art models, like Fable 5 and GPD 5.6. But what's crazy to think about when looking at this index from artificial announces is that this entire section of models
-
1:19
, obre el vídeo en una pestanya nova
were all announced only in the short span of the past 30 days. except for Opus 4.8. Fable 5 was initially released on June 9, GPD 5.6, announced on June 26, and now we have the Grock 4.5 model, all within the past 30 days, which once again shows a brutal competition among Frontier Labs trying to stay ahead. The last time we had a performance lead like this was back in September 2024, when the first reasoning model or one was released by OpenAI as you can see. And it took the rest of the industry nearly six months to catch up. First by Deepseek during the Deepseek moment back in January 2025, and ever since,
-
1:56
, obre el vídeo en una pestanya nova
we had small increments here and there, and yet again, we have a new leap. But this time, started by Enthropate with their mythists and fable model, and the resident industry like OpenAI and Grock are now catching up really fast. I'm not sure why this chart doesn't have GBD5.6 sole added here, but for reference, the 5.6 sole should land right around here. So, we are undergoing yet another leap in model performance in how LLMs are progressing. Now, when we compare the Grock 4.5 models side by side against other state of the R models like Vable 5, Opus 4.8, GBD 5.6 sol, and hopefully Google's Gemini
-
2:32
, obre el vídeo en una pestanya nova
models soon, I hope. The first thing you might have noticed is just how cheap the model is in comparison. Pricing in just at $2 per million in boutokins and $6 per million out boutokins. Now, the phrase that always gets tossed around when it comes to pricing is subsidized plans. Most people don't use credits to buy API costs, but opting to pay monthly subscription plan that subsidizes users to use their models. What's crazy about the GROC 4.5 model though is that GROC 4.5 API costs is comparable to a much lower tier models
-
3:05
, obre el vídeo en una pestanya nova
from other labs like Sonic 5 from Anthropic, GPD 5.6 Lunar from OpenAI and Gemini 3.5 Flash from Google all lower tier models. They're essentially undercutting the entire close labs when it comes to the cost of intelligence. Now when we compare Grockville.5 against open models, they are actually one of the most expensive models available out there. Looking at models like Mini Max M3, KemiK2.6, Deepzig V4 Pro, GLM 5.2, and Q1.3.7 Max, all noticeably cheaper except for Q1.3.7 Max. So really, Grockville.5 is in this wedge not only in performance, but also in pricing, But there are priced better than front-to-labs in the US, but more expensive compared
-
3:48
, obre el vídeo en una pestanya nova
to open models, while being more capable in intelligence than open models, but slightly below the front-to-labs top performance. So the next question here is usefulness. How useful is Grat4.5 model? And what are its use cases? But first, quick work from me exponseoring this video. Agents are what everyone is building today. But there's a huge technical gap when it comes to having your agent be more than just a chat bot.
-
4:12
, obre el vídeo en una pestanya nova
to pull information from databases, make updates to spreadsheets, and update fields in the database. And we rely on MCP to handle all these things, but managing multiple MCPs and making sure they're all working together can get quite complicated. Make offers a super intuitive UI to build your own workflow that supports over 3,000 apps, as you can see, from this drop down list. From here, I can select my super-based database to connect to and also my Google sheets. And now, because these things are added as their own tool inside of make, my agent can vote them through a unified mcp through make on jobs that I give permissions to. As you can see, when I tell my agent to pull information from the database to collect all workup scores, it can do that. And I can also tell my agent to write to my Google Sheets all through
-
4:56
, obre el vídeo en una pestanya nova
make. You can extend your agent through make for free using one month pro subscription that allows 10,000 operations link in the description below. When we read through the release note for for Grockvopin 5, they're signaling that this model is perfect for coding use cases for developing software and office work for knowledge tasks. You can see in their demonstration of how Grockvopin 5, one shot at the solar system app that simulates our solar system in the super interactive way where we have a planets POV like this as they float around in space. But this is just their demonstration, so we got to pull up our sleeves and test it out ourselves.
-
5:33
, obre el vídeo en una pestanya nova
using their own harness called Grok build to start with. And I asked the model to put together a major data centers in the US in a visualized site. And I was surprised by just how fast the model actually got things done. It pulled together a lot of information about data centers but also built the site with code in just a few minutes. And the site looked something like this where it visually shows you a map
-
5:55
, obre el vídeo en una pestanya nova
of all major data centers in the US and I can filter them by companies that are hosting them and go through each one by one. and I can also zoom into them as well. Pretty impressive for how quick it was able to get the job done. Now, to test their office task abilities, I pulled up a different harness since Graph 4.5 is also available on cursor. So I loaded up cursor and picked the model,
-
6:17
, obre el vídeo en una pestanya nova
Graph 4.5 from the dropdown, and I asked it to convert this website we just created into a word file as a corporate briefing and also create a PowerPoint to do presentations on. And within just a minute or two again, you see that Graph was able to generate two extra files here, one for Word document and another for PowerPoint. The document written here shows a briefing outlining all data centers nicely organized and detailed. And the same thing for PowerPoints, very simple and clean as I prompted specifically for
-
6:47
, obre el vídeo en una pestanya nova
Grok to follow. Now beyond this, you can keep trying out the model and for me, I was testing something really technical until I ran out of credits. But my honest feedback is that the model is noticeably fast in output. Looking at their 85 tokens per second SLA, it might not seem all that impressive, especially in comparison to OpenAI's recent GPD 5.6 announcement, running 750 tokens per second as an option, but I think the appeal to a model like this is using less tokens to get the job done, which is a different kind of speed. According to their report, they showed that Grog 4.5 solves the Swedish Probenchmark using less tokens where the model's average
-
7:27
, obre el vídeo en una pestanya nova
output tokens was up to 4.2 times less than Opus. All in all, Grog 4.5 is showing the industry that XAI they're trying to wedge themselves in the market when it comes to pricing and intelligence, but also being a serious competitor in the enterprise adoption of AI models, given its smaller footprint and model size, they're integration with cursor and their large compute availability. And according to Elon, there will be increasing their Context window to 1 million context in just the next few days, which currently the model only offers up to 500,000 tokens.