Gemini 3.6 Flash i els benchmarks: mateixa puntuació no vol dir cap millora
Gemini 3.6 Flash obté 50 punts a Artificial Analysis, però millora en codi, eines i eficiència. Expliquem què amaguen la mitjana, la velocitat i el cost.
El llançament: tres models amb objectius diferents
Google ha presentat Gemini 3.6 Flash, Gemini 3.5 Flash-Lite i Gemini 3.5 Flash Cyber. El vídeo de Miguel Torrez qüestiona el relat del llançament perquè el nou Flash obté una puntuació agregada semblant a 3.5 Flash i queda per sota dels models de màxima qualitat en un rànquing extern.
La crítica assenyala un problema real: un gràfic de benchmark no diu automàticament quin model convé per a una aplicació. Però la conclusió inversa —que no hi ha cap millora— tampoc es desprèn d’una sola puntuació.
Els tres productes no competeixen exactament en la mateixa categoria:
- 3.6 Flash és el model general per a programació, multimodalitat i agents;
- 3.5 Flash-Lite prioritza volum, cost i velocitat;
- 3.5 Flash Cyber és una versió especialitzada que Google oferirà només a governs i socis de confiança mitjançant CodeMender en un pilot limitat.
Google confirma que 3.5 Pro continua en proves amb socis. L’absència d’un Pro públic deixa un buit a la part alta del catàleg, però no demostra per si sola que tota la plataforma hagi quedat enrere.
Puntuació igual no significa producte igual
Artificial Analysis assigna 50 punts a Gemini 3.6 Flash en el seu índex d’intel·ligència. En la captura comentada al vídeo, 3.5 Flash apareix al mateix nivell. L’índex combina proves de coneixement, raonament, ciència, codi, eines i context llarg; una mitjana pot quedar igual encara que canviï la distribució interna.
Google publica millores concretes respecte de 3.5 Flash:
- DeepSWE: 49% davant del 37%;
- MLE Bench: 63,9% davant del 49,7%;
- OSWorld-Verified: 83,0% davant del 78,4%;
- GDPval-AA v2: 1.421 davant de 1.349.
Són xifres seleccionades pel fabricant i s’han de validar en entorns independents. Tot i això, mostren per què no és correcte reduir el llançament a «la mateixa puntuació». Una empresa que necessita editar repositoris o controlar interfícies pot valorar una millora específica encara que l’índex global no es mogui.
També pot passar el contrari: un augment de benchmark no garanteix que el model sigui millor en català, en el domini de l’empresa o amb les eines disponibles. Les proves públiques són mostres, no una descripció exhaustiva.
Velocitat: no és l’únic criteri, però sí que importa
El vídeo considera que, un cop s’arriba a uns 60 tokens per segon, una aplicació prefereix qualitat a més velocitat. És cert en algunes tasques individuals i complexes. En altres casos, la velocitat pot ser determinant:
- assistents interactius amb respostes curtes;
- extracció de milers de documents;
- agents que fan molts passos;
- generació de diverses opcions en paral·lel;
- serveis amb molts usuaris simultanis.
Artificial Analysis mesurava aproximadament 235 tokens de sortida per segon per a 3.6 Flash, molt per sobre de la mediana de la seva classe. També registrava un temps fins al primer token d’uns 19 segons. Aquest contrast recorda que velocitat de sortida i latència inicial són mètriques diferents.
El temps real d’una tasca inclou raonament, crides d’eines, execució de codi, xarxa i reparacions. Un model que escriu molt de pressa però necessita tres intents pot acabar més tard que un altre de més lent i precís.
Preu per token i cost per tasca
Gemini 3.6 Flash costa 1,50 dòlars per milió de tokens d’entrada i 7,50 per milió de sortida. Flash-Lite baixa a 0,30 i 2,50 dòlars respectivament. Aquestes tarifes no determinen soles el cost final.
El càlcul útil és:
cost de la tasca = entrada + sortida + memòria cau + eines + repeticions + infraestructura pròpia.
Google afirma que 3.6 Flash utilitza un 17% menys de tokens de sortida que 3.5 Flash en l’índex d’Artificial Analysis i fins a un 65% menys en DeepSWE. Si es manté la taxa d’èxit, aquesta reducció pot compensar una diferència de tarifa.
El vídeo destaca que Grok 4.5 apareix en el gràfic amb més puntuació i menor cost per tasca. És una comparació vàlida dins d’aquella configuració, però no una regla universal. El cost d’Artificial Analysis és una mitjana ponderada dels seus benchmarks, amb preus, memòria cau i tokens generats. Un flux amb vídeo d’una hora, un repositori privat o una resposta de cent paraules pot invertir l’ordre.
La multimodalitat de Gemini és forta, però no exclusiva
El vídeo diu que Gemini és l’única IA capaç d’analitzar nativament vídeo, documents i qualsevol contingut. La formulació és massa absoluta. Altres famílies també accepten combinacions de text, imatge, àudio o vídeo, segons el model i el proveïdor.
La documentació oficial sí confirma una cobertura àmplia per a 3.6 Flash: text, imatge, vídeo, àudio i PDF com a entrada, amb una finestra de més d’un milió de tokens. També integra cerca, execució de codi, funcions, sortida estructurada i ús de l’ordinador.
L’avantatge potencial no és ser l’únic, sinó oferir aquestes modalitats i eines dins del mateix ecosistema amb gran velocitat. Cal provar la qualitat real d’OCR, taules, vídeo llarg i citacions: admetre un format no garanteix interpretar-lo perfectament.
Flash-Lite pot ser el llançament més rellevant
Miguel Torrez considera Flash-Lite el membre més atractiu per a anàlisi a escala. Google anuncia 350 tokens de sortida per segon, un milió de tokens de context i preus baixos. També incorpora diferents nivells de raonament i l’eina d’ús de l’ordinador.
És una combinació adequada per classificar, extreure, resumir i encadenar subagents quan cada error té un cost baix i es pot verificar automàticament. No és necessàriament la millor opció per a una decisió jurídica, una migració crítica o un problema científic difícil.
Que Flash-Lite quedi lluny dels models de frontera en l’índex d’intel·ligència és esperable. La pregunta de producte és si supera el llindar de qualitat del flux a una fracció del cost.
Com triar sense creure cegament cap benchmark
Una avaluació interna petita és més útil que deu rànquings:
- reunir entre 30 i 100 tasques representatives;
- definir criteris automàtics o revisió cega;
- fixar el mateix context, eines i límit de temps;
- repetir cada tasca per mesurar variabilitat;
- registrar èxit, cost total, latència i intervenció humana;
- separar errors tolerables d’errors crítics.
Per a codi, cal executar tests i analitzadors. Per a documents, comprovar camps i cites. Per a agents, mesurar quants passos acaben sense bloqueig. Una puntuació general pot servir per seleccionar candidats, però la decisió final ha de reproduir el treball real.
Conclusió
La prudència del vídeo davant del màrqueting és saludable: Gemini 3.6 Flash no és el model més intel·ligent de tots els rànquings i la rapidesa no compensa qualsevol error. Tampoc es pot concloure que Google només ha generat soroll.
3.6 Flash ofereix guanys específics en codi, ús de l’ordinador i eficiència de tokens, amb una velocitat de sortida excepcional. Flash-Lite apunta a volum i Cyber a un entorn restringit. El buit d’un nou Pro explica que qui busqui màxima qualitat compari alternatives.
No cal «creure» els benchmarks: cal entendre què mesuren, contrastar fonts i executar una prova pròpia. El millor model no és el primer d’una classificació, sinó el que completa la tasca requerida amb el nivell d’error, temps i cost que el producte pot assumir.
Contrast i context
Fonts consultades
-
01
Miguel Torrez Gemini 3.6 Flash: Don’t Believe the Benchmarks
- 02
-
03
Google AI for Developers Gemini 3.6 Flash model documentation
-
04
Artificial Analysis Gemini 3.6 Flash: Intelligence, Performance and Price Analysis
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Hey, what's up, everyone? My name's Miguel. Now, Google just announced the dropping of three new Gemini models. However, it seems that they're in trouble. Let me break it down for you. We've got the newest models released just today, Gemini 3.6 Flash, 3.5 Flash Light, and 3.5 Flash Cyber. Now, this last one is not really a model per se, it's just a version of Gemini Flash more oriented towards cybersecurity. Of course, if we look at the benchmarks, they're going to tell you that they're
-
0:30
, obre el vídeo en una pestanya nova
performing better and faster than the the last generations. That's something that you should expect from every model, but the true question is, are these models really keeping up with the level of performance that we have now, as well as the costs? And that's where it actually gets tricky. So, for information, Gemini has three different versions of the models. You have the smallest one, which is called Flash Light. This model is probably my favorite model out of all of them because it's extremely affordable, and you can use it for a lot of analysis
-
1:01
, obre el vídeo en una pestanya nova
at scale. You see, the one superpower that's left for all of the Gemini models is that it's still the only AI that can really analyze videos, documents, and anything that you throw at it at a native level. So, it really sees and has incredibly powerful OCR capacities. That means that it can actually visualize the things that you that you have in a document perfectly. Now, Pro, the biggest one out of the three
-
1:31
, obre el vídeo en una pestanya nova
was a contestant when we still had Opus 4.5, I believe. Gemini 3.1 Pro was an excellent model, but we haven't had a release of 3.1 3.5 Pro nor 3.6 Pro in a while now. We don't even know where they are, really. And that makes the question, well, how are these two performing compared to the rest of the market? And well, the news are not really that good. So, as you can see right here on this chart, we have the level of performance of 3.5 flash and 3.6 flash
-
2:12
, obre el vídeo en una pestanya nova
at exactly the same level. So, that means that these model, according to artificial analysis, the source where this graph comes from, doesn't really have a big improvement in terms of performance. Now, not only that, but of course, Gemini 3.5 flashlight, which is the lower end of the models, is still sitting way, way behind in the charts. That's normal. It's a slower model. It's a smaller model. So, it actually makes sense for it to be there. But, where things get actually worrisome is whenever you actually look in detail
-
2:49
, obre el vídeo en una pestanya nova
at what's happening in the back. So, as you can see right here, we have the benchmarks of the state-of-the-art models, Chimi K3, GLM, Fable, Soul, etc. And this is where we see where Gemini actually stands. So, on the intelligence level, we see that 3.6 flash hits a 50, score of 50. However, whenever you actually reach the higher end of the models, so we're going to call this the top five, Fable, say Fable
-
3:23
, obre el vídeo en una pestanya nova
Soul, uh Chimi K3, Grok, and GLM, we see that we have a 10% to 20% increase in levels of intelligence. Of course, Gemini 3.6 flash is still the fastest one out of all of the models, but if you're building applications, really you don't really care about having the fastest model. You care about the best quality at around 60 tokens per second. Now, something that also is quite interesting to see is the cost per task. So, the cost per task is what dictates how much
-
4:00
, obre el vídeo en una pestanya nova
it costs you to actually get stuff done using your AI. And this is where we can actually see that we have models such as Grok 4.5, which actually cost about 40% less and have a higher level of of intelligence. So, for reference, the level of intelligence of Grok 4.5 is very similar to that of Opus 4.8. So, we literally have a model that only performs better, but it's actually also cheaper. So, that begs the question,
-
4:34
, obre el vídeo en una pestanya nova
what do we really do with the Gemini models? Right now, Google has fallen severely behind in the AI race, despite being the company that actually has access to all of the resources that are necessary to keep going. So, the hardware, the software, and the data. We really hope that Gemini 3.5 Pro or 3.6 Pro are going to change this, but things are not looking very good for Google right now. Now, I'm going to say, I hope that you enjoyed this video. I'm not going to recommend the 3.6 model
-
5:11
, obre el vídeo en una pestanya nova
because I firmly believe that you have better options elsewhere via Grok. And if you want to just have your own open-source model, you can also use GLM. It's perfectly fine. Uh using flashlight is good, but even then, nothing has really changed. I hope that you enjoyed this video. I'm going to say, I'll catch you in the next one. Remember stay up to all of the latest AI news. In this case, it was just a bunch of noise. And I'll see you in the next one. See you.