Kimi K3 és millor que Opus 4.8? Deu proves mostren una resposta molt menys simple
Pat Simmons compara Kimi K3 amb GPT-5.6 Sol, Claude Opus 4.8, GLM 5.2 i Kimi K2.7 en deu tasques de codi i treball creatiu. K3 guanya en 3D i gràfics animats, però falla en controls, àudio i textos comercials. És un salt enorme respecte de K2.7, no un vencedor universal.
Kimi K3 arriba amb una promesa ambiciosa: rendiment de frontera, pesos oberts, un milió de tokens de context i un preu inferior al dels models propietaris més cars. Pat Simmons decideix comprovar-ho amb deu tasques que van des d’aplicacions tridimensionals fins a una biografia, una pàgina comercial i un tràiler.
La resposta curta al títol del vídeo és que Kimi K3 no supera Claude Opus 4.8 en tot, però sí que el pot guanyar en treballs concrets i s’acosta molt més a la frontera del que suggereix l’etiqueta de model obert.
En les proves de Simmons, K3 és especialment bo generant 3D i gràfics animats. Queda més enrere quan la funcionalitat importa més que l’aparença, i produeix una de les pitjors versions del text comercial. El salt respecte de Kimi K2.7 és visible, però la classificació canvia a cada tasca.
Què és Kimi K3
La documentació oficial de Kimi situa el llançament el 16 de juliol del 2026 i el descriu com el model més capaç de la companyia. Moonshot anuncia 2,8 bilions de paràmetres, arquitectura híbrida amb atenció lineal, visió nativa i una finestra de context de fins a 1.048.576 tokens.
El model està orientat a programació de llarga durada i treball del coneixement. Funciona sempre amb raonament activat i permet seleccionar nivells d’esforç. La companyia també el presenta com a codi obert; en la pràctica, la possibilitat de desplegar un model d’aquesta mida queda fora de l’abast d’un ordinador domèstic normal i la majoria d’usuaris l’accediran per API o servei allotjat.
Kimi mostra resultats propis molt alts en proves de codi i agents. Simmons els ensenya, però adverteix que no coneix bé diversos dels benchmarks i que els laboratoris oberts tenen incentius per optimitzar les proves que publiquen. Per això prefereix mirar artefactes concrets.
Com es va fer la comparació
El creador posa Kimi K3 al costat de quatre models:
- GPT-5.6 Sol;
- Claude Opus 4.8;
- GLM 5.2;
- Kimi K2.7.
En alguna prova també inclou Claude Fable 5, però no disposa de prou crèdits per utilitzar-lo sempre. Llença els treballs en paral·lel des de Claude Code mitjançant una habilitat pròpia. GLM i K2.7 passen per OpenRouter; GPT i Opus utilitzen les seves subscripcions, i K3 s’executa primer amb la subscripció de Kimi Code i després amb una clau d’API.
Els prompts i resultats es publiquen a SimmonsBench. El nom del model queda ocult mentre Simmons revisa cada generació i es revela després. Aquest petit encegament redueix el biaix de marca, però no converteix la prova en un experiment científic.
No tots els models apareixen a totes les rondes, només hi ha una generació per tasca i la puntuació la decideix una sola persona. A més, alguns resultats es valoren per aparença encara que el prompt exigeixi funcionalitat tècnica.
Prova 1: el centre de comandament 3D
La primera tasca demana un dispositiu tridimensional robust, visible des de qualsevol angle i amb controls operables dins del navegador.
Kimi K3 obté el primer lloc. Genera un objecte amb bona il·luminació, superfícies convincents, botons, retroil·luminació i un procés d’encesa. Simmons considera que és la versió més realista i interactiva.
GPT-5.6 Sol també fa una bona composició, però K3 destaca justament en el camp que Moonshot havia ensenyat al vídeo de llançament: geometria i interfícies 3D generades en una sola passada.
La prova, però, revela petits defectes en tots els models: text invertit, dials difícils de moure i controls amb efectes poc clars. «Visualment impressionant» no és sinònim de producte acabat.
Prova 2: una aplicació meteorològica
El prompt no demana només consumir una API. Vol una aplicació per a Amèrica del Nord que obtingui dades reals i produeixi una previsió pròpia a curt termini.
K3 crea una interfície utilitzable, amb gràfics i cerca de ciutats, però no és la preferida. Simmons penalitza el disseny de degradat blau i alguns elements trencats. GPT-5.6 Sol ofereix la presentació més polida i Opus queda a prop.
La limitació més important és metodològica: la revisió se centra en colors, animació i si els gràfics es veuen bé. No valida la qualitat del model meteorològic, la font de dades ni l’error de previsió. Per tant, aquesta ronda mesura més frontend que meteorologia.
Prova 3: convertir MP3 a MIDI
La tercera aplicació ha de detectar notes, altura, temps i durada d’un fitxer d’àudio, generar MIDI i permetre comparar les dues reproduccions.
K3 queda últim perquè la seva versió no troba correctament el fitxer o no permet completar el flux. Altres models carreguen l’àudio i mostren una visualització semblant a un piano roll.
Simmons reconeix que no sap jutjar la precisió musical. Valora interactivitat i aspecte, però no compara les notes detectades amb una transcripció de referència. Per això tampoc es pot concloure que el model guanyador faci una conversió correcta; només que presenta millor la demostració.
És una bona lliçó per a qualsevol avaluació: si no es pot mesurar el requisit central, la prova acaba puntuant allò que és més fàcil de veure.
Prova 4: un joc de lluita en 2D
Els models reben la instrucció de crear un joc jugable al navegador amb dos combatents, barres de salut, moviments, cops i combos.
K3 queda segon. Produeix un joc que funciona, amb selecció de personatge i controls, i supera clarament K2.7. El guanyador ofereix millor sensació de combat i respostes més coherents.
La ronda mostra una de les conclusions més consistents del vídeo: K3 és molt superior al seu predecessor. La diferència no és un detall estilístic; passa d’una interfície confusa i difícil d’iniciar a una experiència que es pot jugar.
Prova 5: la pàgina d’un rellotge de luxe
Aquí es combinen disseny web i un rellotge fotorealista en 3D que l’usuari ha de poder girar. Els millors resultats separen dues habilitats: modelar l’objecte i construir una pàgina comercial coherent.
K3 no aconsegueix la millor versió. Alguns rivals creen un rellotge més creïble o una composició de producte millor resolta. Simmons situa el model lluny del primer lloc en aquesta ronda.
El contrast amb el centre de comandament és revelador. Guanyar una tasca 3D no garanteix transferència perfecta a qualsevol objecte. Geometria, materials, proporcions i context visual canvien molt entre una consola de ciència-ficció i un rellotge.
Prova 6: una rèplica jugable de Counter-Strike
La sisena prova demana un shooter tridimensional al navegador amb moviment, armes i enemics. K3 produeix gràfics atractius i sembla prometedor abans de jugar.
Quan Simmons prova els controls, troba tecles que mouen en direccions inesperades, velocitat estranya i una arma poc convincent. K3 queda quart. Fable i GPT-5.6 ofereixen experiències més sòlides, i Opus també el supera.
És probablement el resultat més útil de la part de codi. Un agent pot generar una captura espectacular i, alhora, fallar la lògica bàsica. La prova funcional ha d’incloure moviment, col·lisions, recàrrega, dany, pausa i reinici, no només una inspecció visual.
Prova 7: una introducció biogràfica
La primera tasca de coneixement demana una obertura biogràfica sobre el mateix Pat Simmons. El model ha de cercar informació perquè el prompt no aporta dades personals.
GLM 5.2 queda primer i K3 segon. Els models oberts ocupen les tres primeres posicions, mentre GPT-5.6 i Opus queden al darrere segons la preferència de Simmons.
K3 escriu un text més pròxim a una biografia que a un perfil corporatiu i troba informació rellevant. Tot i així, una bona prosa no prova que cada detall sigui correcte. La recerca sobre una persona necessita cites, comprovació de dates i separació entre informació pública i inferències del model.
També és una prova molt subjectiva: una altra persona podria preferir un estil diferent sense que cap text fos factualment millor.
Prova 8: gràfics animats sobre Kimi
Els models han de crear un vídeo explicatiu amb gràfics en moviment sobre Kimi K3. K3 guanya amb transicions, ritme i una composició més cinètica.
És el segon triomf clar del model i reforça el patró de la primera ronda. K3 sembla especialment fort quan ha de combinar codi, SVG, moviment i sensibilitat visual.
Algunes versions inclouen dades de mostra o espais reservats. Abans de publicar un vídeo així cal substituir totes les xifres, comprovar els gràfics i revisar que la narració no converteixi afirmacions promocionals de Kimi en fets independents.
Prova 9: la pàgina d’un curs d’IA
Simmons utilitza un projecte real: la pàgina del seu curs intensiu de quatre setmanes. Demana un redisseny i un text més directe per a fundadors i executius, amb èmfasi en problemes concrets i sense sonar excessivament comercial.
K3 queda cinquè. Genera un titular i un subtítol massa llargs, repeteix fórmules típiques de text d’IA i no capta bé el matís que el creador buscava. GPT-5.6 Sol és la seva opció preferida i K2.7 sorprèn amb una versió millor que K3.
Aquest resultat recorda que una versió nova no domina necessàriament l’anterior en cada estil. L’ajust de dades, la longitud de raonament i les preferències de redacció poden fer que un model menys capaç produeixi un text més adequat.
Prova 10: un tràiler de trenta segons
L’última tasca demana una llista completa de plans per a un tràiler cinematogràfic i després utilitza un model de vídeo per generar les escenes.
GLM 5.2 queda primer i K3 tercer. Els resultats són espectaculars per moments, però alguns clips introdueixen personatges aleatoris o perden la continuïtat. És difícil saber quina part de l’error prové del guió i quina del generador de vídeo.
Quan una prova encadena dos models, la puntuació final barreja les capacitats de tots dos. Per avaluar només el model de text caldria puntuar la llista de plans abans de generar el vídeo i aplicar el mateix procés a cada proposta.
Els límits d’ús van ser un resultat en si mateix
Simmons comença amb un pla mensual de Kimi de 19 dòlars i afirma que només completa dues construccions abans d’arribar al límit temporal. Puja al pla de 39 dòlars, aconsegueix unes quantes generacions més i torna a topar amb la quota. Finalment connecta una clau d’API i paga per ús.
Calcula que la jornada li costa entre 65 i 70 dòlars entre plans i consum addicional. És una experiència individual: l’ús pot variar segons el context, l’esforç de raonament, la mida de les respostes i els límits que Kimi modifiqui.
La tarifa oficial de l’API publicada després del vídeo és de 0,30 dòlars per milió de tokens d’entrada en memòria cau, 3 dòlars per entrada nova i 15 dòlars per sortida. Un context d’un milió de tokens és una capacitat, no una recomanació: omplir-lo i generar respostes llargues pot ser car.
Per què aquest benchmark no corona un guanyador
La prova és útil perquè mostra resultats reals, però té limitacions:
- una sola mostra per prompt;
- valoració d’una sola persona;
- absència d’una rúbrica definida abans de veure els resultats;
- models que falten en algunes rondes;
- accés per serveis i subscripcions diferents;
- cap normalització per cost, tokens o temps;
- funcions centrals que no sempre es comproven;
- tasques creatives amb molta subjectivitat.
Per a una selecció empresarial caldria repetir cada prompt, executar tests automàtics, registrar el cost complet i puntuar cegament amb criteris predefinits. El conjunt també hauria d’incloure tasques pròpies de l’organització, no només demostracions vistoses.
El veredicte
Kimi K3 no és «millor que Opus 4.8» sense afegir «per a què». En aquesta comparació guanya el gadget 3D i els gràfics animats, queda molt bé en la biografia i el joc de lluita, i falla en la conversió d’àudio, els controls del shooter i la pàgina comercial.
La conclusió més sòlida és el salt respecte de K2.7. K3 entra de ple en converses on abans només apareixien models propietaris de frontera. El seu caràcter obert, el context llarg i el preu poden convertir-lo en una opció atractiva per a agents i generació visual.
La decisió pràctica no és substituir tot l’stack perquè ha guanyat dues demos. És donar-li un grup de tasques ben definit, mesurar qualitat i cost, i mantenir un altre model per als casos on els tests mostrin que encara falla. El vídeo no troba un rei universal; troba un competidor que ja mereix una prova seriosa.
Contrast i context
Fonts consultades
-
01
Pat Simmons Kimi K3 Is Here! (Better Than Opus 4.8?)
-
02
SimmonsBench Kimi K3 generations and prompts
- 03
-
04
Kimi Kimi K3 API pricing
-
05
Kimi API model selection
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Alright, well the model releases keep on coming. This time out of the open source world, Moonshot AI has just dropped KimeK3. And on the benchmarks, it's beating its open source competitors and even going totor toe with some of the best front-shar models out there. But you know this is an eventchmarked channel. So in this video, we're putting KimeK3 through the ringer with nine builds. Everything from games and 3D apps to slide decks and creative writing. But we've even got a Quentin Tarantino Directed Movie trailer. Yeah, we're getting a little crazy with it. So by the end of this video, you will see exactly where KimeK3 lines. up against the frontier models. An answer to the question, is this the new open source king? So let's find out, getting right into the builds, we're gonna kick off our fan out of a bunch of different building agents to start running through these tests. I'll just see you in a minute. coating folder here, and then I'm gonna do all of this in-clod code just because it's the easiest way to run all of these across all of these different models. And then I just have this command here that I'm gonna paste in. And this is just telling Claude to across a bunch of different models run these builds. And I really wanted to put this to the test across front here and these open source models. So we're gonna be testing KMK3 against GLM52 against KMK27, it's predecessor, and then Opus48 because I'm really curious. how it holds up against Opus and then GPT-56 soul as well. And then we may do a couple of fabel tests. I'm really running out of usage credits quick on Fable. I don't know how to pay extra for those. So maybe for a couple we'll see how it pairs to Fable as well. I wish I had more usage credits, but well at least get a sense of how it compares to the other frontier models as well. And like I've done previously, I will include this. fan out skill and a second you'll see six item windows pop up running all of these different builds but the way I'm accessing these models is through first open router where available so probably familiar with open router if you've watched any of my videos you see me using this in the past so we're using the open router API key for gil and 5.2 and kmk27 I'm using my subscription for gpfff6 and opus and then on the k3 side this is not available yet in open router so what you need to do is either go to Can be.com and you can interact with K3 here. You can see it's available in the model picker. You'll definitely need to sign up for a subscription. You'll probably get like two generations before you'll run out of the free credits to sign up for a subscription. You just go here. You go upgrade plan. You can see I'm on the model rado right now. Monthly. I'm assuming I'm going to hit usage limits already though. So I might upgrade for this demo. But that's how you do it in. Kame.com and then what you can do as well, this is what I'm doing as I'm using the CLI because I'd prefer to have access to all of these in terminal. So what I'm doing is Kame.com slash code and you can actually go to console and once you sign up for a subscription plan, you can just use your subscription right here. And then all you need to do is just run a bash command. It's just exclamation, Kame login, and then it will just run the authentication, verify, and you're in. And these items popped up on my screen, so we are in progress on all of these different builds. We're going to start with code. then we're going to move into knowledge work. So I'm doing six AI builds here. One is a 3D gadget. I'm going to explain all of these and I'll show you the prompt here in a second. Once these builds complete But that's the first one. The second one is a weather simulator. The third is an MP3 to MIDI visualizer. A fourth is a fighting game. The fifth is Counter Strike 2 demo. Yep. We're running that one back. And then the sixth is a landing page for a luxury watch product. Okay. So all of these have completed. And we're going to go to Simmons Venture in a second. It'll look at these. However, what I noticed is already after only two builds. Just refresh this. Himgy K3 was only able to get two builds done with my moderator, $19 a month subscription plan. And I'm already hit my usage and it resets in two hours. And then I'm at 20% usage. So that is interesting. I mean, they have a lot of different subscription plans. So they probably just try to get you in on this cheap one. I don't know if I'm gonna wait two hours or just upgrade my plan. But Kimmy has a daily driver. I'm already realizing that you probably need, I guess something like the Allegro to regularly using this.
-
3:36
, obre el vídeo en una pestanya nova
I gasped on it. Okay, so I just upgraded to the $39 month plan. Let's see how much usage that will get me. But this would be a good time to say, please like this video, subscribe to the channel if you haven't already. It helps me out a ton and may help in these slightest of ways subsidize some of these AI subscription costs that I'm paying hundreds of dollars a month for. So thank you for doing that. And then let's get back to the builds. And while we wait on those K3 builds to finish up, let's just quickly talk about the benchmarks that promises to be really quick. Just so we can draw some sort of comparisons. You know these open source labs are notorious for benchmarksing. So take this other grain of salts. Hopefully the test that we run will be more indicative of this than the actual benchmarks, but here's where we're looking at. First coding deep sweet can be k coming in third behind Fable 5 and G556 soul. Way ahead of OPEC48, way ahead of GLM 5.2 which is interesting. Frontier sweet, if it coming in seconds, only behind. Fable 5 and a huge jump above 5 6 soul. Then on Kimmy code bench 2.0 which wasn't familiar with that one. They are certainly not bench maxing their own benchmarks. It comes in second. You know, you gotta head to him. They could have just ranked themselves in first behind at Fable 5, terminal bench 2 1, PED of opus and fable and only behind soul. Rogue RAM bench not familiar with that one, but it's at the top and then sweet marathon enough from later. We marathon either, but it's also at the top. I mean just from the coding benchmarks. This is This is not only just crushing Opus for 8, but also giving Fable and 5, 6 Soul Run for the Money. With an Opus Source model, we will see, though, like I said, benchmarks only tell us so much. And then we have General Agents, not familiar with most of these billets we'll just talk about these quickly. I do know GDPVAL, though. This is more for knowledge work. You can see KMK3 coming in third here. There's quite a bit of a differential between Fable 5 and 5, 6 Soul and KMK. There. And then we have a bunch of these other ones not really familiar with these and you can see it coming in at the top with a lot of these. Then visuals too, nothing familiar with this. I don't even know why I'm reading this a lot. I should not be trusted with naming off end of these benchmarks. So take these with the grand assaults. I don't know what I'm reading. I literally just read these a lot. So people can see how it compares. All I care about though is getting back to the build. So that's enough benchmark talk. Let's check in on K3 and see how it's doing. Alright, I just hate the rate limits once again. We got about four more builds done, so it's still not quite done, but I am just going to preview the builds that have completed with K3 and the other models. We'll do a comparison. I have a couple others running. What I did, I just switched to API key, so I'm just running through their CLI with the API key now. and I'm just gonna pay for each generation because I cannot upgrade again to the $100 month plan. But with that, let's go do Simmons mentioning at least compare some of these builds just to start. So we've got, couple of these ready to go. First up, build an interactive 3 command center gadget. Reminder that all of these prompts will be available. in a separate blog post and you'll be able to interact with and see all of these generations live. But let me just quickly gloss over this, build me a visually stunning fully interactive 3D hardware gadget that runs entirely in the browser, a rugged desktop command center device, I can look at from any angle and actually operate. The idea here is in their demos in that little trailer that they had. They were doing a lot of good 3D generations. So I wanted to put it to the test like this is where a lot of these will see a couple more of these where 3D is a big part of it because they emphasize this so much in their trailer. So I really wanted to see how it stacks up against the other models. And just like we've done in the past, I'm going to look at each one of these generations and then reveal which model is which. So let's take a look at this first one.
-
7:13
, obre el vídeo en una pestanya nova
Click in here. Okay, click this. Oh, look at that. All right. I mean, not bad. Oh, all right. Yeah, kind of hard to text a little backwards. But we got some nice dynamic lighting here. This, this is a little off. This little dial here. Yeah. It's all over the place. Kind of turns around. I'm going to lick the glare. Okay, not bad. Not bad. Take a look at this one. Okay, a little glitchy. You can at least read the text. Okay. You can change this even though it's completely wrong. Click these. Okay. Click these, but it's just all over the place. Okay, power's on power's up. Oh, look at nice little power on can't really see it. All right, that one's uh, that one's a little bit better. Even though it's not great third one. All right, we can read it nice. The good is a good sign. Okay. Okay. Okay. Oh, okay. All right, I got it. So I just need to go up. All right, that's really good. That's really realistic. I got to go power. Oh, like that. Nice, I'm going to mess up boot up. Okay. This one's good. Look at that lighting. Ooh, look at this and see if we got sound effects backlight. Net. I don't know what that does. That's a little weird. Since, I mean, this one's the best one for And then we got this last one. Okay. All right. Yeah, readable. Very similar. Okay. Nice. Dial works kind of kind of okay kind of hard to move this. No sound effects unfortunately, but does something. Again, nice little backlight. Again, nice little mode power. Nice little power on. All right. That one's definitely number two. All right. Let's see. Is this Kimmy? No. All right. Kimmy K3 Auto Good Star. Opus 48. No. GpT56 Soul. All right. Wow, all right, all right, we're cooking. Okay, I was for it. And then Kimmy K27. All right, all right, Kimmy K3, you have officially got my attention. Next up, we have build a weather app with its own forecast model. Quickly gloss over the prompts, but it's build me officially stunning weather forecasting web app for North America, pulls real data and produces its own short term forecast, not just a wrapper requirements to do to do to do not read all of these. Again, all these prompts will be available to you in the post below. Let's take a look. Alright, number one. Okay, we just have a search, so Toronto. Okay, pulling live weather. Oh, okay, you're nice little, nice little interface. I don't love the emojis, but in this is kind of broken, but it's some good colors here. Some good design. Get the little sun in the background. You know, it's not bad. It's not bad. Okay. That's number one. Number two. This already looks way too. AI looking. Okay, we're in North America. Okay, Denver. Alright, yeah, yeah, it's not bad either. I don't like the blue gradient, but the chart looks better. These charts are actually not broken. And then we get this nice little thing here. Okay, pretty good. I would say, this one's better than this one. Okay, and then it's like this. Alright, I just, every time I see blue now, I just, I'm triggered, but I mean, this interface looks pretty good. Chart kind of broken. That looks pretty good. And then we can click Chicago show up. Alright, yeah. Um, pretty good. Okay, I would say that was best yet. And then we have this classic cream. Somebody in the comments on my five six soul videos said background cream colors the new purple gradient and I could not agree more. Now I can't unsee that now. So let's do a new gift location permissions. That's fine. Let's do Mexico City. Okay, weird icons. This is actually a good test because none of these really did all that well. Okay, we have some nice one moving SVG. This one's definitely the best, but you can see it's still kind of broken. It's very readable. Char... Okay, it would have been nice to have like a little hover state on this chart here, but that's fine.
-
10:49
, obre el vídeo en una pestanya nova
same with this one but at least it's like not broken. Yeah, okay, so let's see. I want to see like another one. York. Yeah, okay. Yeah, it's nice like move. There's actually some animation with each city. Okay. Yeah, I mean this one is definitely the best K3 again. No way. 5 6 all. Okay. This is my second favorite. Oops for it. Okay. Actually this one is pretty good. Is this K3 K27? All right. K3 down. All right. All right. Fair enough. So here's here's K3. You know, like this wasn't bad. The chart was actually working, but I just do not like these friggin' blue gradients. So I'd say five-six-old, one-opest two. I might even say I prefer K-27 though, just because it. even though I kind of had this nice little background. So there is building a weather app. Next up, we've got an MP3 to MIDI converter. And I'll be honest, I don't even know what this one is. This was an idea that someone left in the comments from my five six-sold video. I believe, which thank you everyone for dropping in those comments. Please, if you have any ideas for future model release demos like this and test, please let me know in the comments. I really appreciate it. And it gives me a time more ideas. I'm really like stressed this because I'm constantly running on out of ideas as these models get better. Anyway, so an MP3 to MIDI converter. Here's the prompt. Build me a browser app that converts an audio file into a MIDI file by actually detecting the notes. Pitch timing and duration and let's me play both back to compare. Okay, let's see. Okay, I need to actually load and I'm P3. I think so. I think there was a demo. I think we gave it some sort of MB3. Oh, it's Tone-Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Cone, Anyway, all right, load input and all this in pitch. All right, so I don't know anything about pitch or music notes or anything like that. So I'm just going to judge this based on the interactivity and how well it visually shows this. So, okay, play original, okay, play MIDI. All right, you can see pretty cool. I'm sure people watching this note what's going on. I have no idea, but it looks like these are piano notes. Somebody please tell me what's going on in the comments. I know someone's watching that understands this, but I guess we can increase the sensitivity. I don't think that didn't anything. BPM. Okay. It's pretty cool. Pretty cool visualizer. Next up we have this one looks very similar. Load input, very similar actually. Okay. I'm loading audio, advanced detection settings, notes. Okay. Original. This is very similar. Look, the MIDI guy kind of just stops. Okay. This one's... a little bit more broken, very similar. Look to the last one, maybe this one as well. Okay. Can you just upload, can you figure out how to upload this audio file once I'm Alright, I can't find the file, so I'm just gonna say didn't do a good job. So these ones were able to load the input. So this one was just last one. The current background again. The new purple gradient looks great. And so we have the input.3. Let's be cool. Okay, it doesn't highly detect the notes. I don't know if I liked how it did in those last builds, but yeah, okay, I mean, all of these are relatively straightforward. I'm gonna guess this is five six sold. This is five six sold written all over it. I'll only have that. This is absolutely gonna become the new. Purple gradient. Okay, so that's five six soul and then this one is last. Can we K3 darn it? You hate to see it. And then we got what? Opus? Nope, can we K27? Opus? Like these don't look so much exactly the same. All right, might have to move away from that test. But next up, we have a fun one. Build a 2D fighting game. Here is the prompt building a playable 2D fighting game in the browser. Thanks Street Fighter to still do two fighters. I help bar for each. Real moves and combos. Then we have a bunch of requirements. Again, prompts are all in the description below. Well look at this first one. Okay, I'll already some text not aligned properly, but okay, with select characters, that's pretty cool. So player one and then computer, okay, two player, oh, interesting, okay, let's go against the computer. Volt versus blaze fight. Okay, I don't know how to punch. Jump, S crouch, V block, here's way better than me. Okay, I just got destroyed. I slowly figured this out. Okay, can I pause this? My goodness, okay. W jump, S crouch, F fight, G heavy.
-
14:26
, obre el vídeo en una pestanya nova
B block. Alright, we're gonna have to move through this quick, but okay, jump F fight. Okay, heavy. Yeah, sounds good. Pretty fun. Pretty fun. Okay, next one. Play over to CPU. Nice. We got some nice little sound effects. Alright, I like this way better. S-Cross J. Light K-Heavy. Okay, this is kind of weird. Geez. Not a gamer. I'm not a gamer. Alright. One more time. Kenji jump, S crouch. Jay, Blight, Kay, heavy. There's no one in the contact one time here. I came to move, okay, jump. There we go. There we go. All right, this one's a little weird. It's like rock, I'm sock, I'm robots kind of. Okay, number three, this looks like another five to six. I bet this is a five to six generation again. Just the fonts of this background, okay. Weird characters, really weird characters, soul. Oh, Wonder what model this is. Gotta be kidding me. Okay, look, can I have a select these? Choose Fighter. I can't even select this. This is crazy. Oh wait. Okay, that's confusing. Okay, P1, CPU. All right. That's a confusing thing. Okay, all right. Don't love this. Don't love this at all. Wey slower. Okay. Don't love that one at all. All right. And then this one, all right. Choose what is going on these are all over the place. Layer 1 select. She's looking at these icons. This is terrible. That'll be Kimming K27. Okay. Can you select this? How do I play her one select? Enter? Nope. This is bad. Okay, that's terrible. All right, that's the worst one. 2-7 and then I would say this one was second-alice. So, like, that did give way. This one is good and then this one... Yeah, this one's definitely most realistic. So I'm going to say second, chemicated three, alright there you go, and then I'll open it for eight. Alright, fair enough, fair enough. Chemicated three, pretty good. I'm going to, you, you just see the jump between, like first of all, it's better than sole here. It's not going to fall. Look at the jump between two seven and three. I mean, it's frid in night and day, like that is a crazy upgrade. Okay, next up. I wanted to do some combination of a landing page and 3D. So we're building this luxury watch landing page, photorealistic rotatable 3D watch of the hero that looks like a ship from a top. So we're gonna take a look at all of these. Okay, this year, this is done. Now, first up, very good 3D. Wow. That's pretty incredible. The alignment can be fixed. I actually like this one. Yeah. Yeah, the actual web design can be used a little bit of work, but I mean, the 3D work is pretty frigging incredible, cheese. I've done some like generations with Fable, which I know I promised that we would. Do a couple of those. I don't know if we're gonna get to them, but this is absolutely fable level. I don't know if this K3, but this is pretty darn good. Okay, next up, Jesus Christ. Okay, yeah, not bad. I mean, the fact that it's just rotating like this is very weird. I do like the dynamic lighting though. Looks good. It doesn't really look like a watch, laying a rest of the landing page. I can actually zoom in. Alright, that's that one. Next one kind of weird. Okay, it's not really weird. rotating properly it does look good if it's rotated properly looks a little low reds too for some reason Again, these models have the dynamic lighting down which is pretty impressive wasn't the case a few months ago. All right, not bad this one. Oh, very nice. Very nice. Okay. Yeah, actually like this cuz it's nice. Okay, kind of weird. I'd wait, but yeah, all right this one between this one and this one kind of like this one better even though the placements are on this one looks a little more realistic
-
18:03
, obre el vídeo en una pestanya nova
People thought oh, I forgot I did film five care that's well Okay, I guess I did have couple of people five's in here alright, film five. Well, there you go. I said this looks a lot like film five and it's film five is this K3? Opus for it interesting. Okay. This has to be K3. Oh five six soul and then K3 interesting interesting K3 didn't do a good job of the watch on the watch test hates it and I forgot to do Kmk27 here, but it probably would have been terrible I'm realizing too. I forgot to include GLM 5 2 and a couple of these my bad my bad For these next builds and we get it in a knowledge work I'll make sure we include GLM5 2. I'm all over the place, but at least we can see how it compares to the frontiers for a lot of these tests. Alright, last up Counter-Strike 2 demo. Counter-Strike 2 demo, you might already be familiar with this one. I have rented a couple times with Fable in 5-6, so if you've seen those videos you will recognize. Fable in 5-6 generations, but we're still going to compare it to K3 into how it does. This is the prompt, build me visually stunning, genuinely playable browser clone of Counter Strike 2, and we just go on to name all these. But let's take a look at first one, sensitivity. Okay, that's a nice touch. Volume. That's new. So this... Might be K3? Oh, nice to take damage. That's okay. Right now. I don't have to reload how to reload. Okay. R's reload. Wow. This is very good. Nice. Nice. Yeah. Okay. And then I think you've seen the rest of these. This. Okay. I think this is. And now I can't remember that might be 5-6-Soul. This was Fable's version, and then this, I'm gonna assume, oh, but no, this was 5-6-Soul. This was 5-6-Soul. There it is. So, okay, this one wasn't bad. Was this, this GLM? Okay, K3, okay. Okay, okay, open-s for it. All right, I had never seen open-s for it. Open-s for it. Open-s for it, it's one. It was really good. I thought this was KMK3. I still think that KMK3, okay, let's actually play this. Give KMK3 a chance here. Okay, terrorist. Employee. Yeah, I mean, the countdown's a little long. W is not going forward either, that's weird. So S is going forward. Okay, that's a weird, that's a weird weapon. You're a little slow too. I am very confused on how to go forward. I'm pressing W now. It's going left. I'm pressing S is going right now. All right, there we go. Okay, this is weird. This is weird controls. So not great from KMK 3. It definitely says third and then I'm sorry. Fable is definitely number one. Still five, six, number two. Opus number three, KMK. 3, number 4. Don't you hate to see it? I thought Kimi would be doing a better job. The graphics look great. The actual shooting not great. The physics and controls is not great. Would it like to see this test on GLM 5.2? My apologies. Yet again, just start rip through these way too fast. Next time I will make sure we don't miss these, but that is counter strike too. Alright, so that concludes coding. And I'm gonna be honest, Kimi at least for these tests, K3. Do I dare I say a little underwhelming? Still a huge jump from K27 and it's still right there. It's still competing with the frontier models. I'll give it that. I'll give it that. Next we're gonna do is put it to the test with some knowledge work and I used knowledge work lightly. I didn't want to do it. We've been doing this in the past where I've been doing copyrighting and LinkedIn posts and stuff like that. I didn't want to do that. So this time where you're going to do kind of a blend of knowledge working coding. We'll have some motion grafting designs, some creative writing, some landing page writing and like a teased in the intro, a quintant internet trailer. Those builds are just wrapping up as well. By the way, I've been paying it. If you had cost this whole time, it'd probably spend another $25 on top of my subscription. So I'm not about 65, 70 bucks on these three generations today. So let's hit a refresh and take a look at some of these and all of the work outputs. First up, we have an Isaacson style biography intro. So I'm asking it to write an opening paragraph of a biography of myself in the style of Walter Isaacson, if you're not familiar with Walter Isaacson, or the recent Elon Musk.
-
21:39
, obre el vídeo en una pestanya nova
by a forced to fame a Steve Jobs bio, he is an all-time writer, and so we're gonna see how these models compare. We're gonna quickly breeze through these because there's a lot of writing, but let's take a look. Ordinary morning in 2022, freelance marketer in New York, but in a little chatbot called GPT3 and found himself as a later recalled laughing at its hallucinations. So by the way, too, I didn't give it any information about me. All I said was Pat Simmons. So it had to go online to research me, find, I don't know. videos or my website, whatever, and actually put this together. So I mean, the fact that I didn't even found this pretty impressive and then yeah, it's already signing it sources. He was an unlikely guy to the new age and he said so himself. No technical background, no formal training and the arcane arts of machine learning. Nice. Nice. Okay. Number two. That wasn't bad. Like I said, writing is it's so subjective that it is tough, but let's let's just say I wanted to try this new kind of unique test here. So next one. Pat Simmons filters the future through the lens of the non-technical operator. He does not lecture about algorithms, instead he sets two rival models to the same bill than let's the artifacts speak for themselves. So this one actually probably found a YouTube video and pulled that transcript to know about that. The channel stated charter promises non-technical AI tools and workflows that can turn a knowledge worker into the AI go-to person on a team and Simmons delivers, okay, this feels more like, I don't know. Forbes article or something. Okay, number three. Pat Simmons is not build the machines he judges. Yeah, um, no kidding. And he has made that limitation his authority rather to the topology. Okay. Little A, I sounding there, but how many people are actually building machines to kind of a weird way to say that. Second sentence. He came to artificial intelligence sideways out of marketing rather than any laboratory. It's like, okay, a freelancer buys on account would follow to board the stamp of the man who trusted his own eyes over the promotional literature. Doesn't seem like this one really understands AI careers, but okay. And a season when artificial intelligence was being sold to the public through leaderboard scores in triumphant benchmarks, Pat someone's built a small, contrary and stage for the opposite idea. Working from the United States with the underhury manner of a practitioner rather than a prophet, this is actually not bad. I like that one. I'm gonna say that one so far as number one. This one, just annoying copied and love it. Let's look at the last one and I will determine. On July 10th, 2026, Pat Simmons opened a 41 minute YouTube review by setting a three new AI models to work, and then he kept going 10 builds. Okay, so we just looked at one YouTube transcript and wrote this, and tried to make it sound like a story. Good for mounting though, that looks cool. All right, I'm gonna say this one was just annoying, so I'm gonna say it's last. This one, because it's more of like a biography, I would say number two, three, this one number four. All right, let's see, which one, what's number one? Okay, GLM 5.2, interesting. Number two, Kamek-3, look at that. There we go. Number three, Kamek-2-7, interesting. Number four, five six-soul, and then number five-opus, interesting. This, by the way, is an interesting finding. You can see how good these open source models are at writing. GLM-5-2, which I've done quite a bit of writing with, is actually yes, very impressive. This is not just an outlier, and then we have the other open source models coming in the top three, and then the frontier models, four, eight, and five-six, coming in last. So this is actually a really good insight here that's worth calling out because coding even though these are the flashier demos, if you're not a software developer, those meet a lot less to you. Most of us are just in these kind of day-to-day things where we need a writing partner where we need something analyzing documents where we need help just getting through basic knowledge work tasks aren't did today. Interestingly, all of the open source models did better. It's something that I've seen time and time again where it really is worth it if you haven't already looking into migrating even just part of your stack over.
-
25:16
, obre el vídeo en una pestanya nova
to an open source model. Rand can plead next up, we've got a motion graphic explorer. And this guy here, Tatsula, has been put in KMK3 to the test today and some of his generations are really cool. What he did here was he... took their launch video here and just had it recreate SVGs. So basically a video clone of Kimmy's video and look how good it is. And it's just so impressive. So I want to do our own test of this by using the hyperframes skill and generating an MP4 with motion graphics that is an explainer of Kimmy K3. So let's take a look at that now. Number one. I'm going to quickly move through these. Okay, some good transitions. Wow, these are getting good. This might be one of the better models, but that's okay. That's really impressive too. It didn't have any benchmarks, which makes sense why I just just price and then all these placeholder elements here. You can see that these post-concraftings are really solid. All right, that's number one. Number two, they all have this dark background, which is interesting. Okay, if you'll go to K3, all right, kind of kind of all over the place. Not bad. Nice little bar chart. Very cool. Okay. Yeah, not bad. What is gimmick a three nice little transition price just all placeholders. All right. This one's less complex still looks pretty good. I'm going to say this one so far as number one this one's the lower the pack. Let's see this one now some good transitions. All right. Kind of boring charts are to read text one model. Okay. Yeah. That one's towards the last and then something didn't generate. I'll see if I can get that generated. But in the meantime, okay. Let's see what this one is. There we go. K three. Yep. darn good at motion graphics. Wow. I mean, look how kinetic and interaction this is. Looks great. Okay. Which one didn't work? Okay. GL2 didn't work for some reason. Five six soul. I would say five six soul is last. And then this one, I would say this one's number two. What is this? Opus 48. This one's number three. Do seven. Yeah. I would say two seven that are better job than. 5-6-Soul on that, and then Jill and 5-2, sorry, I didn't generate. For our next knowledge work test, we have landing-page copywriting, a little bit of landing-page design. This is actually for something that I'm launching. I'm calling it AI Bootcamp. Here is the current website. The link is in the description below if you're interested, but it's essentially a four-week live intensive where each week we get into AI fundamentals, defining the business use case for founders, executives, business owners, on implementing AI in their own businesses, actually building that out, shipping all that kind of stuff, and I've been spending... too long in this landing page just banging my head against the wall. So I figured why not go to the models and see if they can do any better. So that is what this next test is. Here is the prompt, rewrite, redesign the workshop landing page for about seven sharper copy that digs into the avatar. So I really spent a lot of time in the prompt here because this is the problem I was having. I wanted to dig more into pain point and write this in a way that speaks to that pain point a bit more and isn't too sailsy and we're gonna see how these models did first up. Turn we should be using AI into a working tool your business actually uses Not bad. That's not really the focus of but okay build one live AI tool or workflow for your company in four weeks Then lead your teams AI push knowing what works. Okay. That's not terrible. That's not terrible subhead That's kind of a point. And then we have this nice little redesign of the header even though this doesn't make a whole lot of sense in this little graphic here. All right. Wow. Okay. So did
-
28:52
, obre el vídeo en una pestanya nova
And it reads on this whole thing right now as probably costing you twice. Okay, digging into the problem. Nice. Okay, this is actually, this is actually a good problem statements. All right. The software bill keeps climbing. The busy work is still there. You can feel the gap widening. AI is on every agenda. Become the person in the room who actually knows. Okay. Don't love this design. Spot the right problem, build with your eyes open, lead from experience. Okay. But yeah, this is, this is solid. This is solid. Four weeks later that it bet is over. It is running. Okay. Design's not bad. Learn it. Find the case, build it, ship it, small group. Real deadlines, no place to hide. Like, wow, I pulled from a testimonial and then just created this as a polquate. Alright, this one is, I'm impressed. Not sure what it is though. Next one. Build it yourself first, roll it out second for week live intensive for founders and executives. Okay, yeah, they're just going to lean into this. We should be using AI more, which I don't really want to lean into a whole lot. But again, nuance that I didn't probably communicate in the system prompts. The subhead way too long. You're not behind because you don't believe in AI. behind because we've never actually built anything with it. Okay. And then we have these problems here, not bad, kind of kind of hard to read and fall along. Now this is a tool and problem. It's doing problem. Alright, this one's got AI copy written all over it. Don't love that one. Okay, stop renting your operations, start owning them. I think they're talking about software, building software. Before we could live intensive, we're founders and executives build AI systems. Their business actually needs not to create a small code of operators who won't let you quit and dashed it. Give way, your paint for AI twice. Once in software, your team barely uses, again, in the hours lost to busy work. Okay, that's not, that's not bad. quiet dread that you're falling behind quiet the as love using quiet okay nice we've got a little bit of my identity shift copywriting here the operator who stops guessing and starts deciding not bad not bad I think everything else is the same so I'm gonna say this one's number one so far three two let's see what the rest of these are stop reading about AI start building with it just not created it AI copywriting has such a long way to go still. And then this just really long sub headline. A four week live intensive you pick the workflow that's eating your team's time, build the system that runs it. Yada, yada, yada way too long. You know, you should be further along with AI by now, not because you're behind the news. You probably read more about the next one or the team. It's so long. You're paying for AI seeds. Okay, leaning into the problem, not bad. We'll figure it out. AI next quarter is a decision to leave. This is just so... Iki just feels like he's just being talked down to become the leader who actually knows Okay, don't love this the fix isn't more research is one real build way too long copy Okay, I think everything else is the same that one is the last so far and then we've got a way too long of a headline and for we see we'll be you'll have built the thing you keep saying your company needs that's a mouthful way too long of a subhead don't love this design here you've had this meeting you didn't get where you are by being slow you've made harder calls than this one This one keeps sliding. You approve the budget for a ride. It's kind of going like old school copywriters. It feels like a direct mail piece, almost. You bought the seats. Most of them haven't. Okay, this is just all over the place. Gosh, uh, more achy copywriting. You already know the other kind of leader. You don't need to become an engineer. That's not about point one month. Start to ship. You won't be the only one building. Okay, y'all. All right. That one. I'll say that one's fourth. All right. Big reveal. Fifth. Oh no, can we K3? Yeah, hey, to see it. That was really quite bad though if we're being honest. What's number four? Ops 4 8 interesting number three GLM5 2 number 2 is soul really going to be up in the top here. Can we K2 7 is number 2? Wow, and then soul's number 1.
-
32:29
, obre el vídeo en una pestanya nova
I don't know, this could just be an isolated test, but the cover rating was quite terrible. Anyway, that is workshop landing page, last test here, and it's a fun one, wait and tear and tino trailer. All right, so for this one, I really want to put the models creative thinking to the test, and the prompt is right and fully shotless a 30 second trailer for a Quentin Tarantino film, as a complete ready to generate shot plan for an AI text to video model. So the idea is not only writing a script for a 30 second trailer, but also the complete shotless that we then gave to seed ants. to generate. So we're gonna see how all these look and then it looks like it also, at least a couple of them made actual websites too. I'm really like one I'm seeing so far. So let's take a look at this one, the last slice. Okay, not really sure what's going on there. I don't know if it's the video model or the actual script or shot list, just random characters kept showing up. Great shot in the opening here though. Nice stab, a little sh**, but everything else kind of all over the place. I'm not gonna play all of these, but I'm just gonna do this separately and then I'll be sure to add these to the actual post so you can watch each of these if you are really curious. See you down as by the way, it's just an incredible video model. Look at this, it just looks so realistic. But that's not the test. These are so absurd. All right. This one I'm just gonna say is the best one. Which one was it? GL and 5 2. There we go. This very confusing. I'm gonna say fourth. That's one just because the shots prompted. I say number 2. That's one I'll give number 3 because of that establishing shot and then this number 5. Let's see. 4 8. Okay. 5 6. So, number 2. Okay, two, seven, number four, K3, number three. All right, not a great test. I'm just trying to get creative with it. See if we can push these models and do something a little bit different. But there we have it. So with that, let's do one final telly of where K3 landed across the board. So here is what we're looking at. On the coding side, it came in at first, fourth, third, second, fourth, the interactive gadget that 3D rendering was pretty impressive. We kicked it off and I thought we're going to see some better outputs from K3, but these are truly so subjective. And then a knowledge work did a little bit better, second, first, and fifth, third. So it wasn't totally blown away, like I maybe was with Fable or 5.6. But that's also not the point with these open source models. If we looked at some KimmyK27 generations, for example, you can see how much of a step up it is from 27 to not 3. And I mentioned this too, but for the most part, you're not doing crazy hard coding tasks. A lot of this is just trying to get through your day and for an open source model that's much cheaper than any of the frontiers, KimmyK3 seems like a good selection. But, uh, curious what you think, appreciate the position. I know this is a long one. This one kind of got out of control. I'm really trying to work on writing in these model release demos sometimes I get a little craze with it. Continue to just try to test the creativity the kind of day-to-day task especially to see how well these truly do so appreciate you sticking with me as always. If you have any ideas for future demos to run and any of these model release videos, please let me know and I'll see you in the next one.