Ornith 1.0 a prova: el model local de codi de 35B convenç més que el de 9B
Bijan Bowen prova Ornith 1.0 en versions de 9B i 35B amb webs, escenes 3D i correccions agentives. El 35B MoE crea els resultats més atractius, però tots dos deixen errors funcionals i la comparació utilitza maquinari i precisions diferents.
Ornith 1.0 és una família de models oberts especialitzats en programació agentiva. Bijan Bowen en prova dues versions locals: el model dens de 9B i el model mixture-of-experts de 35B.
La versió de 35B produeix el millor resultat global. Construeix un sistema operatiu web atractiu, una mascota d’escriptori i un aparador de rellotges amb models 3D. La de 9B és capaç de reparar parcialment una escena a partir d’una imatge, però deixa més interfícies trencades.
Cap de les dues ofereix fiabilitat plena. Diversos botons no funcionen, els jocs no es poden completar i alguns intents de reparació entren en bucles de raonament o no eliminen l’error.
La conclusió del vídeo és prudent: Ornith 35B sembla potent per la seva mida activa, però encara falta una comparació controlada amb els models Qwen originals per saber quant millora realment l’ajustament.
Què és Ornith 1.0
Deep Reinforce defineix Ornith com una família de models de codi que pot treballar dins d’agents amb terminal, fitxers i eines. La línia inclou:
- 9B dens;
- 31B dens;
- 35B MoE;
- 397B MoE.
En el moment d’enregistrar el vídeo, Bowen troba disponible el 9B i el 35B. El 31B apareix anunciat però encara no es pot descarregar des del lloc que consulta, i el 397B queda fora de l’abast del seu maquinari de prova.
El repositori actual ofereix pesos, variants quantitzades i instruccions per a vLLM, SGLang, Transformers, llama.cpp i Ollama. La llicència anunciada és MIT.
Els models són de raonament: produeixen un bloc intern de pensament abans de la resposta i poden generar crides d’eines. El context anunciat és de 262.144 tokens, però utilitzar-lo complet implica un consum de memòria molt superior al necessari per carregar només els pesos.
L’entrenament també optimitza la bastida
La idea distintiva d’Ornith és que l’aprenentatge per reforç no optimitza únicament la solució final. També aprèn la bastida agentiva que guia el procés.
Una bastida pot decidir:
- com descompondre la tasca;
- quan inspeccionar un fitxer;
- quina eina executar;
- com comprovar un resultat;
- quan corregir el pla;
- com conservar informació entre passos.
En lloc d’aplicar el mateix flux dissenyat per humans a totes les tasques, el sistema permet que la política i la bastida evolucionin conjuntament.
Deep Reinforce menciona GRPO, group relative policy optimization. De manera simplificada, el mètode genera diversos intents per a una mateixa tasca, els compara dins del grup i reforça les trajectòries que obtenen millor recompensa.
Aquesta descripció no demostra per si sola que l’agent sigui més fiable. L’efecte s’ha de mesurar amb la mateixa base, el mateix entorn i prompts no presents durant l’entrenament.
Benchmarks prometedors, publicats pel creador
El repositori atribueix a Ornith 1.0 millores importants en Terminal-Bench, SWE-bench, NL2Repo, ClawEval i SWE Atlas.
Entre les xifres publicades:
| Model | Terminal-Bench 2.1 | SWE-bench Verified | NL2Repo |
|---|---|---|---|
| Ornith 1.0 9B | 43,1 | 69,4 | 27,2 |
| Ornith 1.0 35B | 64,2 | 75,6 | 34,6 |
Són resultats del mateix projecte. La documentació explica harness, temperatura, context i nombre de repeticions, però una puntuació no es trasllada automàticament a qualsevol ús local.
El vídeo aporta un complement útil: tasques visuals i iteratives que deixen veure el producte generat, inclosos els errors que una mitjana pot amagar.
Dos models, però una comparació desigual
Bowen executa el 9B:
- amb quantització Q8;
- en una GPU RTX 5090 de portàtil amb 24 GB;
- mitjançant LM Studio i, més endavant, OpenCode;
- a uns 18 tokens per segon en la primera prova.
El 35B s’executa:
- sense quantitzar;
- en una GPU Blackwell molt més gran;
- mitjançant vLLM i OpenCode;
- amb més memòria disponible.
Per tant, quan el 35B guanya no es pot separar fàcilment l’efecte de:
- més paràmetres totals;
- arquitectura MoE;
- més precisió;
- maquinari superior;
- flux agentiu diferent.
El vídeo és una primera presa de contacte, no un benchmark cara a cara.
Sistema operatiu web: gran diferència entre 9B i 35B
La primera ordre demana un fals sistema operatiu dins del navegador, amb aplicacions i jocs.
Ornith 9B
El resultat del 9B té una estètica peculiar, però les finestres queden amuntegades i diversos controls no responen. El model intenta diagnosticar els errors des de LM Studio i genera una cadena de pensament llarga, sense arribar a una reparació clara.
Bowen atura l’intent perquè sembla atrapat. El problema no és només visual: no es poden tancar o moure elements de manera fiable.
Ornith 35B
El 35B crea una interfície molt més completa:
- menú d’inici i cerca;
- gestor de fitxers, tot i que majoritàriament estàtic;
- bloc de notes que desa un fitxer;
- navegador que obre una pàgina dins del navegador;
- configuració de colors i fons;
- rellotge i data;
- dos petits jocs;
- mascota d’escriptori.
La mascota és el detall que més sorprèn Bowen. Saluda, desapareix i recorda que ha tornat, dorm, menja i conserva estadístiques. En un moment mostra un nivell de felicitat del 0%, un resultat còmic que dona títol a la introducció del vídeo.
El clon de GTA no és plenament tridimensional, però incorpora vianants, vehicles, edificis, animacions i una acció per colpejar. El calabós 3D és visible, encara que el jugador queda atrapat sense arribar als enemics.
És una demo creativa i extensa, no una aplicació acabada.
L’estació de metro i el pas a joc de trets
Tots dos models reben una petició per crear una escena de metro.
El 9B genera alguns elements 3D i cartells, però l’escena és massa fosca. El control de brillantor queda bloquejat quan el punter entra en la vista tridimensional.
El 35B ofereix una estació neta i il·luminada, amb minimapa. També comet errors geomètrics:
- peces del vagó desalineades;
- seients dins d’una estructura tubular;
- moviment vertical acumulatiu en prémer la tecla d’ajupir-se;
- rendiment molt lent.
Bowen demana convertir totes dues escenes en un joc de supervivència amb zombis.
Els dos primers resultats fallen al botó d’inici. El 35B rep l’error dins d’OpenCode i el corregeix ràpidament. Després es pot entrar al joc i els enemics ataquen, però el jugador no els pot fer mal.
El 9B no aconsegueix arreglar la seva versió. Aquesta prova resumeix Ornith: pot construir una gran quantitat d’interfície i lògica, però una ruta essencial queda incompleta.
La web de rellotges: el millor resultat del 35B
Una altra tasca demana una web cinematogràfica per a una marca fictícia de rellotges, amb un model 3D a la capçalera i dues targetes de producte.
El 9B crea el disseny i fins i tot introdueix «2026» a la imatge, però els models tridimensionals no carreguen. Diversos intents de reparació —inclòs un refresc complet i una prova en Firefox— deixen el mateix error.
El 35B sí que presenta:
- una corona que fa de cara del rellotge;
- moviment panoràmic a la capçalera;
- composició de producte coherent;
- dos models a les targetes;
- acabat visual superior al que l’autor esperava.
Bowen remarca que models més grans han tingut dificultats amb aquest prompt. És el cas més convincent a favor d’Ornith 35B, encara que només és una execució.
De fotografia a objecte 3D amb el 9B
El model petit rep una imatge d’una màquina recreativa impresa en 3D i ha de reconstruir-la com a objecte interactiu.
La primera versió no mostra res per un error. Quan el 9B treballa dins d’OpenCode, llegeix el missatge i genera una reparació.
El resultat continua incomplet, però reconeix:
- la pantalla;
- el joystick vermell;
- la base negra;
- part de la distribució de l’objecte.
No reprodueix bé la silueta completa. Bowen ho considera una mostra de potencial perquè combina visió, codi i reparació en un model de 9B quantitzat.
És important distingir «elements recognoscibles» d’una reconstrucció 3D fidel. No es mesuren proporcions ni es compara la geometria amb l’original.
Proves que no arriben a bon port
No tots els encàrrecs mereixen el mateix entusiasme.
Joc de ral·li en C++
El 35B ha de crear un joc de carreres 3D amb estètica retro. El procés es torna lent, barreja noms de carpetes i apunta a un resultat massa elemental. Bowen l’atura abans d’obtenir una versió funcional.
Joc de llanxa
Com a prova final, el 35B quantitzat rep una ordre curta per fer un joc de llanxes 3D. El resultat és una gran decepció i l’autor no inicia una ronda de reparació.
Correccions del 9B
El 9B arregla l’error que impedia veure la recreativa, però no repara:
- el joc de metro;
- els models dels rellotges;
- el sistema operatiu inicial.
La capacitat agentiva existeix, però no és consistent.
Què diu realment aquesta prova
El vídeo dona suport a quatre conclusions.
El 35B té més marge creatiu
La mascota, el sistema web, el clon de GTA i el rellotge 3D mostren capacitat per produir prototips rics a partir d’una sola instrucció.
La presentació pot ocultar errors centrals
Una interfície atractiva no garanteix que el botó principal funcioni. Cal executar cada ruta, inspeccionar la consola i provar l’estat després d’interactuar.
L’agent ajuda, però no sempre convergeix
OpenCode permet donar errors reals al model. Algunes reparacions funcionen i altres només canvien el problema o entren en una explicació circular.
Falta comparar amb la base
Ornith és un ajustament de models Qwen i Gemma. Sense repetir els mateixos prompts amb Qwen3.5 9B i Qwen3.5 35B, no es pot atribuir l’èxit a l’entrenament d’Ornith.
El mateix Bowen reconeix que no recorda amb prou precisió les proves anteriors per declarar una millora definitiva.
Per a qui pot ser interessant
Ornith 9B pot encaixar en una GPU de consum amb quantització i servir per:
- prototips de front-end;
- petites reparacions supervisades;
- assistència offline;
- experiments amb agents;
- generació de primeres versions.
El 35B MoE ofereix més qualitat, però necessita molt més maquinari en precisió completa. Les variants GGUF redueixen requisits, amb una possible pèrdua que el vídeo no mesura.
Cap dels dos s’hauria d’utilitzar per executar canvis sense:
- revisar el diff;
- executar proves;
- validar dependències;
- comprovar seguretat;
- verificar llicències;
- limitar les eines disponibles a l’agent.
Conclusions
Ornith 1.0 35B és el guanyador de la prova de Bowen. Construeix el millor aparador 3D i el sistema web més complet, i resol algun error dins d’OpenCode.
El 9B ofereix una entrada local més accessible. La reparació multimodal de la recreativa és prometedora, però les aplicacions principals continuen trencades.
El títol del vídeo pregunta si són els millors nous models locals de codi. La prova no permet respondre que sí. Permet afirmar una cosa més útil: Ornith pot generar prototips sorprenents, sobretot en 35B, però encara necessita un humà que comprovi cada acció i una comparació directa amb el model base.
Contrast i context
Fonts consultades
- 01
-
02
Deep Reinforce Ornith-1.0: repositori, benchmarks i instruccions
-
03
Deep Reinforce Ornith-1.0-9B
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
It's happiness level is at 0% which is It's just kind of disturbing. Today we're going to be taking a look at Ornith 1.0. This is a family of models that have been getting a lot of praise recently and in my most recent video which was quite a few days ago and I do apologize for the gap. A lot of the recent comments are suggesting that I test these out on the channel. So for today's video we're going to be testing two of the four available. I should say three available versions of these. The 9 billion parameter dense version and the 35 billion parameter mixture of experts version. I save three available because as of the time of this filming, the 31 billion parameter denser version is not currently publicly available at least on their hugging face and the 397 billion parameter version I don't have in available system to run that on. So we're going to be using the two right here, which we can play with for today's video. So before we get into it, please do feel free to subscribe as I do want that 100K plaque and I will continue to focus more on open weight models being that the currently newest model. GPD 5.6 is not something really anyone can access and that does seem to be sadly a point in time where... Restricted access to frontier intelligence may now become the norm which is sad but also shifts focus more on two open weights models like this. However, this is a fine tune of existing models. Mainly the ones we're going to be testing are quen models. The only one here that is Gemma 4 would be this 31 billion parameter dense model which is not currently available so it's listed but not accessible and really the future of open weights At least in terms of the models these are based on which the ones we're testing are based off of Ali Baba quen open weight models. There has not been a lot of noise from them in terms of whether or not they're even going to release any additional quen variants. So seeing what level of performance one can get out of existing models may become something that is more and more popular as time goes on depending on what the open weight situation is in terms of big labs who are actually still putting out models that are accessible and democratized. to those who are not one of the selected Fortune X companies who can access state of the art model. With that, let's take a look at some of the interesting things here. These do seem to benchmark very impressively. Now, I have not at all tested these. I ran one prompt to ensure that my config was set up, so we'll be looking at these results for the first time. So further down in this blog post, they do talk a bit about what was included in the improvements to these models based off of their previously existing versions. And they talk about the self-improving training framework that jointly learns to solve tasks and to construct these scaffolds that guide those solutions. Rather than relying on a fixed human design harness shared across the task category, it treats these scaffold as a learnable object that co evolves with the policy and they do talk about using RL right here with GRPO. Now, if anyone's more interested and I do want to specifically take a moment to show this resource here with all the stuff that's going on with model accessibility, perhaps not being as confident and inspiring as it was a couple of weeks ago. There is a lot of available resource online in specific right here. This is hugging faces LOM courses and they do have a lot of good and palatable information about a lot of the techniques that will appear in blog posts for things like this. So in this case, GRPO is called group relative policy optimization. And if we scroll right here, the core innovation of GRPO is its approach to evaluating and learning from multiple generated responses simultaneously. So multiple responses will have been generated here.
-
3:08
, obre el vídeo en una pestanya nova
And then it compares outputs within the same group to determine which one should be reinforced. So basically the best answer then kind of is used out of that group and then pushes the model towards the correct answer of a group or just Read this they also give some additional information right here and then we get some scary looking math on the page as well as a full benchmark table for the different variations of size that are currently available right here. We've which the Gemma 431B is not hit output it will be very interesting to see assuming these do stack up how that is to note for the benchmarks here These are not one-shot scores. I think they are okay, so they're averaged over five runs and interestingly they ran them at a higher temperature than the suggested sampling parameters for just like general use which I do believe based on the hugging face model card. We're 0.6 so they ran them at one. I don't know why I mentioned that it's probably not as important for a first look in test video. So with that, let's go and take our first look at the model we're going to start by testing, which is the 9 billion parameter dense version of this. I am going to be running this locally right here at a Q8 quantization on a laptop 5090, which is a 924 gigabyte. It's a 24 gigabyte video card. So of course, we're going to be starting with the browser OS test v2.5. This is the one where it needs to create functional 3D games. Now because Because these models are tunes of existing models, I want to be honest about the potential that some of these prompts may have made their way into these specific training set. It can't be discounted especially for a fine tune on it. existing model. So with that, we're going to look at these results for the 9B and the 35B MOE. However, we may try some things that have not really been tested on the channel before, just to get a proper test of how these perform on things that would not have been seen in a dataset. All right, at a screaming 18 tokens per second, we have received the browser OS test for the, all right, this is our browser OS result for the 9 billion parameter model. Interesting. Oh, wow. Oddly, oddly looking is the term that came to mind. Let's see if there's a right click. Okay, there is an a right click which makes me more bullish that it wasn't specifically benchmarksed on this test. Now the main issue where probably going to notice is this somewhat of a cluster of the way that these apps are arranged on the screen and I don't see the ability to actually get rid of them and move them. Let's try with our start menu. Okay. Do seem to be having some functional issues right here. So let's take a look and see if we have any specific errors that could be attributed here. Okay, and I'll just because we're using open code with the 35 B degenerate, the browser OS. I have just instructed it from within the LM Studio chat to fix these issues. So I'm going to be running the same browser OS test just using open code with the 35 B model. This is running through VLM on the 6,000 pro box behind me. It is not
-
6:17
, obre el vídeo en una pestanya nova
quantized so this is like the full performance level of this model. I had to change a setting here because the timeout was set too low so it kept airing out because the generations are going to take longer so I fix that and now I'm re-running this here. So we'll be able to side-by-side compare these. I have also noticed that in initiation of giving this the troubleshooting task with the errors for the browser OS, this is back to the 9B in LM Studio. I am looking through the chain of thought and getting a little concerned we may be entering some four. of thought loop because I keep seeing like a weight exclamation point. Sometimes this happens for a while and then the actual result gets fit out properly so we'll see what happens here but I'm a little concerned. Alright so it did take quite a while but as we saw right there or maybe not depending on when I started refilling this this is like a 17 or 1800 line browser OS test that it just created again. This is the 35B model version. Sadly the 9 billion parameter one. I don't know that this is specifically a thought loop. Maybe if this was running it like a thousand tokens per second, it would have actually found the issue. But as we see right here, this is not looking too good in terms of being able to fix the specific issues we noticed. So for the time being, let's just go ahead and look at the result we received for the 35 B version of this model, which is in its own web OS folder right here. Okay, so far so good. Not bad. Hey, I'm your new desk called pet that is actually kind of cool and I don't believe that is something I've seen from a quen model before I've seen stuff like this from minimax Hello friend. I kind of like it. Can I move it? Oh, no, but there is a right click unintentional right click discovery. Oh look it moves It's kind of cute like to be honest with you. I like that It's a it's a waving at us saving All right, talk a pet. Oh no, all right. Let's Let's leave it there. I like it. I'm back. Good. I like that. It actually has awareness of when it returns to the screen when you retogulate, which is actually some level of depth. I know that seems silly to focus on, but I'm actually happy to see that. Look at the little animations, too. It's like, yeah, I'm not quite sure what it's doing. Okay, it's asleep now. Let's check our start menu. Very good with a search settings. Files. We'll just run through these sequentially. Okay, this is a very well formed and aesthetic. the pleasing non-functional file manager so these are all just static they look good but they don't unfortunately allow us to click anywhere further play with me not right now but maybe later with that hand movement I think not let's turn the pad off for now terminal new fish if it has anything like cool here very good it has a okay next up no pad
-
9:26
, obre el vídeo en una pestanya nova
Save does this save it as a text file? It does Very good. Maybe this one was benchmarks and just not the not the 9B because that was quite terrible So next up we have a browser and it opens to work the PDF Excellent excellent and well done look at this We're in a browser in a browser. All right our settings here geometric Charcoal sky blue sunset fire purple pick a custom background color That seems to override the other options. All right GTA clone time. Oh disgustingly not 3D, but not bad. This is very look there's okay. This is gonna be difficult to see at least. All right. Well, the full screening that did absolutely nothing. There's actually like leg animation movements here. This seems to have a pension for like interesting little sprite style animations. The buildings are interesting. We have a truck right there. The vehicles that actually have visible headlights and windshields as well. This really isn't have bad health money wanted area midtown on foot. The the Pat has come in that may look like a very odd GTA game if someone were to just click to this point in the video and the from my recollection of a quantist if we run over one of these there may be a red. Okay. Good. F is to punch. Oh, okay. They ran inside. So if you press F it just kind of eats them because that's punch downtown midtown uptown. He's side and void. I'm so hungry. Okay, you're going you're only away for now. That was good. Next up we have a dungeon which is like a 3D dungeon. It's actually pretty alright. All right, so unfortunately we're stuck in this room. We can't get to any of our enemies, but the actual game itself was not bad at all. Then finally we have pet. Oh, we can perform actions with the desktop pet. So let's bring it back. It had said it was hungry. So we can feed it. Can we overfeed it? Make it bigger? No, that's messed up. Where's the GOP one button? Nap time. Okay, we'll wake it back up. No, play. That was fun. Pet. Purs loudly. And it actually has pet stats. So it's happiness level is at 0%, which is it's just kind of disturbing from like a empathetic standpoint. Okay, we're gonna close this, but this is definitely a very interesting special feature. I do believe I've only seen something like that with mini-max, and then of course we have the date and time in our local. Was there a special feature? Oh, that would have not really been the pet. Not pet.
-
12:35
, obre el vídeo en una pestanya nova
Definitely definitely a significant improvement over this garbage. I'm going to go out on the limb here and say this is still not made any progress in terms of actually fixing the issues we saw. Is it still thinking? Yeah, all right, so I'm going to stop this because it's just not. Next up we're going to be doing a version of the beautiful static subway scene. I don't know if this is the exact same prompt that I tend to run in all of the other model tests, but I'm okay with some slight differences because it can throw things a little off, which we do want when testing something. So I have run this both through open code for the 35 B version and as well as the 9 B version right here, just running locally on this laptop. Alright, let's take a look at our subway results. These are just a beautiful static subway scene. This is what what we got with the 9 billion parameter model. Okay, now, if I can get to the lighting slider, which I can't, this is always a gacha that stumps less capable models where we do have the adjustable brightness slider. Unfortunately, the only way to get to it is to escape out of the movable scene where it is then locked by the inter-scene station. portion so that is kind of regrettable, but nonetheless we're still able to see at least some artifacts here and perhaps without the bright studio lights pointing at me. This will be a bit more visible in the actual video than what I see right here. We have some elements of a 3D scene and I would say just from what is visible which is really not too much visible. This is not a bad 3D result for a model of the size, although it is a dense model. This one's tough to judge but we do have some individual elements and some interesting artistic posters as well as some red and some other things. So here's the 35B version. Okay. Oh, okay good. I was about to say okay the 9B did better. Wow, this is incredibly incredibly laggy. We do have a mini map in the top left which is interesting as well as some other. This is like Really really really clean Very sterile. I am not even gonna begin to question what these things are these almost look like they were supposed to be conduit on the ceiling, but they're oriented in the wrong way which is you know it happens We have a tube there as well, which I'm wondering if we can go in it if that is the train and Space is to rise see is to crouch interesting. Oh Okay, so those are actually not like you can't crouch and then uncrowch it just as you press C you just continue to go down so Okay, so these do seem to be elements of the subway car that are unfortunately just not 100% properly put together if we go into this tube right here There are seats and things like this so it has some of the elements available just not necessarily oriented in the correct way Interesting nonetheless and very
-
15:44
, obre el vídeo en una pestanya nova
Clean and sterile as I said a bit laggy lights look good on the ceiling's Interesting now I've instructed both of these to turn their results into an FPS with zombie enemies weapon recoil sound effects, et cetera, and we'll see what we get from that all right the results for turning these subway stations into FPS games have been completed Let's start with the 9 billion parameter model. Oh Click to survive, and then you just can't actually click That's you know Now we also have the 35B version, okay subway survival, enter the station. Did both of these fail? That's just a bit frustrating. I'm giving the 35B one through open code the issues we're facing where the enter the station button doesn't do anything, as well as the specific error we got here in the developer tool. I don't really know what to do because this will take so long and I'm worried it will get into a thought loop just because we are using the 9 billion parameter 1 through LM Studio. Alright, so I am glad to see that the 35B subway station results seem to be fixed very quickly, so let's refresh it and see if we can now actually enter the game. Okay. Very interesting, a bit significantly more laggy than it was previously, but alright. This sounds like a... Maybe that's not an enemy. Oh, I don't think that was an enemy. All right, these definitely are yes, so my mistake. All right, let's interesting grouping of sound effects that chose to use. All right, I don't think we can actually attack any of these enemies. Oh, hello. All right, dry again. And then try again is just like the most like W3 schools tutorial looking botany you've ever seen in your life. Alright, let's do it. Let's yeah, I don't think we can actually do any damage to these things. Which is alright. Well, it was in improvement because it did fix the issue we were having. So that's good to see. Now in lieu of just continuing this stellar result, I'm probably going to exit this and run the same front end test. Now here that we are. currently running with the 9 billion parameter model. So I'm giving the 9B a new front-end test that I've been running recently in some of our newer videos where it needs to create a beautiful website for slap-ish watch company. The point of the website is it needs to... create these assets in 3D and the hero section needs to have a very cinematic panning, good-looking shot of the 3D watch. Then it should have two cards featuring 3D models of the watches as well, just for pricing sales card. So it'll be interesting to see what it does for one 3D model and capability and two overall front end design capability with this test. So we will also get our Slap-Swatch company result here from the 35B. Alright, so here's the front end watch website result from Arn.
-
18:53
, obre el vídeo en una pestanya nova
9 billion parameter model. Okay, unfortunately the assets are just not loading properly. Okay, does say 2026 in the photo which is good. However, I wonder if there's a trivial error here maybe like a missing import map or something. Okay, there isn't unfortunately. I'm going to let this agentically try to fix some of the results that it has generated that have not properly worked at. later point because we're going to use it through open code and right now I'm waiting for the 35B to generate this result, but I will not just leave the 9B 1 through LM Studio entirely, so that's something we're going to do as well. Is see how it can agentically in a coding agent like open code fix some of its results. In the meantime, we have our preliminary result from the 35B. Now this yes, it looks a bit silly. I'm going to say for a mixture of experts model at this size, although it is not at all quantized. even in a big quantity probably have knocked this out pretty similarly. This is not bad. As a matter of fact, that crown there is actually the face everything. This is a better result than I would have expected to receive here. This is something I've been testing with a lot of new frontier models, at least until we stop having access to them. But I'm gonna say this is actually really not bad. This is an impressive result. The 35 B1 definitely seems to have some proper chops. With that said, I don't have full recollection of my general QN-35B MOE testing, so it is possible that it's not a hugely... but... For what I'm seeing here independently judging this, I'm impressed. Now this is where the tests are going to start to differ a bit from within open code for the 35B model and I'm just starting it from within build mode not in plan mode. I'm giving this a C++ test however, not the skateboard one because that is probably pretty decently known at this point for a C++ test. This is to generate a 3D racing game with retro rally game style aesthetics to it. Okay, it is asking for permission now, which is good. So it has created or written some of the script for our skate game. Oh, it's not a skate game. The problem is that I put it in a folder named skate. So this is a 3D racing game. Keep that in mind. I had just tripped myself up. So this really took this prompt and the combination of what I'm trying to say here is this is about as like bare bones. in a way that it's trying to make it that it could have. To the point where I'm not 100% confident that it's actually going to produce a successful result, especially because for some reason there has... skake or skate. I don't think that's a mistake that I made. I should probably check, but... yeah. Okay, that's... alright, we'll see what happens. Alright, unfortunately the C++ racing game just needed to be stopped. It wasn't... likely to work and it was taking quite a while so unfortunately regrettably we're probably going to have to omit that from this Pacific test So I'm giving this very simple instruction here with an image of course as well to create an interactive functional 3D model of this image I would imagine this will probably just end up using 3.js or something looks like a plywood box. No, that is high quality PLA plastic 3D printed, but we can excuse that all right, so here is our
-
22:02
, obre el vídeo en una pestanya nova
image 2 3D model replication test from the 9B. Okay, unfortunately, and keep in mind this is a pretty difficult prompt for this to have knocked out. So at this point, I think I'm probably going to hold off on running additional tests through LM Studio from the 9B because I'd like to see how... capable will be in fixing some of its initially generated but broken results. But something I would like very much to do now is swap this so the 9 billion parameter model is in open code and will give it a chance to fix some of the issues that had with the scripts that were generated just. through LM Studio. All right, so we now have the 9 billion parameter model from the system working in open code. And I want to start with having it try to fix the image to 3D model generation test. So I will just say fix this issue and paste in the specific error that we had. It is in a directory with only this script. We now supposedly have a fix to at least that specific error that was happening. I should also make note that I moved it into its own specific folder, which is why the file had disappeared there. Let's see if we You know what? That's not half bad because yes, okay. It's half bad The reason that I say it's not half bad is because it did do a decent job of actually replicating some of the things that were on screen the vision capabilities in the bottle that this is based on which is the quen three point five nine B-dense We're always pretty good unfortunately we don't have the full picture here of the shape of the arcade machine, but there are elements of it that are included, such as the red joystick and the screen. So I'll say that this is actually 4-9 billion parameter model at a Q8, especially trying to replicate something from an image into 3D. I'm okay with this, which may seem weird because it's not a great result, but it shows some promise. Next up, I'm going to try to tackle the subway FPS game, so I've just given it the script in its own directory and said the game is too dark, and the game can't be initiated when clicking the button to start it, so we'll see if we get a functional result. Alright, we now supposedly have a fixed subway result as well, so this was the 9 billion parameter model. Okay, well unfortunately this one did not work. This was a more complex fix, likely. It is a little disappointing as I was kind of feeling excited and hopeful that this would work. But regardless, it's just interesting to see to what level of completion we can get some of these. The last one that I would like to attempt to have this fix is the watch website where it needed to create the 3D models and unfortunately those just never really showed up. So we do have a specific error to give it here. So we're going to have it try to fix the watch website where basically none of the 3D models were appearing and we're interesting to see if we get it because the 35 B
-
25:11
, obre el vídeo en una pestanya nova
result for this was actually quite good all things considered. That was a very quick fix so if we look at the reloaded version here, all right, unfortunately we're still not getting our 3D models, however it is possible that a new error has appeared. Interesting, it's saying the fix is correct, you may need to hard refresh, which I did do prior to sending it that. So let's refresh it one more time. We'll hard refresh it. It is still giving me the same issue. I want to verify just that like we're actually working from within the same directory, which we are, okay, let's try opening it and Firefox then. And we still are not getting functional watch models. Interesting. So I'm giving it some pushback saying it's a console issue. It still shows after hard refresh and the watch models don't appear at all. Alright, unfortunately this just didn't quite work out. However, I do have one final thing that I did do as a bonus on the machine behind me. So I gave it a very simple prompt the 35B version at a Q8 quantity. I said make me a 3D speedbook game. And we now have that file. I've not looked at it and I'm excited to see what it did. So let's see if we got a decent speedboat game or not. Okay, so that is huge let down and unfortunately I don't really have the desired try to fix this right now It was just like a one final thing so that is going to lead us into probably the closing thoughts for these models They do seem good, but I want to pre-face saying that with I don't fully have recollection in my mind freshly how the 9B or the 35B A3B for these quen models performed that these are based on so I can't definitively give a judgment like it significantly better, it significantly worse. I will say I wasn't pressed with mainly the 35B model that I saw. Some of the things it generated, especially for a mixture of experts with such few active parameters, were really quite decent. Things like this watch website, not that one. The one that actually worked, this 3D watch model right here that this knocked out as well as the proper spinning hero section was really not have bad considering that this has proven to be quite a challenge for model significantly larger than this. So I was quite happy to see that right there. Additionally, let's see what else do we have. So I had moved some of these one just trying to have the 9B fix them. The arcade machine that was multi-modal coding test base, that the 9B model initially failed to produce anything at, and then subsequently fixed through open code. It did a really good job replicating the screen and everything on there because the vision component to the model this is based on was very strong and still is. It also had some aspects or elements of it like the joystick and the black base. Not great, but something that started out as not even showing anything and we were able to fix it with that model which was nice to see. Unfortunately the 9B never really got the subway station or the watch website working properly. Additionally to that, the web OS that the 35B model created was really actually quite good. This little like desktop pad was very cool except for when it's like
-
28:20
, obre el vídeo en una pestanya nova
satisfaction or happiness just stated 0% with a little concerning. The 3D game that it implemented here was quite neat. Really something I was quite enamored with even though it wasn't fully 3D was this GTA clone. It had a lot of interesting immersiveness, even to the walking animation in these little sprites. And then the ability to punch was kind of funny. So I very much liked this result. This was a very proper result. Unfortunately, the 9B1 had some interesting hope or potential to it with the aesthetics that we see here. It just never really quite worked properly and I didn't think it was worth trying to have it. through open code because it's more of a simple test. The subway game that was created with the 35B again. This was really very clean. Now, oh, look at that. Some of the logic unfortunately didn't really work where we couldn't actually put damage on any of these enemies but they could put it on us. Nonetheless, this was a very sterile and clean result in terms of the 3D items that were created and this was a follow up test to turn it into this survival game with zombie enemies. So that was cool to see. I would say just in a light testing here, the 35B definitely seems to be the winner of the two with the caveat that it was run on the blackwell card at full precision, so it wasn't quantized at all. So that'll give it a little boost in strength as well versus running the 9B as we did in a Q8. But the 9B definitely seems decent. If you have a card that is a good fit for that 9V preeminent or model, I would definitely say it's worthy of trying. Really the thing to do would probably be a head-to-head just to see how the results stack up compared to the model. It is based off of as sometimes the fine tunes can make things better or make things worse. In this case, I don't believe it made things worse. I could definitely see that there's a genuine improvement here just based off of some of the performance we saw on a few of these tasks. So that is probably going to conclude our first look and test of the get the name right so I don't put your it. The Deep Reenforce AI or Nith 1.0 models mainly the 9B and the 35B MOE. I am very interested to see two things. One, if they come out with the 31B that is based off of Gemma4, that could potentially be a pretty potent option. Additionally to that, I would love to see a version of this that is created from the Dense 27B Quand model, which still is probably one of the best local coding models that exists. And I am slightly concerned it may continue holding that position for deforeseeable future, just depending on the way open weights goes and access to front-tier models. So there's definitely a change going on and I would imagine a lot of folks are going to be more interested in open weights models and things that run locally like this. So I absolutely wanted to test these and that's going to wrap it up. So if you have any questions, please feel free to leave them in the comments and thanks for watching.