Qwen3.8 Max Preview a prova: brillant en interfícies, irregular en codi i 3D
Una prova pràctica de Qwen3.8 Max Preview mostra un gran frontend, però resultats febles en C++, CAD, microcontroladors i coherència espacial.
Qwen3.8 Max Preview és el nou model de raonament allotjat d’Alibaba que Bijan Bowen prova amb interfícies web, videojocs, CAD i un microcontrolador. El resultat és molt desigual: crea un sistema operatiu al navegador visualment original, però falla en coherència espacial i no completa la tasca de maquinari.
La paraula important és Preview. Quan es va gravar el vídeo no hi havia un informe tècnic, una targeta de model ni una taula pública de benchmarks. Les afirmacions sobre 2,4 bilions de paràmetres, rendiment només inferior a Claude Fable 5 i futura publicació dels pesos provenen del breu anunci que veu el creador, no d’una documentació completa que permeti auditar-les.
1. Què està confirmat i què continua sent una promesa
Al començament, 00:08, Bowen explica que és la primera menció pública de la sèrie Qwen3.8 i que l’accés encara és preliminar. La documentació de QwenCloud confirma l’identificador qwen3.8-max-preview, una finestra de context d’un milió de tokens, raonament, ús de funcions i eines integrades.
La plataforma també avisa que durant el període de prova el model pot canviar i que després serà retirat o substituït per una versió de producció. Això impedeix tractar els resultats del vídeo com una avaluació definitiva.
La interfície mostrada anuncia:
- 2,4 bilions de paràmetres totals;
- un nivell de capacitat “segon només a Fable 5”;
- la intenció de publicar-ne els pesos;
- millores contínues abans de la disponibilitat general.
Sense arquitectura, paràmetres actius, llicència i benchmarks reproduïbles, aquestes dades s’han de presentar com a afirmacions d’Alibaba. “Pesos oberts” tampoc significa necessàriament codi obert: encara caldrà llegir la llicència i saber què es publica realment.
2. Un model obert enorme encara pot donar més opcions
A 01:45, el creador anticipa una objecció: si el model és tan gran que gairebé ningú no el pot executar a casa, per què importa que els pesos estiguin disponibles?
La resposta és la competència entre proveïdors. Diferents empreses podrien desplegar el mateix model, ajustar-ne el preu, oferir més velocitat o allotjar-lo en una regió determinada. Amb un model tancat, l’usuari acostuma a dependre d’un únic operador i de les seves condicions.
No obstant això, executar 2,4 bilions de paràmetres exigiria una infraestructura extraordinària. Una quantització agressiva podria reduir la memòria, però també la qualitat, i executar-lo principalment amb CPU seria molt lent. Els pesos oberts beneficiarien sobretot centres de dades, universitats i proveïdors especialitzats, no un ordinador domèstic convencional.
3. La millor prova: un sistema operatiu dins del navegador
La primera tasca pràctica comença a 04:05. El model ha de construir una simulació de sistema operatiu al navegador amb finestres, menús, jocs i aplicacions.
És el resultat que més convenç Bowen. En destaca:
- una estètica pròpia, allunyada dels dissenys genèrics que veu sovint;
- finestres redimensionables i menú contextual;
- rellotges funcionals;
- dos petits jocs amb el mateix llenguatge visual;
- un monitor del sistema i opcions de tema;
- un seqüenciador musical interactiu;
- una funció que registra la sessió i permet retrocedir.
El creador reconeix que la valoració visual és subjectiva. També adverteix que els seus prompts són públics i coneguts, de manera que no pot descartar que una empresa els hagi inclòs en proves prèvies. Una demostració impressionant no equival a un benchmark independent.
4. Metro i FPS: detall funcional amb problemes de llum
A 11:47, Qwen genera una estació de metro i després la transforma en un joc de trets en primera persona. L’escena està massa fosca, fins i tot amb la brillantor al màxim, i no inclou cap tren.
La versió FPS, en canvi, conté més detall del que esperava el revisor: arma animada, retrocés de la corredora, recàrrega, passos, so ambiental, enemics i canvi de postura quan el jugador corre. La lògica bàsica funciona, encara que una part de la interfície de llum continua fallant.
Bowen la considera bona però inferior a la que havia obtingut de Kimi K3. És una comparació informal: els models poden variar entre execucions i no es mostren múltiples intents amb les mateixes condicions.
5. El joc en C++ revela persistència, però no qualitat visual
La prova de C++ comença a 15:53. El model ha de crear en un únic fitxer un joc de monopatí amb ambient de passeig marítim californià.
Qwen troba molts errors quan intenta escriure el fitxer a través de l’agent. La solució és enginyosa: genera un script de Python que, al seu torn, crea el codi C++. El programa compila i el joc funciona, però els gràfics són molt pobres i alguns moviments queden mal resolts.
La part més útil arriba quan el creador li passa una captura de pantalla. El model identifica diversos defectes visuals i els corregeix sense una llista detallada d’instruccions. No transforma el joc en un producte acabat, però demostra una capacitat multimodal interessant: observar una sortida, diagnosticar-la i iterar.
6. CAD i microcontrolador: dos fracassos importants
A 19:32, la tasca és dissenyar en OpenSCAD un motor V8 imprimible en 3D que pugui allotjar un petit motor elèctric. Algunes peces —com les tapes i els col·lectors— tenen aparença recognoscible, però l’orientació i la coherència global són incorrectes. Bowen conclou que no val la pena imprimir-lo.
El test del microcontrolador, a 22:24, és encara més problemàtic. El model ha d’identificar una placa USB amb ESP32‑C6 i mostrar-hi l’ús de la GPU. Primer la confon amb un ESP32‑C3, explora fitxers aliens a la tasca i no aconsegueix escriure correctament a la pantalla.
Després dels intents, la placa només mostra una pantalla negra i deixa de comunicar-se amb l’ordinador. El vídeo no demostra que el dispositiu hagi quedat danyat físicament, però sí que la prova falla i que executar ordres sobre maquinari real pot tenir conseqüències més difícils de revertir que un error dins d’un navegador.
7. La línia temporal d’una ciutat: obediència sense criteri
L’últim test, a 25:54, demana una ciutat que evolucioni entre diferents èpoques. La primera versió conté caixes geomètriques simples. Després d’una crítica explícita, millora alguns vehicles, el clima i les transicions, però continua lluny d’un resultat d’alta qualitat.
Quan Bowen exigeix almenys 5.000 línies, el model treballa obsessivament per superar la xifra. Això mostra persistència, però també un risc dels agents: compleixen una mètrica fàcil de comptar encara que no sigui una bona mesura del resultat. Més codi no implica una ciutat millor.
Conclusions
Qwen3.8 Max Preview impressiona en generació d’interfícies i en la capacitat de continuar treballant després d’errors. El sistema operatiu web és creatiu i funcional, i el FPS mostra un nivell de detall notable.
La resta rebaixa l’entusiasme. El joc de monopatí és lleig, el model CAD no és utilitzable, la placa ESP32 queda en un estat problemàtic i la ciutat necessita instruccions artificials per millorar. Per a un model anunciat amb 2,4 bilions de paràmetres, Bowen esperava més.
El principal interès no és proclamar-lo “el millor model obert”, perquè els pesos encara no s’havien publicat i faltava documentació. És observar una versió primerenca amb punts forts clars en frontend i raonament llarg, però també amb errors espacials, excés de deliberació i poca fiabilitat quan toca maquinari.
Contrast i context
Fonts consultades
- 01
-
02
QwenCloud Text generation models
-
03
QwenCloud Token Plan overview
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:05
, obre el vídeo en una pestanya nova
I didn't click stop. So Ali Baba has announced a new model, which is the Quen 3.8 series, and this is the first public mention of a Quen 3.8 model. Now, before we get into it a couple of things. One, please do feel free to subscribe, so I can get that 100k plaque, assuming you like my content. But two, this is pretty much going to serve as the entirety of our announcement post, because this is currently still in preview. So everything we know about this model is contained within these few sentences right here. I would imagine in the coming days or weeks when this moves to a more general availability or whatever the terminology is, there will be a big announcement release post, there will be benchmarks and things of the sort. However, in a lieu of benchmark jpegs, we do have a very, very bold claim right here saying that this is second only two fable five. Beyond that though, there's some other really exciting things here. One, this is going to be open weight. Now historically, Alibaba Cloud, their quen family of models, the Max ones were always closed source and proprietary. No previous quen max model has actually been open weight and that's very exciting that they've decided they're going to open weight the max model. I was surprised to see this but positively surprised I guess could be said. Additionally to that the size. This is 2.4 trillion parameters. That's massive and previous quen max models where I believe somewhere in the neighborhood of around a trillion parameters. So 3.6 max 3.7 max, etc. It's a significantly larger than those, though it is a bit smaller than KMK3, which is 2.8 trillion parameters. It's kind of like keeping track of all the different sizes gets a little cumbersome. Now something I want to touch upon here is that, okay, this is open weight, but like no one's going to be able to run this locally at home. So what is exciting about that? If it's open weight, but like no one can actually use it. It is true. Someone who has a system that is not five figures will have no hope of being able to run this. And even if you had a 5-figure plus system, you would either have to run this at such a extreme quantity that the intelligence would be degraded to the point where it would not compete with Fable 5 or even much smaller models. Or it would be so slow that you'd get like half a token or a token per second using a big CPU and system ram only server, which I am in the process of planning out to build because I really do want to run some of these locally. But the exciting part is that this being open weight means a bunch of different providers can download these weights and serve them themselves from their own infrastructure from their own physical locations. And that is exciting as a consumer who pays and uses these models because when we use anthropic or open AI we use their model from them and we pay them directly. There's no alternative options except for maybe some niche like more business use cases where oh it's being served from bedrock or something like that. But for the vast majority of people we have one option to use those models. When a big model like this that is performant is open weight or it's not able to be run locally on hardware we own in our own home. However there's a bunch of different providers that are going to be serving this. Some will have it being served faster. Some will have it being served cheaper and it really opens up a lot of different options
-
3:16
, obre el vídeo en una pestanya nova
for folks who want to use this model as opposed to the flip side of that coin where you have one option and you pay this one price. So even though it's not necessarily able to be run locally, I sometimes see the sentiment expressed like why should I care about these models if they're so big that there's no hope of running them. And I just think that the democratization of them actually being served by more people, giving the consumer more choice is still exciting even if you can't use it at home. So with that we don't really have much else to say about this. We know it's size. We know it's claimed performance and we know it is currently in preview. So if if there are some oddities to it. Hopefully there'll be ironed out by the time that this is no longer in preview, but I'm not gonna wait till this is officially released to do a video, because I'm very, very excited to test this. And of course we're going to begin just with the tried and true browser OS test V2.5, which we can see has now begun to actually be written. This thought for quite a while, I do believe that I sent it this initial prompt if we can scroll up all the way. Probably 10 to 15 minutes ago, and it thought for quite a while, If I click on this, does it show us specifically how long it thought for? It doesn't, but look as I scroll up. I don't know if it's easy to see, but the scroll bar is there. It's very small and it's still very close to the bottom of the page. So there was a large, large amount of thought that had gone through this result. And now we just reached at the top of that chain of thought. So very, very, very intricate amount of thinking. All right, so we have our browser OS test from the 2.4 trillion for MRE, 1 3.8 max preview. This is a very different UI splash screen that I've ever seen for this result ever. It said the only external thing was like some fun pack. I don't think that would be tripping it up. You're silly silly model. All right. So I'm actually happy to see that it just gave us the patch because it shows some behavioral behavior in that like it new like I'm not going to regenerate this entire script very good. Oh my I. Or it is subjective. Personally, I love it. I absolutely dig this UI aesthetic. It's just very, very, very different to anything I've seen before doing this test seriously. They're kind of all the same at this point, not this one. This is a very, very interesting style. Now, good. We can resize. I love this. I really like this. Okay. We do have a clock that is showing the correct time in our local. Both in the top right and bottom right of the screen so it really went hard with the clocks. Do we ever right click? Yeah we do. Okay and again, we're showing just this UI is really, really novel. I can say definitively that this UI is not showing any signs of distillation from anything that currently exists so that is good as well just for those of you who keep in a high on those distillation attacks. Nonetheless, let's check out our start menu. Okay, we have restart and we don't have the search.
-
6:11
, obre el vídeo en una pestanya nova
I would like to see a search from a big beefy model. Nonetheless, let's just run through them. I do have the sound on. Okay, maybe we can click on them from our start menu. Did you, gee, this car. Yeah, it's, you have to, being a popular influencer, the possibility does exist that these tests can be, of course, like properly created to benchmarks is the term that comes to mind. I'm not saying this is, I'm not saying it. I'm simply saying the possibility exists when doing some of these tests. So keep that in mind. However, a model of this size and this stated performance should definitely, very good crunch. That's disgusting. Should definitely have the capacity to pull out a result like this. Again, the one thing I'm noticing here that's just totally different is the UI aesthetic of this is entirely unique. And I personally like it. There's definitely some issues here. Oh, the police are seemingly having a drift session in the city commons. Nonetheless, we could probably stay away from them, but we want to check to see if the actual logic works to catch us. Okay, our health is going down, and we're wasted. Very, very different. I really, I don't know what to say. I like it. Void Patrol. Okay. These usually bore me. I'm not really into like space games. Oh, yeah. That's actually kind of cool. What is the planet with the rings would be Saturn? That is one singular ring though, so I don't know what, uh, oh, this is actually kind of cool and it's hard to, it's got that same style though in the way that it drew everything for the GTA game. If it keeps this aesthetic for some of our later prompts, this could be a very interesting review of Pete Lab. Let's go back to that one because I'll never to waste my entire much time and it and have to time lapse. It's another list. She's do help. I also don't like having a guest account in my own browser OS, but I suppose
-
8:13
, obre el vídeo en una pestanya nova
640k memory ought to be enough for anybody. That's actually like a meta computer science reference and that's clever. I like seeing that and Dressing so the matrix the matrix has you follow the hourglass. Is that from the movie? I don't think it is. Kelsey. Good. I know this probably premature to make this claim, but I think this model might be quite good. press Ctrl Z on the desktop trust me. Oh, that's the special feature. I don't know, the special feature is saved for last. All right. I should just delete this video and come out with this aesthetic because no one's ever seen this. This isn't like a really like AI looking thing. And I came out with a new anti-slop AI aesthetic, supposed to get up and then this will come out. No bug. Oh, they ripped stuff. I'm sorry that was a tangent. All right. System monitor. Very good. We have our frames per second. Our load and our windows. Settings. Accent. So this would change it to I should have started with that. I'd like to see some purple instead of you know what? It bodes well for how much I like this. The fact that I haven't been raging about the orange color tone here. Look at this though, like, night circuit. Is there any movement? No, but coral reef. Alright, that one's kind of like my first S.V.G. orbit. And this was the background for one of our games. Daybreak. It's just so like hipster. And we have the about which, wow, this is quite the, I did change the theme. Let's just check our special feature and then we'll go back to that beat lab. Oh, Control Z. Normally, I've seen special features like this before the difference here is that it's been recording from the entirety of us using this. Most of them like you take a static snapshot and then you can go back to that point in time. This one's just full on like Microsoft copi letting us. So that is definitely kind of interesting. And then the final thing which I'll turn the speaker up for a little, the beat lab, it made it like a sequencer. Oh, I didn't click stop. That was just like really like a, I can't say that on camera, but that was like a gigantic like, oh you're excited, like, you're all right. Stop stopping. Oh, and we can turn more on. That's the whole point here. The one thing missing here is a loop button. This was, honestly this was one of the best ones I've seen and that may be subjective because I like the aesthetics. So keep that in mind. I spent eight minutes on that, I believe. It's not that bad. All right, that was gold, yo, that slept. It's gonna be like, I'm glad you liked that. Let me know if you want any other changes like I could have this in, yep, of course. We're gonna do the subway test, and we're gonna do this in the classic way. Being that one, we're sending it the first prompt, which is just to make a good looking subway scene, two, we're doing it from within the website interface. All right, so we've received the result for our beautiful subway scene.
-
11:59
, obre el vídeo en una pestanya nova
I haven't even looked at it yet, but I've sent the follow up prompt, turn it into an awesome FBS. So here's our beautiful static subway scene, we're really in Avenue. All right, click to step onto the platform. That is extremely dark to the point where it's basically impossible to see anything that's not illuminated. And it makes, so this is a simple mistake that gets made sometimes. So I have an amendment to make to my statement that the brightness slider didn't work. It's actually not mouse control. It uses the plus and minus. So the problem is it's just still significantly too dark, even at the highest setting. But in looking at the code, something that I do want to say, I don't normally mention this because we don't look at the scripts. Maybe we should start doing that. It's very, very well commented. So there's just a lot of contract pit back wall behind the platform, far wall across the tracks, pipes and cable tracks. So everything is just very well commented here, which I like to see. All right, I made this a bit brighter just manually because I do want to actually be able to see some of this. So is that graffiti? I think it is graffiti but I don't know. It's still too dark. And this is the brightest setting. All right cool, we can see down the track. There is no train here, mind the gap. Inverted wet floor sign but still that's kind of cool to see. It's interesting that a lot of these are putting puddles on the floor of these results and that's not something that has happened until like all the subnetures did and now everything's doing it and then we have a nice exit here danger and then we also have the tunnel. All right pretty cool. No train though. It says it in the top right. No trains are running tonight. The platforms yours. Seems clever. All right so here's our subway scene that was turned into an FPS deadline. Interesting it's putting some lore in the bottom just with this like marquee scrolling thing. Platform lighting it's just turn that up all the way it's probably We still gonna be too dark. All right, let's check it. All right. Their ability to do weapon models is definitely improved in this generation of models can't ask for more than that. I'm a little let down at the fact that the sound seems to not be working, but that's a be-gen issue. So the speaker was off, and that's my fault. Very good. And it is possible, of course, that these results are benchmarks at this point because this is like a pretty well-known test due to my influencer status. Look at the actual graphics on the shirt, though, and stuff.
-
14:32
, obre el vídeo en una pestanya nova
Hello. Oh, oh, and we got a points for headshot. Look at the weapon model. We actually see the slide going back and forth when the ammunition's being fired. I don't normally see that. I don't think I've ever seen that in a matter of fact. Goodbye. We have footsteps. We have ambient noise and sound effects. reload is cool. Now the K3 result, they were like crawling out from under the track. I was like a freak of nature result. This is still not bad. Like, okay, so it just puts like a dark, like element over, still though. I mean imagine like a year and a half ago seeing this, it would be quite tough. We've cleared. That's interesting. So when I press shift, I'm about to press shift and run. It actually like tucks the weapon in a different orientation when we're running. Just interesting things to notice of thoughts that it implements into these sorts of things. So let's say that was actually, can we trim a lighting up more? No, okay, it's just a UI glitch. Still very good. Not K3 good, but good. All right, so we're going to also do, of course, the C++ skateboard game test. Now in the KMK3 video, I had changed this where the location needs to be a mall. Folks didn't like that. The reason I changed it is just in case the boardwalk California aesthetic one was benchmarksed on, but nonetheless we're changing it back. So this is the California boardwalk aesthetic, single file C++ gate game. In addition to that, I have just told it don't use Ray Lib, but everything else that it had in its plan was perfectly acceptable, and it has given us its plan right here. So I'm just going to, from within build mode, tell it to build. So this was just really odd and partially because it's in preview, but because I don't really have a lot of good documentation to go on and how specifically they should be configured. From within Open Code, it had a large, large, large amount of errors in trying to write the file. They were just happening one after another. So then it opted to write a Python file to generate the C++ file.
-
16:51
, obre el vídeo en una pestanya nova
I'm trying to scroll up to see if we can keep back to the point that will show us what specifically this did. Oh wow. Yep, and this did do it hardcore. I'll write a Python script to generate the C++ file instead. And if we scroll up, we can see some of the issues that we're occurring here. So partially an issue, maybe with Open Code, maybe partially an issue with the model being in preview. Nonetheless, it did find a way to actually generate this game for us. It seems to have compiled it without any errors. So good. And now we're met with the completed result. Let's go take a peek at it. Why do you? frustrating, very, very frustrating. I'm not happy with this and something I'm going to go on on a limb here and say is, I think if we do a kick flip right now, it's going to rotate the player and the board as one. Okay, I'm going to eat my words on that one and I'm happy with that. I just based off of some of the things I saw here, I was not very bullish. This is just, this is not good. I'm not happy with this and I want to give it a photo of this. I don't know that it's 100% properly configured to be able to see things visually here. So I'm going to just ask it first and foremost. Can you actually see the screenshot? That was an incredibly, incredibly long bow to thinking. I don't know if I'd call it overthinking, but I'm going to inevitably just have time lapsed it. So folks can make their own judgment. It wasn't as bad as the Kimmy K3 overthinking, but it definitely was going back and forth a lot. So nonetheless, it did finally come to a conclusion, and it's now going to try to fix all of these issues. made all the fixes that it decided it needed to. And again, before we take a look at this, I wanna reiterate the result of everything that we just saw was simply me saying to it, can you see the screenshot in this folder? I didn't know if I had everything correctly set up in terms of its multi-modal capability. So it went, it looked at the screenshot, identified a whole bunch of issues. And then the fixes that it made give us this. Okay good.
-
18:44
, obre el vídeo en una pestanya nova
Now, this is still like, it's not very impressive visually, But it did properly fix basically every single issue that we noticed, just based off of a screenshot. So that I'm going to say, I'm satisfied with, the kick flip seems to have become a bit messed up unfortunately. I will say it's actually done a better job with the text than a lot of models we see when given this specific test. Though, unfortunately, the graphics here are quite poor. I'm going to say this is not up to the level of what I would have hoped for a model of this size in this performance. So keep that in mind, we did go rail the rail there, that was pretty cool. And then we have these weird, like, I don't know what they are. So yeah, but it showed some interesting multi-modal coding capability or whatever, nonetheless, nothing more. All right, next up, I'm going to be giving this a CAD test where I'm telling it, I have a DC motor, a 280 size one, for my RC car project, and I need a 3D printable model of a V8 engine. That this little DC motor will fit inside. So when you create this, make sure that it's properly able to be printed and does fit the motor. I've begun this from within thinking mode here. Excuse me, from within plan mode. And something I want to bring up is I noticed a comment or two where someone had asked, does this system still have remnants of other examples of these tests being run? No. Every single result from the KimiK3 model test that was the most recent test is no longer on the system. Does not exist no logs, no nothing like that. So this won't be able to look into other models answers. it's coming up with all of this entirely on its own with no prior art on the system. I guess could be said. Okay, I'll just tell it. It is that one open-scad recommended is fine, recommended, multi-part. I'm gonna go with that because that's what K3 did, and it was really just like, quite something. All right, and we've answered the questions that will come up with its plan inevitably, probably quicker than where we're doing like a
-
20:40
, obre el vídeo en una pestanya nova
skate game, fiasco, and then we'll tell it to build it. All right, we got a nice plan. It was very succinctly done, and it seems like it will have a lot of cool elements to this motor. Also, don't know why cursors open. I must have missed clicked that. Alright, we've received our 3D printed V8 motor design. So let's take a peek at it. We'll just start with taking a look at the renders. Okay, there are no renders. That's all right. We'll just look at the entire thing V8 engine model. Okay. Tarn it. Ah, you've. It's made some mistakes. I've noticed quen models, especially the big state of the R1s have over the Time like they've struggled spatially This is definitely exhibiting that behavior now the sad part is it extruded a V shape But it's in the wrong orientation now again the parts are Separate so it's possible that this may print nicer than it's looking right here Nonetheless, I'm gonna probably have to say this is not worth printing Unfortunately, I will take a look at some of the parts individually though just to see. So let's just look at the block by itself. Yeah, that's that's unfortunately just a trocice because this would in no way be usable for now. This is something's gone quite wrong here unfortunately and that is pretty much the story when we look at the entire thing. So not quite right. There are elements here that do look good like the valve covers. covers the headers in particular which are these things right here look pretty nice. There is even seemingly an oil pan down there but it falls apart when it comes to like just general coherence for the entirety of the shape. A little disappointing. Next up I want to just try like a random piece of hardware test so I have this it's a little display development board that I bought a few of these for micro center just because it's always fun to have stuff like this flying around. So I'm going to plug this into the computer and I'm going to say there's a USB device plugged in. Figure out what it is. The thing is this has a little screen on it so then we'll have it maybe write some unique or intelligent animation to be displayed on the screen for this little microboard. That's all I'm going to say and we'll see what it does. This is still leaving a lot to interpretation. So it's going after probe the ports figure out what the hack is is exactly plugged in here, try to get some dimensions for what the screen is showing. All right, slight change of plans. I've put it in its own little folder, and I've changed what I wanted to do. I've told it that I wanted to show system GPU utilization on this tiny little screen, while this is connected to the system. I said stay in your folder, because it went in with searching and finding information from stuff
-
23:23
, obre el vídeo en una pestanya nova
that is not at all pertinent to this, like other types of displays that I've worked with, not for model testing, It was like assuming it was one of those. So it just wasn't quite right. So the specific name for this is an ESP32-C6. It would be that expressive USB serial unit is the one that it specifically. OK, it's calling it a C3. Good. It's an ESP32-C6, not C3, recompiling with the correct target. OK, now the error seemingly gave it that bit of information. But however it gets to that solution, it's fine with me. Easy as fixed. Have it join your Wi-Fi and receive GPU data over UDP. I like that one of the options that it gives us here is try harder. That is what I'm going to opt to do, but I also applaud it for giving that as an option and not just trying to use the simplest one. So good. I'm getting a little tired of holding this up, but nonetheless we'll do it for science. I must be honest that I'm starting to lose a little bit of hope that this actually maybe something we can get done with this current setup. I think these little boards are probably fairly new, at least I've never seen them before, and I try to keep tabs on some of this stuff, but it may have been around for a while, though it does not seem like it's able to make any progress actually writing to this or doing anything. It's just now stuck on the same screen that it was the first time I plugged this into power before even trying to attempt to do anything with it, so we may be getting to the point of an impasse, I would say. Alright, so I can't get it to do anything now. I've tried plugging it into multiple different computers, trying a bunch of different button presses while plugging it in. Well, it's plugged in. And something has gone awry, so I'm going to call this test to fail for now. And maybe we'll revisit this when this is out of its preview. So I got the screen to just light up. It's totally black, but you can tell that this thing's actually on just from another different computer I tried it on, but nothing shows up at all. So I think it may have actually properly wiped the flash on this. It's still, I can't get this computer to actually communicate with it right now. Again, we'll probably save the remainder of this for the dedicated general availability test. But it seems like it may be an interesting test because it posts them difficulty so we can perhaps keep this around. Next up we're going to do the city timeline test where it just shows a city over a number of different year periods and there has to be nice transition effects between them and things like that. This is always fun to do because they sometimes will put funny or cool niche detail in specific time periods. So apparently in four minutes we've received our city timeline prompt. That significantly quicker than I believe I've seen from any model when giving them this test. So we'll see what we get. Make sure my speaker on. definitely checks out. This is something is wrong here. This is I'm pretty sure
-
26:31
, obre el vídeo en una pestanya nova
Quinn 27B would have knocked this out. So we have people, we have old cars. The core is there. Okay, the transitions between the scenes do work. I almost wonder would have started this in plan mode made much of a difference. Okay, 255 we have more hovering things and a lot of neon. So let's see like that school, it's just the rest of this is, you know, and folks have VR headsets on, they do still use traffic lights in 255. Alright, let's let's give it some feedback. So I'm going to give it some genuine honest feedback and we'll see what it does. The user is very unhappy with my previous attempt. They wanted to polished high end thing and what I delivered was essentially a bunch of primitive boxes and so on. There's very basic geometric shapes that barely represent a city. I did see someone online just mentioning that they had significantly better luck with this when telling it like you need to think very, very intricately. I believe the terminology they used was like just getting very angry with it in all caps. So it's possible we'll respond well to negative feedback. Okay, so So it's still not very good. I guess it's the takeaway. Ration books here. It's the models in preview. And that's why I'm gonna perhaps hold some of my judgment for now. But okay, I do like the transitions between them. I think that's pretty cool. Do the cars have fins, they do. Only one model neglected to do that. I think it was 5.6 sold, but don't quote me on that. Okay, that's kind of cool. 2005 looks better. The cars do have LED strips. Is it raining? I think it's raining. And then 255. Cool, and we have, oh wow. Very, very, all right, just out of curiosity and because why not, I yelled at it after this result and said that I wanted to have something that's like at least 5,000 lines of code. It worked very, very hard to ensure that it was over 5,000 to the point where it was like, okay, I have 4,100 lines, and then it would make an edit and it would have like 4,150. And I'd be like, okay, now I have 4,150. It did show me that it seems to want to pursue a goal as ridiculous as the goal may be. So this will either be incredible or complete disaster. Okay, it's somewhat better. It's not as intricate as I would have expected the building density seems to be a bit more. We can see differences though unlike the vehicle models actually seem a bit more detailed here. I'm not gonna spend a lot of time on this, but I just found it interesting that,
-
29:37
, obre el vídeo en una pestanya nova
okay, that scene right there looked cool for about a fraction of a second. It was just very interesting how this worked really, really, really hard at ensuring that it was that 5,000 lines of code. It did add weather and things like that. So we see the clouds floating here, interesting. We have taxis there. So this is perhaps some form of a New York aesthetic. We also have a bus. Yes, the vehicle movement is definitely a big issue. There were news stands there as well. Okay, the cars are very glossy and they do have LED strips for lights as well. The street lights have changed. It does seem like it's raining in 2025. Maybe this is the model, perhaps giving its opinion on the current state of things, or it's just raining. And then, 255, we have blue Christmas trees and still the, oh wow, okay. All of the buildings now have, all right, interesting. But that's gonna conclude our first look in test of Quen 3.8 Max Preview. And I'd very much like to emphasize the preview part of that name. I was not overall very impressed with the results we've seen, especially considering the raw size of this model, which is gigantic. However, this is still in a preview state, so it is very likely and very possible that once it becomes generally available and it is out of preview, there will be some pretty significant changes. And part of the reason this is in preview is to get feedback and perhaps see things like we noticed today, we're okay like the city time resolved. It just was like irritatively bad and maybe there's something causing it to not want to go as hard as a can on stuff like that. Additionally, the little micro controller test that we did was perhaps a waste of time being that nothing interesting really happened but nonetheless it gives us a basis and reference point to be able to test this. If not in the same exact test something similar in scope once it's out of review just to see how good it becomes. I think really the takeaway and forget doing like a results overview or something like that. This is very exciting because it's going to be open-weight. It is the first max model from Ali Baba and the Quint family that will be. It's 2.4 trillion perimeters, so it's absolutely massive and it represents a continued, at least commitment to open-weight from Ali Baba. A lot of folks have been wondering, is this kind of the end, like they haven't really done anything in terms of updated open weight models. So seeing another one coming out is very, very bullish for the future of open weight and accessibility to models. This browser OS test was fantastic. Again, it may be subjective because I really liked the aesthetic, but I will say the UI style here was definitely its own unique and very different. So I like that and I'm interested to see what folks experience in testing this model. So with that, that is going to conclude first look in test of Glenn 3.8 Max Preview. If you have any questions please feel free to leave them in the comments and thanks for watching.