Vuit IA construeixen un generador de vídeo local: el resultat amb 12 GB de VRAM
WeeklyHow posa diversos assistents d’IA a treballar en una aplicació local de text a vídeo amb narració. El resultat funciona, però revela els límits de qualitat, maquinari i coherència dels models oberts.
WeeklyHow proposa un experiment ambiciós: posar diversos assistents d’intel·ligència artificial a escriure una aplicació local capaç de generar clips de vídeo i narracions a partir de text. La màquina disponible té 12 GB de VRAM, una limitació que condiciona tot el projecte i obliga a distingir entre entrenar una IA nova i construir una eina amb models ja existents.
El resultat és una aplicació funcional amb interfície web, diagnòstic de maquinari, generació de clips i síntesi de veu. Ara bé, les proves també deixen al descobert errors d’anatomia, text i coherència visual que els models locals encara no resolen bé.
1. El repte real: crear una aplicació, no entrenar un model des de zero
El vídeo comença preparant amb ChatGPT una instrucció detallada per als agents de programació. L’objectiu és obtenir una aplicació local de qualitat de producció que generi vídeos narrats a partir d’un prompt, amb text a vídeo i veu sintètica.
La primera conclusió de la mateixa IA és decisiva: 12 GB de memòria gràfica poden sostenir una canalització local competent, però no permeten entrenar des de zero un model fundacional modern de vídeo. Per tant, malgrat el reclam de “fer que una IA creï una IA”, el projecte no fabrica pesos nous. Els assistents escriuen el programari que coordina models oberts prèviament entrenats.
Aquesta distinció és important: la IA generativa actua aquí com a desenvolupadora i integradora, no com a substitut del procés caríssim de recopilació de dades i entrenament d’un model de vídeo.
2. Diversos agents construeixen el backend i la interfície
El desenvolupament es divideix en tres fases i s’executa dins de Cursor amb diversos models d’IA. Després de la primera fase ja hi ha un paquet de Python instal·lable, un backend i una interfície web més completa del que el creador esperava.
L’aplicació final separa les funcions en diverses pantalles:
- generació de clips;
- narració de text;
- selecció de models i configuració;
- diagnòstic de maquinari;
- generació simulada per comprovar codificadors i dependències.
El diagnòstic detecta la GPU, PyTorch i FFmpeg. Durant la demostració cal corregir una dependència de PyTorch que apareix com a absent, un exemple dels problemes pràctics que continuen requerint intervenció humana encara que bona part del codi l’hagin escrit agents.
3. La primera prova: cinc segons de vídeo vertical
Per avaluar el sistema, WeeklyHow utilitza el prompt “Will Smith eating spaghetti”, una referència habitual als primers vídeos generats per IA que mostraven figures humanes molt deformades. Configura un clip de cinc segons, relació vertical 9:16 i el perfil estàndard de 480p.
La primera sortida recorda precisament aquells resultats antics. La persona no s’assembla a l’actor, apareix sostenint un plat de manera estranya i el moviment de la boca, les mans i els dits és poc coherent. El model entén la composició general —una persona, menjar i una acció—, però falla en els detalls fins i en la continuïtat temporal.
Una segona prova amb text confirma una altra limitació: el model tampoc no representa bé les lletres dins del vídeo. Els ajustos del prompt milloren una mica la imatge, però persisteixen artefactes com dos polzes, plats duplicats i reflexos o persones que apareixen sense una causa clara.
4. El salt al model de 14.000 milions de paràmetres
L’agent proposa dues vies de millora: afinar els prompts o canviar el model base. Després d’hores de proves, una actualització del sistema i dels controladors permet executar la variant més gran de Wan 2.1, de 14.000 milions de paràmetres.
Amb el mateix prompt, el vídeo guanya detall i context. La taula incorpora una ampolla, el menjar sembla més creïble i apareix una segona persona, tot i que no s’havia demanat. La figura principal també s’aproxima més a alguns trets generals sol·licitats pel prompt.
La millora, però, no elimina el problema principal: els dits i les mans continuen deformant-se. El vídeo mostra que augmentar la mida del model pot millorar la riquesa visual i la comprensió de l’escena, però no garanteix una anatomia estable ni un control exacte de tots els elements.
5. Narració local amb Kokoro-82M
L’altra peça del projecte és una pantalla de síntesi de veu basada en Kokoro-82M. La interfície permet triar veus agrupades per accent i gènere i ajustar-ne la velocitat. El motor es carrega, sintetitza l’àudio i es descarrega després de cada petició, de manera que no ocupa la GPU mentre s’està generant un vídeo.
La primera narració de prova dura vuit segons i sona prou natural. El creador prova després diverses veus i considera que una d’elles és especialment consistent. Aquesta part demostra que és viable incorporar veu local amb un model relativament petit.
Hi ha, tanmateix, una expectativa que no es compleix: la pantalla no narra automàticament un vídeo carregat ni sincronitza els llavis. Funciona com un generador de veu independent, comparable en concepte a un servei de text a veu, i encara caldria unir explícitament el clip i l’àudio dins del flux de producció.
6. Què han creat realment les IA?
El resultat és valuós com a prova d’enginyeria: els agents han aixecat una aplicació amb frontend, backend, configuració, comprovacions del sistema i integració de models locals. També han accelerat la resolució de problemes i han proporcionat alternatives quan els primers resultats no eren satisfactoris.
Però no és un model de vídeo entrenat de zero. És una capa d’orquestració sobre Wan 2.1, Kokoro, PyTorch i FFmpeg. Aquesta arquitectura és molt més realista per a un ordinador personal, perquè reutilitza models oberts i concentra l’esforç en la interfície i el flux de treball.
El projecte també evidencia els límits de l’automatització: el creador ha de respondre preguntes sobre el maquinari, instal·lar dependències, actualitzar controladors, interpretar errors i decidir si prioritza qualitat o consum de recursos.
7. Un prototip funcional, però encara lluny dels millors generadors
Els clips finals funcionen i demostren que la generació local és possible, però WeeklyHow admet que la qualitat queda per sota dels models de vídeo més recents. L’equip disponible tampoc no permet afegir fàcilment entrenament o ajustament avançat sobre el model gran.
Els punts forts del projecte són la privadesa local, el control del programari i una interfície que fa accessibles eines complexes. Els punts febles són la lentitud, les exigències de memòria, la fragilitat de la instal·lació i els errors visuals persistents.
Conclusions principals
L’experiment demostra que diversos assistents d’IA poden construir una aplicació multimèdia local força completa, però també obliga a rebaixar el significat de “crear una IA”. Amb 12 GB de VRAM és possible coordinar un generador de vídeo i un sintetitzador de veu; no és realista entrenar un model fundacional modern des de zero.
Wan 2.1 ofereix una base funcional: la variant gran millora clarament el detall respecte de la petita, però les mans, el text i la coherència continuen sent punts febles. Kokoro-82M aporta una narració local convincent, encara separada de la generació del clip. El producte final és, sobretot, una demostració útil del que avui poden fer els agents de programació quan treballen sobre models oberts i maquinari de consum.
Contrast i context
Fonts consultades
-
01
YouTube · WeeklyHow I Made AI Make AI (Video Gen Model)
- 02
-
03
Hugging Face · hexgrad Kokoro-82M
-
04
PyTorch Documentació de CUDA a PyTorch
-
05
FFmpeg Sobre el projecte FFmpeg
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
A lot of you guys have been asking me to make a video gen model with AI, and so in today's video, I'm going to be using all of these AI models like GPT 5.6 Sol, Fable 5, and Grog 4.5 I guess. And they're the best models on the market right now, and I'm pretty sure with all of these models, we'll be able to make an AI video generation model. So without further ado, ladies and gentlemen, let's start building the prompt using chat GPT. Okay, so the first question is, what hardware will run the model? Okay, so I answered all the questions about my hardware and what video generation approach. We will do text to video from prompts, and the audio generation approach is AI voice narration. The key constraint is now clear. 12 gigabytes is enough for a capable local pipeline, but it's not enough for trading a modern video foundation model from scratch. Yeah, I knew it's not going to be possible. Okay, so here's the prompt,
-
0:54
, obre el vídeo en una pestanya nova
build a production quality local application that generates complete narrated videos from a text prompt. Okay, let's just copy this prompt and then go back to cursor and paste here. All right, let's send this. All right, guys, this is going to take a while. So while we're waiting, let's just talk about what I know about image and video gen models. Again, I am not an AI engineer. I'm not an expert, but from what I know, image gen models uses some sort of algorithm called Gaussian denoicing or something. And for example, there's this image. And for each time step, the model adds noise to the image until it's just a bunch of noise. And then it will denoise that image to make sense out of it. That's all I know about image models. But when it comes to video gen models, from what I know, video gen models do the same thing, like instead of a single model, they have multiple images that need denoicing or that need this algorithm, right? That's all I really know. So you guys can
-
1:52
, obre el vídeo en una pestanya nova
fact check me in the comments below. All right, let's just move on. All right, so it is finally done. What do we have here? So we have the backend already. So it's an installable Python package with this and that we have all these files looks good. We also have a front end. Okay, I did not expect that I thought we just need to run like a Python script. And that's it. But we are not done yet, because we still have faces two and three. So we just continue to face number two. And this time, I'll use GPT 5.6 all. Let's proceed to phase two. Okay, day two guys, because yesterday I was related to I had to get some sleep and but I were back, we're back. What we're going to do next is proceed to phase number three. So we can just say, let's go. Let's do phase three. I'm gonna teach the fable five once again. All right, guys, looks like it is almost done. So phase number three is complete and verified and to end on your machine co-coron narration works fully offline, placed in the browser and acceptance script confirms TTS runs cleanly after one 2.1 video model is
-
3:00
, obre el vídeo en una pestanya nova
unloaded. Okay, so we got a couple of things here. We got the backend fully built and we also have the front and fully built. We have this new narration screen with a voice speaker grouped by accent or gender. Oh, you know what, let's stop reading and let's get to it. And here is the app. This looks super basic. We got a couple of things here. We got generate clips. We got narration, we got models and settings, hardware diagnostics and mock generation. Let's take a look at the models and settings first. Okay, looks like this all we got here. What about the hardware diagnostics? Okay, so I have the GPU, I have the PyTorch looks like we still need to install this. I thought I already installed it, but it's still saying missing. And there we go. We have FFmpeg. Cool. I think we are ready. Okay, so if you're curious about this mock generation page, it's literally just a page where you can test your encoders. There's nothing really in here. Now let's just proceed to generate clip. And now we can just say will Smith eating spaghetti. If you guys have no idea what this is,
-
4:01
, obre el vídeo en una pestanya nova
it's where everything started. I think it was around 2022 and they released the very first video gen model benchmark. And that video is will Smith eating spaghetti. Yeah, the more you know, for the duration, we'll just set it to five seconds for the aspect ratio. We'll just set it to nine by 16. The preset standard is for ADP. We'll just set it to standard or hold on standard for ADP high quality for ADP. What? Okay, let's do this generate local clip. Here we go. Okay, so it's finally done. First of all, that is not real Smith and looks like it's holding a plate. So maybe it's about to eat spaghetti. Okay, so let's just play it. Here we go. All right, there we go, guys. We got the classic spaghetti video. Yeah, this does look very similar to those AI videos or old AI videos that just look bizarre. Like at this point, the model still doesn't really know how humans work or, you know, our anatomy. I mean, it knows we have a head, a face and two arms, but still goes wrong all the time with the fingers and the mouth. So like, those tiny details are still very hard
-
5:14
, obre el vídeo en una pestanya nova
for the model to arm. Perfect. And the only thing we can do, I think, is to change the model. Let me try and generate another video. Let's do naked gradaway. I tried generating a video with a text. And yeah, it's also not capable of making videos with text. I mean, I'm not really surprised to be honest. If it can't perfect this spaghetti video, there's a high chance that it won't be able to make the text. So yeah, I went back to cursor and asked the agent to improve this project and it gave me a couple of options. First is to make adjustments in my prompts. And I did that and here is the result. Okay, I'm gonna say this looks a little better, but still kind of the same, because there are a couple of issues I noticed. First is if I take a look here, there are two thumbs and his point fingers just doesn't look good. And of course, we can see that there are two plates. This guy right here must be very hungry. But wait, hold on. Maybe this makes sense, because this one right here is just the spaghetti. But this one right here kind of looks like a
-
6:11
, obre el vídeo en una pestanya nova
different dish. Another issue I noticed is something happened here. Look, there was a reflection right over there. That is a person right there. See, the video is generated by this model is not that they're not as good as the videos generated by the recent models, you know, like nano banana. This is absolutely delicious. I think I nailed the sauce this time. But I guess I'm just asking too much. Okay, so what we can do next is upgrade the backbone. This time, we're going to use one 2.14 billion parameter. I've been working on this project for hours last night and I couldn't make it to generate a video. But today, I managed to fix it, but I wasn't recording my screen. So yeah, honestly, I'm not really sure what fixed the problem. But after updating my system and the drivers, it just worked. So here I generated a video using the same Will Smith eating spaghetti prompt. And this one doesn't look a little better. First of all, there are far more details compared to the previous ones. We can see that there is a bottle here, like it understood that
-
7:20
, obre el vídeo en una pestanya nova
we are dining right now, so there must be something to drink. There is also another person here, which really surprised me. I didn't really mention in my prompt to include another person, but for some odd reasons, the model just decided to add a woman. Anyway, the food looks great. This time it does look appetizing. I would honestly eat that. Now here is what really surprised me. In my prompt, I said quote unquote, Will Smith eating spaghetti. And you guys see what I'm seeing, right? The person is the same race as Will Smith. I guess this confirms that the model is much better than the previous one. But as much as I like the result, the same problems still exist. I'm talking about the fingers. This right here. It just looks really weird. So yeah. But it all serious as though this looks phenomenal. Now there's actually one more page that we haven't seen yet. And that's the narration. And in this page, we have this. So generate speech locally with cockro 82 million parameters, the engine loads, synthesizes and unloads per request.
-
8:26
, obre el vídeo en una pestanya nova
It never runs on the GPU while a video job is active. Okay, let's give it a try. Okay, so we have an example here at sunrise, a small explorer entry to forgotten greenhouse. Among the dust and broken glass, it found something still alive. And underneath of that, we have voice. Oh, we got plenty. We also have speed. Okay, let's give this a try. Here we go. It's synthesizing. Okay, so here is the generated narration. It's eight seconds long. Not bad. This is gonna be good. Let's find out. At sunrise, a small explorer entered the forgotten greenhouse. Among the dust and broken glass, it found something still alive. That's actually really good. To be fair, I'm a little bit disappointed because I thought we would be able to upload a video and then the video will be narrated or, you know, lip synced, but it doesn't look that way. It just generates audio like 11 labs or something. Anyway, let's try more voices and see how they sound. Your mom, sometimes she would tell stories. Subscribe to weekly how if you haven't yet.
-
9:28
, obre el vídeo en una pestanya nova
That's a massive car. She said a car that's right, a massive car. She didn't mean to say that like brother. Sagas, Kajus, geeks, ricks, ricks, ricks, ricks. Six, seven, six, seven, six, seven. What the fuck is going on here? Shut the fuck up, George. What did you say? You fucking guy stop fighting. Let me get some of the jokes aside. I think they call us the best one here. She sounds weird, but she's the most consistent out of all the voices I've tried. So yeah, guys, that is the narration. Anyway, I went back to generating more videos and here are the results. We can clearly see that it's not as good as the more recent models and as much as I'd like to improve it or, you know, train something on top of it. I just don't think my system is capable of handling that. But let me know the comments below what you'd like me to do with this project or what you want to see next. Hey, Nicole. Yes, weekly how do you want to say something? Give me your.