quantització vídeo amb IA Wan 2.2 LTX-2.3 ComfyUI FP8 Q4

Quantitzar vídeo amb IA: Q4 aguanta, però l’àudio cau abans

Una comparació de Wan 2.2 i LTX-2.3 mostra que Q4 sol ser el punt dolç, Q8 pot superar FP8 i l’àudio es degrada abans que la imatge.

Executar un generador de vídeo amb IA en una GPU domèstica sovint obliga a reduir la precisió dels seus pesos. La pràctica habitual és descarregar una versió Q4, perquè ocupa aproximadament la quarta part que un model de 16 bits, però aquesta comoditat obre una pregunta: què es perd realment en comprimir-lo?

Alex Ziskind compara vuit nivells de quantització en dos models diferents, Wan 2.2 i LTX-2.3. Manté el mateix prompt, llavor i configuració, i canvia només el format dels pesos. Els resultats deixen tres idees pràctiques: el format importa tant o més que el nombre de bits, l'àudio pot degradar-se abans que la imatge i no hi ha un únic llindar vàlid per a totes les escenes.

1. Una comparació controlada de FP16 fins a Q2

A 00:35, Ziskind presenta els dos casos de prova. Wan 2.2 genera vídeo, mentre que LTX-2.3 produeix imatge i àudio sincronitzats. La documentació oficial de LTX defineix LTX-2.3 com un model audiovisual DiT que genera ambdues modalitats dins del mateix sistema; per tant, la quantització pot introduir errors visuals, sonors o de sincronització.

Totes les versions s'executen en una RTX Pro 6000 Blackwell amb 96 GB de VRAM, prou capacitat per carregar els models de precisió completa. El muntatge evita que una versió petita guanyi només perquè la gran pateix descàrrega a RAM. Per a Wan s'utilitzen cinc escenes, des d'un got d'aigua fins a persones i mercats nocturns. LTX afegeix diàleg, text a la imatge i coherència de veu.

La comparació combina inspecció humana amb SSIM, LPIPS, alineació amb el prompt, taxa d'error de paraula i distància entre espectrogrames. A 02:03, el vídeo adverteix que cap mètrica substitueix veure i escoltar els clips. També és una prova exploratòria amb pocs prompts, no una certificació universal de cada quantització.

2. Q8 supera FP8 tot i utilitzar els mateixos bits

La primera sorpresa arriba amb un got d'aigua. FP8 redueix Wan d'uns 27 a 14 GB i, vist sol, el resultat sembla acceptable. Tanmateix, a 05:29, la distància LPIPS respecte a FP16 és 0,19. Q8_0, també amb vuit bits per pes i una mida semblant, baixa fins a 0,07: s'acosta molt més a la referència.

La mitjana dels cinc prompts repeteix el patró. FP8 s'allunya aproximadament el doble que Q8. En l'escena d'un cotxe, a 06:37, FP8 fins i tot canvia la direcció del moviment i fa que el vehicle circuli enrere. No passa en les altres variants.

LTX-2.3 mostra una diferència similar. A 08:45, Q8 ocupa uns 23 GB davant dels 46 GB de BF16, però conserva millor la imatge i l'àudio que FP8. Aquesta observació no significa que FP8 sigui sempre pitjor: depèn de la implementació, l'escalat, el maquinari i el model. Sí demostra que «vuit bits» no identifica una qualitat concreta. La representació numèrica i la manera d'agrupar els pesos són part de la decisió.

3. Q4 és un bon compromís per a moltes escenes

Entre Q6 i Q5 la degradació és gradual. Quan arriba Q4, Wan ocupa uns 9 GB i ja pot encaixar còmodament en moltes GPU de consum. A 10:28, la cara és menys nítida i canvia una mica la identitat del personatge, però els paisatges, el cotxe i les escenes generals encara es mantenen.

Per això Ziskind descriu Q4 com el punt dolç: redueix molt la mida sense convertir automàticament el vídeo en inservible. És especialment raonable per provar prompts, ajustar un flux de ComfyUI o generar material on la identitat exacta no és crítica.

El repositori oficial de Wan 2.2 ajuda a entendre per què el model complet és exigent. La variant A14B utilitza dos experts d'uns 14.000 milions de paràmetres: un treballa les primeres fases de soroll i l'altre refina els detalls. Hi ha uns 27.000 milions de paràmetres totals, encara que només 14.000 milions estiguin actius a cada pas. Comprimir els pesos facilita l'execució local, però qualsevol error pot propagar-se durant moltes passes de denoising.

Q4 no és una recomanació cega. Per conservar el rostre d'una persona, text petit o moviments precisos, Q5, Q6 o Q8 poden justificar més memòria. Per a previsualitzacions o objectes simples, Q4 pot ser més que suficient. El tipus de contingut fixa el mínim acceptable.

4. En LTX, l'àudio cau abans que la imatge

La descoberta més fàcil de passar per alt apareix en LTX. De Q5 a Q4, la distància de l'espectrograma mel puja de 9,9 a 46,9, gairebé cinc vegades. En canvi, a 11:56, les mètriques de vídeo continuen prop dels nivells anteriors: SSIM al voltant de 0,88 i LPIPS de 0,07.

Si l'avaluador només mira fotogrames, pot concloure que Q4 ha preservat la qualitat. En reproduir el clip, però, les veus incorporen reaccions estranyes, canvis de timbre i paraules que no existien a la referència. La taxa d'error de paraula ja havia detectat en FP8 un riure afegit que Whisper no trobava en BF16.

A 12:33, Ziskind formula una regla útil: l'audiència tolera una imatge imperfecta més temps que un so molest o inintel·ligible. En models audiovisuals, la prova final ha d'incloure auriculars, transcripció i sincronització labial. Una miniatura nítida no garanteix un clip publicable.

5. Q2 trenca persones i moviment complex

A dos bits, Wan baixa fins a uns 5 GB, però les escenes humanes perden identitat i consistència. LPIPS arriba a 0,54 en el personatge i a 0,57 al mercat nocturn, aproximadament un 60% més que Q4. A 13:05, cabells, cares i moviment comencen a desfer-se.

LTX segueix el mateix camí. A 13:59, la imatge es deforma i la veu sona robòtica; en altres clips, la persona ja no mira a càmera o sembla un subjecte diferent. Q3 també pot travessar el llindar quan hi ha rostres, roba detallada i diàleg simultani.

No obstant això, un got d'aigua estàtic conserva una puntuació SSIM de 0,74 fins i tot a Q2. D'aquí neix la tercera conclusió: no existeix un precipici universal. Un objecte simple amb poc moviment resisteix més compressió que una cara. I a 18:00, una bola que baixa per una rampa té una física incorrecta fins i tot en precisió completa; aquell defecte és del model base, no de la quantització.

Conclusions

La prova avala Q4 com a punt de partida pràctic, no com a garantia. Estalvia molta memòria i manté resultats convincents en diverses escenes, però cal pujar de precisió quan importen la identitat, el text, la física o una veu neta.

La lliçó més general és escollir per evidència: comparar formats, no només bits; mirar els clips complets, no només fotogrames; i escoltar l'àudio per separat. Q8 pot superar FP8 amb la mateixa mida nominal, Q4 pot conservar la imatge mentre deteriora la veu i Q2 pot funcionar amb un objecte estàtic mentre destrossa una persona. La quantització adequada depèn del model, del contingut i de l'error que el projecte pot tolerar.

Contrast i context

Fonts consultades

4 fonts
  1. 01
  2. 02
  3. 03
  4. 04

Font de treball

Transcripció amb marques de temps

12 fragments
Consulta la transcripció
  1. 0:00 , obre el vídeo en una pestanya nova

    Here's a thing that nobody tells you. If you're running an AI video model locally, you're running it quantized. You just... R. Come for you, I probably the most popular tool for generating videos. It has your GGUF version, and most workflows assume you grab the Q4 on day one and never looked back. Basically means the model weights are quantized down to 4 bits. But nobody actually tells you what you give up to get there. So, I wanted to check it out. I took two video models, and I ran each one all the way down and eight step ladder. From full precision at FP16 down all the way to 2 bits. Some of these look okay into bits, actually. through that. When 2.2, 14 billion parameters text to video and LTX 2.3, 22 billion parameter model, this one is a bit spicier, it's because it generates the video and the audio together, which means now the picture and the sound can break separately. I ran the same prompt, the same seed, the same settings, every single time. And as we walked out later, only one thing actually changes just the quantization. And all of this on a custom Linux box that I built running in RTX Pro 6000 Blackwell. That's 96 gigs of RAM, which comfortably fit all the precision models that I have. Now, here's what I didn't expect. Two completely different architectures built by two completely different teams and they eventually kind of fall apart in the same way. Let's go through it. So this is the reference for when I got Fp16.

  2. 1:33 , obre el vídeo en una pestanya nova

    Looking good, looking good. For LTX, it's basically the same thing as BF16, just slightly different quantization, but they're both full format or full precision. And this is what everything else we're gonna compare, gets compared against. For when, I ran five prompts, some are simple objects, some are humans, some busy scenes for LTX. I measured a couple more things. Whether the words state correct, does this support the new Phantom 5090? And whether the audio still sounded intelligible. try turning it off and on again. Then I scored the videos with some common benchmarks like SSIM, L-Pips and Promp the Lime and Test. Plus the test of my actual eyeballs and ear holes, but enough about my holes. Let's see the next level so we can actually get something to compare it again, shall we? So here's when Fp16 the whole mile is 27 gigabytes, quite chunky. And I do see... Pretty good character consistency looks pretty good except the very beginning two frames or so where it looks like an old I don't know TV CRT screen but after that it balances out and the character remains pretty consistent throughout the leaves looked normal. Let's see LTX. Does this support the new Phantom 5090? Oh yeah 24 lanes wait Does this support the new Phantom 50? That's crazy. Yeah my prompt had him doing something funny but uh that was just nuts. That's me and Dan at Microsenter, but our voices are just not at all the same. Although...

  3. 3:03 , obre el vídeo en una pestanya nova

    The voices are very clear, very consistent, and after analyzing both the image quality and the audio quality on this one, it's kind of hard to see, but these are our bass lines, and I will go through these very soon here. This tiny thing has quietly become one of the most useful tools I carry. As someone who constantly is bouncing between meetings, conferences, and content, I need a better way to keep track of the details. The hard part for me is I cannot stay fully present, listen, and get footage, and take clean notes all at the same time. So now I use... plot, note pin S, like a second brain for meetings, interviews, and event days. Most note taking setups still leave me doing the worst part afterwards, sorting out who said what and what actually matters. This thing is actually really simple to use. One click starts the recording, the physical button helps mark key moments, and then it organizes everything for me after. It can capture up to 20 hours nonstop, and then turns that into transcripts with speaker labels, clean summaries, and actionable to do lists, instead of one giant recording that I never revisit. I also really like ask plot. It lets me pull up past information or think through next steps without digging through everything manually. The big win for me is less mental load and better follow through. And I'm saying that as someone who's been using plots since 2024, my went through all the different versions. Obviously, use it where recording is appropriate and always get permission. But as a workflow tool, this thing has really been useful for me. Use my code, Alex 15 off for 15% off. There's also a prime day deal. and plus 30 day free return policy. Check the link in the description. Now we cut to half the bits and this is where it gets weird. Eight bits per weight, but specifically the FP8 format, so floating points, half the bytes of FP16. FP8 is natively supported on this GPU. So that should be a slam dunk. FP8 is 14 gigabytes on disk, how's the quality? So here I started with the simplest thing that I rendered and this is just a glass of water, just sitting there.

  4. 4:59 , obre el vídeo en una pestanya nova

    cameras moving around just a little bit. Whatever changes you see here is just the model basically rendering things differently. Even though the seat is exactly the same, the FP60 version looks just a little bit more realistic to me except for what are these lines walking across the table. I don't know it might be raining or something or something else is moving in the room. On the right it's a little chopper and a little bit more blurry, but if it wasn't next to the FP60 image it would just pass. Here's how things change. FP8 is already off the baseline here. This is the static detail video. Point one 9 on L-Pips and L-Pips is basically like my eyeballs in software. If it was zero that mean it's identical and the higher the numbers it means it's more different from the original. Now here's the surprise. The very next level down it's not really down. It's parallel. I'd say it's Q80. So it's still 8 bits per weight but in a different format it's integer 8 with. K-Quant grouping. On the glass video Q8 lands at 0.07. It's almost identical to full precision while FP8 is about twice that much. And yeah, look at that, the glass actually looks almost identical. to the FP16 in the video itself, according to my balls and I. Q8 pretty much the same disc size as FP8. But better fidelity. And guess what? The average across all the 5 prompts here, FP8 is still about twice as far from baseline as Q8. So, FP8, the format with hardware acceleration, everybody said was the future. It drifts further and further from full precision than the older int8 approach. Same number of bits, the difference is just where it's spent

  5. 6:36 , obre el vídeo en una pestanya nova

    that precision. And here is the kicker, the red car. A simple motion prompt. There's FP16. And FP8, the car drives backwards. Not in any other quantization does this happen. Just FP8. And this time, it's not just slightly different rendering. It's just wrong. So more bits is not the same as better. The format does more work than the bit count does. Remember that, because we're about to watch the same exact surprise happen in a completely different model. FPA again. Same hardware acceleration, same expected slam dunk. Does this support the new Phantom 5090? Looks about the same, right? Watch the face and listen at the end. Oh yeah, 24 lanes. Wait. It wasn't yet! Did you catch that? Whistper picks it up word error rate on FPH is 0.18 and basically word error rate is just how many words drifted from the baseline transcript. Zero is identical. This is the only quantization between BF16 and Q3KM with a non-zero WR word error rate. The model added a laugh that does not exist in the baseline audio. That's the say I'm 0.87 L-pips is at 0.07 Pretty good. By the way, the top three are for video quality comparisons. And the bottom two for LTX only are for audio. And here we have Mel Spectrogram MSC 26.7. And for video, by the way, the SSIM is the PixelMatch.

  6. 8:12 , obre el vídeo en una pestanya nova

    So one would be identical. Obviously BF16 is identical to itself, so that's one. And here on FPA, we're sliding down 2.8. And Mel MSC down here is the audio fingerprint. So zero would be identical in this case. Anything higher than that is drift. Does this support the new Phantom 5090? Oh, yeah, 24 lanes. Wait. So compare that to Q801 level down on the ladder, sort of more like... horizontal, same 8 bits per weight of course different format though. Look at the difference between BF16 and Q80. There's a huge difference in the amount of space it takes, 46 gigabytes versus 23, but they look identical. Even the motion blur in those exact moments in those frames is exactly the same, but FPA, not. And Q8 is better on every single metric, both video and audio. So FPA underperformed in one, When one, I don't know how to say that properly, don't ask me, and it also underperformed in LTX. Two different model architectures, two different teams, two different training runs. We are right? Maybe not. So if you've got a choice between FB8 and Q8 for a video model right now. Take Q8. I'm gonna skip over levels between Q6 and Q5 because they do get little bit worse, but there's no real big jumps here until you get to Q4. In fact, with these glasses, Q4 looks pretty good to me too. Same thing with the car. I see slight differences with the detail of the car itself, but overall, the motion, the clarity looks...

  7. 9:59 , obre el vídeo en una pestanya nova

    Pretty good. Let's take a look at the lady here. Q6 and Q5 both 12 and 11 gigabytes respectively on this. It definitely looks like the same exact lady. So what happens at Q4? Can we save more space and actually get away with it? Q4 9 gigabytes on disk. This is the kind of thing you'd run on a 24 gig consumer GPU. And this is where most of the community lives. And if you watch this in isolation, you probably think, It's fine. Sure, the face is a little bit fuzzier and a little bit not as crisp, but it's possible. The forest still moves, but if you compare it, El Pips says we are... Point three four. That's further away than FPH. For a complex scene, we're at point three five. We're about that glass of water. Not so bad here. Point two. For L-Peps that character consistency is getting there. Let me show you. Notice anything different about her hair and also it kind of doesn't look like the same woman anymore. Here's the complex scene. It's a little bit harder to tell here. Tokyo Night Market. Quarter disk size of FB16. But... It still holds together. This is kind of like a sweet spot over here and you can stop here if you don't have any reason to get any smaller. Red car looks fine and do comment down below if you notice any weirdnesses that I didn't notice. We save a ton of space in LTX 2.3. The Q4 is only 14 gigabytes here. There's some extra weirdness going on here with the face there and my eyes and dance eyes and the unexpected reaction makes its reappearance.

  8. 11:36 , obre el vídeo en una pestanya nova

    I'm gonna be like a man! But look at the Mell Spectrogram here, which is jumped to 46.9. One level up at Q5, it was just 9.9. So the audio fidelity just dropped roughly five times in a single step. Now check the top row, the video metrics SSIM is 0.88. Basically, right where it was, L-Pips is 0.07. The picture is right about where it was for FPH. So it's definitely gradually getting worse here. And Q4 shows a big change at least in the objective measurements. But not as much as audio. So that's finding number two. Audio degrades before video here. Even though it might not be perceived as such or might not be as noticeable as video. If you're only watching the picture, you're gonna miss that. And what does this matter? The picture degrading is something you might catch on a rewatch, but the audio degrading is something the viewer hears immediately. You ever watch a YouTube video with terrible audio? I hope. I hope my audio is actually decent here. It can be perfect, but I try. But you can put up with pretty bad video. However, if you hear terrible audio, people will just click away right away. Very different tolerance levels. So models like LTX, the ones that support audio, have to take extra special care. Alright, two bits per weight. This is five gigabytes for when. This is the bottom of the ladder. And it shows that glass is...

  9. 13:05 , obre el vídeo en una pestanya nova

    Pretty bad looking. There's even differences frame to frame. Look at the lady. That's terrible and Q3's very similar to Q4 where it does mess with the hair quite a bit from the original and doesn't look at all like the original but Q2 is a whole different level that's completely unusable here at this point. Look at this character consistency chart. Elpips is 0.54 Tokyo market to complex scene We're at 0.57, even worse. Both up about 60% from Q4. Now, LTX at the same two bits does the exact same thing. Look at that first frame looks pretty good, right? This is image to video by the way in case you didn't know. Small model, 8.3 gigabytes, but look what happens if I play this. Does this support the new Phantom 5090? Oh yeah, 24 lanes. Wait. They lied to us again! Oh my god, that would just like tugs at the hard strings. The audio is very different. It sounds robotic. Until he screams, then it sounds like more of somebody trying to get an Oscar. But look how a terrible the video quality is at every frame. Anything that's moving is basically completely destroyed and smashed. Whoa. Look at the transcript here. Word error rate here for Q3 and Q2. If when from daylight was again to they lied to us again. caps in the middle of the sentence. These are not the prompts by the way. These are the transcripts for that whisper test that are generated from the video. And for LTX is not just this one skit. LTX again, this is BF16, FPA, looking like a very different person. Q8, Q6, Q5, look like BF16 pretty much the same person. But the leaves are wrong, even in BF16.

  10. 14:55 , obre el vídeo en una pestanya nova

    Those are not maple leaves. I don't know what that is like a starfish in the shape of a leaf or a leaf in the shape of a starfish I guess. The shirt is also very difficult here. Every single one of these has a different shirt, Q5 and Q4 are kind of the same. And look what happens with Q3. We have a different person here. Can you tell which one of me is the real one? She's not even looking at the camera. Forget about Q2. Can you tell which one of me is the real one? Yeah, it's not you. Actually it's not of a butt. Some of these could pass for real one. Can you tell which one of me is the real one? Did you notice the audio difference between BF16 and Q2? Huge difference. This one sounds real. Here's Q2 again. Can you tell which one of me is the real one? Totally fake at this point, right? This is what her charts look like. Pretty gradual SSIM, L-Pips, also pretty gradual, but, you know, it gets up there. They all have a little bump on FB8, which makes it equivalent to Q4 or slash Q3. Let's check out the tech support scene, BF16, listen to the audio too. Did you try turning it off and on again? I am the IT department! Not sure what's going on over there something's going on with FPA the lady is totally different try turning it off and on again I Am the IT department Look at her. She's not impressed and the lady looks different in pretty much every single one of these her shirt color is different the wall decorations are different What's on the computer screen is different but once we get to the 11 gigabyte Q3 things just start melting down check this out

  11. 16:49 , obre el vídeo en una pestanya nova

    Did you try turning it off and on again? I am the IT department! Kind of looks like Mr. Smith from the Matrix. You know she's trying to push his buttons and she's winning. Q2. Did you try turning it off and on again? I am the IT department. Did you? Wow, that one is way off the rails. First of all, the lady sounds robotic. The video is just totally destroyed. So faces are still the hardest part. Whisper for Q2 and Meld Spectrogram for Q2. All just are off the chart. Carable so by now it looks like two bits just Rex everything, but here's what I did not expect the glass of water seen SSIM is that 0.74 a two bits per weight. It doesn't look great when I'm examining with my balls of eye but the numbers are saying it's okay. So take these tests also with a grain of salt. I didn't come up with these tests. These are pretty much standard tests. So yeah I haven't shown you this one yet, the marble rolling down the ramp. And even though each one of these frames by themselves might look okay, the physics of this thing is just terrible at every single weight level. Even if full precision here, the ball just doesn't roll like a real ball and then two balls get merged. That's just the model, not the quantization, it's just wrong.

  12. 18:25 , obre el vídeo en una pestanya nova

    not necessarily worse. Hello. At Q2, you can pretty much say it's worse. So finding number three I'd say is that there is no universal cliff and that's true in both models. If the thing you're making is a single heart object in motion, you can probably ride this all the way down. But if you're making humans in it, the floor is a lot higher. Bottom line. Run Q4 unless you've got a reason not to and the sweet spot doesn't move when you switch model architectures But also if your output has audio listen to it at Q4 the picture will trick you but the voice will not and if you remember one thing from this whole video Let it be this Format matters more than bit count the full ladders are linked down below every clip in every chart and thanks to this rig I was able to generate these videos pretty fast. If you want to know whether this 96 gigabyte RTX Pro 6000 rig was actually worth it, that video is over here. Thanks for watching and I'll see you next time.