DeepSeek V4 Flash en AMD Strix Halo: proves i límits
Donato Capitella prova DeepSeek V4 Flash amb DwarfStar sobre un i dos equips Strix Halo: quantitzacions, velocitat, qualitat i configuració.
Un model de 284.000 milions de paràmetres en un PC
Donato Capitella prova DeepSeek V4 Flash en equips AMD Strix Halo amb 128 GB de memòria unificada. El model oficial és una mescla d’experts de 284.000 milions de paràmetres totals, però només n’activa 13.000 milions per token i admet fins a un milió de tokens de context. Aquesta arquitectura redueix el còmput, però no elimina la necessitat d’emmagatzemar una quantitat enorme de pesos.
La demostració utilitza DwarfStar, o ds4, un motor especialitzat en aquesta família. A diferència d’un carregador general com llama.cpp, pot implementar de manera molt específica l’atenció comprimida, el format del model, la memòria cau KV i les crides d’eines. El cost d’aquesta especialització és una compatibilitat més estreta i un projecte encara qualificat de beta pels seus mantenidors.
El vídeo està patrocinat per AMD. Capitella, però, no es limita a una prova comercial: documenta errors inicials, variants de pesos, mesures de rendiment i una configuració distribuïda reproduïble.
El port a ROCm havia de ser correcte abans de ser ràpid
Les primeres versions per a AMD van adaptar el backend CUDA a HIP i van aconseguir carregar el model. El problema va aparèixer en tasques llargues de programació: l’agent repetia crides d’eines indefinidament. Segons el relat, una implementació incompleta de l’indexador d’atenció feia que el model perdés informació anterior quan creixia el context.
La comunitat va corregir aquesta part, va comparar la deriva numèrica amb els altres backends i va reescriure el camí ROCm incorporant les optimitzacions provades pels col·laboradors. El suport actual de DwarfStar inclou Strix Halo al repositori principal. La lliçó és important: veure tokens a la pantalla no demostra que un port conservi el comportament del model.
Capitella valida primer amb un agent de codi i només publica quan desapareixen els bucles. És una disciplina més útil que perseguir una xifra de tokens per segon: una optimització que altera l’atenció o la selecció d’experts pot accelerar un resultat incorrecte.
Q2, una variant híbrida i Q4 per a dues màquines
La quantització Q2 amb matriu d’importància ocupa aproximadament 80,8 GB. No redueix tots els tensors de manera uniforme: comprimeix sobretot els experts encaminats i conserva altres parts crítiques amb més precisió. La matriu d’importància es calibra amb codi i raonament per identificar quins pesos convé degradar menys.
Una segona variant ocupa prop de 97 GB i manté en Q4 les capes expertes 37 a 42. Entra en 128 GB, però deixa menys marge per al context, els buffers de ROCm i la resta del sistema. El repositori de les eines recomana un node dedicat i lots de prefill més petits si apareixen errors de memòria.
La quantització Q4 completa ronda els 153 GB i ja no cap en una sola màquina de 128 GB. El vídeo la reparteix entre dos Framework Desktop. “Q2” i “Q4” són abreviatures: els fitxers combinen precisions diferents segons el tensor, de manera que no s’ha d’estimar la mida multiplicant simplement paràmetres per dos o quatre bits.
Velocitat: el context penalitza sobretot el prefill
En el banc de proves del vídeo, la Q2 comença al voltant de 218 tokens per segon processant el prompt i baixa fins a uns 123 a 64.000 tokens. La variant híbrida és una mica més lenta. En generació, totes dues se situen prop de 15 tokens per segon amb context curt i cauen cap als 12 quan creix.
La Q4 distribuïda manté el prefill aproximadament en 50 tokens per segon i genera entre 13 i 11. La corba més plana no significa que escali millor: la xarxa i el pipeline ja s’han convertit en el coll d’ampolla. Per a un agent de codi, processar un repositori i un historial llarg pot pesar molt més en la latència total que mostrar la resposta final.
Són resultats d’un muntatge concret —Framework Desktop, Fedora 43, una versió determinada de ROCm i els pesos indicats—, no una promesa per a qualsevol Strix Halo. Canvis de kernel, backend, context, refrigeració o memòria disponible poden moure les xifres.
El model més lent pot acabar abans una tasca
Capitella també executa una selecció de 50 problemes de SWE-bench Verified amb el seu agent Pi. La Q2 resol 35 casos, un 70%; la híbrida, 38, un 76%; i la Q4 distribuïda, 45, un 90%. La sorpresa és que la Q4 triga de mitjana uns 21 minuts per tasca, semblant a la Q2 malgrat generar més lentament.
La seva interpretació és que una quantització de més qualitat necessita menys passos, evita decisions errònies i compensa part del cost d’inferència. És una observació valuosa: el rendiment d’un agent s’ha de mesurar fins al resultat correcte, no només en tokens per segon.
Cal llegir els percentatges amb el seu abast real. És un subconjunt “mini” de 50 tasques, executat amb una configuració d’agent i una metodologia pròpies que l’autor diu haver revisat per reduir falsos positius. No és el marcador oficial de les 500 tasques de SWE-bench Verified ni permet comparar sense més amb resultats d’altres laboratoris.
Com s’executa en un sol node
La guia ofereix una imatge de contenidor amb DwarfStar compilat per a ROCm i una interfície de terminal, DS4 Cockpit, per descarregar pesos i iniciar el servidor. En Fedora es pot fer servir Toolbox; en altres distribucions, Distrobox, Docker o Podman. El servidor exposa endpoints compatibles amb OpenAI i Anthropic, útils per connectar-hi un xat o un agent de programació.
El sistema ha de permetre que la GPU integri una part molt gran de la memòria compartida. La documentació actual d’AMD explica que Strix Halo no té una VRAM física separada: el límit GTT controla quanta RAM pot mapar cada procés. També exigeix versions de kernel amb correccions específiques. Per això és preferible seguir la matriu ROCm vigent en lloc de copiar cegament paràmetres d’arrencada d’un vídeo.
Algunes guies desactiven l’IOMMU per guanyar rendiment, però això anul·la l’aïllament contra accessos DMA i pot inutilitzar l’NPU. No és un ajust innocu. Per a un equip d’ús general, cal valorar seguretat i compatibilitat abans de reservar gairebé tota la memòria a la inferència.
Dos Strix Halo sumen memòria, no eliminen la xarxa
Per executar Q4, cada node conserva una còpia dels pesos al disc però només carrega el seu tram de capes. Un coordinador processa les primeres capes i un treballador les següents; tots dos han de compartir context i paràmetres compatibles. El muntatge del vídeo utilitza Ethernet de baixa latència i ample de banda elevat.
Això és paral·lelisme de pipeline, no una suma transparent de dues GPU. Les activacions viatgen entre màquines i la generació autoregressiva espera que cada token completi tot el recorregut. S’obté capacitat per carregar un model millor, però no el doble de velocitat.
L’elecció final depèn de l’ús. La Q2 deixa més context i respon abans en un sol equip; la híbrida millora la qualitat amb menys marge; la Q4 necessita més maquinari i paciència, però va destacar en aquesta prova de codi. El vídeo demostra que DeepSeek V4 Flash ja és viable localment en maquinari de consum d’alta gamma, sempre que “viable” inclogui configuració, validació i límits molt concrets.
Contrast i context
Fonts consultades
-
01
YouTube — Donato Capitella DeepSeek V4 Flash Inference on Strix Halo: ds4, Quantizations, Distributed Inference and Benchmarks
-
02
YouTube Canal de Donato Capitella
-
03
DeepSeek Fitxa oficial de DeepSeek V4 Flash
- 04
-
05
Strix Halo DS4 Toolbox Contenidors, pesos i configuració de DeepSeek V4 Flash
- 06
- 07
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Just over a month ago, DeepSeek released V4 Pro, a 1.6 Trillion Parameter model, and alongside it came the Flash version, a mixture of experts model with 284 billion parameters. In this video, I will show you how to get the Flash version hub and running on Streaks Halo devices, such as the Framework Desktop, and if you have two of them like me, I will also show you how to cluster them together to run higher quality conversations. And of course, we'll look at the benchmarks and what you can expect from the model on this platform. Before showing all of these, I want to give you a peek into the background work that the open source community puts into making these models actually work properly on the hardware. You have the timestamps, so feel free to skip this section if you just want to see the benchmarks or a tutorial on how to get DeepSeek before Flash up running on your Streaks Halo box. But I really want to take the time to highlight the great work that the open source community has been doing without which we would not have any of these. Also, a quick shoutout to AMD for sponsoring this video and for helping me bring content like this to the community, not just reviews, but in depth analysis and practical tutorials to get the best out of our local AI hardware. So when the model dropped, my first instinct was to wait. As you know, I don't rush to make a video every time a new model is released. Specifically for DeepSeek before Flash, I was a bit skeptical. Because this has around 300 billion parameters running the full 16-bit weights requires over 600 gigabytes of memory. Even on the highest Streaks Halo configurations with 128 gigabytes, that means compressing most of the weights down to 2 bits just to make it fit into the memory. And to beat contentizations are generally not great for coding or reasoning tasks, the precision loss is high enough that the model tends to make more mistakes and that was a major concern for me. Another concern I had was Lama CPP. I typically use Lama CPP as an inference engine, because it has the most compatible and mature model and quantization ecosystem. It is perfect for local inference.
-
2:48
, obre el vídeo en una pestanya nova
However, the integration work for DeepSeek before inside Lama CPP is still ongoing. The main reason is that DeepSeek before introduces architectural innovations like its compressed attention, which drastically reduces the memory needed to handle long context. And it is what makes their 1 million token context window very practical, but it also means that a general purpose inference engine has to implement these specific new operations and architectures without breaking the compatibility for the hundreds of other models it supports. Then somebody in the Streaks Halo Discord community pointed me to a project called DS4 or the Warf Star 4 by Salvatore Sanfilipe. Salvatore is better known online as Antirates and among other things is the original author of Redis, a very popular in memory database that I'm sure many of you have used before. The idea behind DS4 is a dedicated inference engine built from scratch for this specific model family rather than a general purpose from time to come to every single model architecture. I think originally it targeted upper silicon and was then extended to Kuda which naturally opened the path to a rock-haming heap implementation for AMD GPUs. And that was the point where I started paying close attention. When I looked at the DS4 GitHub repository there was already work underway to bring it to AMD hardware and specifically to Streaks Halo. A user by the name of Alanceve and I have no idea if this is the way the user name is meant to be pronounced, but essentially Alanceve had done the initial port. He took the existing Kuda backhand and ported it to heap using compatibility headers that map Kuda API calls today rock-hame equivalent. He got the model loading and running on Streaks Halo providing the first working baseline. Then Nick, whose GitHub username is AJPIA, another user name that I have no idea how it's meant to be pronounced, Nick forked the repo and built a native rock-hame implementation the performance number he posted were meaningfully faster. So I took this and built a toolbox to experiment. At this point I could have stopped. I had a working build and an optimized branch showing solid performance numbers.
-
5:43
, obre el vídeo en una pestanya nova
I could have set it up, record a quick benchmark and publish the video. But that is not what I want to do with this channel. Before I tell people something works, I want to spend the time using the model and checking it actually performs well and it's not just more hype. So I run the model through SWE Bench verified mini a software engineering benchmark that tests an RLAMS ability to work as a coding agent and that did not look good. The model fell into infinite loops on every single task. It called the same tools with the same arguments over and over again without making any progress. So working through the GitHub issues read on the DS4 repository, Alan Sav identified the root causes. The main one was the indexer, which is the attention component that allows the model to track long context was not implemented properly for the rock-hame part. Past a certain context length, the model was not correctly attending to earlier tokens. But even after those fixes, Salvadorism Philippe realized that there was still a noticeable drift between the rock-hame implementation and the other. But, hence, numerically it was not as close to baseline as it could be. So he decided to restart from scratch and rewrote the entire rock-hame implementation incorporating all of the correctness fixes and pulling in all of the performance ideas from Nick and Alan Sav while trying to measure and keep the result numerically close to baseline. After a few days of testing, we now have the rock-hame support merged into the main branch of DS4, the looping is gone, the results on coding benchmarks look good and the performance on strict sale is solid. You check all the history on the GitHub issue, it will be clear that this is the kind of work that takes time and it is exactly why I did not want to rush into publishing a video before I was sure that these had been tested thoroughly. So now that we have a solid rock-hame implementation, what can we actually run on it? There are a few choices. The standard Q2I matrix quant takes around 80 gigabytes. To make this model fit into memory, the majority of the weights are compressed down to 2.
-
8:30
, obre el vídeo en una pestanya nova
But however, some critical parts of the model are kept at a higher precision and this helps preserve accuracy. This model is also calibrated using an i matrix or importance matrix. During calibration, the model was run against millions of tokens of coding tasks and reasoning data to track which specific weight parts were demost active. The quantization process then compressed those highly active parts less aggressively. When you compress a model this way, that makes a practical difference. The i matrix version makes fewer mistakes, so in reasoning tasks and coding tasks, then if you did a uniform to beat quantization on all the weights. There is also a hybrid variant that keeps layers 37 to 42 at 4-bit precision sitting at around 97 gigabytes. These gives better accuracy, where those later expert layers are more relevant. But it is a bit slower and leaves less room for context. Beyond those, there is also a 4-bit quantization that runs at around 153 gigabytes. That of course does not fit on a single 128 gigabytes straight-sale or node, so I'll show you how to configure a multinode setup if you have two of these devices. I should also mention that DS4 also has an SSD streaming mode where experts that exceed the available memory are loaded from this on the map, allowing to run larger quantization, trading some speed for significantly more capacity. I'll rock out these not implemented yet, but I know that people are already trying to port this feature, so it should land very soon. So before we look at tutorials and demos, I think you are most interested in the benchmarks. You can find the benchmarks here, if you go on strict sale to boxes and benchmarks, these are the dwarf star benchmarks. There are three quantizations that we can run, the Q2 or the 2-bit quantization, and another one which has some layers in 4-bit, some of the last layers from 37 to 42, and then we have the Q4 quant, which you need to run over 2-strikes sale. The devices, and again, we'll see in a second out, to run all of these. As you sure we have two categories that we are interested in, pre-fill or prom processing, which is how many tokens per second it can process in our prompt.
-
11:32
, obre el vídeo en una pestanya nova
This is very important for coding agents, which have to process large code bases, and as you can see, as the context grows, typically the performance drops, and the faster model around 218 tokens per second at the beginning, from processing, obviously, is the smaller quant, the 80 gigabyte quant, and you can see the curve as the context grows, and then we have the same amount of time. The context grows, the performance goes down to around 123 tokens per second at 64K, context which is actually a pretty large context, but this is still usable. The same goes for the other larger quant, the 93 gigabyte quant, it is obviously a little bit slower, but not drastically slower, and I still find that it is useful if you have a little bit of patience, and what you see here, instead, much slower, is obviously the distributed inference, and you can see that, from processing states the same 50 tokens per second, regardless of the context, because in this case, you have the pipeline parallelism, and all the networking that's really the bottleneck, and that states kind of constant in these particular setup, so you don't see the performance go down, because the bottleneck is somewhere else. Obviously, we have token generation, how many tokens per second it can generate, and as you can see, the two quants are quite close to each other, roughly 15 tokens per second when you start, and then it goes down to around 12 tokens per second, and the performance is pretty much the same, and even the distributed inference is not far off that we start from 13 tokens per second, and maybe we go down to 11 tokens per second. And which is actually useful, so if you don't have a very large context, this is still kind of viable, not extremely viable, but kind of viable. If you have a moderate context, you really need to use any of the other two quants though, because obviously prompt processing is the most important part there. So, of course, speed is very important, but the quality of the output and the ability of the model to correctly complete tasks is equally very important, and this is why I've started benchmarking models in the pie coding agent against the SWE bench verified, meaning benchmark, this is a very common software engineering benchmark with real tasks that come from real,
-
14:29
, obre el vídeo en una pestanya nova
so GitHub projects and it allows us to figure out how well the model performs in a real world task, so to access that you can go on strict zero tool boxes, and that's the current dashboard. So you might have seen these in one of my previous videos on coding agents, I have since improved the way I run this benchmark, I considerably improved the way I run this benchmark to reduce false positives and make it a little bit more accurate, so you will see that the numbers actually changed from the previous time you saw this in a video, and this is how run much better. So you can see here, DeepSeek V4, flush, this is these molar 80 gigabyte version of the model, achieve 70% which is a very, very good score on this benchmark, and it was able to complete 35 tasks in the average time was 21, means then you also have the NTP version of the same model of the NTP run of the same model, obviously it did achieve the same performance, but it was 1 minute faster, so not massive difference, and then I also benchmarked the larger model, the 93 gigabyte model where the layer 37 to 42 are in Q4 instead of being 2 bits, and this did complete 38 tasks instead of 35, and so obtaining 76% which is the highest score, I've been able to get on a strict sale of devices with any of these models, even when 3.5122 billion parameters which used to top my leaderboard is now just behind these by 1 task, so this is actually measurable good quality in completing tasks. Now, I know what you're thinking, what about the Q4 model, what about the larger ones that I'm running on 2 strict sale of devices? Through these, this benchmark takes a long time, especially when a model is that slow, so I will run it, but it's going to take probably a couple of days, do not throw from the future, while I made it in this video, the benchmark is completed on the dual strict sale with the Q4 quantity, and you can see the results here, 8.90%, 45 out of 50 tasks, so this is obviously a much better quality quantity,
-
17:11
, obre el vídeo en una pestanya nova
and Q4 tends to be just about the correct quantity to retain enough ability, enough precision from the original model, but the other striking thing I want to point your attention to is the time it took to complete the tasks 21 minutes, which is essentially the same as the Q2 quantity, so what's going on there? Well, this is obviously a much better quality quantity, it means that it completes a particular task taking way less steps, so even if it is half as fast as inference, because it makes better decisions, it takes it roughly the same time as the other model, and this is a interesting finding, because maybe this is more viable in real world tasks than as more quant. To get the basic upper running on strict sale, I prepared a toolbox, which you can get as usual from strict sale toolboxes.com. Here we go into that, take a look at the host configuration, if you watched my other videos, you are already familiar with this, but I'm running all of these on a framework test-top with 128 gigabytes of unified memory. My base operating system is Fedora 43 updated to the latest kernel, actually this is a bit old in this documentation, I have the latest kernel, and most importantly, there are some parameters that you have to pass to the kernel, and there are other ways of setting this now, but essentially you want to make most of the memory available dynamically to the GPU, and this is particularly important for the PC, because it does need a lot of memory. After you've done that, you can head to the toolbox, which is this one over here, the wall, star. Now, I feel like I always repeat the same things in my videos, essentially a toolbox is just a container, but it's integrated in a way that you can enter that environment and then you feel like a shell imabs your home directory and your host network environment directly into this pseudo shell, and so it's much easier if you want to experiment with things. And you can load that particular container, or toolbox, just like this toolbox, create, you pull the image, and then obviously you have to expose all of the devices, the GPU device, and the right permissions, so that the GPU can be accessed from within the container,
-
20:07
, obre el vídeo en una pestanya nova
and then you can toolbox enter that container, it's very, very simple. Now, this works if you are on Fedora, if you are on DiBian or Arch Linux, toolbox doesn't work very well, and you want to use something called these throw box. Now, another way to get up and running is to use something that I call DS4 cockpit, which you can install with this particular PPEX command, and so we are just going to do that. I'm going to head to my framework desktop, and I'm just going to run this command, which is going to install the DS4 cockpit, that's all installed, and then you can say DS4 cockpit and there it is. You can see a list of the different toolboxes or containers available. We have the DS4 rock and 7.2.4 which tracks the main branch of antirate's repo, and this is the one you should be using, and then we have an experimental one which is a fork I made, which I'm going to show you later, that has multi-node support for strict sale, so we're going to run essentially a larger model across two strict sale devices. But essentially, this allows you to download a container, so if you select it, you can create or update the container, and once it's there, you can enter the container, as I said, it drops you into a shell, and you have, for example, the DS4 command, and you can specify all of the parameters manually, and you can find all of the CLI parameters inside the readme. So this is how you can run inference, for example, and this is how you run the server. But typically I prefer to run the server mode. So if you go here, you can see that you can choose between podman or doc, so you don't have to use toolbox or distrobox. You can just use the container, you select the image that you want. Again, in this case, it's the standard one, but later on we're going to use the multi-node one, and then you can select a model that you want to use. We have a model manager, and this uses a directory, you can change, and then allows you to select one of the models. So we have the Q2, I matrix model, the 80 gigabyte one, we've got the Q2 Q4 model that we discussed, we have the Q4, I matrix, and we also have the model for multi-token prediction, speculative, the coding.
-
22:50
, obre el vídeo en una pestanya nova
You can just select a model and click download, and this will go on AginFace and download the model for you. Now on this particular host, actually I happen to have the models on an external disk, I think it's under here. In fact, you can see that these are already downloaded. So then let's say I want to run the basic Q2, I matrix, quant, this molest one. Available, I can set a context size, I can set a cache on disk for the KV values, select the port, I can decide if I want to use mdp or not, and then I can just click start the s for server. You can see here the entire command line, it's copying the weights in memory, and actually if we do a mdp top, you will see the memory being allocated and the weights being copied in memory, there it is. This is now up and running on port 8000, and you can use it with your favorite chatbot. So what I'm going to do here, I'm just going to open Visual Studio code, and once it opens, we can go into the Pi extension, and this is a project I've been working on, but essentially, if you look here, I can now connect it to deep seek before flush, this is a configuration that I added to the Pi that you can see from here, so I added these DS4 provider with some of the parameters pointing it to obviously what I'm running, the server and configuring the model, you can alter this configuration. So now I can jump into a new chat and I can ask what is this project about, and obviously it takes its time, and we can go and check that this is actually working and look how much memory this is using, I wouldn't recommend trying to run a UI on top of this. But you can see the speed at which this is working, it understands exactly what disease you can see all of the tool usage, all of the thinking, and the speed is fairly acceptable if you have to work locally on a project. So because I have two framework test options, I thought it would be nice to try distributed inference. Now, for that to work, I had to fork the DS4 repository and add some custom code because the implementation on CUDA doesn't really work in the way that memory is being managed on strict sale.
-
25:44
, obre el vídeo en una pestanya nova
So you are going to need to use the toolbox with my fork, but it is actually very easy to do, let me show you how to do that. So this is one of the two nodes, and I'm bringing up DS4 cockpit and here I will select first of all the multi node container with my patches, and then obviously I'm going to select the Q4 model and you can see that this is actually a pretty high quality quant, some of the ways are actually in F16 and some are in eight beats, so this is not just four beats. And what we're going to do here, this is a perfect candidate for MTP actually, so let me just select the MTP model, and then we have to select the role, this is going to be the co-ordinator, and I'm going to put here the IP address of my co-ordinator on the network interface that's connecting the two devices. Now if you look at my previous video on the strict sale of cluster, you will see that I have it connected with a very low latency Ethernet connection and high bandwidth, so that's the link I am using for this. And you will also see that as soon as I set up the co-ordinator role, it pre-selected the first 22 layers to be allocated to this, so when we set up the worker, the other layers will be loaded there. You will also see that selecting the co-ordinator changed the default context size, and that's again because we have so much more memory to be honest, you could go much higher than that at this point, but this is all set up and I can start the server. So this is now starting to allocate the memory on the co-ordinator node, and as it does that we are going to do the same on the other strict sale node, so selecting the correct toolbox, and obviously you need to have a full copy of the weights here. And you need to select the worker role, the context here must match exactly the one you set for the worker. You can see the layers that are selected here, and then we need to put here the IP address of the coordinator, and I don't think for this, I don't think we need to select MTP because that only is needed on the coordinator.
-
28:30
, obre el vídeo en una pestanya nova
So now that's gonna start. And if we go here on the coordinator, we see that this is now up and running and pretty soon, hopefully after the layers are allocated in memory, this should be able to connect, as you can see, it connected to the coordinator, and there we go. So we can now use this, and I'm going to show you maybe again in a pie-coding agent, so let me open it up. So there it is again, I can get into pie, and maybe once again I can ask it, what is this project about? And we can see this is now working as we saw from the benchmarks before, this is not the fastest, but it is also much better quality. This quality as a cost which is the performance, but we can see the two nodes actually working, and there it is. Again, remember this is done with pipeline parallelism. I don't think this is tense so parallelism, so there is obviously a network overhead that we already discussed when we saw the benchmarks. But ultimately you can see that it is working and it is doing the tool calling and the evaluation, it is simply going to be quite slow, but, reason why I'm showing you this and I'm getting this up and running and also benchmarked on SWB bench is because when the new devices with AMD Gorgon Halo come out, they will have an option for 190 to gigabytes of RAM, which will fit a model like this perfectly, but because they don't have the obviously network latency and pipeline parallelism and everything can be run from the same host, I expect, although they will be slower than the Q2 quant, the performance will still be much better than this and fairly useful, and this model is exceptional quality, this Q4 quant as we saw in SWB bench is really really good. Thank you for watching until the end. Remember this channel is a hobby project of mine and I put a lot of time and I've fought into testing, debugging and documenting all the stuff to make sure it's accurate and reproducible. If you find any of these useful, please support the channel in all the usual ways by liking, commenting and subscribing.
-
31:12
, obre el vídeo en una pestanya nova
You can also support my work directly via via via my coffee, I leave a link in the description. All the links to the tool boxes, setup commands, etc are available in the description and add streaks hello toolboxes.com Thank you for watching and I'll see you in the next one.