Intel·ligència artificial Maquinari llama.cpp ASUS NUC Intel Panther Lake LLM local clúster

Tres ASUS NUC poden executar un model d’IA de 70B?

Tres ASUS NUC 16 Pro reparteixen un Llama 3.3 70B i comparen CPU, GPU, NPU, Ethernet i Thunderbolt. El clúster suma memòria, però no sempre velocitat.

Alex Ziskind connecta tres miniordinadors ASUS NUC 16 Pro amb processadors Intel Core Ultra Series 3 per respondre dues preguntes diferents: poden repartir-se un model d’intel·ligència artificial que no cap en cap màquina individual? I, si el model ja hi cap, tres equips el poden servir més ràpid?

La resposta curta és sí a totes dues, però no amb la mateixa arquitectura. Dividir un model permet sumar memòria i executar un Llama 3.3 de 70.000 milions de paràmetres, encara que sigui molt lent. Replicar un model més petit a cada node, en canvi, multiplica el rendiment agregat perquè cada ordinador atén peticions independents.

1. Tres NUC, 192 GB i tres motors d’IA per node

El vídeo ensenya el resultat abans d’explicar-lo: a 00:00, els tres NUC participen en l’execució d’un model que no cap en cap d’ells per separat. Cada unitat de prova disposa de 64 GB de memòria, de manera que el conjunt suma 192 GB bruts.

El maquinari apareix amb detall a 02:34. Són NUC 16 Pro d’ASUS amb CPU Intel Core Ultra Series 3, gràfics Intel Arc integrats i una NPU dedicada. La configuració concreta del vídeo no representa totes les variants comercials: ASUS anuncia diferents processadors i memòries segons el model.

La fitxa oficial d’ASUS confirma que la família pot incorporar Core Ultra Series 3, GPU Arc, NPU de fins a 50 TOPS, doble Ethernet de 2,5 Gb i dos ports Thunderbolt 4. Intel, per la seva banda, identifica Series 3 com la primera plataforma basada en Intel 18A.

2. La GPU duplica el prompt, però no la generació

Abans de formar el clúster, Ziskind mesura els tres motors d’un sol NUC. A 05:06, activar la GPU fa passar el processament del prompt d’una mica més de 1.000 a més de 2.200 tokens per segon. Llegir el context inicial —per exemple, els fitxers que envia un assistent de programació— millora aproximadament al doble.

La generació posterior queda clavada al voltant de 46 tokens per segon tant amb CPU com amb GPU. El vídeo ho atribueix al «mur de la memòria»: en una GPU integrada, CPU i gràfics comparteixen la memòria del sistema, i generar cada token obliga a moure moltes dades dels pesos. Quan el límit és l’amplada de banda de memòria, afegir unitats de càlcul no accelera aquesta fase.

Les xifres són resultats d’aquesta màquina, model i configuració; no s’han de generalitzar a qualsevol LLM. La distinció sí que és útil: el processament del prompt tendeix a aprofitar millor el paral·lelisme, mentre que la generació pot quedar limitada per la velocitat amb què es llegeixen els pesos.

3. La NPU consumeix menys, però el programari encara limita

L’experiment amb la NPU comença a 06:30. llama.cpp no la pot utilitzar en aquesta configuració, així que el creador passa a OpenVINO. Un model petit funciona, però un de més gran exigeix reconstruir els artefactes perquè fins i tot alguns models preparats no carreguen directament.

Quan la prova arrenca, la GPU és la més ràpida, la NPU supera la CPU en alguns models petits i el temps d’inicialització de la NPU és visible. A 08:06, Ziskind compara la mateixa GPU amb dos entorns: llama.cpp sobre Vulkan ronda 34 tokens per segon, davant d’uns 14 amb OpenVINO en aquella prova.

La NPU recupera terreny en consum instantani. A 09:04, mesura uns 17 W per a la NPU, 24 W per a la GPU i gairebé 30 W per a la CPU. Però la GPU completa la feina més ràpid i acaba oferint millor energia per token. Consum més baix durant l’execució i eficiència total no són exactament la mateixa mètrica.

4. Dividir un model petit fa el clúster més lent

Amb els tres equips connectats per Ethernet de 2,5 Gb, el primer clúster reparteix un model Qwen de 35.000 milions de paràmetres. El resultat de 10:15 va en direcció contrària a la intuïció: una sola màquina genera uns 35 tokens per segon, mentre que les tres juntes baixen aproximadament a 17.

La causa és el recorregut de cada token. Una part del model resideix en un node, la següent en un altre i el càlcul ha de creuar la xarxa repetidament. No s’estan sumant tres generadors independents, sinó inserint comunicació entre etapes que abans compartien memòria.

Al mur de memòria s’hi afegeix, doncs, un impost de xarxa. El model de 35B ja cabia en un NUC; repartir-lo no resol cap limitació de capacitat i només introdueix latència. Sense una interconnexió i un protocol pensats per a aquest patró, més ordinadors poden significar menys velocitat.

5. El model de 70B funciona perquè el clúster suma capacitat

El cas que justifica dividir pesos arriba a 11:30. El Llama 3.3 70B quantificat ocupa uns 75 GB segons el vídeo, massa per als 64 GB d’un node. Distribuït entre els tres, carrega i genera text a prop d’1,4 tokens per segon.

No és una velocitat còmoda per a un xat, però demostra la idea central: el clúster permet executar una mida de model inaccessible per a una sola unitat. La suma de memòria serveix per capacitat, no converteix automàticament els nodes en una màquina més ràpida.

Ziskind prova després una connexió triangular Thunderbolt anunciada a 20 Gbps. A 13:03, el 70B continua pràcticament igual: 1,43 tokens per segon abans i després. El coll d’ampolla no era només l’amplada de banda màxima del cable, sinó la latència, la memòria i milers de transferències petites.

6. Replicar el model sí que augmenta el throughput

La segona arquitectura apareix a 14:03. En comptes de tallar un model entre màquines, cada NUC carrega una còpia completa d’un model que sí que cap en 64 GB. Un distribuïdor envia cada petició a un node diferent.

Sota càrrega, un equip processa al voltant de 196 tokens per segon agregats. Tres còpies arriben a prop de 500, una millora aproximada de 2,5 vegades. Una conversa individual no necessàriament rep respostes 2,5 vegades més ràpides; el benefici és atendre més sol·licituds simultànies.

La regla pràctica queda clara: dividir pesos amplia la mida màxima, però penalitza la velocitat; replicar-los amplia la capacitat de servei, però exigeix que cada node pugui allotjar tot el model. Per a una oficina amb diversos usuaris, la rèplica pot tenir sentit. Per a un únic usuari amb un model petit, un sol NUC evita cost i complexitat.

També cal protegir la xarxa. La documentació oficial del backend RPC de llama.cpp el descriu com una prova de concepte fràgil i insegura i adverteix que no s’ha d’exposar en una xarxa oberta ni en un entorn sensible. Un clúster experimental s’ha d’aïllar, actualitzar i tractar com a infraestructura, no com un servei públic per defecte.

Conclusions

Els tres ASUS NUC 16 Pro aconsegueixen executar un Llama 3.3 70B que no cap en cap node individual. Ho fan a només uns 1,4 tokens per segon, i ni una xarxa Thunderbolt més ampla elimina el cost de comunicar fragments del model.

L’experiment més útil és la comparació entre objectius. Si cal memòria, es divideix el model i s’accepta la penalització. Si cal servir més usuaris, es replica un model que càpiga a cada màquina i es reparteixen les peticions; així el vídeo passa de 196 a gairebé 500 tokens per segon agregats.

El maquinari Panther Lake és capaç, sobretot en processament de prompts, però GPU, NPU i clúster depenen d’una pila de programari encara irregular. La pregunta correcta no és simplement si tres miniordinadors «sumen», sinó què es vol sumar: memòria per executar un model més gran o throughput per atendre més feina.

Contrast i context

Fonts consultades

4 fonts
  1. 01
  2. 02
  3. 03
  4. 04

Font de treball

Transcripció amb marques de temps

15 fragments
Consulta la transcripció
  1. 0:00 , obre el vídeo en una pestanya nova

    Right now this little machine is running a 70 billion parameter AI model. And so is this one and this one all three together. But here's the part that doesn't add up. Not one of them is big enough to hold this model. Not even close. By the way, this is Intel's newest silicon, a brand new chip, a brand new GPU, a dedicated AI chip sitting on top of that, the NPU. And I got three of them to try some the crazy. So how is it running a model that none of them can actually fit? Okay, so the dream here is simple. Many PCs are getting pretty ridiculous. And if you're a developer, this is pretty convenient stuff here. Not only will it run your IDEs and code smoothly now, it used to be kind of a joke just a couple years ago. Now they're serious. These things can also run AI, locally. Your own models, your own data. No API built showing up at the end of the month. I have it next to my Mac Mini to show you the relative size of this new one. It's shorter, it's narrower.

  2. 0:56 , obre el vídeo en una pestanya nova

    it's a little bit longer than a Mac Mini. But it is pretty small. The question is, can you take a few of these and bolt them together into one machine that runs stuff that you couldn't run on just one? The part that people forget about a clean local setup is what happens when you need storage somewhere else. For anyone dealing with servers, data sets, checkpoints, or large project files, storage gets complicated fast. You want redundancy and remote access, but you don't want to give up privacy just to get it. Most services make uploading easy, but they're not really built around the idea that your data should stay private. That is where internet comes in. It's a privacy-first cloud storage platform built around keeping your data under control. Your files are protected into N10, zero-knowledge encryption, so only you can access them.

  3. 1:41 , obre el vídeo en una pestanya nova

    Not even internet can access them. It also uses post-quanom cryptography to help you protect against current and future threats. And it fits real workflows too. With web access, desktop, and mobile apps. Plus, CLI, WebDev, ArcLone, and NAS support. Yeah, I can sync my NAS right to it. I like this more as a companion to a logo setup, not a replacement for it. On top of storage, you also get privacy tools. File versioning for supported file types, and the lifetime ultimate plan gives you five terabytes for a one-time payment. No subscription. So the pitch here is simple. Private cloud storage that actually fits the way technical people already work. This is not just vague security language. Internet is open source and independently ordered by secure them. It's also GDPR compliant and ISO 20701 certified. Use my link in the description or go to internet.com slash al-exist-in and use code al-exist-in to get 87% off the lifetime 5 terabyte plan. So in case you're not familiar,

  4. 2:40 , obre el vídeo en una pestanya nova

    ais is just released these, these are the Nuck 16 Pro. This particular design actually varies quite a bit. You get WiFi 7 and Bluetooth 6 in these, and they carry the pro name because they have redundancy for pretty much everything. Two Thunderbolts, two HDMI's, four USBAs, dual Ethernet ports, a Granible RAM, but they can come with Intel Core Ultra 5, 325, for 559, Core Ultra 7356, and these are Core Ultra X7, 358H, so this is like the high end once. Here's my favorite part though. If you wanna get inside, it's just a little pull tab like that. More of any PC issues should be doing this kind of thing. Look how easy it is to get in and get out. Done. I wasn't side of that just now. They can actually get up to the Intel Core Ultra X9,

  5. 3:27 , obre el vídeo en una pestanya nova

    which I don't have, but I'd imagine those would be pretty crazy. They also have the new Arc B390 GPUs inside, and a separate GPU that's the dedicated AI engine. Now my version also comes with 64 gigs of memory. This one on Nuick has 32, and it's 30, at $1,700, so you can imagine how much these cost. Check the current pricing, because that's always changing. Now, even though there's definitely no question that these will make incredible little dev machines, I'm gonna be doing something you probably shouldn't be doing with them, making it AI cluster out of them. And that dual Thunderbolt that I was talking about, keep an eye on that, because it comes back later in a big way. So if you've been watching this channel at all recently, you know that I have a few machines here that I've been clustering including these Dell GB Tens, which are basically the same thing as the DJX Spark and the Mac Studios and all those machines can be clustered using RGMA. Which basically means you can go back

  6. 4:23 , obre el vídeo en una pestanya nova

    and watch the videos, but basically that means that the more machines you add, the faster it gets, generally speaking, this is basically for things like large language models, all the lamps. When you're generating text, generating images, if you decrease the network latency between the machines, you're gonna get faster generation. And when you're doing text generation or image generation, you would not, there's always two stages. One is prompt processing. If you're working with code, for example, you've got your IDE talking to an LLM for code generation or debugging or whatnot. You're gonna be sending all that code as context to the LLM. That's where prompt processing comes in handy.

  7. 5:02 , obre el vídeo en una pestanya nova

    So prompt processing has to be fast. I won't get into the details of that setup In this video, because I have a bunch of other videos showing that, now I just want to test how fast these things do that. And the first thing I want to know on a brand new chip, can that new Arc GPU actually speed up and LLM. So I ran the same model on the CPU and then on the GPU. Turning the GPU on essentially doubles the speed that it reads your prompt. Just over a thousand tokens per second up to 2200 tokens per second. Boom. Nice. The second part of LLM's inference is token generation. After it's done, it's prop processing. It switches over to the generating part. And that, here my friends, is about 46 tokens per second on the CPU and about 46 tokens per second on the GPU. It didn't move. And this is the thing I want you to remember because it comes back to haunt us. The rest of this video. This is the memory wall. Talking Generation relies on memory bandwidth,

  8. 6:03 , obre el vídeo en una pestanya nova

    prompt processing relies on that brand new hot GPU chip that's really fast. Both sides are very important, and on these particular chips, the patho link chips, the GPU shares the same memory as the CPU, and generating text is limited by that memory speed. So the GPU's extra muscle does nothing for it. For the generation tokens per second, memory bandwidth is the boss. So the GPU helps, but it's not magic, and they're still that third engine that I hadn't touched. The NPU. If the GPUs cap by memory may be a dedicated AI chip is the real story, right? So let's get into the game. First problem. The tool that we can use on these machines, the most popular tool for doing inference is Lama CPP. It can't even talk to the NPU. So I switch over to Intel's own OpenVeno. This is Intel's own software and small model on the NPU actually works cool. By the way, if you're not familiar with OpenVeno, it's supposed to be able to use the GPU, the CPU, and the NPU and basically automatically decide for you which is best. An OpenVeno has their own toolkit in models on hugging face. These are specifically openVeno models with the OV extension so you can know what they are. because hugging face model naming conventions are so conventionny. Right. And then I try a bigger model and it fails. Intel's own pre-built models will not run on Intel's own NPO. So I have to go and rebuild the model myself just to get it to load. That's the bleeding edge for you. But we know these days what happens with these chips. They come out. They're great. But if you want to do any kind of bleeding edge AI stuff,

  9. 7:48 , obre el vídeo en una pestanya nova

    The software has to get shop. However, once I got it running, here's all three engines across a few models. Small-ish models GPU is fastest everywhere. But the NPU beats the CPU, it's just not a speed demon, and it's a bit slow to spin up. Now hold on. I just used two completely different pieces of software. Lama CPP for the GPU and open Vino for the NPU. But I did mention that OpenVeno also works on the GPU. So on the exact same GPU, which one is actually faster? OpenVeno or Lama CPP? Well, I ran that too. Same model, same chip. And I did not expect this. The free Open Source Tool beat Intel's own software by about two and a half times. 34 tokens per second with Vulcan. Vulcan is just the API that Lama CPP was using in this case. Versus Intel's 14 tokens per second. on Intel's own GPU, come on. So far that away, because when we get to the cluster, we're gonna be sticking with Lama CPP here. It's faster and it's really the only way to get clustering going anyway, on this particular setup. But there's one more thing that I wanted to settle

  10. 8:55 , obre el vídeo en una pestanya nova

    before leaving one single box. Speed was never NPU's whole pitch anyway. Efficiency is, everyone says their NPU is the most efficient one and everybody's got an NPU now. Okay, let's actually measure it. Real Watts while it's generating and the answer is both things are true. The NPU actually draws the least amount of power 17 watts versus 24 for the GPU versus almost 30 for the CPU and the power draw the GPU has the lowest draw. It's the coolest and the quietest that part is real. However the GPU is actually more efficient Porto can that's that red line right there because it's twice as fast It finishes the work sooner and uses less total energy per token. And usually you're going to be doing these kinds of burst the operations on a machine like this, unless you're running agents 24-7. So the NPU Sips the least power, but the GPU gets the most AI done per dual. The CPU? It loses on both, but that's kind of always been the case. However, the CPU's making it come back and I'll talk more about that in future videos. So the whole lesson from one box, memory bandwidth runs the show. So I squeezed everything I could out of one machine. The only way to go bigger is more machines. I brought in the other two knocks,

  11. 10:17 , obre el vídeo en una pestanya nova

    wired them all together, and split one model across all three. More machines, more speed, right? No. That was kind of hopeful, but now I'm a little disappointed. But we've seen this story before, with my framework video cluster situation before I got RDMA working on that. And my minis forum cluster video, it got slower. The same model on one machine was about 35 tokens a second and across all three, 17. It got essentially cut in half. This is Quint 3.635 billion by the way, which easily fits into one machine. Once you see why this happens though, it all makes sense, even if it stinks. This kind of clustering splits the model into chunks and every single token has to hop from one machine to another machine to the third machine over the network. Over these Ethernet wires go into the switch, and then back, you're not adding power, you're adding traffic. So we have the same memory wall as before, but now we've just added the network tax on top of it. So if clustering doesn't make a model faster? What's it actually for? With a caveat that I've mentioned before, Cluster and Kamek, uh, inference faster. If you use the right technology, which is RDMA, which I'm showing in other videos, and I'll link those down below. But the answer is the whole reason I did it. It's not for speed, it's for size. Let's find the model that no single one of these machines could ever run on its own. Ah, yes. The good old Lama 3.370B.

  12. 11:52 , obre el vídeo en una pestanya nova

    This is a dense model. It's large. It's been around for a little while, but all the models have been either mixture of experts models, which means they have a certain number of parameters, like 200 something, but they only have a small number of active parameters. So they've been bigger in parameter count than the 70B or much lower in parameter count in the 70B. They're not really dense. They're mixture of experts. This one is dense and the one I ran is 75 gigabytes. A single machine has 64 gigs of memory. It basically cannot load this period. But three of them together that's 192 gigabytes in the pool. So I split the model across all three and it works. It actually works folks. Three little machines not one of them big enough on its own to run the 70 billion per hour model. It's slow. Okay, 1.4 tokens a second, but that's not the point. The point is it ran. The fun part is that I asked it what is remarkable about three small computers running one big AI model. And it described its own situation right back to me. Okay, so it works, but it's slow. And every one of those tokens is hopping between the machines over playing 2.5 gig Ethernet. The fix seems obvious, right? Make the network way faster. Each one of these machines has two Thunderbolt ports on it. Remember that second port I told you about? 20 gigabit, like eight times the bandwidth. So I wear all three together in a triangle.

  13. 13:18 , obre el vídeo en una pestanya nova

    Every machine straight to every other one. This has to fix it. And it did basically nothing. Loading the model got a little faster on the small one, but the big 70 billion parameter barely budged. And generating tokens on a 70 billion parameter, it was identical. 1.43 before 1.43 after. This model model actually got worse and it kept crashing. Now here's why. It's the same wall from that very beginning. The bottleneck was never the cable or the bandwidth speed. It's memory speed and thousands of tiny little messages flying back and forth every second. A wider highway doesn't fix the traffic jam that isn't on the highway. Well play physics. Splitting one model across machines is a dead end for speed. But there is a completely different way to use these three as a cluster. And I've shown kind of a demo of this in a previous video. We don't split the model at all. We put a full copy on every machine and just send each request to a different one.

  14. 14:18 , obre el vídeo en una pestanya nova

    This of course will only work for models that will fit inside the RAM over one machine. Buy 70 billion parameters. Aha! This one. This one scales. One machine handled about 196 tokens a second under load. Three machines, almost 500 tokens a second, two and a half times faster, finally. So here's the rule, and it's the whole thing in one line. You cluster one way for size, you split the model, and a completely different way for speed. And for throughput, copy the model and spread the work. Don't mix them up. So step back and look what actually happened here. Same lesson at two different scales. There's three AI engines in one box, for three boxes in one cluster. Now I think that this could be solved at a software level to some degree. I'm not 100% sure about that, but Apple did it just magically updated software and suddenly RDMA works. Do not look up a tele in his videos. By the way, his channel somewhere down below. Also, he got it working with his strict halo boxes. I did the same thing following his example, so maybe there's a way to unlock that in software here. Is the newest Intel stuff ready with the software stack? The hardware is real and it's impressive, but the bleeding edge means rough edges. Intel's own software wasn't ready for Intel's own chip. I'm rooting for them.

  15. 15:44 , obre el vídeo en una pestanya nova

    They made some really nice leaps, especially with their discreet professional GPUs, and I've made a video about the B60, B70, so you can check all those out. So if I were you, for model that fits on one machine, the cluster is kind of pointless. One box wins every time. The cluster like this would be better served as a proxmox serving virtual machines, for example. This whole setup only earns us keep and it's price tag when the model is too big for any single machine. Or you've got an office, so people, and you wanna serve them. I'm curious, if you were to run a cluster like this with many PCs that are APUs, for example, like this one, Would you run it for AI or would you run it for proxmox or virtual machines? What would you do? Let me know in the comments down below. Very curious to read that and I read all your comments. Do check out that cluster setup right over here and the framework cluster setup over here. Thanks for watching and I'll see you next time.