Dues Claude Skills per crear vídeos amb IA: el pipeline de JOEY amb Higgsfield
JOEY comparteix dues Claude Skills per convertir idees en prompts coherents per a Nano Banana Pro i Seedance 2.0. El seu flux separa personatge, escena, imatge de referència i direcció de vídeo per reduir proves inútils a Higgsfield.
Fer una imatge atractiva amb IA és relativament fàcil; mantenir el mateix personatge, l’estètica i la lògica visual durant diversos plans és molt més difícil. JOEY proposa resoldre aquest segon problema amb dues Claude Skills creades durant la producció de CTRL, el seu projecte de ciència-ficció i K-pop fet amb eines generatives.
Les skills no generen el vídeo directament. Funcionen com una capa d’instruccions per convertir una idea en prompts detallats i repetibles abans d’enviar-los a les eines de Higgsfield. Una prepara personatges, vestuari i imatges de referència per a Nano Banana Pro; l’altra tradueix aquestes decisions a llenguatge de càmera per a Seedance 2.0. L’objectiu és malbaratar menys crèdits, no prometre una generació perfecta al primer intent.
El problema no és un prompt, sinó un sistema repetible
JOEY divideix bona part del contingut sobre vídeo amb IA en dues categories: demostracions d’una sola imatge espectacular i cursos que ensenyen receptes aïllades. El que troba a faltar és el procés que connecta personatge, escena, referències i múltiples plans sense començar de zero en cada conversa.
Un xat nou no coneix automàticament les decisions visuals preses abans. Si cada petició torna a explicar estil, lent i color, les instruccions poden variar. La seva resposta és conservar aquest criteri dins de skills reutilitzables perquè Claude apliqui una estructura consistent.
Això encaixa amb la definició oficial de les Skills de Claude Code: carpetes amb un SKILL.md d’instruccions que es carrega quan és pertinent i pot tenir material de suport. En el cas de JOEY, el valor és haver-hi codificat decisions nascudes de producció real.
Banana Pro Director construeix la base visual
La primera skill s’anomena Banana Pro Director i s’ocupa dels prompts d’imatge. JOEY diu que pot ajudar a preparar fitxes de personatge, referències de vestuari i scene plates, és a dir, bases visuals de l’escenari. També incorpora el seu vocabulari sobre textura de pell, càmeres i aparença hiperrealista.
Aquesta fase busca fixar elements abans d’animar-los. Primer es defineix qui apareix: trets, roba i referències coherents. Després es construeix l’espai i s’hi col·loca el personatge. El resultat és una imatge de referència que resumeix la composició desitjada i que servirà d’entrada per al vídeo.
Nano Banana Pro és el nom comercial del model d’imatge Gemini 3 Pro Image. Google el presenta com una opció orientada a més fidelitat, control creatiu i edició basada en referències. Això no elimina les variacions pròpies de la generació, però explica per què una fitxa i una referència ben preparades poden ser més útils que una descripció improvisada cada vegada.
Cinema World Builder converteix la imatge en direcció
La segona skill, Cinema World Builder, prepara els prompts de vídeo. Segons el creador, inclou cinc modes cinematogràfics, cadascun amb llenguatge de càmera, lent, estoc visual i tractament de color. Així, una indicació com rodar de nit en un aparcament amb reflexos anamòrfics es transforma en una especificació més tècnica.
La separació de responsabilitats és important. Banana Pro Director decideix l’aspecte del fotograma de partida; Cinema World Builder descriu com es mouen la càmera, els subjectes i l’acció al llarg del temps. No és el mateix demanar una escena estàtica que indicar un pla seqüència o una successió de plans amb durades diferents.
Seedance 2.0 admet instruccions i referències multimodals d’imatge, vídeo i àudio, i ByteDance en destaca la generació de seqüències amb diversos plans i més control sobre moviments i muntatge. La skill de JOEY no substitueix el model: organitza la informació perquè aquesta capacitat rebi una direcció més clara.
El flux pràctic: personatge, escena, referència i vídeo
El tutorial final del vídeo redueix el pipeline a quatre passos. Primer cal definir un personatge i una escena. A continuació es crea la fitxa del personatge, amb prou detalls perquè les decisions no canviïn accidentalment. El tercer pas és situar-lo dins l’escenari i generar una imatge de referència.
Aquesta imatge s’adjunta a Claude perquè la segona skill redacti el prompt de moviment. L’usuari ha d’indicar explícitament si vol una sola presa o una seqüència de diversos plans. També convé estimar quants segons ocuparà cada part: la durada no és només una decisió narrativa, perquè afecta el consum de crèdits de la plataforma.
El sistema, per tant, no consisteix a prémer un botó que produeix una pel·lícula. És una cadena d’artefactes intermedis que permet detectar errors aviat. Si el personatge no és coherent, es corregeix la fitxa; si la composició falla, es refà la referència; si el moviment no comunica la idea, es modifica la direcció de vídeo sense reinventar tota la identitat visual.
CTRL és la demostració, però el pipeline és el producte
JOEY ensenya un fragment de CTRL per demostrar què ha pogut construir en unes dues setmanes: personatges, canvis de vestuari, escenes i seqüències creats amb Higgsfield, Seedance, Nano Banana Pro, Suno, Claude i Adobe Premiere. No presenta el videoclip com una obra perfecta. De fet, remarca que continua aprenent i que ja canviaria diversos resultats.
Aquesta autocrítica ajuda a entendre la proposta. El videoclip és una prova visible; el producte reutilitzable és el procés que l’ha fet possible. Les skills provenen del gust i de les decisions de JOEY, però ell anima a modificar-les perquè cada creador hi codifiqui la seva pròpia manera de treballar. Copiar les instruccions no hauria de produir automàticament l’univers de CTRL.
La descripció del vídeo enllaça tant la versió utilitzada a la demostració com una actualització posterior, identificada com a 2.0. Per tant, qui les descarregui ha de comprovar quina versió està adoptant i conservar una còpia dels canvis propis, especialment si converteix la base compartida en una eina habitual de producció.
Com adoptar les skills sense perdre el criteri propi
La documentació de Claude Code situa les skills personals a ~/.claude/skills i les de projecte a .claude/skills. Cada skill té una carpeta i un SKILL.md; els recursos addicionals poden quedar al costat. Abans d’instal·lar fitxers de tercers, és prudent llegir-ne les instruccions i comprovar quines ordres, eines o recursos demanen.
Un bon punt de partida és provar-les amb una escena curta i mesurable. Cal comparar el prompt generat amb l’objectiu, anotar quines paraules produeixen resultats útils i eliminar les preferències que no encaixin amb el projecte. Amb el temps, la skill hauria de reflectir exemples propis, convencions de noms, limitacions de pressupost i una gramàtica visual concreta.
Menys intents inútils no vol dir zero errors
La promesa raonable és millorar la qualitat inicial dels prompts i fer més consistent el procés. JOEY adverteix explícitament que un bon vídeo encara necessita múltiples preses i generacions, i que les skills no eviten gastar centenars o milers de crèdits. La coherència també depèn dels models, de les referències, del muntatge i de la revisió humana.
Hi ha altres límits. Una instrucció molt detallada pot perpetuar una mala decisió amb gran consistència; un model actualitzat pot interpretar de manera diferent un prompt antic; i una estimació de durada no garanteix el ritme final. Cal tractar les skills com a documentació executable del criteri creatiu, no com una garantia artística.
La lliçó més útil del vídeo és passar de col·leccionar prompts a construir un pipeline. Personatge, escena, referència, moviment, durada i cost formen un sistema que es pot revisar i millorar. Les dues skills de JOEY ofereixen una base gratuïta per començar, però el salt important arriba quan cada creador les adapta al seu llenguatge visual i aprèn dels resultats que no funcionen.
Contrast i context
Fonts consultades
- 01
- 02
-
03
Claude Code Extend Claude with skills
-
04
Higgsfield Higgsfield AI
- 05
-
06
ByteDance Seed Seedance 2.0
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:04
, obre el vídeo en una pestanya nova
I'm about to do something no one in the AI space is doing right now. I spent the last 2 weeks building Claude skills to help us prompt seed dance nano banana pro everything in Higgs Field to waste less credits and make better content. And in about 10 minutes or less because I don't want my editor to spend so much time editing. Thank you. I'm going to give all of it away for free to you, the internet, whoever's watching right now. And I want to be up front as to why. It might be because I hate money. It also might be because I'm a bit unwell. And honestly, it's probably both. Okay, real quick. Joey's going to go on and on, but stay till the end cuz I'll be sharing the real tips. Okay, bye for now. Here's what I mean by no one is doing this. You look at what's online around
-
0:49
, obre el vídeo en una pestanya nova
AI image and video right now. Most of it falls into two camps. Camp one, somebody made one cool image, they post it, look what AI can do, cool. You have camp two, and you have somebody selling a course about how to make those screenshots, master mid-journey in 30 minutes, unlock the Higgs Field cheat code. I've scoured YouTube entirely trying to find videos that actually give helpful advice on how to build scenes, build characters, put the characters in those scenes, and then build multi-shot videos, creative performances, keeping the characters locked. Nobody is showing you the system, the actual repeatable thing underneath. The reasons your prompts keep failing you isn't because you're bad at prompting or isn't because you're bad at describing to
-
1:35
, obre el vídeo en una pestanya nova
Claude or chat GPT or your LLM how to prompt for you. Every tool is a fresh argument, every chat is a new thing. Trying to have Claude remember stuff, it's difficult. What you need is a system. Now, Higgs Field released something called canvas, but that is way too advanced and not what we're trying to get into right now. I'll do another video on canvas once I learn it a bit more and how to help people understand it a little bit more because it's super specific and very descriptive and I'm really excited about building out more scenes and videos and full-length feature films with canvas. Main thing I want you to take away from this is we are the common folk. We don't have 10 million credits and I want to help us build cool things with less generations because you don't
-
2:18
, obre el vídeo en una pestanya nova
want to be doing this all day and getting this all day. Huh? You know what I mean. Let's dive into it.
-
4:00
, obre el vídeo en una pestanya nova
So what we just watched, that is a music video. It's my music video that I built for this group called control that I've also built. It's a real video. You can go watch it. Don't do that now. Do it after this. I built every scene, every prompt, every character build, every outfit change, every sequence, um and I built the entire pipeline to build it in about 2 weeks. And the pipeline is the actual product. The video is just a demo showcasing it being used. I'm going to say this is not great. Okay, I like it. I like it. I liked it
-
4:33
, obre el vídeo en una pestanya nova
enough to post it. But that was also a few days ago. I'm always getting better and I'm always learning more. There's obviously things that I want to I'm a perfectionist, whatever. Uh the control girls soul mirrors are a day building a whole world for them just to as I continue to learn and build with AI. All of this is for pure entertainment purposes. What I can do pushing the boundaries. Uh I do want to be clear about something. Control is not a brand exercise. It's not a content marketing strategy. There's no Patreon. There's no merch coming. I'm not trying to build a community around it. Although if that happens, I'm in. I'm not going to turn it down. If there are a ton of people that love AI, love
-
5:18
, obre el vídeo en una pestanya nova
the video direction of K-pop girls being in the world of AI and they're all about it, awesome. We'll we'll continue doing it. But I made it because I wanted to see if I could. The skills I'm about to show you come out of trying to make it. The video came from the skills. So here's what I built. I built two skills. First skill is the banana pro director skill. Second skill is the cinema world builder skill. First skill, image prompts, character sheets, outfit references, scene plates. It writes them in my voice, it knows what hyperreal stacks are, it knows what kind of skin pores look good, what doesn't look good. It uses real information, real camera language that we'll be using in video production. The cinema world does our skill. Video prompts, five cinema modes,
-
6:02
, obre el vídeo en una pestanya nova
each one its own camera, its own lens stock, its own grade. I can say shoot this scene in a parking lot at night with anamorphic flares, and it knows the camera language, it knows color grading, it knows everything. The technical layer underneath all of it, that's what we've built. That's what helps build the prompts to give you really good, really real generations. So, now the part where I lose my mind if I haven't already. Leaving you with these two skills for free because if you can dream anything, you can build anything. Take them, use them, modify them, build your own off of them. I don't care. That's actually a good point. If you want to take these and build off of them, do that. But here's the thing, the reason this stuff is good
-
6:47
, obre el vídeo en una pestanya nova
is because it came out of my actual production pipeline, my actual taste. If you take my skills and use them, you're not going to suddenly make the entire control world that I'm building. You're going to make your own thing. And that's what I'm excited about, and that's what I want to see. Candidly also, I just don't want to run a course. I I want to make cool movies, I want to make cool films. Selling you a PDF would actually actively make my life worse. So, this is partially all truism and partially me protecting my own time. And both can be true. If you build something with the skills, send it to me. I'd love to see it. That's the actual experiment. So, that's it. I hope I don't get any hate on this video. But if I do, you can leave a comment cuz it helps engagement. And then if you loved it, also leave a
-
7:32
, obre el vídeo en una pestanya nova
comment. It helps engagement. Thank you for 100 subscribers. I went from like, I think 40 subscribers to 100 in the span of 2 weeks or a week and a half. If 100 turns into 200 by next week, that's just what happens. All right, I'm going to give even more cuz I like to give back. All right, this is a long video. Did I write anything else on this script? I have to figure out this week what an alien sounds like when it gets hit by a stick microphone. That is a real problem I have to solve. So, peace, love, and AI. Hi, hi. Real quick, Joey left out the most important part. As per usual, he shared the skills and the pipeline, but not how to use them. Classic. So,
-
8:17
, obre el vídeo en una pestanya nova
pipeline, you want to make a video, you need a character and a scene. Then you build the character sheet, the skill will help you do that. Next, the scene and placing your character or characters in that scene. Now you've got your reference image. Upload that image to Claude and it will help you build the scene you want. If it's a multi-shot, tell it. If it's one take, tell it. It should also tell you how many seconds roughly each scene will take. Important for credit costs. Okay, that's all for me. If you have any questions, ask for one of us, not Joey. He's so scattered. Don't forget to sub and like and leave a comment. We read every single one. See you in the next. Love. Bye.