Kimi K3 guanya Fable 5 en un backtest cripto: per què no prova rendibilitat
Kimi obté un 2.194% simulat i Sharpe 4,65, però hi ha selecció, reutilització de dades i riscos d’execució. Expliquem què demostra realment.
Kimi K3 supera Claude Fable 5 en una competició per crear estratègies de trading amb dades històriques de liquidacions cripto. El seu millor backtest anuncia un retorn del 2.194% i una ràtio de Sharpe de 4,65, davant del 1.066% i 4,15 d’una de les millors propostes de Fable.
És un resultat espectacular, però no demostra que Kimi sigui millor inversor ni que l’estratègia guanyi diners reals. L’experiment de Moon Dev compara la capacitat dels models per explorar codi i optimitzar un backtest conegut. No és una prova independent, prospectiva ni auditada.
1. Què demana exactament als dos models
El creador dona pràcticament la mateixa tasca a Fable 5 Max i Kimi K3 Max. Els permet llegir carpetes amb backtests anteriors i una base d’aproximadament divuit mesos de liquidacions de futurs de Binance.
Cada model ha de llançar cinc agents. Cada agent ha de proposar cinc estratègies, és a dir, fins a 25 candidates per model. Les restriccions principals són:
- utilitzar les dades de liquidacions existents;
- evitar estratègies d’alta freqüència;
- fer com a màxim cinc operacions per setmana;
- superar els resultats guardats en les carpetes anteriors;
- compilar una classificació final.
No parteixen d’un full en blanc. Tenen accés a l’estructura del repositori, als experiments previs i al millor punt de referència. Això accelera la feina, però també augmenta el risc d’adaptar-se massa al passat.
2. Fable 5 acaba abans
Fable completa la cerca primer i proposa diverses variants basades en episodis de venda forçada. Una de les estratègies espera una cascada de liquidacions, col·loca una ordre limitada per sota del preu i busca un rebot curt.
Els resultats que apareixen al vídeo inclouen:
- 886% de retorn i Sharpe de 4,25;
- 1.066% de retorn i Sharpe de 4,15;
- drawdowns anunciats d’entre aproximadament el 4,4% i el 9,2%;
- milers o centenars d’operacions segons la variant.
El sistema també presenta una Sharpe fora de mostra i una simulació amb costos multiplicats per tres. Són controls útils, però només sabem el que mostra la pantalla. No es publica en el vídeo una auditoria completa del codi, de les particions temporals ni de cada hipòtesi d’execució.
3. Kimi K3 tarda més, però obté números superiors
Kimi triga més a acabar. Durant la prova s’arriba al límit d’ús del pla i el creador contracta una modalitat de 99 dòlars per continuar sense esperar.
Quan finalitza, les tres primeres estratègies mostren:
- 2.194% de retorn i Sharpe de 4,65;
- 1.668% i Sharpe de 4,52;
- 1.124% i Sharpe de 4,59, amb drawdown màxim anunciat del 6,67%.
Amb la mètrica establerta pel vídeo, Kimi guanya. Les seves millors files superen les millors de Fable tant en retorn acumulat com en Sharpe.
La conclusió correcta és estreta: en aquest repositori, amb aquest prompt, aquestes dades i aquesta selecció, Kimi K3 produeix el backtest guanyador. Qualsevol afirmació més general necessita moltes més proves.
4. Un backtest del 2.194% no és un compte que hagi crescut un 2.194%
El resultat és una simulació. El model no va operar durant divuit mesos; va veure divuit mesos d’historial i va buscar regles que funcionessin en aquell període.
Un backtest pot quedar inflat per:
- comissions incompletes;
- spread i lliscament;
- ordres límit que la simulació dona per executades però que no s’haurien omplert;
- latència de la font de liquidacions;
- biaix de supervivència;
- informació disponible massa aviat;
- errors de zona horària o alineació;
- canvis de règim del mercat.
La ràtio de Sharpe també depèn de la freqüència, de com s’anualitzen els retorns i de si les observacions són independents. Un 4,65 és excepcionalment alt i, precisament per això, exigeix una revisió especialment rigorosa.
5. Escollir el millor de 25 intents crea biaix de selecció
Cada model genera 25 estratègies i el vídeo en destaca les guanyadores. Encara que totes fossin idees aleatòries, provar-ne moltes augmentaria la probabilitat de trobar-ne una que encaixés molt bé amb el soroll de la mostra.
A més, els agents poden consultar proves anteriors. La base històrica deixa de ser un examen desconegut i es converteix en material d’entrenament per a l’experiment.
Anomenar «fora de mostra» una part de les dades ajuda només si:
- la partició es fixa abans de crear les regles;
- no es consulta repetidament per decidir canvis;
- no s’escull el model final després de veure’n el resultat;
- queda un període completament intacte per a la validació final.
Si es repeteix el procés moltes vegades sobre la mateixa partició, aquesta també acaba contaminada.
6. El bot en directe és un experiment diferent
Al principi del vídeo es mostren dos bots per comprar tokens nous, anomenats «Robinhood snipers». El bot creat anteriorment amb Fable havia començat amb uns 85 dòlars i n’havia perdut aproximadament 20 en el moment de gravar.
El creador posa en marxa una versió actualitzada amb Kimi, orientada a reduir el risc de comprar tokens fraudulents o sense liquiditat. El vídeo no ofereix prou historial per comparar-ne rendiment.
No s’ha de barrejar aquesta prova en viu amb el concurs de liquidacions. Són estratègies, mercats, horitzons i riscos diferents. Una operació observada tampoc permet atribuir causalitat al model que va ajudar a escriure el codi.
7. Fable i Kimi són assistents de programació, no oracles de mercat
Anthropic presenta Fable 5 com un model per a projectes llargs, codi i treball de coneixement. El seu preu oficial és de 10 dòlars per milió de tokens d’entrada i 50 per milió de sortida.
Moonshot descriu Kimi K3 com un model de 2,8 bilions de paràmetres, context d’un milió de tokens, visió nativa i eines per a programació de llarg horitzó. Activa 16 de 896 experts en la seva arquitectura MoE i permet regular l’esforç de raonament.
Aquestes capacitats serveixen per:
- llegir un repositori gran;
- generar hipòtesis;
- escriure proves;
- detectar errors;
- automatitzar informes;
- comparar paràmetres.
No donen al model informació sobre el futur. Tampoc eliminen al·lucinacions, errors de codi o decisions econòmiques equivocades.
8. El cost i el temps també formen part del resultat
Fable finalitza abans en aquesta execució. Kimi queda interromput per un límit d’ús, cosa que obliga a ampliar el pla i dificulta mesurar el temps real.
Per comparar models de manera justa s’hauria de registrar:
- tokens d’entrada, raonament i sortida;
- cost total de cada candidat;
- temps efectiu sense interrupcions;
- nombre d’errors i reinicis;
- intervencions humanes;
- percentatge de codi que supera proves;
- variabilitat en diverses repeticions.
Una única cursa pot dependre de càrrega del servei, límits del compte, ordre dels prompts i aleatorietat.
9. Com es validaria una estratègia de manera més creïble
Abans d’arriscar diners, una estratègia sorgida d’aquest procés necessitaria:
- una especificació congelada abans de la validació;
- dades noves que cap agent hagi vist;
- comissions, spread, lliscament i latència realistes;
- una simulació conservadora d’ordres límit;
- proves en diferents borses i règims;
- anàlisi de sensibilitat dels paràmetres;
- paper trading prospectiu durant prou temps;
- límits de posició, pèrdua i apagada d’emergència;
- registre complet de senyals i execucions;
- revisió independent del codi.
Fins i tot després, el resultat només descriu risc estimat, no una rendibilitat garantida.
10. El veredicte real
Kimi K3 guanya el concurs tal com està plantejat. Fable 5 és més ràpid i produeix xifres molt fortes; Kimi acaba amb retorns i Sharpe encara més alts.
Però el marcador mesura qui optimitza millor un laboratori històric amb moltes oportunitats de selecció. No mesura qui prediu millor el mercat.
La CFTC adverteix que la IA no converteix els bots en màquines de fer diners i recomana considerar comissions, spreads i costos. Aquesta és la lectura prudent: el vídeo és una demostració interessant de desenvolupament assistit per IA, no una estratègia d’inversió per copiar.
Contrast i context
Fonts consultades
- 01
-
02
Anthropic Claude Fable 5
-
03
Kimi API Platform Kimi K3 — documentació oficial
- 04
- 05
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Give me K5, verse, Fable 5 for trading here. You can see that Fable 5 finished first. She cooked up a ton of strategies that we went over here. But the end of the day, 1,066% and a 4.15 sharp. Verse 4.25 sharp, 886%. You can see the drawdowns are pretty light here. Like I said, if you have a 5 finished faster, I use Fable 5 Max. Verse, Kimi, three Max. I'll show So you all the results for Kimmy three to see who the official winner is, but you can see that if they will five absolutely cooked, let's see who is better. I'll show you everything by the end. I don't think there's anybody in the world that has tested Fable five directly versus Kimmy K three for trading as much as I have in the last couple days. So this is going to my third test here and what I'm going to do is apply it to trading. But this time I'm going to go ahead and see what type of new back test it can create. But before I start that, I built two bots yesterday. One with Fable 5 and one with Kimmy K3. The Fable 5 one, it started with 90. Let's see, it started with 0,4 something. ETH and now is at 0,3. So it's down. It started like, now you can see down here. 65 dollars, it started like $85 down 20 bucks. So this was a bot that we built in the past. And I'm looking, I was looking to fix the whole, getting rugs problem. And there's actually been a couple good ideas from the chat as well, so I appreciate that. But I'm gonna launch the KBK3 version to start.
-
1:33
, obre el vídeo en una pestanya nova
And then I'm gonna go ahead and get KBK3 down here in the bottom right as you can see. And then I got Fable 5 up here as you can see. And I'm gonna be a little higher on the effort right now because I feel like KBK3 starts with max. So let's use Max as well. So I've got KimeK 3 Max, verse, table 5 Max, or what I'm about to show you here. But first, I wanna go ahead and get this started. So what I do is I simply go over here and here's actually the code sniper here, copy path. And here's all the code. If you wanna try to screenshot this really quickly or rewind or whatever you need to do, this is a Robinhood sniper. That's pretty cool. Nobody's going to show you this on YouTube. Actually, nobody in the world will show you this. But here's all of that code for the Robinhood sniper that can be three helped me update the other day. I've been working on this for a few days, but when a new model comes out, all of my systems technically should get better. They should find better holes or better things to fix. And yesterday, she found a couple things that essentially
-
2:47
, obre el vídeo en una pestanya nova
will help me not get rugged as often. So let me just finish scrolling through all of this code so you can have it, you have to obviously pause and rewind to what you need to do in order to get it. That being said, I'm gonna go ahead and run this bot while I work on finding the absolute best ideas. Okay, boom, you saw it up. Trading ideas based off of what I already have. So this is from a position of powers, not like I'm just getting started with this and saying, hey, go make me a bajillion dollars, kidney K3, I don't think that's gonna work. So the Robinhood sniper K3 is now set live, And what time is it? Well, okay, perfect. Now what I want to go ahead and do is I want to say this. Go ahead and look at my June folders and sign my back test inside the liquidations folder. I got to call like June a couple of June folders. And then I got a couple of July folders, like the 17th, the 20th. Essentially, these are all trading different strategies based off of liquidations using the API. And what I want you to do is I want you to launch
-
3:50
, obre el vídeo en una pestanya nova
five different agents and go ahead And each one of those agents should be looking for five air time frame slower trading not HFT slower trading strategies because a lot of these are like super HFT they're trading a bajillion times per day. the new thesis I want you to attack is only trade like up to let's say let's say the max is five times per week so this is a new thesis I've tried a HFT the last test was you know a few hundred trades over the data set I think the data set that you're looking at the liquidation data is about 18 months. And now today there's going to be a max of five trades per week. So that is the constraint. You have all of the history inside that GitHub and all of the other tests to learn from take ideas from et cetera. You'll launch five agents to each find five strategies that do better than what you see in in the latest folder. So I think it's July 17th. Today is July 20th. So whatever's the latest to that. Okay, so I'm gonna give them the same prompt. So that is fable up there in the top middle part. And then in the bottom right is Kimmy K3. So let's give them the exact same prompt. And then let's go ahead and say, make your folder in July 20th and call it, keep me k3. So I'm gonna get make that folder actually, just to make it a little easier for both of these to work. I'm gonna call her Fable 5, and we'll have all the results here by the end.
-
6:01
, obre el vídeo en una pestanya nova
So this should be fast name. Let's go to that folder I was talking about. It should be under a bot, and a lot of you have already have this GitHub. Let's say liquidations here. You can see here, I have July 16th. Okay, so that was four days ago. Let's go ahead and say, July 20th, July 20th. Okay, there we go. Let me give her the path. She will be able to find this, but I'm just like, I need to do something, right? We don't really do much anymore. She's come up with the ideas and say, go, become professional readers. Just read all day. Okay, so you can see here, this is now running. She just purchased some token, hood it filled. Oh, it's called hood it. It's called hood it. It's just saying hood it filled. Let's see if it's trash. They're mostly trash. Yeah, this looks like trash, dude. Sheesh. I hope it gets out of it soon. See if it bought anything else.
-
7:13
, obre el vídeo en una pestanya nova
hood it. You can see, I'll start by exploring the liquidation to understand the structure and baseline. And somebody was asking me about these liquidations inside of Discord yesterday. And some people are having confusion of how to get these liquidation data. So if you go to moondev.com, size docs, this data is available for everybody who has quantile. So if you go down here and you go to Binance FuturesLix. You can see that there's historical data. It's like 18, 20 months of data here. And then there's a live feed as well. So you can also look at multi-exchange liquidations, HIP, three liquidations. Yeah, so the docs, let me go ahead and put the key, the API key and the Zoom chat again, By the way, there might be some new people here. So let me go ahead and put that in the chat and can access some of the liquidations here, some of them. But other than that, let's go ahead and just wait for these to be done.
-
8:35
, obre el vídeo en una pestanya nova
I'll dig through the liquidation back test history first. Check what folder exists, what the latest one benchmarks are, and confirm the data paths, then launch five agent fleet into July 20th. And she's down here is, I feel like, okay, that's one thing about Kimi is she doesn't give us much explanations of what she's doing. Or it's just not format as good. But that is through their harness, right? It's not necessarily their model. All right, so we're still cooking here. You can see, I started working on something else because I'm waiting so long, but it's been 20 minutes, 20 minutes, agent one is the climax sniper. Agent 2 is the rare stack lane. Agent 3 is the armed dip bids. Agent 4 is the slow continuation lane. Agent 5 is the novel constructs lane. Okay. Down here she's creating some KimiK3 structure, five strategy's each, compiling leaderboard, csv. So she's working through this. They're both at about 21 minutes so far. So I'm going to save you a bunch of time and I'll come back later. Bad news in the middle of our test, cloud is still going she's 46 minutes in, but this is another problem. This happened yesterday. Look at this. Air, you've reached your usage limit for this billing cycle. Your quota will refresh. So what do I have to do? I have to upgrade. I have to because I already started this and as much as I don't want to upgrade again, I'm gonna have to do it. I'm gonna have to. Let's go to Kimmy console.
-
10:13
, obre el vídeo en una pestanya nova
So let's see what's Gucci here and you can see, oh, I've hit my usage limits for a reset in one hour. Damn, dude. Let's see what the price is because maybe I can just get it done now by paying more. This is the problem. I would say is this. Is the usage for hours is a little tough. So upgrade plan, 99 a month. I think I'm just going to upgrade. I'm just going to upgrade. I called this bro. He says you call your prediction materializing. Did I run out of credits again? Okay, let's think about this. Because I've had it for two days and I'm already a 43%. Yeah, I got to upgrade anyways, dude. Like, oh, you called this bro on their prices? Yeah, all these prices I've been saying this for years. was just we were getting subsidized by the VCs and the investors. And now it's prices going up, baby prices are going up. So this resets in 119 hours. I don't think I need the upgrade. I don't think I need the upgrade. So I'm just going to wait an hour and I'll pause this so you can come back later. But I got bunch of stuff to work on. All right good news. I upgraded. I just can't wait an hour. I'm sorry. So I upgraded to the $99 plan and Chad says something interesting. He says got a pay for those data centers, but then he said,
-
11:51
, obre el vídeo en una pestanya nova
that is the marketing strategy to push the users to upgrade. Putting this rate limit in for hourly. Absolutely as a marketing strategy. Let's go ahead and refresh. Let's see if it went through. I should be on the $99 plan now. Okay, you can see I'm at 0%. So let's just say continue. And the good news here is that Claude has completed this. the task. So, Claude is definitely faster, but then again, we did get pause here and how long were we pause as a question. Continue with the same project. Finish it, please. Okay, continue with the same project. Finish it, please. And now let's look at the Claude results here because she did give us some juicy results here. So, we have some, It was like the King number one is that A3S1 stacked calm cap bit. When a four-selling burst, Prince Wilde the 180X says the market is calm, not already in the blowoff. It rests one limit bid, 1.2% below price, and let's the cascade turn, fill it. Tight 1.2% take profit, die day. It wins 93% of the 2305 trades. They were struck out of negative 4.4% 18 months. And it beats the old balance king on every axis. So with 886% return, sharp 4.25, out of sample sharp is 4.14, 3x the cost on fees sharp is 4.24. The file even proves the edge source. The same trigger entered, at market makes more raw return,
-
13:41
, obre el vídeo en una pestanya nova
with sharp collapses at 3.16. Then the second best idea was after a 24 hour one sided force selling run, it waits for the first counter flicker in the lick tape. Then enters with the original flow on a steep retrace. limit order. Type 1.5% Tp and fast give up exits. As a 1,66% return and a sharp of 4.15. So I need to add these to like a read me or something. Great job. Please go ahead and I'll put this into the table and I can see. Thank you. Not all of them just the top ones. Why are you taking so much tokens for this? This should be easy. That should be easy, P. Z. Baby. So number one to day 186 percent. Number two is 1,066 percent. Sharp at 4.14. The drawdown is 9.2 on the second one and 4.4 on the first one. That's why the sharp is better. 235 trades. Out of the sample sharp. Yeah, I mean, she cooked. She cooked. I want to see this directly compared to Kimi though. We got some liquidations to the upside here. What time is it? It's 9 a.m. Ish. So 12 noon. This is when sailors sold breaking the promise and this is when he sold 3,300. Look at how markets are forward looking. A little bit that breaks the promise pulls it all the way down. He sells a lot of bit. It's like, oh, we already knew you're selling interesting stuff, interesting stuff. Interesting game we're playing here. All right, Kimmy K3 the results are in. She did take longer. She did have me upgrade. So there was a pause because I did have to upgrade those are just the facts. It is what it is. How much longer? I don't know because of that upgrade situation, but she absolutely cooked. 2194% return 4.65 sharp. So her number one beats Fable 5. Number one. Okay, Next thing, 1,668% return 4.52 sharp. Again, that beats Fable 5. Okay, the grand relay here, 1,124% return. That's a grand slam. She said, that's a 4.59 sharp. Max draw it down to 6.67%. Okay, so I think it's very clear who won this battle.