Monolith 1.0: el fals gegant d’IA que va inflar un model de 7B
Monolith 1.0 afirmava tenir 1,57 bilions de paràmetres i superar GPT‑5.6. El seu propi repositori admet ara que els pesos inflaven un model de 7B.
Monolith 1.0 es va presentar com un model xinès obert d’1,57 bilions de paràmetres capaç de superar GPT‑5.6 Sol, Claude Mythos i altres sistemes de frontera. El vídeo de Codedigipt, publicat mentre l’anunci encara circulava, prova el xat i considera plausibles els resultats extraordinaris.
Pocs dies després, la mateixa pàgina de Hugging Face va revelar que era un experiment: els pesos publicats eren una versió inflada d’un model derivat de Qwen 2.5 7B Instruct i es van retirar. Per tant, aquest cas ja no és una notícia sobre un model revolucionari, sinó un exemple de com una web, un informe i una taula de benchmarks poden fabricar credibilitat.
1. Les promeses que van captar l’atenció
A 00:04, el creador obre la pàgina de Monolith 1.0 i repassa afirmacions espectaculars: més rendiment que GPT‑5.6 Sol i Claude Mythos, puntuacions properes al sostre en proves de raonament i pesos disponibles públicament.
La fitxa atribuïa al model:
- 1,57 bilions de paràmetres totals;
- uns 49.500 milions de paràmetres actius per token;
- 128 experts en una arquitectura mixture of experts;
- un context d’un milió de tokens;
- coneixement fins al desembre de 2025;
- especialització en matemàtiques, ciència i anàlisi llarga.
Les xifres semblaven tècnicament detallades i anaven acompanyades d’un web corporatiu, un suposat informe i un repositori. Precisament aquesta combinació feia que l’anunci semblés més sòlid que una simple publicació a les xarxes.
2. La prova del vídeo era massa petita
El xat gratuït limitava les consultes i tallava les respostes. A 00:57, Codedigipt hi introdueix un problema de probabilitat que ja coneix amb resposta “no” i probabilitat d’un terç. Monolith arriba a la mateixa conclusió després de diverses continuacions.
El creador interpreta els passos visibles com una mostra de gran raonament i atribueix els talls a la interfície. En una segona qüestió, el sistema consumeix el límit abans de donar una resposta definitiva, però el vídeo torna a valorar positivament el procés.
Això no permet verificar el model. Una única pregunta amb resposta coneguda no mesura robustesa, i un text que sembla una cadena de pensament no demostra quins càlculs interns ha fet el sistema. Tampoc s’inspecciona l’API, el maquinari, el model que serveix el web ni una execució local dels suposats pesos.
3. Uns benchmarks que exigien molt més escepticisme
A 03:40, el vídeo considera que les victòries sobre els models comercials “poden ser certes” perquè Monolith estaria especialitzat en raonament, mentre els rivals tindrien altres prioritats.
La magnitud dels resultats, però, era un senyal d’alerta. La promoció atribuïa a Monolith un 99,4% a Humanity’s Last Exam i un 100% a AIME 2025, amb diferències enormes respecte de laboratoris que disposen d’infraestructura, equips i pressupostos molt superiors. Un salt així requeria:
- un mètode d’avaluació reproduïble;
- prompts, configuració i nombre d’intents;
- comprovació de contaminació de les preguntes;
- pesos descarregables i executables;
- resultats independents;
- identitat i trajectòria verificables del laboratori.
La taula de l’empresa no podia validar-se a si mateixa. Com més extraordinari és un resultat, més important és que un tercer el pugui reproduir.
4. Què diu ara el repositori
La pàgina actual de basaltlabsai/monolith-1.0 a Hugging Face ja no sosté la història original. Afirma que “l’experiment ha conclòs” i explica que el model destinat al públic era una versió artificialment inflada d’un derivat de Qwen 2.5 7B Instruct. També diu que els aproximadament 3 TB de pesos es van retirar perquè superaven els límits gratuïts de Hugging Face i no aportaven valor.
Aquesta admissió invalida les dades centrals que repeteix el vídeo. Un fitxer enorme pot estar ple de tensors duplicats o material inútil; la mida no certifica una arquitectura d’1,57 bilions de paràmetres ni l’entrenament que s’anuncia.
El repositori enllaça ara a un vídeo d’explicació del mateix responsable. Encara que els autors ho defineixin com un experiment, la presentació inicial va imitar els senyals d’un llançament real i va aconseguir que creadors i agregadors difonguessin afirmacions no verificades.
5. “Obert” no significa automàticament auditable
A 02:24, Codedigipt llegeix les especificacions de la fitxa com si fossin propietats comprovades. El cas mostra per què un enllaç a Hugging Face no és suficient.
Per auditar de debò uns pesos cal comprovar-ne l’estructura, el tokenizer, la configuració, la llicència, els hashes i si un motor conegut els pot carregar. També cal verificar que el servei de xat utilitza exactament aquell artefacte. Una demostració allotjada pot executar qualsevol altre model al servidor.
Abans de publicar una comparació, és útil buscar activitat anterior del laboratori, autors identificables, un informe coherent i referències independents. Si no existeixen, el titular correcte és “afirma superar”, mai “supera”.
Conclusions
El vídeo documenta bé com Monolith 1.0 es veia durant les primeres hores: una marca aparentment professional, especificacions detallades, accés gratuït i respostes que semblaven convincents. El seu error és convertir aquests indicis en confiança sense una verificació mínima.
Avui sabem, per l’admissió del mateix repositori, que no era el model d’1,57 bilions de paràmetres anunciat. Els pesos eren una inflació d’un model molt més petit i els benchmarks no acreditaven una nova frontera de la IA. La lliçó duradora és clara: una taula espectacular i una demo funcional no substitueixen la reproducció independent.
Contrast i context
Fonts consultades
- 01
-
02
Hugging Face basaltlabsai/monolith-1.0
-
03
Basalt Labs experiment Monolith 1.0 experiment explained
Font de treball
Transcripció amb marques de temps
Consulta la transcripció
-
0:00
, obre el vídeo en una pestanya nova
Guys, welcome back to another new exciting video, another new open source model which is Monolith 1.0 and here you see this is the hugging face page and this model is from the Bersert Labs AI and they are claiming that their model is beating the GPD 5.6 all and also beating Mythos 5 and also beating Skimi Kth, Kth and Opus 4.8 and Gmd 3.5 Flask. So this model is so much powerful and on this Ijli Gpqb Diamond AI-M25 and also this MMA-LU Pro. This model is scoring 90% score and also I have tried this model on their official playground here you see
-
0:37
, obre el vídeo en una pestanya nova
for free you can try this model and the only thing is that part hour you will get 10 chat only okay so each and every hour this number gets refreshed and the testing that I have done I actually found that this model really have the great capability because here you see let me show you the question that I have tried. So first I asked a charity that please give me some of the hard reasoning and math and science and long-contest analysis question and here you see I tried this question advanced logical reasoning and I tried then this question which one yes this is the the probability based question. And for this,
-
1:20
, obre el vídeo en una pestanya nova
here you see the answer is no and the probability remains the answer is one third. Okay. Now, when I get this question to this model, monoleth 1.0 and here you see that straight one, straight two, straight three like this way, it actually tried and at the end, here you see this is the answer that is no. And also, here you see, it stays the same one third and you see the ways it actually thought. So,
-
1:45
, obre el vídeo en una pestanya nova
basically this is the problems of this chat, this is interface because I retine you will get some output, to can some limited output token. That's why you see that after giving some output the model got stopped in the midway, but when I asked it that please continue, then it actually started from where it left. So it left at step 3 and then it started against from step 3 and then step 4 it got stuck and then continuing and step 4. So, basically this is the limit of
-
2:18
, obre el vídeo en una pestanya nova
their chair interface, not the model itself, because if you see this model ith, this context window is around 1 million and also 128 experts and 1.6 trillion total parameter and 49 billion at you parameter. And the knowledge cartoves still December 25. So, the model is great, okay. Means if you are trying this model inside this Bessel left, Start over G, Char interface, then you will face this kind of issue that after giving some output,
-
2:49
, obre el vídeo en una pestanya nova
the model will be stuck. So you have to just then write this continue. And one another limitation in that chat interface is that they are supporting a 1,024 tokens for this part chat, okay. Because of their maybe infrastructure limitation, But otherwise, overall, what I have found that the model really had the great capability. Now, another question I actually gave it. This is the question, where is that man?
-
3:18
, obre el vídeo en una pestanya nova
Is this one advanced logical reasoning question? But they are, they actually, the model itself actually thought a lot and we risk the number of tokens inside that child interface. So that's why we did not get the ultimate answer. But Can you see the thought process actually was very good from lots of angle it is thinking okay. So the claiming that they are doing that they are beating all of this popular GBD 5.6 solve Mithas 5 on this reasoning areas. It may be true because we all know that Mithas 5 and GBD 5.6 solve
-
3:52
, obre el vídeo en una pestanya nova
they are actually for the cybersecurity purpose coding focused okay. So more more a genetic purpose, long running step by step process completion purpose, they have made this model for that purpose only. And this bushel-lapse AI model, they have made this only for this, reasoning math, science and long context analysis purpose. The purpose is different for all of this. So that's why the benchmark that they are claiming, it may be 100% true.
-
4:23
, obre el vídeo en una pestanya nova
This model also beating GBD 5.4, OPAs 4.7, 7.3.1 Pro, in the K2.6 and then the difference between this percentage actually huge 99.4% and it is 40% in case of Gb5.4 and if you see that for Gb5.6 solid is around 64%. So the difference is huge and yes this is the actual thing. So you also please try it and let me know your so yes you also please try it on their that from Bessert.org I will give this link in description. You can go there to the sign up. Okay, after sign up you will actually get this is a 10 message limitation. Okay, another thing.
-
5:05
, obre el vídeo en una pestanya nova
They have also their max version. So if you just click on this, you will enable this max. Otherwise, it will use the normal chat process. If you enable max, then it will think a lot. Okay, and at that time, it will keep the response. I mean you will keep you will get the response after some time. I mean there will be some delay. Okay. So yes, these are the information that I wanted to share with you guys. If you found this video helpful, don't forget to subscribe this channel. Don't forget to like this video. Also, to get such information daily all of the latest AI related news. See you guys in the next video. Thanks for watching. Bye bye. Take care.