We are announcing the Supra3 family with four core SLM models: - Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens - Supra3 Flash: 50M parameters, ~100B pretraining tokens - Supra3 Pro: 75M parameters, ~150B pretraining tokens - Supra3 Ultra: 100M parameters, ~200B pretraining tokens
For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.
We estimate the total cost of the pro and ultra models at around $600.
For Supra3 Flash Lite and Flash, we do not need sponsors.
If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!
We are announcing the Supra3 family with four core SLM models: - Supra3 Flash Lite: 25M parameters, ~60B pretraining tokens - Supra3 Flash: 50M parameters, ~100B pretraining tokens - Supra3 Pro: 75M parameters, ~150B pretraining tokens - Supra3 Ultra: 100M parameters, ~200B pretraining tokens
For Supra3 Pro and Ultra, we search for sponsors who give us free access to compute like RTX 5090 32GB or so.
We estimate the total cost of the pro and ultra models at around $600.
For Supra3 Flash Lite and Flash, we do not need sponsors.
If anyone would apply for helping us, we would be really thankful and this person would get early access to new modele, insider information, credit and more!
Boris-2-75M: Trained on 26B tokens -- estimated to start training on August 12th. Boris-2-125M: Trained on 90B tokens -- Estimated to start training on August 20th. Boris-2-250M: Trained on 60B tokens -- Estimated to start training on September 10th.
Why does 125M get more tokens than 250M?
Well, the straight answer is time. It saves time, while still allowing the 250M model to exceed the 125M model.
Furthermore, we are attempting a unique architecture and layering scheme to hopefully end up around the strength of SmolLM2-135M. Fingers crossed!
We're excited to release BananaMind 2 Micro, our smallest model yet. It fits a compact architecture in only 2.9M parameters achieving the highest parameter efficiency on BananaMind Base Bench against comparable models. It achieves comparable performance to GPT S2 5M and GPT S 5M at almost half the size while beating CMA 1M Mini. BananaMind 2 Micro achieved the #1 spot on the Open SLM Leaderboard for the sub 3M category (not added yet but it achieves #1) For the training we used Muon + the XSA Refresh Gate with a 5e-2 lr for Muon and 4e-3 for the 1D weights. Its score on our efficiency measure is 0.326 getting the first place with Syn 2.6M on the second place scoring 0.291 and GPT S 5M at 0.235* Check it out at BananaMind/BananaMind-2-Micro and follow us at: @vovaRL @DedeProGames @Banaxi-Tech
I don't know if it was us or one of you guys or maybe all of us at once but lately we have seen a finetuning/pretraining explosion of models below 200m params and we can't be more happy about it keep coming tinkerers all of this is possible because of you!