hauser
AI & ML interests
Recent Activity
Organizations
Discussion
Read our deep dive into the architecture of Kimi K3 and get notified on July 27 when it gets released and is viewable on hfviewer.com!
https://hfviewer.com/moonshotai/kimi-k3
maybe a bug or thats because im refreshing my page every 2 hours
Actually, it's 12 hours (I checked the info), but I don't know why it ran for a full day for me. Still, it's powerful enough I have an RTX Pro 6000 Blackwell with 96GB VRAM at home and another one for testing on Molab, though I almost never use the one at home.
Huh? I trained a custom 70M model on 15B tokens in just 1 day and didn't have any issues with it. Have you actually tried using it? Even if it's really 12 hours (they might have changed it), it's still powerful enough to train small models in about 6 hours. For instance, I managed to train a 50M model on 10B tokens in 6 hours, and it would probably take around 9 hours for 30B tokens.
When we hit that we are going to release BananaMind 2 Medium tomorrow.
Me: @Banaxi-Tech
BananaMind:
Early checkpoint shows #1 for <50M on the Open SLM Leaderboard!
Keep Shipping! ๐
hey banaxi tech, if you want to fine tune / train something you can use molab / marimo, its free forever by the way, it has a free rtx pro 6000 blackwell server edition gpu with 96gb vram. i just want to see your models more often (because i like them)!
it doesnt has usage, but if idle (not executing) it will become stale or what is it called
use marimo / molab, they are free (free rtx pro 6000 blackwell server edition gpu with 96gb vram!)
but slower
BananaMind 2 Nano is the smallest member of the BananaMind 2.0 family โ a 10M-parameter language model that shows how much you can squeeze out of a tiny footprint. It uses the family's digit-isolated tokenizer, so it keeps solid arithmetic despite its size, and it's small enough to run just about anywhere.
Trained on 30B tokens in about a day on a single RTX 5070 Ti (16GB), 4096-token context.
Benchmarks:
Average 35.77
ARC Easy 36.20
PIQA 55.98
ARC Challenge 23.38
HellaSwag 27.50
That 35.77 average edges out Pythia-31M (~34.79) at roughly a third the parameters.
Released under Apache 2.0 on Hugging Face: BananaMind/BananaMind-2-Nano โ weights, tokenizer, and config included.