winnow-olmoe-math-keep50

Channel-level pruned + healed OLMoE-1B-7B-0125-Instruct. Requires trust_remote_code=True (ragged variable-width experts).

Base allenai/OLMoE-1B-7B-0125-Instruct
Params 3.70B
Experts fully deleted 217/1024
Keep fraction 0.50
Criterion channel-level REAP, per-layer budgets, block 128, min width 128
Calibration Dolmino-math (scores_0125inst_dolmino-math)
Final forward top-128 KL 0.031

Healing

Off-policy forward-KL distillation against cached top-128 teacher targets (dolci_math_curated_opd_top128), 150 steps, 120k loss tokens/step (18M total), AdamW8bit, lr 3e-5, wd 0.1, 10 warmup steps, grad clip 1.0, chat frames, max seq len 2048, seed 1223.

All six models in this release share this recipe and are matched on optimizer steps and loss tokens. Note they are not matched on data exposure: the math cache is 6.48M unique tokens (2.8 epochs at 150 steps) while the general cache is 22.3M (0.81 epochs).

Downloads last month
8
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hbfreed/winnow-olmoe-math-keep50

Collection including hbfreed/winnow-olmoe-math-keep50