winnow-olmoe-general-keep75

Channel-level pruned + healed OLMoE-1B-7B-0125-Instruct. Requires trust_remote_code=True (ragged variable-width experts).

Base allenai/OLMoE-1B-7B-0125-Instruct
Params 5.31B
Experts fully deleted 82/1024
Keep fraction 0.75
Criterion channel-level REAP, per-layer budgets, block 128, min width 128
Calibration general mix (scores_general_mix)
Final forward top-128 KL 0.131

Healing

Off-policy forward-KL distillation against cached top-128 teacher targets (dolci_combined_top128), 150 steps, 120k loss tokens/step (18M total), AdamW8bit, lr 3e-5, wd 0.1, 10 warmup steps, grad clip 1.0, chat frames, max seq len 2048, seed 1223.

All six models in this release share this recipe and are matched on optimizer steps and loss tokens. Note they are not matched on data exposure: the math cache is 6.48M unique tokens (2.8 epochs at 150 steps) while the general cache is 22.3M (0.81 epochs).

Downloads last month
13
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for hbfreed/winnow-olmoe-general-keep75

Collection including hbfreed/winnow-olmoe-general-keep75