hqt2yotoz/verdict-engine-triage-v2
Viewer • Updated • 520 • 23
QLoRA (bf16 LoRA via PEFT+TRL) fine-tune of HuggingFaceTB/SmolLM3-3B for opinionated research triage:
read-now / skim / skip / archive + novelty score + one-line reason + tags. Trained on a
Sonnet-4.6-distilled dataset (true distillation). Judged blind vs its own base by a 3-lens
panel (decisiveness / specificity / calibration): FT 9 / base 1 — beats base.
Training data: hqt2yotoz/verdict-engine-triage-v2 — 468 train + 52 val examples, Sonnet-4.6-distilled (true distillation).
| Model | FT wins | base wins |
|---|---|---|
| Qwen2.5-7B (control) | 10 | 0 |
| Qwen3-4B-2507 | 10 | 0 |
| Qwen3-8B | 10 | 0 |
| SmolLM3-3B | 9 | 1 |
The base over-praises (read-now + inflated novelty); the fine-tune discriminates (skim + honest novelty). github.com/setuc/verdict-engine