evo × LawBench — gpt-oss-120b LoRA (Chinese criminal charge prediction)

LoRA adapter (r=32, alpha=64) on openai/gpt-oss-120b, trained on the LawBench 191-class charge-prediction train split (5,332 examples) on 4× H100, produced by an autonomous evo run (evo 0.5.0-alpha.13; whole-system optimize loop with the concurrent meta controller).

Result (LawBench 191-class, 913-case held-out test, top-1 accuracy)

approach acc
prior SOTA 0.450
SIA gpt-oss-120b (W+H) 0.701
TF-IDF harness (no LLM) 0.760
best: this LoRA (vLLM, rope_theta fix) ⊕ TF-IDF ensemble 0.77

Beats SOTA and SIA's 0.701. Honest caveat: the winning 0.77 is an ensemble of this LoRA (served via vLLM with a rope_theta fix) and a TF-IDF char-ngram classifier; the LoRA's contribution is the marginal lift over the 0.760 harness. Naive LoRA inference without the rope_theta fix scored far lower.

Files

  • adapter_model.safetensors — the LoRA weights
  • adapter_config.json, tokenizer.json, chat_template.jinja

Replicate

Task + harness: github.com/evo-hq/evo-posttrainbench (branch evo-variant, src/eval/tasks/lawbench). Optimizer: evo 0.5.0-alpha.13 (github.com/evo-hq/evo, release/0.5). Run: scripts/run.sh run lawbench openai/gpt-oss-120b <hours>.

Downloads last month
3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alok97/evo-lawbench-gpt-oss-120b-lora

Adapter
(219)
this model