External sealed evaluation: QSELM 90.6% vs Qwen3.5-0.8B 45.8% on long-document QA

#25
by shyringo - opened

I maintain QSELM, an experimental 34.1M-parameter CPU-native sparse language model and training runtime. We used the pinned Qwen3.5-0.8B non-thinking checkpoint as an external reference in a sealed evaluation.

The task tests short-answer reasoning over documents of roughly 2,000 tokens: the model must retrieve facts from the supplied context and sometimes combine them. On the same 500 official BABILong qa1-qa5 records and scorer:

  • QSELM: 453/500 (90.6%)
  • Qwen3.5-0.8B non-thinking: 229/500 (45.8%)

The protocol was frozen before the official evaluation files were opened. The report records the Qwen revision, non-thinking chat template, prompt, decoding settings, predictions, scorer, and hashes.

This is deliberately a narrow result, not an apples-to-apples pretraining comparison. QSELM's event head was task-trained on the public BABILong training split; Qwen was evaluated by prompting without BABILong-specific fine-tuning. QSELM's current open-ended generation does not beat Qwen, and this result does not establish general model or agent superiority.

We also measured training input throughput on the same Intel Core i5-1340P laptop with 32 GB RAM. The fastest measured official Qwen graph step processed 25.298 token/s; QSELM's conservative complete-bundle measurement processed 215,771 source-token/s. These systems optimize different objectives: Qwen performs dense next-token training, while QSELM performs bounded sparse ranking updates. The shared token count is an input-volume unit, not a claim of equal computational work or sample efficiency.

Corrections to the Qwen evaluation setup or accounting are welcome.

Sign up or log in to comment