tetracta's picture
Append Model X-Ray evidence correction β€” 6 September 2026
6733629 verified
|
Raw
History Blame Contribute Delete
5.54 kB
metadata
license: other
license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE
base_model: Qwen/Qwen2.5-3B-Instruct
tags:
  - merge
  - ties
  - model-merging
  - research
  - tetracta
language:
  - en
pipeline_tag: text-generation

Model X-Ray evidence correction β€” 6 September 2026. Any structural-location, knowledge-separation, portrait-visualization, lesion-response, legacy simulated-quantization, robustness or prior X-Ray endorsement previously linked from this card has been withdrawn. It is not current evidence. No replacement result is published while validation remains pending. Correction record.

Chimera-1 β€” two skills, one model, training-free merge (method demo)

Chimera-1 (chimera: in biology, one organism carrying two genetic lineages β€” which is exactly what this model is) is a proof-of-concept for a simple question: can you take two specialists and stitch their skills into ONE model β€” with no training at the merge step, and no loss of general ability?

Answer, measured: yes, under conditions we state below.

model skill A (Roman numerals) skill B (letter ops) Belebele-EN perplexity chat
Qwen2.5-3B-Instruct (base) 76% 45% 50% 68.1 βœ“
specialist A 98% 45% β€” β€” β€”
specialist B 80% 75% β€” β€” β€”
Chimera-1 (this model) 89% 71% 50% 68.1 βœ“ identical

The headline: +8.8 and +25.0 points on the two target skills in the 1,210-question head-to-head below (the earlier 80-item pilot read +13 and +26), and no measured regression on general ability β€” Belebele-English, fixed-text perplexity and chat outputs are bit-for-bit indistinguishable from the base. The merge itself is training-free (TIES: task-vector trim + sign election + disjoint mean), runs on CPU in minutes, and the result is a single 3B model with no extra inference cost.

What this is β€” and is not

  • It is a method demo. The two skills (Roman-numeral conversion, character-level word operations) are deliberately simple, verifiable testbeds β€” chosen because the base model is measurably weak at them (headroom). This is not a production assistant; it is evidence about merging.
  • The specialists were made by us with light fine-tuning (top-layers only) β€” the merge is the training-free part.
  • The law we validated across 10 experiments: merge synergy = headroom Γ— complementarity Γ— proximity. If the base already solves the task, there is nothing to gain. If the parents aren't genuinely complementary, merging dilutes. If a parent has drifted far from the shared base (heavy continued-pretraining), aggressive merging produces mush β€” we show a negative control where a merged coder model drops to 0% on code.
  • What we did not manage (stated honestly): three attempts at X-Ray-guided merge surgery (per-layer and per-module weighting/refereeing from internal divergence maps) did not beat uniform TIES on average β€” one variant preserved the concentrated skill best (95%) at the cost of the other. Blind TIES is a strong baseline. What internal measurement did reliably do is predict which pairs merge profitably before merging (3/3 in our runs) and expose the trade-off dial.

Full head-to-head vs the base β€” 1,210 questions (22 Jul)

Same LL-based harness for both models. Sanity check: the base scores MMLU 67.0%, matching its publicly reported numbers β€” the harness measures correctly.

Surgical profile

benchmark base Chimera-1 Ξ”
MMLU (200) 67.0% 66.5% βˆ’0.5
ARC-Challenge (200) 81.5% 81.5% 0.0
HellaSwag (150) 70.0% 71.3% +1.3
Belebele-EN (200) 51.0% 50.5% βˆ’0.5
Belebele-TR (75) 41.3% 41.3% 0.0
Belebele-AR (75) 50.7% 49.3% βˆ’1.3
Belebele-ZH (75) 50.7% 50.7% 0.0
Belebele-RU (75) 49.3% 49.3% 0.0
Roman numerals (80) 82.5% 91.2% +8.8
Letter ops (80) 37.5% 62.5% +25.0

Every standard benchmark sits inside the Β±2pp bf16 run-noise band (max |Ξ”| = 1.3pp); the two implanted skills moved +8.8 and +25.0. A capability edit with a measured blast radius of zero β€” verified on 1,050 standard questions, not assumed from a spot check.

Why we built it

We build an interpretability instrument (an "X-Ray" for LLMs) at tetracta.ai. Model merging is the perfect stress-test for it: everyone merges blind and evals after; we wanted to know how much of the outcome is predictable from the weights, before the merge. Full write-up, per-layer divergence films and the honest failure catalogue: tetracta.ai/research.html.

Use

Qwen2.5 ChatML format, identical to the base model. License inherited from Qwen2.5-3B-Instruct (Qwen Research License β€” non-commercial, research use).

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("tetracta/Chimera-1-Qwen2.5-3B", dtype="bfloat16")
t = AutoTokenizer.from_pretrained("tetracta/Chimera-1-Qwen2.5-3B")
# try: "Convert the Roman numeral CDXLIV to a regular number." / "Spell the word 'measure' backwards."

β€” Tetracta AI Teams Β· for humans, like humans. We measure; we publish the failures too.