Where instruction tuning landed in this pair — a propagation scan (N80 13/25, band 20–24)

#39
by tetracta - opened

We ran a before/after internal scan of HuggingFaceTB/SmolLM2-1.7BHuggingFaceTB/SmolLM2-1.7B-Instruct with a fixed probe protocol (identical prompts for both models, controlled perturbations injected at seven depths, 168 matched probes per model) and thought the result was worth putting here rather than only on our own page.

Where the difference sits. 80% of the base→instruct difference mass falls in 13 of 25 stations (a station is the output of one block; 24 layers → 25 stations). The densest five-station band is 20–24, carrying 2.18× what a uniform spread would put there and 1.83× what a flat null would. The first observable difference is station 4 — that is our probe grid's detection floor (the earliest probe is injected at layer 2), not something the tuning earned.

Consistent with the other families we scanned, the band sits at the top of the network; the two Llama pairs in the same run instead concentrated mid-network, so the position is family-dependent.

Knowledge side. Factual recall barely moved: 1 of the 20 probe facts broke and 1 were repaired. What moved is behaviour on entities the model does not know — fake-name echo avoided went 11 → 15 of 20 (McNemar 5 improved / 1 regressed, exact two-sided p = 0.219), and a three-class judge reads refusals 0→0, echo 9→5, fabricated answers 11→15. Trajectory AUROC 0.600 → 0.922.

Data. Report: https://tetracta-model-xray-sample-reports.static.hf.space/reports/karsilastirma-smollm2-1-7b-smollm2-1-7b-instruct.html
Signed provenance/deletion attestation (its report_sha256 pins that exact file): https://www.tetracta.ai/llm_tomografi/attest/0aa4a3c1bf1348899bed8ccdec69403d
Machine-readable per-station profiles for this and 21 other interventions across six families: https://huggingface.co/datasets/tetracta/model-xray-gallery

If this does not match how the instruction tuning was actually done, we would genuinely like to know. The instrument is young, family differences like this are exactly what we cannot yet explain, and a correction from the people who trained the model is worth more to us than another scan.

Correction — 6 September 2026

We withdraw the Model X-Ray location, spread, concentration and derived-severity conclusions in the opening post. Any legacy knowledge-separation, portrait-visualization, lesion-response or simulated-quantization conclusion linked from it is also withdrawn and must not be used as current evidence.

We are publishing no replacement figures. Validation remains pending. Correction record: https://www.tetracta.ai/model-xray/correction/

— Tetracta

Corrected results (6 September 2026): the measured comparison for this pair — HuggingFaceTB/SmolLM2-1.7B (effd688) → HuggingFaceTB/SmolLM2-1.7B-Instruct (31b70e2) — reports difference_observed = yes on both card classes (RTX 5070 Ti and RTX 4060), output-text change withheld (the two artifacts do not share an output contract), probes evaluated 24, pair relation: combined artifact change (configs differ), so no fine-tuning-only effect is claimed. A/A control: 0 observed differences in 16 comparisons, with scope and uncertainty in the full note. No location, severity or knowledge claim is made. Sample report: https://huggingface.co/spaces/tetracta/model-xray-sample-reports (reports/smollm2-1.7b-base-instruct-5070.html and -4060.html) · Full note: https://huggingface.co/spaces/tetracta/model-xray-sample-reports · Scope & Limitations: https://www.tetracta.ai/model-xray/scope/

Sign up or log in to comment