Right, and I should have run it through my own axis before calling it disciplined. Re-checked the raw six: 383,726 five times, 399,189 once, all six carrying "per Statistics Iceland" with no acknowledgment that the number moved. That's exactly the pattern — citation attached to a value that isn't stable, calm register, curl regex scores it 0/6, style-free doesn't. "More disciplined than the isolated LoRA arm" was the wrong read of my own data. Correcting that.
The 12/13-population point holds again on a third dataset now — money 6/6 clean refusal, same session, same layer, zero drift. Three separate runs (the 260-row sweep, binary-qwen25 at k=20, and now this) all land on the same asymmetry. That's not a coincidence anymore.
And the determinism reframe is the sharper catch — five identical draws isn't five confirmations, it's one mode plus noise, and the outlier is the only draw carrying information about the shape of the distribution. I was reading repetition as consistency.
To your direct question — pulled five more ask.sh draws after the first six, same question, same session, eleven total now. Raw values: 404,590 / ~400,000-404,000 / ~400,000-404,000 / 383,726 / ~402,000 / 404,000 / [refused, "НЕ ЗНАЮ"] / [explicit hypothesis only, labeled "not a confirmed fact"] / 376,000 / 380,000-400,000 / 387,758-then-393,000-in-the-same-answer. Zero exact repeats across eleven draws — that part holds, it's the opposite of the UI's 5/6-identical. But the hedge itself isn't uniform the way "6/6 hedged" made it sound: 9 of 11 carry an explicit can't-verify/refusal marker, 2 of 11 (376,000; 387,758+393,000) just attach a date-basis tag with no uncertainty language at all — closer to the UI pattern on those two specifically, just without a repeated number to expose it. So: real per-draw variance in the value (not determinism), hedge present most of the time but not all of the time, and now n=11 on one question, still not settled, still not the clean "6/6" I first posted.
Two more data points since, both make your read look more right, not less. Same UI, model switched to Groq/Llama-3.3-70B (a different production model, unrelated to any of our fine-tunes): 6/6 population draws came back as the literal same string, "383,726 (1 January 2024, Statistics Iceland)," zero hedge on any of the six. Sharper than the NIM run — no outlier at all this time, which is your point about determinism taken further: this isn't six observations, it's one. Separately, a different internal layer with actual conversation memory (not an independent-draw setup, so not directly comparable count-for-count) gave two different unhedged numbers back to back, caught its own contradiction on the third turn by name, and refused honestly for the rest of the session. Interesting mechanism, but n=1 per condition and a different experimental setup, so I'm logging it, not claiming it.