Oracle Lens (AO) β RL final checkpoint iolens.final.ddp600.s0 (iter_000600)
A LoRA adapter on Qwen/Qwen3.6-27B that turns the model into an oracle lens: feed it one
of its own residual-stream activation vectors and it verbalizes the content of that activation
as a bullet list of concepts. This is the on-policy GRPO checkpoint (600 steps, DDP-8) warm
started from the iolens.final SFT bullet student β the first positive RL result on the AO
line (held-out joint FVE 0.120 β 0.155 against a frozen activation-reconstructor reward).
Load contract (read this before using β a wrong contract measures nothing)
| knob | value |
|---|---|
| base model | Qwen/Qwen3.6-27B (text decoder) |
| injection site | the embedding slot of marker token γ (id 158983), flanked by <activation>/</activation> (neighbor ids 29, 510) |
| injection | replace the marker's embedding with 16000 Β· h / βhβ (unit transform, Ξ±=16000; scale-free) |
| supported layers | residual stream layers 20β60, step 4 (11 layers) |
| prompt | the checkpoint's own concepts_raw template (in nla_meta.yaml): "An activation vector from layer 20 of a language model is enclosed in activation tags: <activation>γ</activation>. Produce distinct concepts that encode this activation, each as a '- ' bullet on its own line." rendered with the chat template, add_generation_prompt=True |
| parse | one concept per - bullet line |
| eval sampling used in our benchmarks | k=1, T=1.0, top-p 0.95, top-k 64, max-new 256 |
Usage sketch
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.6-27B", torch_dtype=torch.bfloat16, device_map="cuda")
tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.6-27B")
model = PeftModel.from_pretrained(base, "<this repo>")
TEMPLATE = ("An activation vector from layer 20 of a language model is enclosed in activation "
"tags: <activation>γ</activation>. Produce distinct concepts that encode this "
"activation, each as a '- ' bullet on its own line.")
ids = tok.apply_chat_template([{"role": "user", "content": TEMPLATE}],
tokenize=True, add_generation_prompt=True)
mp = ids.index(158983) # the marker slot
h = ... # your residual vector, layer in {20,24,...,60}
v = 16000.0 * h / h.norm()
embed = model.get_input_embeddings()
inputs_embeds = embed(torch.tensor([ids], device="cuda"))
inputs_embeds[0, mp] = v.to(inputs_embeds.dtype).to("cuda") # REPLACE the slot embedding
out = model.generate(inputs_embeds=inputs_embeds, do_sample=True, temperature=1.0,
top_p=0.95, top_k=64, max_new_tokens=256)
print(tok.decode(out[0], skip_special_tokens=True))
Benchmarks (workspace-bench, strict-judged, 2026-09)
Judged baseline macro 0.785 over the 12 bank families (association 0.88, multihop 0.89, multihop-mt 0.87, multilingual 1.00, typo 1.00); hillclimbing: conjunctive 0.60, role-bound association 0.85, moral-committed 0.81, moral-deliberative 0.62, user-modeling 0.44 (k=1). Notable qualitative reads: never-written arithmetic intermediates, both poles of moral dilemmas from the empty assistant prefix, doctrine retrieval (Tarasoff, Double Effect) at layer 20, and latent user-attribute inference (see the workspace-bench records).
Provenance
- Training: on-policy GRPO (self-contained trainer), 600 steps, Ξ²_KL 0.02, lr 5e-5,
group size 64, reward = per-bullet joint FVE against a frozen activation reconstructor
(whitened, unit concept dirs, kβ€4); warm start
ckpts/ao/distill/final.s0/step105. - Files:
adapter_config.json,adapter_model.safetensors(PEFT LoRA),meta.json(training summary),nla_meta.yaml(the prompt/injection sidecar β the source of truth for the load contract),run_config.yaml(full trainer config). - Canonical archive:
agu18dec/local-workspaceβckpts/ao/rl/iolens.final.ddp600.s0/.
Model tree for agu18dec/oracle-lens-ddp600
Base model
Qwen/Qwen3.6-27B