Decoder-at-Layer Sweep for OpenVLA-OFT — Technical Report
Question studied: when OpenVLA-OFT decodes robot actions from its final LLM layer, is the final layer's advantage due to better features or to feature–decoder compatibility? The answer determines whether the model's depth can be cut at inference.
Summary of findings: attaching and training the real action head at intermediate layers recovers 94–100/100 closed-loop success from layer 16 upward (half the network); an 8.4M-parameter adapter in front of the frozen original head recovers most of it; and tuning as few as 4 transformer blocks recovers full baseline success from layer 12 — i.e. with 51% of the deployed parameters removed. The final layer holds no privileged action information; its advantage is a small, learnable compatibility transformation.
Setup: OpenVLA-OFT 7B with the L1 regression action head
(moojink/openvla-7b-oft-finetuned-libero-spatial), LIBERO-Spatial,
libero_spatial_no_noops data, seed 7. Offline evaluation: 1,491 action-chunk samples
(first 10 episodes per task). Readout training: 20 disjoint episodes per task.
Closed-loop evaluation: 10 trials × 10 tasks per configuration, fixed benchmark initial
states. Hardware: A100-SXM4-40GB, torch 2.2.0+cu121 (exact pins: env/pip_freeze.txt).
All analysis code: code_snapshot/.
1. Experimental-stack validation
| Check | Result |
|---|---|
| Model + headless-simulator smoke tests | pass |
| Offline final-layer L1 (1,491 samples) | 0.008120 (reference value 0.0081, tolerance ±5e-4) |
| Live bitwise invariant (instrumented layer-32 decode == stock output) | 0.0 on every sample of every run |
Stock predictions vs reference artifacts (baseline_reference/) |
bitwise identical (max dev 0.0) |
| Control: layer-16 linear probe retrained on the fresh extraction | L1 0.1733 vs reference 0.1724 (tolerance 0.01) |
The evaluation stack reproduces the reference predictions exactly, so differences
reported below are attributable to the interventions.
Replicate: python code_snapshot/gate2_check.py, python code_snapshot/gate3_probe_check.py;
recomputed independently by code_snapshot/verify_results.py (checks 1–2).
2. Main result: success vs attach layer
The real L1RegressionActionHead (151M params) is fresh-initialized per layer and
trained on cached action-token hidden states with the LLM frozen (lr 3e-4, 60 epochs,
cosine schedule, early stopping on a 10%-episode validation split; the recipe was tuned
at layer 32 only and then applied unchanged to all layers).
| Attach layer | 1 | 2 | 6 | 12 | 16 | 20 | 24 | 28 | 30 | 31 | 32 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Offline eval L1 | .205 | .189 | .168 | .153 | .129 | .095 | .048 | .024 | .019 | .017 | .013 |
| Rollout success /100 | 11 | 24 | 22 | 77 | 97 | 94 | 97 | 99 | 100 | 99 | 100 |
Reference points: full model with the original head 98/100; a 200k-parameter linear
probe at layer 16 achieves 77/100; the original head reused frozen at layer 31
achieves 70/100 (probe and reference records in baseline_reference/).
Findings:
- Reusing the pretrained head at intermediate layers fails for compatibility reasons, not information reasons. With a per-layer retrained head, offline L1 is strictly monotonic in depth (0.205 → 0.013); the error spike that the frozen pretrained head shows mid-network (peaking at layer 24) vanishes. Closed-loop, layer 24 rolls out at 97/100.
- No cliff at the final layer. Retrained heads at layers 16–32 all land at 94–100/100. Verdict: the final-layer advantage is a compatibility artifact, fully repaired by retraining the head at the attach layer.
- The information boundary for static readouts sits between layers 12 and 16 (77 → 97). Below it the curve collapses (22/100 at layer 6; the 11–24 floor reflects tasks solvable by trivial policies). Early exit at half depth costs ~1 point of success (97 vs the 98 full-model reference; the retrained head at layer 32 itself reaches 100).
- Readout capacity matters closed-loop: at layer 16, the real head reaches 97 vs the linear probe's 77 on identical training data.
- Offline action metrics do not predict closed-loop success — the offline L1 ordering across layers 16–24 does not match the rollout plateau, and §4 shows a starker case. All accept/reject conclusions here rest on rollouts only.
Representation analysis (linear CKA, action-token positions, all 33 layers):
adjacent-layer similarity is 0.95–0.99 through the mid-network, with three landmarks: a
large early reorganization at layers 2–3 (CKA 0.11–0.17), a dip at 27–28 (0.94), and
0.92 at 31→32 (the final entry includes the last RMSNorm). Similarity to the final
layer rises smoothly from 0.33 (layer 3) to 0.91 (layer 28+): representations converge
gradually — consistent with the smooth success curve and inconsistent with any single
"decision layer". Values: cka_results.json.
Replicate: every number above recomputes from the raw artifacts in this repository
via python code_snapshot/verify_results.py --results-dir . --ref-dir baseline_reference --skip-cka
(exit 0 = pass). Sweep table: sweep_results.json; per-episode records:
rollout_head_layer*.json.
3. Compatibility mechanism: adapters into the frozen original head
Adapter = token-wise residual MLP, LayerNorm→4096→1024→GELU→4096 (+x), 8.4M params (1/18 of the head's size), same training recipe. Variant A trains ONLY the adapter and feeds the frozen original checkpoint head; variant B trains adapter + fresh head.
| Rollout success /100 at layer | 12 | 16 | 30 |
|---|---|---|---|
| A: adapter + frozen original head | 77 | 92 | 98 |
| B: adapter + trained head | 81 | 96 | 99 |
| Retrained full head (151M, §2) | 77 | 97 | 100 |
| Offline eval L1 (A) | .149 | .119 | .014 |
Findings:
- The compatibility artifact is a small, learnable transformation. An 8.4M projector suffices for the untouched original head to read layer-30 features at 98/100 (= full-model reference) and layer-16 features at 92/100.
- At layer 12, readout capacity is irrelevant: adapter+frozen head (77), adapter+trained head (81), and the 151M retrained head (77) are statistically indistinguishable (SE ≈ 4). No static readout of frozen layer-12 features exceeds ~80/100 — §4 shows this ceiling belongs to the frozen features, not the information.
- Offline, the adapter variants match or beat the retrained head at every layer — another instance of offline metrics compressing real closed-loop differences.
Replicate: adapter_results.json + rollout_adapter*_layer*.json; recomputed by
the verify script. Training code: code_snapshot/train_adapters2.py.
4. Tuned attention blocks below layer 12
Attach at layer 12 and jointly train the head plus the last k decoder blocks below it (blocks 12−k .. 11; block lr 2e-5, head lr 3e-4, 12 epochs, early stopping). Training and deployment share an identical block-recompute path from a layer-8 full-sequence cache, gated functionally: the recompute with original weights reproduces the layer-12 sweep result exactly (offline L1 0.1531) before any training.
| k tuned blocks | 0 (frozen, §2) | 1 | 2 | 4 |
|---|---|---|---|---|
| Trainable params | 151M (head) | +202M | +404M | +808M |
| Offline eval L1 | .153 | .142 | .142 | .140 |
| Rollout success /100 | 77 | 93 | 95 | 98 |
Findings:
- The layer-12 ceiling is shallow. One tuned block lifts closed-loop success 77 → 93; four tuned blocks reach 98/100 = full-model reference while exiting at layer 12 (20 of 32 blocks, the final norm, and the LM head removed).
- Contrast with §3: at layer 12, more readout capacity never beat 81 — but letting attention over the full sequence adapt recovers everything. What limits layer 12 is cross-token integration, not information in the action-token features.
- The sharpest offline↔rollout dissociation observed: offline L1 improves 7% (0.153 → 0.142, flat in k) while rollout success moves 77 → 93 → 95 → 98. Offline metrics cannot rank these configurations at all.
Replicate: unfreeze_results.json + rollout_unfreeze_k*.json; checkpoints in
unfreeze/; training code code_snapshot/train_unfreeze.py, rollout code
code_snapshot/rollout_unfreeze.py. Note: the LLM uses bidirectional attention
(not causal) — code_snapshot/unfreeze_common.py documents the verified re-run
semantics and the bf16 kernel-noise bounds.
5. The assembled minimum-parameter model
min_model/model.safetensors packages the best size/accuracy point (§4, k=4): vision
encoders + projector + embeddings + LLM blocks 0–11 (8–11 tuned) + proprio projector +
action head.
| Params | Size | LIBERO-Spatial success | |
|---|---|---|---|
| Full OpenVLA-OFT stack | 7.709 B | ~15.5 GB | 98/100 |
| min_model (exit@12, 4 tuned blocks) | 3.772 B (48.9%) | 7.9 GB | 98/100 |
Parameter count, not latency, is the claim: the rollout above ran the full forward with
the tuned blocks applied on the side (rollout_unfreeze_k4.json); a truncated forward
implementation is required to realize wall-clock savings.
Replicate/verify: every tensor's SHA256 and source (public base checkpoint or
unfreeze/unfreeze_k4.pt) is recorded in min_model/manifest.json; check with
python code_snapshot/verify_min_model.py --bundle min_model --unfreeze-ckpt unfreeze/unfreeze_k4.pt
(add --check-base to byte-compare all frozen tensors against the base repo).
6. Included measurement data
Every rollout logged, per policy query, the executed normalized action chunk and the
stock (original head, layer 32) chunk (rollout_*.queries.npz). Mean |executed −
stock| grows with attach distance from the final layer (0.010 at layer 32 → 0.084 at
layer 16) and rises modestly within episodes (layer 16: 0.073 in the first third of an
episode → 0.087 in the last third) — bounded drift, consistent with the high success
rates. Complete run logs are in logs/.
7. Replication guide
Three tiers, detailed in REPRODUCING.md:
- Verify every reported number (CPU, ~2 min): the two verify commands above.
- Re-run rollouts with the released checkpoints (1 GPU, hours):
env/setup_env.shrebuilds the exact environment;code_snapshot/rollout_eval2.py/rollout_unfreeze.pyre-run any configuration with fixed seeds and benchmark initial states. - Re-derive everything from public sources (~1 GPU-day): base checkpoint and dataset are public; extraction, training, and rollout scripts include built-in gates (bitwise stock-reproduction, probe-reproduction control, sweep-reproduction check) that re-validate the chain at each stage.