| # Decoder-at-Layer Sweep for OpenVLA-OFT — Technical Report |
|
|
| **Question studied:** when OpenVLA-OFT decodes robot actions from its final LLM layer, |
| is the final layer's advantage due to *better features* or to *feature–decoder |
| compatibility*? The answer determines whether the model's depth can be cut at inference. |
|
|
| **Summary of findings:** attaching and training the real action head at intermediate |
| layers recovers 94–100/100 closed-loop success from layer 16 upward (half the network); |
| an 8.4M-parameter adapter in front of the *frozen original* head recovers most of it; |
| and tuning as few as 4 transformer blocks recovers full baseline success from layer 12 |
| — i.e. with 51% of the deployed parameters removed. The final layer holds no privileged |
| action information; its advantage is a small, learnable compatibility transformation. |
|
|
| **Setup:** OpenVLA-OFT 7B with the L1 regression action head |
| (`moojink/openvla-7b-oft-finetuned-libero-spatial`), LIBERO-Spatial, |
| `libero_spatial_no_noops` data, seed 7. Offline evaluation: 1,491 action-chunk samples |
| (first 10 episodes per task). Readout training: 20 disjoint episodes per task. |
| Closed-loop evaluation: 10 trials × 10 tasks per configuration, fixed benchmark initial |
| states. Hardware: A100-SXM4-40GB, torch 2.2.0+cu121 (exact pins: `env/pip_freeze.txt`). |
| All analysis code: `code_snapshot/`. |
|
|
| ## 1. Experimental-stack validation |
|
|
| | Check | Result | |
| |---|---| |
| | Model + headless-simulator smoke tests | pass | |
| | Offline final-layer L1 (1,491 samples) | **0.008120** (reference value 0.0081, tolerance ±5e-4) | |
| | Live bitwise invariant (instrumented layer-32 decode == stock output) | 0.0 on every sample of every run | |
| | Stock predictions vs reference artifacts (`baseline_reference/`) | **bitwise identical (max dev 0.0)** | |
| | Control: layer-16 linear probe retrained on the fresh extraction | L1 0.1733 vs reference 0.1724 (tolerance 0.01) | |
|
|
| The evaluation stack reproduces the reference predictions exactly, so differences |
| reported below are attributable to the interventions. |
| **Replicate:** `python code_snapshot/gate2_check.py`, `python code_snapshot/gate3_probe_check.py`; |
| recomputed independently by `code_snapshot/verify_results.py` (checks 1–2). |
|
|
| ## 2. Main result: success vs attach layer |
|
|
| The real `L1RegressionActionHead` (151M params) is fresh-initialized per layer and |
| trained on cached action-token hidden states with the LLM frozen (lr 3e-4, 60 epochs, |
| cosine schedule, early stopping on a 10%-episode validation split; the recipe was tuned |
| at layer 32 only and then applied unchanged to all layers). |
|
|
| | Attach layer | 1 | 2 | 6 | 12 | 16 | 20 | 24 | 28 | 30 | 31 | 32 | |
| |---|---|---|---|---|---|---|---|---|---|---|---| |
| | Offline eval L1 | .205 | .189 | .168 | .153 | .129 | .095 | .048 | .024 | .019 | .017 | .013 | |
| | **Rollout success /100** | **11** | **24** | **22** | **77** | **97** | **94** | **97** | **99** | **100** | **99** | **100** | |
|
|
| Reference points: full model with the original head 98/100; a 200k-parameter linear |
| probe at layer 16 achieves 77/100; the original head reused *frozen* at layer 31 |
| achieves 70/100 (probe and reference records in `baseline_reference/`). |
|
|
| **Findings:** |
|
|
| 1. **Reusing the pretrained head at intermediate layers fails for compatibility |
| reasons, not information reasons.** With a per-layer retrained head, offline L1 is |
| strictly monotonic in depth (0.205 → 0.013); the error spike that the frozen |
| pretrained head shows mid-network (peaking at layer 24) vanishes. Closed-loop, |
| layer 24 rolls out at 97/100. |
| 2. **No cliff at the final layer.** Retrained heads at layers 16–32 all land at |
| 94–100/100. **Verdict: the final-layer advantage is a compatibility artifact, fully |
| repaired by retraining the head at the attach layer.** |
| 3. **The information boundary for static readouts sits between layers 12 and 16** |
| (77 → 97). Below it the curve collapses (22/100 at layer 6; the 11–24 floor reflects |
| tasks solvable by trivial policies). Early exit at half depth costs ~1 point of |
| success (97 vs the 98 full-model reference; the retrained head at layer 32 itself |
| reaches 100). |
| 4. **Readout capacity matters closed-loop:** at layer 16, the real head reaches 97 |
| vs the linear probe's 77 on identical training data. |
| 5. **Offline action metrics do not predict closed-loop success** — the offline L1 |
| ordering across layers 16–24 does not match the rollout plateau, and §4 shows a |
| starker case. All accept/reject conclusions here rest on rollouts only. |
|
|
| **Representation analysis (linear CKA, action-token positions, all 33 layers):** |
| adjacent-layer similarity is 0.95–0.99 through the mid-network, with three landmarks: a |
| large early reorganization at layers 2–3 (CKA 0.11–0.17), a dip at 27–28 (0.94), and |
| 0.92 at 31→32 (the final entry includes the last RMSNorm). Similarity to the final |
| layer rises smoothly from 0.33 (layer 3) to 0.91 (layer 28+): representations converge |
| gradually — consistent with the smooth success curve and inconsistent with any single |
| "decision layer". Values: `cka_results.json`. |
|
|
| **Replicate:** every number above recomputes from the raw artifacts in this repository |
| via `python code_snapshot/verify_results.py --results-dir . --ref-dir baseline_reference --skip-cka` |
| (exit 0 = pass). Sweep table: `sweep_results.json`; per-episode records: |
| `rollout_head_layer*.json`. |
|
|
| ## 3. Compatibility mechanism: adapters into the frozen original head |
|
|
| Adapter = token-wise residual MLP, LayerNorm→4096→1024→GELU→4096 (+x), **8.4M params** |
| (1/18 of the head's size), same training recipe. Variant A trains ONLY the adapter and |
| feeds the **frozen original checkpoint head**; variant B trains adapter + fresh head. |
|
|
| | Rollout success /100 at layer | 12 | 16 | 30 | |
| |---|---|---|---| |
| | A: adapter + frozen original head | 77 | 92 | 98 | |
| | B: adapter + trained head | 81 | 96 | 99 | |
| | Retrained full head (151M, §2) | 77 | 97 | 100 | |
| | Offline eval L1 (A) | .149 | .119 | .014 | |
|
|
| **Findings:** |
|
|
| 1. **The compatibility artifact is a small, learnable transformation.** An 8.4M |
| projector suffices for the *untouched* original head to read layer-30 features at |
| 98/100 (= full-model reference) and layer-16 features at 92/100. |
| 2. **At layer 12, readout capacity is irrelevant:** adapter+frozen head (77), |
| adapter+trained head (81), and the 151M retrained head (77) are statistically |
| indistinguishable (SE ≈ 4). No *static readout* of frozen layer-12 features exceeds |
| ~80/100 — §4 shows this ceiling belongs to the frozen features, not the information. |
| 3. Offline, the adapter variants match or beat the retrained head at every layer — |
| another instance of offline metrics compressing real closed-loop differences. |
|
|
| **Replicate:** `adapter_results.json` + `rollout_adapter*_layer*.json`; recomputed by |
| the verify script. Training code: `code_snapshot/train_adapters2.py`. |
|
|
| ## 4. Tuned attention blocks below layer 12 |
|
|
| Attach at layer 12 and jointly train the head plus the last k decoder blocks below it |
| (blocks 12−k .. 11; block lr 2e-5, head lr 3e-4, 12 epochs, early stopping). Training |
| and deployment share an identical block-recompute path from a layer-8 full-sequence |
| cache, gated functionally: the recompute with original weights reproduces the layer-12 |
| sweep result exactly (offline L1 0.1531) before any training. |
|
|
| | k tuned blocks | 0 (frozen, §2) | 1 | 2 | 4 | |
| |---|---|---|---|---| |
| | Trainable params | 151M (head) | +202M | +404M | +808M | |
| | Offline eval L1 | .153 | .142 | .142 | .140 | |
| | **Rollout success /100** | **77** | **93** | **95** | **98** | |
|
|
| **Findings:** |
|
|
| 1. **The layer-12 ceiling is shallow.** One tuned block lifts closed-loop success |
| 77 → 93; four tuned blocks reach **98/100 = full-model reference** while exiting at |
| layer 12 (20 of 32 blocks, the final norm, and the LM head removed). |
| 2. Contrast with §3: at layer 12, more *readout* capacity never beat 81 — but letting |
| *attention over the full sequence* adapt recovers everything. What limits layer 12 |
| is cross-token integration, not information in the action-token features. |
| 3. **The sharpest offline↔rollout dissociation observed:** offline L1 improves 7% |
| (0.153 → 0.142, flat in k) while rollout success moves 77 → 93 → 95 → 98. Offline |
| metrics cannot rank these configurations at all. |
|
|
| **Replicate:** `unfreeze_results.json` + `rollout_unfreeze_k*.json`; checkpoints in |
| `unfreeze/`; training code `code_snapshot/train_unfreeze.py`, rollout code |
| `code_snapshot/rollout_unfreeze.py`. Note: the LLM uses **bidirectional** attention |
| (not causal) — `code_snapshot/unfreeze_common.py` documents the verified re-run |
| semantics and the bf16 kernel-noise bounds. |
|
|
| ## 5. The assembled minimum-parameter model |
|
|
| `min_model/model.safetensors` packages the best size/accuracy point (§4, k=4): vision |
| encoders + projector + embeddings + LLM blocks 0–11 (8–11 tuned) + proprio projector + |
| action head. |
|
|
| | | Params | Size | LIBERO-Spatial success | |
| |---|---|---|---| |
| | Full OpenVLA-OFT stack | 7.709 B | ~15.5 GB | 98/100 | |
| | **min_model (exit@12, 4 tuned blocks)** | **3.772 B (48.9%)** | **7.9 GB** | **98/100** | |
| |
| Parameter count, not latency, is the claim: the rollout above ran the full forward with |
| the tuned blocks applied on the side (`rollout_unfreeze_k4.json`); a truncated forward |
| implementation is required to realize wall-clock savings. |
| |
| **Replicate/verify:** every tensor's SHA256 and source (public base checkpoint or |
| `unfreeze/unfreeze_k4.pt`) is recorded in `min_model/manifest.json`; check with |
| `python code_snapshot/verify_min_model.py --bundle min_model --unfreeze-ckpt unfreeze/unfreeze_k4.pt` |
| (add `--check-base` to byte-compare all frozen tensors against the base repo). |
| |
| ## 6. Included measurement data |
| |
| Every rollout logged, per policy query, the executed normalized action chunk and the |
| stock (original head, layer 32) chunk (`rollout_*.queries.npz`). Mean |executed − |
| stock| grows with attach distance from the final layer (0.010 at layer 32 → 0.084 at |
| layer 16) and rises modestly within episodes (layer 16: 0.073 in the first third of an |
| episode → 0.087 in the last third) — bounded drift, consistent with the high success |
| rates. Complete run logs are in `logs/`. |
| |
| ## 7. Replication guide |
| |
| Three tiers, detailed in `REPRODUCING.md`: |
| |
| 1. **Verify every reported number** (CPU, ~2 min): the two verify commands above. |
| 2. **Re-run rollouts with the released checkpoints** (1 GPU, hours): |
| `env/setup_env.sh` rebuilds the exact environment; `code_snapshot/rollout_eval2.py` |
| / `rollout_unfreeze.py` re-run any configuration with fixed seeds and benchmark |
| initial states. |
| 3. **Re-derive everything from public sources** (~1 GPU-day): base checkpoint and |
| dataset are public; extraction, training, and rollout scripts include built-in |
| gates (bitwise stock-reproduction, probe-reproduction control, sweep-reproduction |
| check) that re-validate the chain at each stage. |
| |