Stage 3 — UR5 teleop VLA (Qwen2.5-VL-3B), current-view variant

Training run complete: 60,000 steps. This is the ablation partner of Ngseo/stage3-ur5-actheavy: same objective, same data, same hyper-parameters — the frozen V-JEPA2 targets are computed on the current frame instead of on future frames.

act-heavy     :  V-JEPA sees frames t+4 … t+32   (1.07 s ahead)
current-view  :  V-JEPA sees frame t, repeated 8x  ← this model

Both branches move together: the task branch that z_a is aligned to and the domain branch that z_b is decorrelated from. Nothing in this model's loss looks ahead in time.

The question it answers: does the V-JEPA target have to predict the future, or is shaping the representation against the current frame enough?

The setup this was trained for

Each of the 7 tasks in Ngseo/ur5_teleop_multitask is only ever shown from one of the 4 cameras, so viewpoint alone almost determines the task. Object colour is deliberately crossed between the two task families so that "colour ⇒ camera" is not a valid shortcut on its own:

camera tasks
camera_0 Point at the red cup · Pick up the blue die → basket
camera_1 Pull a tissue out of the box · Close the laptop · Stand the shoe upright
camera_2 Point at the blue cup · Pick up the red die → basket

Architecture

Qwen2.5-VL-3B + LoRA ─ forward_full_hidden ─┬─ Head A → z_a ─┐
                                            └─ Head B → z_b ─┴─ concat(8192) → ResNetActionHead → 30×7

L = 1.0·L1(action) + 0.02·InfoNCE(z_a, z_task_CURRENT) + 0.002·SIGReg([z_domain_CURRENT ; z_b])

Frozen target encoder: Ngseo/stage1 disentangled V-JEPA2 ViT-L. Both heads are AttentiveLatentHead (proj 4096, 8 queries, depth 2, 167.8M each).

LoRA r=32, α=64 — LLM q_proj/v_proj (7.4M) and vision tower qkv (5.2M)
Trainable 386M of 3.77B
Inputs 1 RGB frame @224 + task string + 7-D joint state
Output 30-step action chunk (absolute joint positions, 1.0 s @ 30 fps)
Optimiser AdamW, lr 1e-4, 1000-step warmup + cosine to 0, batch 32, bf16
Augmentation ColorJitter/SharpnessJitter ×2 + DomainRandomization p=0.7

Final metrics vs. the future-frame variant

current-view (this) act-heavy (future)
action L1 @ 10k 0.0826 0.0836
@ 30k 0.0438 0.0442
@ 50k 0.0326 0.0305
@ 60k (final) 0.0319 0.0300
InfoNCE (chance 3.466) 1.86 1.79
cos(z_a, z_target) 0.233 0.292
cos(z_b, z_domain) −0.002 −0.000

On in-distribution action accuracy the two are within ~6% of each other, i.e. looking ahead buys almost nothing here. That is expected: the InfoNCE term carries weight 0.02, so it barely competes with the action loss.

Reference points on the same normalised scale, none of which use vision or language: dataset mean 0.834, copying the input state across all 30 steps 0.155. The per-step copy error grows from 0.019 at k=0 to 0.291 at k=29.

Not evaluated on a robot. Everything above is a training-set loss. The question this study is actually about — what happens when a task is requested from a camera it was never trained on — is not answered by these numbers, and is exactly where the two variants might diverge.

Caveat specific to this variant: the current-view clip is built from the same augmented context frame the VLM is shown, so its InfoNCE aligns two encodings of identical pixels. The future-frame variant aligns against a different, independently-augmented clip. The two therefore differ in more than just "future vs current", and the comparison should be read with that in mind.

The 7-D joint state is also an input, and a probe on that state alone recovers which of the 7 tasks is running with 94% accuracy (chance 14%).

Contents

epoch_6.pt holds model_state_dict (full VLM incl. LoRA), latent_head_state_dict (Head A), free_latent_head_state_dict (Head B), stage2_action_head_state_dict, optimiser state, and the run config. config.yaml is the exact training config.

Downloads last month
-
Video Preview
loading

Model tree for Ngseo/stage3-ur5-currentview

Finetuned
(840)
this model

Dataset used to train Ngseo/stage3-ur5-currentview