Pass the Baton: Trajectory-Relayed On-Policy Distillation Paper • 2607.26057 • Published 5 days ago • 30
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published 5 days ago • 148
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes Paper • 2607.24516 • Published 6 days ago • 7
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 11 days ago • 72
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Paper • 2607.19139 • Published 12 days ago • 74
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 17 days ago • 205
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 12 days ago • 306
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 14 days ago • 92
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning Paper • 2607.10966 • Published 20 days ago • 5
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation Paper • 2607.15686 • Published 16 days ago • 16
KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation Paper • 2607.14202 • Published 18 days ago • 42
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 20 days ago • 84
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published 19 days ago • 108
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published 22 days ago • 71
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published 19 days ago • 102
Dual Latent Memory in Vision-Language-Action Models for Robotic Manipulation Paper • 2607.07608 • Published 25 days ago • 57
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 25 days ago • 64
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning Paper • 2607.02963 • Published about 1 month ago • 29