Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 7 days ago • 36
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 12 days ago • 99
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 9 days ago • 70
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published 11 days ago • 37
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report Paper • 2607.18367 • Published 17 days ago • 59
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 21 days ago • 44
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Paper • 2607.13125 • Published 20 days ago • 138
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 28 days ago • 87
TurboServe: Serving Streaming Video Generation Efficiently and Economically Paper • 2606.19271 • Published Jun 17 • 38