DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence Paper • 2606.19348 • Published Apr 26 • 33
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI Paper • 2608.10915 • Published 3 days ago • 179
Wan-Animate-2: Pushing the Application Boundaries of Character Animation Paper • 2608.06009 • Published 6 days ago • 4
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 9 days ago • 58
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 10 days ago • 92
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 11 days ago • 155
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 19 days ago • 105
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model Paper • 2607.22083 • Published 18 days ago • 7
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis Paper • 2607.28618 • Published 15 days ago • 298
PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 15 days ago • 167
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 16 days ago • 71
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 15 days ago • 302
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 16 days ago • 139
Inflect v2 Collection Complete local text-to-waveform speech models at 3.96M and 9.36M parameters, with official PyTorch and ONNX Runtime releases. • 5 items • Updated 18 days ago • 10