SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published 7 days ago • 124
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published 4 days ago • 25
RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks? Paper • 2609.05324 • Published 10 days ago • 28
TANGO: Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model Paper • 2609.09158 • Published 6 days ago • 23
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow Paper • 2609.03563 • Published 11 days ago • 19
LatentPress: Context Compression Beyond Text and Vision Paper • 2609.01507 • Published 13 days ago • 121
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 27 days ago • 159
DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-NEO-MAX-MTP-GGUF Image-Text-to-Text • 27B • Updated 21 days ago • 958k • 2.34k
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark Paper • 2608.05850 • Published Aug 6 • 23
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published Aug 6 • 102
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published Aug 4 • 26