OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 13 days ago • 216
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 3 days ago • 8
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published 8 days ago • 43
Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills Paper • 2608.01851 • Published 11 days ago • 11
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published 8 days ago • 39
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 14 days ago • 96
WorldExam: Benchmarking World Models from Apparent Appearance to Inherent Reactivity Paper • 2608.02603 • Published 11 days ago • 34
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications Paper • 2607.28617 • Published 15 days ago • 37