StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling Paper • 2608.15089 • Published 14 days ago • 445
Modular TTT: Rethinking Test-Time Training as Composable Modules Paper • 2608.07110 • Published 22 days ago • 9
Characterizing the Quality Profile of AI-Generated C++ in Production Paper • 2608.06640 • Published 23 days ago • 11
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows Paper • 2608.06714 • Published 22 days ago • 10