RevengeBench: Reverse Engineering Code-Space Policies from Behavioral Experiments Paper • 2606.26094 • Published about 1 month ago • 1
A Practitioner's Guide to Continual Multimodal Pretraining Paper • 2408.14471 • Published Aug 26, 2024 • 1
Fake It Till You Make It: Face analysis in the wild using synthetic data alone Paper • 2109.15102 • Published Sep 30, 2021
A high fidelity synthetic face framework for computer vision Paper • 2007.08364 • Published Jul 16, 2020
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published 29 days ago • 52
DataComp-VLM: Improved Open Datasets for Vision-Language Models Paper • 2606.28551 • Published 29 days ago • 52
QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents Paper • 2606.32034 • Published 25 days ago • 12
Great Models Think Alike and this Undermines AI Oversight Paper • 2502.04313 • Published Feb 6, 2025 • 32
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Paper • 2412.06745 • Published Dec 9, 2024 • 6