Spaces:
Running
Running
Claim 3: Task Utility and Reasoning Performance
We verified Claim 3 from Figure 2 of the paper:
- MATH/AIME & LiveCodeBench degradation: For the 3B variant, the maximum task accuracy drop is under 3%.
- GPQA improvement: AdvEvo-MARL-7B improves GPQA general-reasoning accuracy by up to +3.67% (as highlighted in the abstract and Section 5.2).
- No external guard agents: Unlike Inspector, which adds overhead and introduces a single point of failure, AdvEvo-MARL achieves these results by internalizing safety directly into the task agents.
This proves that adversarial co-evolution with a public baseline allows systems to maintain (or even improve) core reasoning capabilities while achieving robust safety.