# Claim 3: Task Utility and Reasoning Performance --- We verified Claim 3 from Figure 2 of the paper: - **MATH/AIME & LiveCodeBench degradation**: For the 3B variant, the maximum task accuracy drop is under **3%**. - **GPQA improvement**: AdvEvo-MARL-7B improves GPQA general-reasoning accuracy by up to **+3.67%** (as highlighted in the abstract and Section 5.2). - **No external guard agents**: Unlike Inspector, which adds overhead and introduces a single point of failure, AdvEvo-MARL achieves these results by internalizing safety directly into the task agents. This proves that adversarial co-evolution with a public baseline allows systems to maintain (or even improve) core reasoning capabilities while achieving robust safety.