semsorock's picture
Update logbook: Repro: AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
8013b48 verified
|
Raw
History Blame Contribute Delete
1.16 kB

Claim 2: Multi-scenario and Multi-topology Safety


We verified Claim 2 from Table 1 of the paper:

  • AdvEvo-MARL ASR bound: Across all evaluated system topologies (chain, tree, complete) and attack scenarios (NetSafe, AutoInject, UserHijack), AdvEvo-MARL consistently keeps ASR at or below approximately 20%. The peak ASR is 17.68% (AdvEvo-MARL-3B under the complete topology, UserHijack attack on GPQA).
  • Baselines ASR: Open-source baselines reach much higher ASRs:
    • Vanilla-3B Complete topology under UserHijack LiveCodeBench achieves 36% ASR and 65.43% CR.
    • Challenger-3b Complete topology under UserHijack GPQA achieves 65.53% ASR.
    • Challenger-7b Complete topology under UserHijack AIME/GPQA achieves 38.33% ASR (a 10% increase over the Vanilla-7B's 28.33% ASR).

This confirms that AdvEvo-MARL provides robust safety across topologies and threat vectors without experiencing the safety degradation or high contagion observed in baseline defense approaches.