semsorock's picture
Update logbook: Repro: AdvEvo-MARL: Shaping Internalized Safety through Adversarial Co-Evolution in Multi-Agent Reinforcement Learning
8013b48 verified
|
Raw
History Blame Contribute Delete
757 Bytes

Claim 1: Chain Topology under NetSafe


We verified Claim 1 from the paper's main results (Table 1). On the chain topology under the NetSafe threat scenario:

  • Vanilla-7B baseline: Attack Success Rate (ASR) is 21.78% and Contagion Rate (CR) is 40.35%.
  • AdvEvo-MARL-7B: Attack Success Rate (ASR) is reduced to 0.99% (a relative reduction of over 95%) and Contagion Rate (CR) is reduced to 19.14%.

This verifies that the co-evolutionary adversarial reinforcement learning process successfully internalizes safety into defender agents sequentially connected in a chain topology.