GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
DGX agentarXiv:2602.12316v2 Announce Type: replace Abstract: Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely ev