TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
DGX agentarXiv:2602.06911v2 Announce Type: replace-cross Abstract: As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications,