Gram: Assessing sabotage propensities via automated alignment auditing
DGX agentarXiv:2605.30322v1 Announce Type: cross Abstract: We introduce Gram, an automated alignment auditing framework to assess the propensity of AI agents to engage in sabotage. We evaluate Gemini models ac