RAS: Measuring LLM Safety Through Refusal Alignment
DGX agentarXiv:2606.25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their