AgenticEval: Toward Agentic and Self-Evolving Safety Evaluation of Large Language Models
DGX agentarXiv:2509.26100v2 Announce Type: replace Abstract: The rapid integration of Large Language Models (LLMs) into high-stakes domains necessitates reliable safety and compliance evaluation. However, exis