Beyond Benchmark Islands: Toward Representative Trustworthiness Evaluation for Agentic AI
DGX agentarXiv:2603.14987v2 Announce Type: replace Abstract: Agentic AI systems increasingly act through tool-augmented, multi-step workflows whose failures (unsafe tool use, unauthorised actions, social harm)