REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
arXiv:2608.10669v1 Announce Type: new Abstract: Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interact