ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
DGX agentarXiv:2607.04686v1 Announce Type: cross Abstract: Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never calls a ne