PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
arXiv:2608.08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face three struct