Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
DGX agentarXiv:2608.06329v1 Announce Type: cross Abstract: Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor