SAGE: A Service Agent Graph-guided Evaluation Benchmark
arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Ex