RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
DGX agentarXiv:2605.20204v1 Announce Type: cross Abstract: LLM-based user simulation is the primary mechanism for end-to-end agent evaluation, yet simulated users are poor proxies for real humans: unconstraine