Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
arXiv:2604.06132v2 Announce Type: replace Abstract: Large language models are increasingly deployed as autonomous agents for multi-step workflows in real-world software environments. However, existing