BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
DGX agentarXiv:2605.29225v1 Announce Type: new Abstract: Self-evolving agents improve over time by reflecting on past failures, but existing evaluation is limited in two ways: it measures only task scores, lea