ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
DGX agentarXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend