Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
DGX agentarXiv:2608.13417v1 Announce Type: new Abstract: Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation. To understa