Local Ai
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
arXiv:2607.17100v2 Announce Type: replace-cross Abstract: Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged
arXiv:2607.17100v2 Announce Type: replace-cross Abstract: Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judged by its terminal pipeline. A terminal score cannot reveal which technical decision produced a gain or distinguish a reusable discovery from a change adapted to development feedback. We introduce intervention-centered Auto Research, which validates research decisions rather than only final artifacts and makes their reliability measurable. Feature, Model, Representation, and Data axes are searched independently with inner five-fold feedback. Each axis winner is frozen before an outer-holdout matrix compares all alternatives on evidence the loop never sees. Across 701 agent-executed attempts spanning ten Matbench endpoints, outer evidence confirms the selected intervention on nine of ten endpoints and preserves 89.3% of non-tied intervention orderings. It also rejects an aggregate Representation gain that inner feedback endorsed. The resulting matrix reveals an information-dependent hierarchy. Composition-only tasks support several routes to improvement, whereas structure-informed tasks favor local geometry features and complementary tree ensembles. A subsequent compatibility test combines already frozen Feature and Model code without further search or tuning and raises mean outer-holdout improvement from 19.0% to 26.3%. By validating decisions rather than only artifacts, this design turns adaptive search into reusable evidence wherever agents propose executable alternatives against a fixed evaluator.
Related
- Can Open-Source LLM Agents Replace Static Application Security Testing Tools? An Empirical Assessment
- When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents
- Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Source: arXiv cs.AI | 2026-07-31