Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction
DGX agentarXiv:2607.12835v1 Announce Type: new Abstract: Rubric-based evaluation is a promising approach for assessing open-ended outputs from LLM-based research agents, particularly in paper reproduction, whe