Research
AutoVerifier: Residual-Guided Non-Parametric Optimization for Reference-Based Answer Verification
arXiv:2608.25637v1 Announce Type: new Abstract: Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable
arXiv:2608.25637v1 Announce Type: new Abstract: Reference-based verifiers are important for evaluating reasoning models and providing accurate outcome rewards in reinforcement learning with verifiable rewards. To improve verification accuracy, prior work has explored rule-based, model-based, and tool-augmented verifiers for checking answer equivalence across diverse answer forms. However, the equivalence of answer forms such as 1+3.14 and 1+pi may depend on the question and scoring criterion. We frame such implicit assumptions as verifier inductive biases. To address this challenge, we propose AutoVerifier, a residual-guided non-parametric optimization method that learns these biases from recurring verifier errors. Specifically, AutoVerifier records these biases in rule cards and promotes them to code modules or prompt guidance only after replay validation detects no direct regressions, keeping accepted updates auditable, editable, and reusable. Experiments on four verifier benchmarks demonstrate that AutoVerifier outperforms state-of-the-art verifiers by a large margin.
Related
- From Reasoning Chains to Verifiable Subproblems: Curriculum Reinforcement Learning Enables Credit Assignment for LLM Reasoning
- LEEPS: Latent-Guided Explore-Exploit Prompt Sampling for Efficient RLVR in Large Language Models
- Outcome Rewards Do Not Guarantee Verifiable or Causally Important Reasoning
- Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
- Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards
Source: arXiv cs.CL | 2026-08-27