Safety
Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~
arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, and there is no unified theoretical framework for process-level advantage estimation. To bridge this gap, we propose extbf{S}elf-Guided extbf{P}rocess extbf{R}eward extbf{O}ptimization~(extbf{SPRO}), a novel framework that enables process-aware RL through two key innovations: (1) we show that process rewards can be derived intrinsically from the policy model itself, and (2) we redefine step-wise advantage by introducing well-defined Cumulative Process Rewards~(extbf{CPR}) and extbf{M}asked extbf{S}tep extbf{A}dvantage~(extbf{MSA}), which facilitates rigorous step-wise action advantage estimation within shared-prompt sampling groups. Our experimental results show that SPRO outperforms vanilla GRPO with 3.4x higher training efficiency and a 12.9% test accuracy improvement. Furthermore, SPRO maintains a stable and elevated policy entropy throughout training while achieving a considerable reduction in response length, evidencing sufficient exploration and prevention of reward hacking. Notably, SPRO incurs no additional computational overhead compared to outcome-supervised RL methods such as GRPO, which benefit industrial implementation.
Related
- Breaking the Self-Confirming Loop: Diagnosing and Mitigating Systemic Reward Bias in Self-Rewarding RL
- Asymmetric Advantage Modulation Calibrates Entropy Dynamics in RLVR
- ProcessThinker: Enhancing Multi-modal Large Language Models Reasoning via Rollout-based Process Reward
- DVAO: Dynamic Variance-adaptive Advantage Optimization for Multi-reward Reinforcement Learning
Source: arXiv cs.CL | 2026-07-27