Safety

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~

DGX agentpaper
safetyarxiv-cs-cl

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, and there is no unified theoretical framework for process-level advantage estimation. To bridge this gap, we propose extbf{S}elf-Guided extbf{P}rocess extbf{R}eward extbf{O}ptimization~(extbf{SPRO}), a novel framework that enables process-aware RL through two key innovations: (1) we show that process rewards can be derived intrinsically from the policy model itself, and (2) we redefine step-wise advantage by introducing well-defined Cumulative Process Rewards~(extbf{CPR}) and extbf{M}asked extbf{S}tep extbf{A}dvantage~(extbf{MSA}), which facilitates rigorous step-wise action advantage estimation within shared-prompt sampling groups. Our experimental results show that SPRO outperforms vanilla GRPO with 3.4x higher training efficiency and a 12.9% test accuracy improvement. Furthermore, SPRO maintains a stable and elevated policy entropy throughout training while achieving a considerable reduction in response length, evidencing sufficient exploration and prevention of reward hacking. Notably, SPRO incurs no additional computational overhead compared to outcome-supervised RL methods such as GRPO, which benefit industrial implementation.

Related

Source: arXiv cs.CL | 2026-07-27

Loading related sources…