Safety
Rethinking Reinforcement Fine-Tuning in LVLM: Convergence, Reward Decomposition, and Generalization
arXiv:2604.19857v1 Announce Type: cross Abstract: Reinforcement fine-tuning with verifiable rewards (RLVR) has emerged as a powerful paradigm for equipping large vision-language models (LVLMs) with ag
arXiv:2604.19857v1 Announce Type: cross Abstract: Reinforcement fine-tuning with verifiable rewards (RLVR) has emerged as a powerful paradigm for equipping large vision-language models (LVLMs) with agentic capabilities such as tool use and multi-step reasoning. Despite striking empirical successes, most notably Visual Agentic Reinforcement Fine-Tuning (Visual-ARFT), the theoretical underpinnings of this paradigm remain poorly understood. In particular, two critical questions lack rigorous answers: (i)how does the composite structure of verifiable rewards (format compliance, answer accuracy, tool executability) affect the convergence of Group Relative Policy Optimization (GRPO), and (ii)2}). Third, we establish a PAC-Bayes generalization bound for tool-augmented policies that explains the strong out-of-distribution transfer observed in Visual-ARFT (extbf{Theorem~3}).why does training on a small set of tool-augmented tasks transfer to out-of-distribution domains? We address these gaps by introducing the Tool-Augmented Markov Decision Process (TA-MDP), a formal framework that models multimodal agentic decision-making with bounded-depth tool calls. Within this framework, we establish three main results. First, we prove that GRPO under composite verifiable rewards converges to a first-order stationary point at rate O(1/sqrt{T}) with explicit dependence on the number of reward components and group size (extbf{Theorem1}). Second, we derive a Reward Decomposition Theorem that bounds the sub-optimality gap between decomposed per-component optimization and joint optimization, providing a precise characterization of when reward decomposition is beneficial (extbf{Theorem
Related
- RL-PLUS: Countering Capability Boundary Collapse of LLMs in Reinforcement Learning with Hybrid-policy Optimization
- Perception-Aware Policy Optimization for Multimodal Reasoning
- Rethinking Token-Level Credit Assignment in RLVR: A Polarity-Entropy Analysis
- Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
- From P(y|x) to P(y): Investigating Reinforcement Learning in Pre-train Space
- Visually-Guided Policy Optimization for Multimodal Reasoning
Source: arXiv cs.CL | 2026-04-23