Research

R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning

arXiv:2602.05940v2 Announce Type: replace Abstract: Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reas

DGX agentpaper
researcharxiv-cs-cl

arXiv:2602.05940v2 Announce Type: replace Abstract: Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reasoning in the question language. Even with the same reasoning language, semantically equivalent English and non-English questions still exhibit a clear performance gap. Together, these phenomena reveal two distinct bottlenecks: target-language question understanding and target-language reasoning. Existing methods typically optimize only one of these capabilities. However, simply combining them may not be sufficient to optimize both effectively, as answer correctness alone cannot distinguish failures in question understanding from those in reasoning. We propose R3S, a reinforcement learning framework that disentangles the optimization of the two capabilities. R3S refines translation rewards derived from downstream reasoning accuracy through English-solvability filtering and recovers target-language RLVR signals using self-generated English hints. Together, these designs require neither external model feedback nor external multilingual training data. Experiments across three backbone models and five languages show that R3S improves language-consistent accuracy over the target-language RLVR baseline on MMATH by an average of 10.3 percentage points, while maintaining near-perfect language consistency. Consistent gains on MMLU-ProX further demonstrate its generalization beyond math problems.

Source: arXiv cs.CL | 2026-08-11

Loading related sources…