R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning
arXiv:2602.05940v2 Announce Type: replace Abstract: Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reas