STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
DGX agentarXiv:2602.15620v4 Announce Type: replace Abstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic