LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
DGX agentarXiv:2607.14952v3 Announce Type: replace Abstract: Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prom