Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL w…
Most people training agentic LLMs with RL right now have a silently broken training loop and have no idea. Here's the trap: single-turn RL works beautifully. Clean curves, sane rewards, everything con