Bilevel Optimization of Synthetic Trajectories for Multi-Turn LLM Fine-Tuning
DGX agentarXiv:2605.24743v1 Announce Type: cross Abstract: While LLMs excel at single-turn generation, they struggle with long-horizon, multi-turn interactions. Offline reinforcement learning (RL) offers a sca