State2State: Environment-Derived Mid-Training for LLM Agents
arXiv:2608.04934v1 Announce Type: new Abstract: Training LLM agents commonly relies on supervised fine-tuning from expert trajectories or online reinforcement learning over human-specified tasks with