Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
arXiv:2608.04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on