Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation
arXiv:2608.02087v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM