Safety
Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation
arXiv:2608.02087v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM
arXiv:2608.02087v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM action-space structure introduces challenges distinct from classical RL, with implications for inducing exploration. New methods are required that leverage the broad knowledge and flexibility of pre-trained LLMs to deliberately generate diverse experience at training time. We propose Instruction-Conditioned Exploration (ICE), which supplements task prompts during training with one of several distinct instructions, increasing the coverage of behaviours attempted. To facilitate ICE, we propose Asymmetric-RL/SD, a combined Reinforcement Learning and Self-Distillation training objective, to transfer explored behaviours to the unconditioned test-time policy. ICE with the Asymmetric-RL/SD objective improves Qwen3-1.7B held-out pass@1 performance at 4K response length on mathematical reasoning tasks by 5.0% relative to training with DAPO, with improvement persisting at a longer 8K context.
Related
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance
- Learning from the Self-future: On-policy Self-distillation for dLLMs
- When Importance Sampling Misallocates Credit: Asymmetric Ratios for Outcome-Supervised RL
- Trust Region On-Policy Distillation
Source: arXiv cs.CL | 2026-08-04