Weak-to-Strong Elicitation via Mismatched Wrong Drafts
DGX agentarXiv:2605.17314v1 Announce Type: cross Abstract: We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-tuning (e.g.