PIRL: From Open-Loop Exploration to Closed-Loop Reinforcement Learning [R]
TL;DR: Most RL post-training algorithms optimize the current batch and move on. But after an update, did the new policy actually become better? We introduce Policy Improvement Reinforcement Learning (