Information-Directed Offline-to-Online Reinforcement Learning
DGX agentarXiv:2605.29405v1 Announce Type: new Abstract: Decision-making from offline datasets typically warm-starts a policy or score model from fixed offline data and then refines it with limited online inte