Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
DGX agentarXiv:2605.05102v1 Announce Type: new Abstract: We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a di