Tutorials
Variance-Reduced Q-Learning over Static and Time-Varying Networks
arXiv:2607.21876v1 Announce Type: new Abstract: We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The a
arXiv:2607.21876v1 Announce Type: new Abstract: We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel epoch-based distributed Q-learning algorithm called VRDQ, where within each epoch, agents locally estimate the Bellman optimality operator and diffuse information using a consensus-based protocol. For both static and time-varying networks, we establish high-probability finite-time convergence rates for VRDQ that enjoy linear speedups from collaboration. Crucially, we prove that such speedups in sample-complexity require only ilde{O}(1) communication, substantially improving upon the communication costs in prior work.
Related
- Accelerating Optimization and Machine Learning through Decentralization
- Central Limit Theorem for Two-Time-Scale Approximate Distributionally Robust RL
- Provable Benefits of RLVR over SFT for Reasoning Models: Learning to Backtrack Efficiently
Source: arXiv cs.LG | 2026-07-27