Model Releases

Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry

arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmet

DGX agentpaper
model-releasesarxiv-cs-lg

arXiv:2608.12753v1 Announce Type: new Abstract: We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions with common rewards, (B) observed actions with independent rewards, and (C) unobserved actions with independent rewards. Players cannot communicate during learning but may agree on a protocol a priori. For Problems A and B we propose exttt{mQ-learning} and exttt{mQ-learning-intervals}, achieving ilde{O}(sqrt{H^4 S A_{ext{joint}}, T}) regret, where H is the horizon, S the state count, T = KH the total steps, and A_{ext{joint}} = prod_{i=1}^M |A_i| the joint action space across M players. For Problem C we give exttt{mEXC} and exttt{mEXC-Bellman}, two-phase explore-then-commit algorithms with regret ilde{O}(H (S A_{ext{joint}})^{1/3} T^{2/3}). Against the centralized joint-action benchmark, decentralized learning under information asymmetry matches the single-agent Q-learning rate of ite{jin2018q} up to logarithmic factors. Because A_{ext{joint}} grows exponentially in M, the bounds are most meaningful for small M or small per-player action sets.

Related

Source: arXiv cs.LG | 2026-08-14

Loading related sources…