Model Releases
Unveiling High-Probability Generalization in Decentralized SGD
arXiv:2605.10205v1 Announce Type: new Abstract: Decentralized stochastic gradient descent (D-SGD) is an efficient method for large-scale distributed learning. Existing generalization studies mainly ad
arXiv:2605.10205v1 Announce Type: new Abstract: Decentralized stochastic gradient descent (D-SGD) is an efficient method for large-scale distributed learning. Existing generalization studies mainly address expected results, achieving rates limited to Oleft(frac{1}{elta sqrt{mn}}right), where elta is the confidence parameter, m the number of workers, and n the sample size. When m=1, D-SGD reduces to traditional SGD, whose optimal high-probability generalization bound is Oleft(frac{1}{sqrt{n}}log (1/elta)right). This discrepancy reveals a gap between high-probability guarantees for SGD and those for D-SGD. To close this, we develop a high-probability learning theory for D-SGD, aiming for the optimal Oleft(frac{1}{sqrt{mn}}log (1/elta)right) rate. We refine bounds for D-SGD using pointwise uniform stability in distributed learning-a weaker notion than uniform stability-and analyze them across convex, strongly convex, and non-convex settings. We also provide high-probability results for gradient-based measures in non-convex cases where only local minima exist, and derive optimization error and excess risk bounds. Finally, accounting for communication overhead, we analyze generalization bounds for local models within time-varying frameworks.
Source: arXiv cs.LG | 2026-05-12