Research

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

arXiv:2601.02022v2 Announce Type: replace Abstract: We prove that Thompson sampling exhibits ilde{O}(sigma d sqrt{T} + d r sqrt{Tr(Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(mu_

DGX agentpaper
researcharxiv-cs-lg

arXiv:2601.02022v2 Announce Type: replace Abstract: We prove that Thompson sampling exhibits ilde{O}(sigma d sqrt{T} + d r sqrt{Tr(Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(mu_0, Sigma_0) prior distribution on the coefficients, where d is the dimension, T is the time horizon, r is the maximum ell_2 norm of the actions, and sigma^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent burn-in'' term d r sqrt{Tr(Sigma_0)} decouples additively from the minimax (long run) regret sigma d sqrt{T}. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.

Source: arXiv cs.LG | 2026-07-07

Loading related sources…