Practical and Optimal Algorithm for Linear Contextual Bandits with Rare Parameter Updates
DGX agentarXiv:2606.00984v1 Announce Type: cross Abstract: We study linear contextual bandits under rare parameter updates: the learner may incorporate reward feedback into its parameter estimate only at a sma