Best-of-Both-Worlds Multi-Dueling Bandits: Unified Algorithms for Stochastic and Adversarial Preferences under Condorcet and Borda Objectives
DGX agentarXiv:2603.18972v3 Announce Type: replace Abstract: Multi-dueling bandits, where a learner selects m geq 2 arms per round and observes only the winner, arise naturally in many applications including r