Safety

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

arXiv:2607.15067v2 Announce Type: replace Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator,

DGX agentpaper
safetyarxiv-cs-lg

arXiv:2607.15067v2 Announce Type: replace Abstract: This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including weighted importance sampling), particularly under behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of weighted importance sampling with the linearity of vanilla importance sampling.

Source: arXiv cs.LG | 2026-08-05

Loading related sources…