Kernel weighted importance sampling for off-policy evaluation in contextual bandits
It provides a more robust off-policy evaluation method for researchers and practitioners working with contextual bandits, particularly when the behavior policy is misspecified.
The paper introduces Kernel-WIS, a new estimator for off-policy evaluation in contextual bandits that is asymptotically consistent and empirically outperforms strong baselines, especially under behaviour policy misspecification.
This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weighted importance sampling), particularly under complex conditions including behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of vanilla weighted importance sampling with the linearity of vanilla importance sampling.