Fetching the paper…
Reading the bibliography…
Contextual bandit algorithms are increasingly replacing non-adaptive A/B tests in e-commerce, healthcare, and policymaking because they can both improve outcomes for study participants and increase the chance of identifying good or even best policies.
Nathan Kallus and Masatoshi Uehara · 1909
Earlier work this paper cites.
Estimation of regression coefficients when some regressors are not always observed
James M Robins, Andrea Rotnitzky, and Lue Ping Zhao · 1994
Earlier work this paper cites.
Weak Convergence and Empirical Processes
A. van der Vaart and J. Wellner · 1996
Earlier work this paper cites.
Asymptotic statistics
Aad W van der Vaart · 2000
Earlier work this paper cites.
Unified methods for censored longitudinal data and causality
Mark J van der Laan and James M Robins · 2003
Earlier work this paper cites.
The offset tree for learning with partial labels
Alina Beygelzimer and John Langford · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and learning
Miroslav Dudík, John Langford, and Lihong Li · 2011
Earlier work this paper cites.
Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
On the minimal penalty for Markov order estimation
R. van Handel · 2011
Earlier work this paper cites.
Estimation bias in multi-armed bandit algorithms for search advertising
Min Xu, Tao Qin, and Tie-Yan Liu · 2013
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, Lihong Li, et al · 2014
Cited alongside, same era.
Statistical inference for the mean outcome under a possibly non-unique optimal treatment strategy
Alexander R. Luedtke and Mark J. van der Laan · 2016
Cited alongside, same era.
Dynamic pricing with demand covariates
Sheng Qiang and Mohsen Bayati · 2016
Cited alongside, same era.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Bernd Bischl, Giuseppe Casalicchio, Matthias Feurer, Frank Hutter, Michel Lang, Rafael G Mantovani, Jan N van Rijn, and Joaquin Vanschoren · 2017
Cited alongside, same era.
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Later among the works it cites.
Balanced policy evaluation and learning
Nathan Kallus · 2018
Later among the works it cites.
Why adaptively collected data have negative bias and how to correct for it
Xinkun Nie, Xiaoying Tian, Jonathan Taylor, and James Zou · 2018
Later among the works it cites.
Confidence intervals for policy evaluation in adaptive experiments
Vitor Hadad, David A Hirshberg, Ruohan Zhan, Stefan Wager, and Susan Athey · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unbiased estimation for response adaptive clinical trials
Jack Bowden and Lorenzo Trippa · 2017
Cited alongside, same era.
Estimation considerations in contextual bandits
Maria Dimakopoulou, Zhengyuan Zhou, Susan Athey, and Guido Imbens · 2017
Cited alongside, same era.
From ads to interventions: Contextual bandits in mobile health
Ambuj Tewari and Susan A Murphy · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık · 2017
Cited alongside, same era.
A sequential and adaptive experiment to increase the uptake of long-acting reversible contraceptives in cameroon, 2018
Susan Athey, Sarah Baird, Julian Jamison, Craig McIntosh, Berk Özler, and Dohbit Sama · 2018
Cited alongside, same era.
Ae: A domain-agnostic platform for adaptive experimentation
Eytan Bakshy, Lili Dworkin, Brian Karrer, Konstantin Kashin, Benjamin Letham, Ashwin Murthy, and Shaun Singh · 2018
Cited alongside, same era.
Intrinsically efficient, stable, and bounded off-policy evaluation for reinforcement learning
Nathan Kallus and Masatoshi Uehara
Cited in the paper.
Simon Quinn, Alex Teytelboym, Maximilian Kasy, Grant Gordon, and Stefano Caria · 2019
Later among the works it cites.
On the bias, risk and consistency of sample means in multi-armed bandits
Jaehyeok Shin, Aaditya Ramdas, and Alessandro Rinaldo · 2019
Later among the works it cites.
Cab: Continuous adaptive blending for policy evaluation and learning
Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims · 2019
Later among the works it cites.
Dynamic assortment personalization in high dimensions
Nathan Kallus and Madeleine Udell · 2020
Later among the works it cites.
Risk minimization from adaptively collected data: Guarantees for supervised and policy learning
Aurelien Bibaut, Maria Dimakopoulou, Antoine Chambaz, Nathan Kallus, and Mark van der Laan · 2021
Closest in time.
Adaptive treatment assignment in experiments for policy choice
Maximilian Kasy and Anja Sautmann · 2021
Closest in time.