Fetching the paper…
Reading the bibliography…
We consider off-policy evaluation in the contextual bandit setting for the purpose of obtaining a robust off-policy selection strategy, where the selection strategy is evaluated based on the value of the chosen policy in a set of proposal (target) policies.
Doubly robust off-policy evaluation with shrinkage
Y. Su, M. Dimakopoulou, A. Krishnamurthy, and M. Dudík · 1907
Earlier work this paper cites.
Efron-Stein PAC-Bayesian Inequalities
I. Kuzborskij and C. Szepesvári · 1909
Earlier work this paper cites.
A note on importance sampling using standardized weights
A. Kong · 1992
Earlier work this paper cites.
Weighted average importance sampling and defensive mixture distributions
T. Hesterberg · 1995
Earlier work this paper cites.
Empirical bernstein stopping
V. Mnih, C. Szepesvári, and J.-Y. Audibert · 2008
Earlier work this paper cites.
Truncated importance sampling
E. L. Ionides · 2008
Earlier work this paper cites.
Empirical bernstein bounds and sample variance penalization
A. Maurer and M. Pontil · 2009
Earlier work this paper cites.
Doubly robust policy evaluation and learning
M. Dudík, J. Langford, and L. Li · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Counterfactual reasoning and learning systems: the example of computational advertising
L. Bottou, J. Peters, J. Quiñonero Candela, D. X. Charles, M. Chickering, E. Portugaly, D. Ray, P. Y. Simard, and E. Snelson · 2013
Earlier work this paper cites.
Monte Carlo theory, methods and examples
A. B. Owen · 2013
Cited alongside, same era.
Concentration inequalities: A nonasymptotic theory of independence
S. Boucheron, G. Lugosi, and P. Massart · 2013
Cited alongside, same era.
Doubly robust policy evaluation and optimization
M. Dudík, D. Erhan, J. Langford, and L. Li · 2014
Cited alongside, same era.
Pareto smoothed importance sampling
A. Vehtari, D. Simpson, A. Gelman, Y. Yao, and J. Gabry · 2015
Cited alongside, same era.
Effective evaluation using logged bandit feedback from multiple loggers
A. Agarwal, S. Basu, T. Schnabel, and T. Joachims · 2017
Cited alongside, same era.
Optimal and adaptive off-policy evaluation in contextual bandits
Y.-X. Wang, A. Agarwal, and M. Dudik · 2017
Rethinking the effective sample size
V. Elvira, L. Martino, and C. P. Robert · 2018
Later among the works it cites.
Offline a/b testing for recommender systems
A. Gilotte, C. Calauzènes, T. Nedelec, A. Abraham, and S. Dollé · 2018
Later among the works it cites.
More robust doubly robust off-policy evaluation
M. Farajtabar, Y. Chow, and M. Ghavamzadeh · 2018
Later among the works it cites.
Balanced policy evaluation and learning
N. Kallus · 2018
Later among the works it cites.
A. Bietti, A. Agarwal, and J. Langford · 2018
Later among the works it cites.
Deep learning with logged bandit feedback
T. Joachims, A. Swaminathan, and M. de Rijke · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
UCI machine learning repository, 2017
D. Dua and C. Graff · 2017
Cited alongside, same era.
Computing nonvacuous generalization bounds for deep (stochastic) neural networks with many more parameters than training data
G. K. Dziugaite and D. M. Roy · 2017
Cited alongside, same era.
Policy optimization via importance sampling
A. M. Metelli, M. Papini, F. Faccio, and M. Restelli · 2018
Cited alongside, same era.
Batch learning from logged bandit feedback through counterfactual risk minimization
A. Swaminathan and T. Joachims
Cited in the paper.
High-confidence off-policy evaluation
P. S. Thomas, G. Theocharous, and M. Ghavamzadeh
Cited in the paper.
High confidence policy improvement
P. Thomas, G. Theocharous, and M. Ghavamzadeh
Cited in the paper.
Later among the works it cites.
Empirical likelihood for contextual bandits
N. Karampatziakis, J. Langford, and P. Mineiro · 2019
Later among the works it cites.
Code to reproduce the results in the paper ’empirical likelihood for contextual bandits’
P. Mineiro and N. Karampatziakis · 2020
Closest in time.
Policy learning with observational data
S. Athey and S. Wager · 2021
Closest in time.