Fetching the paper…
Reading the bibliography…
Off-policy evaluation (OPE) aims to estimate the performance of hypothetical policies using data generated by a different policy.
Eligibility Traces for Off-Policy Policy Evaluation
Doina Precup, Richard S. Sutton, and Satinder Singh · 2000
Earlier work this paper cites.
Marginal mean models for dynamic regimes
Susan A Murphy, Mark J van der Laan, James M Robins, and Conduct Problems Prevention Research Group · 2001
Earlier work this paper cites.
The offset tree for learning with partial labels
Alina Beygelzimer and John Langford · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
A Contextual-bandit Approach to Personalized News Article Recommendation
Lihong Li, Wei Chu, John Langford, and Robert E Schapire · 2010
Earlier work this paper cites.
Learning from Logged Implicit Exploration Data
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade · 2010
Earlier work this paper cites.
Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
Counterfactual Reasoning and Learning Systems: The Example of Computational Advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson · 2013
Earlier work this paper cites.
Doubly Robust Policy Evaluation and Optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li · 2014
Earlier work this paper cites.
Rich feature hierarchies for accurate object detection and semantic segmentation
Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik · 2014
Earlier work this paper cites.
Offline policy evaluation across representations with applications to educational games
Travis Mandel, Yun-En Liu, Sergey Levine, Emma Brunskill, and Zoran Popovic · 2014
Earlier work this paper cites.
Batch learning from logged bandit feedback through counterfactual risk minimization
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims · 2015
Earlier work this paper cites.
High confidence policy improvement
Philip Thomas, Georgios Theocharous, and Mohammad Ghavamzadeh · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning
Nan Jiang and Lihong Li · 2016
Earlier work this paper cites.
Large-scale validation of counterfactual learning methods: A test-bed
Damien Lefortier, Adith Swaminathan, Xiaotao Gu, Thorsten Joachims, and Maarten de Rijke · 2016
Cited alongside, same era.
Data-efficient Off-policy Policy Evaluation for Reinforcement Learning
Philip Thomas and Emma Brunskill · 2016
Cited alongside, same era.
Effective evaluation using logged bandit feedback from multiple loggers
Aman Agarwal, Soumya Basu, Tobias Schnabel, and Thorsten Joachims · 2017
Cited alongside, same era.
Off-policy Evaluation for Slate Recommendation
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni · 2017
Cited alongside, same era.
Optimal and Adaptive Off-policy Evaluation in Contextual Bandits
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik · 2017
Cited alongside, same era.
Cab: Continuous adaptive blending for policy evaluation and learning
Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims · 2019
Later among the works it cites.
On the design of estimators for bandit off-policy evaluation
Nikos Vlassis, Aurelien Bibaut, Maria Dimakopoulou, and Tony Jebara · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue · 2019
Later among the works it cites.
Tengyang Xie, Yifei Ma, and Yu-Xiang Wang · 2019
Later among the works it cites.
Benchmarking graph neural networks
Vijay Prakash Dwivedi, Chaitanya K Joshi, Thomas Laurent, Yoshua Bengio, and Xavier Bresson · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alberto Bietti, Alekh Agarwal, and John Langford · 2018
Cited alongside, same era.
Adapting multi-armed bandits policies to contextual bandits scenarios
David Cortes · 2018
Cited alongside, same era.
More robust doubly robust off-policy evaluation
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh · 2018
Cited alongside, same era.
Offline a/b testing for recommender systems
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé · 2018
Cited alongside, same era.
Policy evaluation and optimization with continuous treatments
Nathan Kallus and Angela Zhou · 2018
Cited alongside, same era.
Offline evaluation of ranking policies with click models
Shuai Li, Yasin Abbasi-Yadkori, Branislav Kveton, S Muthukrishnan, Vishwa Vinay, and Zheng Wen · 2018
Cited alongside, same era.
Representation balancing mdps for off-policy policy evaluation
Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A Faisal, Finale Doshi-Velez, and Emma Brunskill · 2018
Cited alongside, same era.
Closest in time.
Open graph benchmark: Datasets for machine learning on graphs
Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec · 2020
Closest in time.
A practical guide of off-policy evaluation for bandit problems
Masahiro Kato, Kenshi Abe, Kaito Ariu, and Shota Yasui · 2020
Closest in time.
Counterfactual evaluation of slate recommendations with sequential reward interactions
James McInerney, Brian Brost, Praveen Chandar, Rishabh Mehrotra, and Benjamin Carterette · 2020
Closest in time.
Doubly robust off-policy evaluation with shrinkage
Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dudík · 2020
Closest in time.
Off-policy evaluation and learning for external validity under a covariate shift
Masatoshi Uehara, Masahiro Kato, and Shota Yasui · 2020
Closest in time.
Offline Evaluation of Multi-Armed Bandit Algorithms Using Bootstrapped Replay on Expanded Data
Jin Dai · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, Ziyu Wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, et al · 2021
Closest in time.
Finite-time analysis of globally nonstationary multi-armed bandits
Junpei Komiyama, Edouard Fouché, and Junya Honda · 2021
Closest in time.
Debiased off-policy evaluation for recommendation systems
Yusuke Narita, Shota Yasui, and Kohei Yata · 2021
Closest in time.
Counterfactual learning and evaluation for recommender systems: Foundations, implementations, and recent advances
Yuta Saito and Thorsten Joachims · 2021
Closest in time.
Evaluating the robustness of off-policy evaluation
Yuta Saito, Takuma Udagawa, Haruka Kiyohara, Kazuki Mogi, Yusuke Narita, and Kei Tateno · 2021
Closest in time.
Causal combinatorial factorization machines for set-wise recommendation
Akira Tanimoto, Tomoya Sakai, Takashi Takenouchi, and Hisashi Kashima · 2021
Closest in time.