Fetching the paper…
Reading the bibliography…
Off-policy Evaluation (OPE), or offline evaluation in general, evaluates the performance of hypothetical policies leveraging only offline log data.
On the likelihood that one unknown probability exceeds another in view of the evidence of two samples
William R Thompson. 1933 · 1933
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup. 2000 · 2000
Earlier work this paper cites.
The offset tree for learning with partial labels. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 129–138
Alina Beygelzimer and John Langford. 2009 · 2009
Earlier work this paper cites.
Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade. 2010 · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. 2014 · 2014
Earlier work this paper cites.
The self-normalized estimator for counterfactual learning. In Advances in Neural Information Processing Systems , Vol. 28. 3231–3239
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning. In International Conference on Machine Learning , Vol. 48. PMLR, 652–661
Nan Jiang and Lihong Li. 2016 · 2016
Earlier work this paper cites.
Effective evaluation using logged bandit feedback from multiple loggers. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 687–696
Aman Agarwal, Soumya Basu, Tobias Schnabel, and Thorsten Joachims. 2017 · 2017
Earlier work this paper cites.
UCI machine learning repository
Dheeru Dua and Casey Graff. 2017 · 2017
Earlier work this paper cites.
Optimal and adaptive off-policy evaluation in contextual bandits. In International Conference on Machine Learning , Vol. 70. PMLR, 3589–3597
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık. 2017 · 2017
Cited alongside, same era.
More robust doubly robust off-policy evaluation. In International Conference on Machine Learning , Vol. 80. PMLR, 1447–1456
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018 · 2018
Cited alongside, same era.
Offline a/b testing for recommender systems. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining . 198–206
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Cited alongside, same era.
Using cumulative distribution based performance analysis to benchmark models. In NeurIPS 2018 Workshop on Critiquing and Correcting Trends in Machine Learning
Scott M Jordan, Daniel Cohen, and Philip S Thomas. 2018 · 2018
Cited alongside, same era.
Behaviour policy estimation in off-policy policy evaluation: Calibration matters
On the design of estimators for bandit off-policy evaluation. In International Conference on Machine Learning , Vol. 97. PMLR, 6468–6476
Nikos Vlassis, Aurelien Bibaut, Maria Dimakopoulou, and Tony Jebara. 2019 · 2019
Later among the works it cites.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue. 2019 · 2019
Later among the works it cites.
Implementation Matters in Deep RL: A Case Study on PPO and TRPO. In International Conference on Learning Representations
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Firdaus Janoos, Larry Rudolph, and Aleksander Madry. 2020 · 2020
Later among the works it cites.
Evaluating the performance of reinforcement learning algorithms. In Proceedings of the 37th International Conference on Machine Learning . PMLR, 4962–4973
Scott Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip Thomas. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aniruddh Raghu, Omer Gottesman, Yao Liu, Matthieu Komorowski, Aldo Faisal, Finale Doshi-Velez, and Emma Brunskill. 2018 · 2018
Cited alongside, same era.
Intrinsically Efficient, Stable, and Bounded Off-Policy Evaluation for Reinforcement Learning. In Advances in Neural Information Processing Systems , Vol. 32. 3325–3334
Nathan Kallus and Masatoshi Uehara. 2019 · 2019
Cited alongside, same era.
Triply Robust Off-Policy Evaluation
Anqi Liu, Hao Liu, Anima Anandkumar, and Yisong Yue. 2019 · 2019
Cited alongside, same era.
Efficient counterfactual learning from bandit feedback. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 4634–4641
Yusuke Narita, Shota Yasui, and Kohei Yata. 2019 · 2019
Cited alongside, same era.
Cab: Continuous adaptive blending for policy evaluation and learning. In International Conference on Machine Learning , Vol. 97. PMLR, 6005–6014
Yi Su, Lequn Wang, Michele Santacatterina, and Thorsten Joachims. 2019 · 2019
Cited alongside, same era.
Doubly robust off-policy evaluation with shrinkage. In International Conference on Machine Learning , Vol. 119. PMLR, 9167–9176
Yi Su, Maria Dimakopoulou, Akshay Krishnamurthy, and Miroslav Dudík. 2020a
Cited in the paper.
Adaptive Estimator Selection for Off-Policy Evaluation. In Proceedings of the 37th International Conference on Machine Learning , Vol. 119. PMLR, 9196–9205
Yi Su, Pavithra Srinath, and Akshay Krishnamurthy. 2020b
Cited in the paper.
Off-Policy Evaluation and Learning for External Validity under a Covariate Shift. In Advances in Neural Information Processing Systems , Vol. 33. 49–61
Masahiro Kato, Shota Yasui, and Masatoshi Uehara. 2020 · 2020
Later among the works it cites.
Off-policy Bandit and Reinforcement Learning
Yusuke Narita, Shota Yasui, and Kohei Yata. 2020 · 2020
Later among the works it cites.
Doubly robust estimator for ranking metrics with post-click conversions. In Fourteenth ACM Conference on Recommender Systems . 92–100
Yuta Saito. 2020 · 2020
Later among the works it cites.
Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation
Yuta Saito, Shunsuke Aihara, Megumi Matsutani, and Yusuke Narita. 2020 · 2020
Later among the works it cites.
Optimal Off-Policy Evaluation from Multiple Logging Policies. In Proceedings of the 38th International Conference on Machine Learning , Vol. 139. PMLR, 5247–5256
Nathan Kallus, Yuta Saito, and Masatoshi Uehara. 2021 · 2021
Closest in time.