Fetching the paper…
Reading the bibliography…
Efficient methods to evaluate new algorithms are critical for improving interactive bandit and reinforcement learning systems such as recommendation systems.
Semiparametric regression estimation in the presence of dependent censoring
Andrea Rotnitzky and James M Robins. 1995 · 1995
Earlier work this paper cites.
Eligibility Traces for Off-Policy Policy Evaluation, In Proceedings of the 17th International Conference on Machine Learning
Doina Precup, Richard S. Sutton, and Satinder Singh. 2000 · 2000
Earlier work this paper cites.
A Contextual-bandit Approach to Personalized News Article Recommendation, In Proceedings of the 19th International Conference on World Wide Web
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Earlier work this paper cites.
Learning from Logged Implicit Exploration Data, In Advances in Neural Information Processing Systems 23
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade. 2010 · 2010
Earlier work this paper cites.
Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms, In Proceedings of the Fourth ACM International Conference on Web Search and Data Mining
Lihong Li, Wei Chu, John Langford, and Xuanhui Wang. 2011 · 2011
Earlier work this paper cites.
Doubly Robust Policy Evaluation and Optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. 2014 · 2014
Earlier work this paper cites.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. 2016 · 2016
Earlier work this paper cites.
Doubly Robust Off-policy Value Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning . 652–661
Nan Jiang and Lihong Li. 2016 · 2016
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30 . 3146–3154
Guolin Ke, Qi Meng, Thomas Finley, Taifeng Wang, Wei Chen, Weidong Ma, Qiwei Ye, and Tie-Yan Liu. 2017 · 2017
Cited alongside, same era.
Off-policy Evaluation for Slate Recommendation, In Advances in Neural Information Processing Systems 30
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni. 2017 · 2017
Cited alongside, same era.
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudik. 2017 · 2017
Cited alongside, same era.
Double/debiased machine learning for treatment and structural parameters
Victor Chernozhukov, Denis Chetverikov, Mert Demirer, Esther Duflo, Christian Hansen, Whitney Newey, and James Robins. 2018a · 2018
Cited alongside, same era.
Locally Robust Semiparametric Estimation
Victor Chernozhukov, Juan Carlos Escanciano, Hidehiko Ichimura, Whitney K. Newey, and James M. Robins. 2018b · 2018
Representation balancing mdps for off-policy policy evaluation. In Advances in Neural Information Processing Systems 31 . 2644–2653
Yao Liu, Omer Gottesman, Aniruddh Raghu, Matthieu Komorowski, Aldo A Faisal, Finale Doshi-Velez, and Emma Brunskill. 2018 · 2018
Later among the works it cites.
Cross-Fitting and Fast Remainder Rates for Semiparametric Estimation
Whitney K. Newey and James M. Robins. 2018 · 2018
Later among the works it cites.
Off-Policy Evaluation via Off-Policy Classification. In Advances in Neural Information Processing Systems 32
Alex Irpan, Kanishka Rao, Konstantinos Bousmalis, Chris Harris, Julian Ibarz, and Sergey Levine. 2019 · 2019
Later among the works it cites.
Efficient Counterfactual Learning from Bandit Feedback, In Proceedings of the 33rd AAAI Conference on Artificial Intelligence
Yusuke Narita, Shota Yasui, and Kohei Yata. 2019 · 2019
Later among the works it cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library, In Advances in Neural Information Processing Systems 32
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
More Robust Doubly Robust Off-policy Evaluation. In Proceedings of the 35th International Conference on Machine Learning . 1447–1456
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018 · 2018
Cited alongside, same era.
Offline A/B Testing for Recommender Systems, In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Cited alongside, same era.
Data-Efficient Off-Policy Policy Evaluation for Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning . 2139–2148
Philip Thomas and Emma Brunskill. 2016a
Cited in the paper.
Data-efficient Off-policy Policy Evaluation for Reinforcement Learning, In Proceedings of the 33rd International Conference on Machine Learning
Philip Thomas and Emma Brunskill. 2016b
Cited in the paper.
Later among the works it cites.
Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes, In Proceedings of the 37th International Conference on Machine Learning
Nathan Kallus and Masatoshi Uehara. 2020 · 2020
Closest in time.
Minimax Weight and Q-Function Learning for Off-Policy Evaluation, In Proceedings of the 37th International Conference on Machine Learning
Masatoshi Uehara, Jiawei Huang, and Nan Jiang. 2020 · 2020
Closest in time.