Fetching the paper…
Reading the bibliography…
In recommender systems (RecSys) and real-time bidding (RTB) for online advertisements, we often try to optimize sequential decision making using bandit and reinforcement learning (RL) techniques.
Eligibility traces for off-policy policy evaluation. In Proceedings of the 17th International Conference on Machine Learning . 759–766
Doina Precup, Richard S. Sutton, and Satinder Singh. 2000 · 2000
Earlier work this paper cites.
The offset tree for learning with partial labels. In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining . 129–138
Alina Beygelzimer and John Langford. 2009 · 2009
Earlier work this paper cites.
Learning from logged implicit exploration data. In Advances in Neural Information Processing Systems , Vol. 23. 2217–2225
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade. 2010 · 2010
Earlier work this paper cites.
Doubly robust policy evaluation and optimization
Miroslav Dudík, Dumitru Erhan, John Langford, and Lihong Li. 2014 · 2014
Earlier work this paper cites.
Doubly robust off-policy value evaluation for reinforcement learning. In Proceedings of the 33rd International Conference on International Conference on Machine Learning , Vol. 48. 652–661
Nan Jiang and Lihong Li. 2016 · 2016
Earlier work this paper cites.
Data-efficient off-policy policy evaluation for reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning . 2139–2148
Philip Thomas and Emma Brunskill. 2016 · 2016
Earlier work this paper cites.
Real-time bidding by reinforcement learning in display advertising. In Proceedings of the 10th ACM International Conference on Web Search and Data Mining . 661–670
Han Cai, Kan Ren, Weinan Zhang, Kleanthis Malialis, Jun Wang, Yong Yu, and Defeng Guo. 2017 · 2017
Earlier work this paper cites.
Optimal and adaptive off-policy evaluation in contextual bandits. In Proceedings of the 34th International Conference on Machine Learning , Vol. 70. 3589–3597
Yu-Xiang Wang, Alekh Agarwal, and Miroslav Dudık. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning for list-wise recommendations
Xiangyu Zhao, Liang Zhang, Long Xia, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2017 · 2017
Earlier work this paper cites.
Offline a/b testing for recommender systems. In Proceedings of the 11th ACM International Conference on Web Search and Data Mining . 198–206
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Earlier work this paper cites.
Real-time bidding with multi-agent reinforcement learning in display advertising. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management . 2193–2201
Junqi Jin, Chengru Song, Han Li, Kun Gai, Jun Wang, and Weinan Zhang. 2018 · 2018
Earlier work this paper cites.
David Rohde, Stephen Bonner, Travis Dunlop, Flavian Vasile, and Alexandros Karatzoglou. 2018 · 2018
Earlier work this paper cites.
Budget constrained bidding by model-free reinforcement learning in display advertising. In Proceedings of the 27th ACM International Conference on Information and Knowledge Management . 1443–1451
Di Wu, Xiujun Chen, Xun Yang, Hao Wang, Qing Tan, Xiaoxun Zhang, Jian Xu, and Kun Gai. 2018 · 2018
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau. 2019a · 2019
Earlier work this paper cites.
Offline evaluation to make decisions about playlist recommendation algorithms. In Proceedings of the 12th ACM International Conference on Web Search and Data Mining . 420–428
Alois Gruson, Praveen Chandar, Christophe Charbuillet, James McInerney, Samantha Hansen, Damien Tardieu, and Ben Carterette. 2019 · 2019
Earlier work this paper cites.
RecSim: A configurable simulation platform for recommender systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier. 2019 · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction. In Advances in Neural Information Processing Systems , Vol. 32. 11784–11794
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019 · 2019
Earlier work this paper cites.
Batch policy learning under constraints. In Proceedings of the 36th International Conference on Machine Learning , Vol. 97. 3703–3712
Hoang Le, Cameron Voloshin, and Yisong Yue. 2019 · 2019
Cited alongside, same era.
Empirical study of off-policy policy evaluation for reinforcement learning
Cameron Voloshin, Hoang M Le, Nan Jiang, and Yisong Yue. 2019 · 2019
Cited alongside, same era.
Deep reinforcement learning for search, recommendation, and online advertising: a survey
Xiangyu Zhao, Long Xia, Jiliang Tang, and Dawei Yin. 2019 · 2019
Cited alongside, same era.
Lixin Zou, Long Xia, Zhuoye Ding, Jiaxing Song, Weidong Liu, and Dawei Yin. 2019 · 2019
Cited alongside, same era.
MOPO: Model-based offline policy optimization. In Advances in Neural Information Processing Systems , Vol. 33. 14129–14142
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. 2020 · 2020
Later among the works it cites.
Sim-to-real transfer in deep reinforcement learning for robotics: a survey. In 2020 IEEE Symposium Series on Computational Intelligence (SSCI) . 737–744
Wenshuai Zhao, Jorge Peña Queralta, and Tomi Westerlund. 2020 · 2020
Later among the works it cites.
Model-based offline planning. In International Conference on Learning Representations
Arthur Argenson and Gabriel Dulac-Arnold. 2021 · 2021
Closest in time.
Benchmarks for deep off-policy evaluation
Justin Fu, Mohammad Norouzi, Ofir Nachum, George Tucker, Ziyu Wang, Alexander Novikov, Mengjiao Yang, Michael R Zhang, Yutian Chen, Aviral Kumar, et al · 2021
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
D4RL: Datasets for deep data-driven reinforcement learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
RL Unplugged: A Collection of Benchmarks for Offline Reinforcement Learning
Caglar Gulcehre, Ziyu Wang, Alexander Novikov, Thomas Paine, Sergio Gómez, Konrad Zolna, Rishabh Agarwal, Josh S Merel, Daniel J Mankowitz, Cosmin Paduraru, et al · 2020
Cited alongside, same era.
Dynamic knapsack optimization towards efficient multi-channel sequential advertising. In Proceedings of the 37th International Conference on Machine Learning . 4060–4070
Xiaotian Hao, Zhaoqing Peng, Yi Ma, Guan Wang, Junqi Jin, Jianye Hao, Shan Chen, Rongquan Bai, Mingzhou Xie, Miao Xu, Zhenzhe Zheng, Chuan Yu, Han Li, Jian Xu, and Kun Gai. 2020 · 2020
Cited alongside, same era.
MOReL: Model-based offline reinforcement learning. In Advances in Neural Information Processing Systems , Vol. 33. 21810–21823
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. 2020 · 2020
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning. In Advances in Neural Information Processing Systems , Vol. 33. 1179–1191
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
Offline reinforcement learning: tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2020
Cited alongside, same era.
Off-policy learning in two-stage recommender systems. In Proceedings of The Web Conference 2020 . 463–473
Jiaqi Ma, Zhe Zhao, Xinyang Yi, Ji Yang, Minmin Chen, Jiaxi Tang, Lichan Hong, and Ed H Chi. 2020 · 2020
Cited alongside, same era.
Caglar Gulcehre, Sergio Gómez Colmenarejo, Ziyu Wang, Jakub Sygnowski, Thomas Paine, Konrad Zolna, Yutian Chen, Matthew Hoffman, Razvan Pascanu, and Nando de Freitas. 2021 · 2021
Closest in time.
Deployment-efficient reinforcement learning via model-based offline optimization. In International Conference on Learning Representations
Tatsuya Matsushima, Hiroki Furuta, Yutaka Matsuo, Ofir Nachum, and Shixiang Gu. 2021 · 2021
Closest in time.
Improving long-term metrics in recommendation systems using short-horizon offline RL
Bogdan Mazoure, Paul Mineiro, Pavithra Srinath, Reza Sharifi Sedeh, Doina Precup, and Adith Swaminathan. 2021 · 2021
Closest in time.
NeoRL: A near real-world benchmark for offline reinforcement learning
Rongjun Qin, Songyi Gao, Xingyuan Zhang, Zhen Xu, Shengkai Huang, Zewen Li, Weinan Zhang, and Yang Yu. 2021 · 2021
Closest in time.
Evaluating the Robustness of Off-Policy Evaluation
Yuta Saito, Takuma Udagawa, Haruka Kiyohara, Kazuki Mogi, Yusuke Narita, and Kei Tateno. 2021 · 2021
Closest in time.
Interpretable performance analysis towards offline reinforcement learning: A dataset perspective
Chenyang Xi, Bo Tang, Jiajun Shen, Xinfu Liu, Feiyu Xiong, and Xueying Li. 2021 · 2021
Closest in time.
A general offline reinforcement learning framework for interactive recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 4512–4520
Teng Xiao and Donglin Wang. 2021 · 2021
Closest in time.
Representation matters: Offline pretraining for sequential decision making
Mengjiao Yang and Ofir Nachum. 2021 · 2021
Closest in time.
COMBO: Conservative offline model-based policy optimization
Tianhe Yu, Aviral Kumar, Rafael Rafailov, Aravind Rajeswaran, Sergey Levine, and Chelsea Finn. 2021 · 2021
Closest in time.
Xianyuan Zhan, Haoran Xu, Yue Zhang, Yusen Huo, Xiangyu Zhu, Honglei Yin, and Yu Zheng. 2021 · 2021
Closest in time.
DEAR: Deep reinforcement learning for online advertising impression in recommender systems. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 35. 750–758
Xiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiwang Yang, Xiaobing Liu, Hui Liu, and Jiliang Tang. 2021 · 2021
Closest in time.
Off-policy deep reinforcement learning without exploration. In Proceedings of the 36th International Conference on Machine Learning , Vol. 97. 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019b · 2062
Closest in time.