Fetching the paper…
Reading the bibliography…
Offline reinforcement learning (RL), a technology that offline learns a policy from logged data without the need to interact with online environments, has become a favorable choice in decision-making processes like interactive recommendation.
Q-learning
Christopher JCH Watkins and Peter Dayan. 1992 · 1992
Earlier work this paper cites.
Actor-critic algorithms
Vijay Konda and John Tsitsiklis. 1999 · 1999
Earlier work this paper cites.
Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. In International Conference on Machine Learning (ICML ’15) . PMLR, 814–823
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Earlier work this paper cites.
Generative Adversarial Imitation Learning. In Advances in Neural Information Processing Systems (NeurIPS ’16, Vol. 29) , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.)
Jonathan Ho and Stefano Ermon. 2016 · 2016
Earlier work this paper cites.
The LFM-1b Dataset for Music Retrieval and Recommendation. In Proceedings of the 2016 ACM on International Conference on Multimedia Retrieval (New York, New York, USA) (ICMR ’16) . 103–110
Markus Schedl. 2016 · 2016
Earlier work this paper cites.
DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia) (IJCAI’17) . 1725–1731
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017 · 2017
Earlier work this paper cites.
What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?. In Proceedings of the 31st International Conference on Neural Information Processing Systems (Long Beach, California, USA) (NeurIPS ’17) . 5580–5590
Alex Kendall and Yarin Gal. 2017 · 2017
Earlier work this paper cites.
Visual foresight: Model-based deep reinforcement learning for vision-based robotic control
Frederik Ebert, Chelsea Finn, Sudeep Dasari, Annie Xie, Alex Lee, and Sergey Levine. 2018 · 2018
Earlier work this paper cites.
Offline A/B Testing for Recommender Systems. In WSDM ’18 . 198–206
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Earlier work this paper cites.
Self-attentive sequential recommendation. In International Conference on Data Mining (ICDM ’18) . IEEE, 197–206
Wang-Cheng Kang and Julian McAuley. 2018 · 2018
Earlier work this paper cites.
Algorithmic framework for model-based deep reinforcement learning with theoretical guarantees
Yuping Luo, Huazhe Xu, Yuanzhi Li, Yuandong Tian, Trevor Darrell, and Tengyu Ma. 2018 · 2018
Earlier work this paper cites.
Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining (Marina Del Rey, CA, USA) (WSDM ’18) . 565–573
Jiaxi Tang and Ke Wang. 2018 · 2018
Earlier work this paper cites.
Quantitative analysis of Matthew effect and sparsity problem of recommender systems. In 2018 IEEE 3rd International Conference on Cloud Computing and Big Data Analysis (ICCCBDA) . 78–82
Hao Wang, Zonghu Wang, and Weishi Zhang. 2018 · 2018
Earlier work this paper cites.
Top-K Off-Policy Correction for a REINFORCE Recommender System. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (Melbourne VIC, Australia) (WSDM ’19) . 456–464
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H. Chi. 2019 · 2019
Earlier work this paper cites.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau. 2019a · 2019
Earlier work this paper cites.
Stabilizing off-policy q-learning via bootstrapping error reduction
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019 · 2019
Earlier work this paper cites.
A Simple Convolutional Generative Network for Next Item Recommendation. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining (Melbourne VIC, Australia) (WSDM ’19) . 582–590
Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. Jose, and Xiangnan He. 2019 · 2019
Earlier work this paper cites.
Reinforcement Learning to Optimize Long-Term User Engagement in Recommender Systems. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Anchorage, AK, USA) (KDD ’19) . 2810–2818
Lixin Zou, Long Xia, Zhuoye Ding, Jiaxing Song, Weidong Liu, and Dawei Yin. 2019 · 2019
Earlier work this paper cites.
An optimistic perspective on offline reinforcement learning. In International Conference on Machine Learning (ICML ’20) . PMLR, 104–114
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. 2020 · 2020
Earlier work this paper cites.
Algorithmic Effects on the Diversity of Consumption on Spotify. In Proceedings of The Web Conference 2020 (Taipei, Taiwan) (WWW ’20) . 2155–2165
Ashton Anderson, Lucas Maystre, Ian Anderson, Rishabh Mehrotra, and Mounia Lalmas. 2020 · 2020
Earlier work this paper cites.
Keeping Dataset Biases out of the Simulation: A Debiased Simulator for Reinforcement Learning Based Recommender Systems. In RecSys ’20 . 190–199
Jin Huang, Harrie Oosterhuis, Maarten de Rijke, and Herke van Hoof. 2020 · 2020
Cited alongside, same era.
MOReL: Model-Based Offline Reinforcement Learning. In Advances in Neural Information Processing Systems (NeurIPS ’20, Vol. 33) , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). 21810–21823
Rahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, and Thorsten Joachims. 2020 · 2020
Cited alongside, same era.
Conservative Q-Learning for Offline Reinforcement Learning. In Advances in Neural Information Processing Systems (NeurIPS ’20, Vol. 33) , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). 1179–1191
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Cited alongside, same era.
Offline reinforcement learning: Tutorial, review, and perspectives on open problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2020
Disentangling User Interest and Conformity for Recommendation with Causal Embedding. In Proceedings of the Web Conference 2021 (Ljubljana, Slovenia) (WWW ’21) . 2980–2991
Yu Zheng, Chen Gao, Xiang Li, Xiangnan He, Yong Li, and Depeng Jin. 2021b · 2021
Later among the works it cites.
Reinforcement Learning based Recommender Systems: A Survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far. 2022 · 2022
Later among the works it cites.
Off-Policy Actor-Critic for Recommender Systems. In Proceedings of the 16th ACM Conference on Recommender Systems (Seattle, WA, USA) (RecSys ’22) . 338–349
Minmin Chen, Can Xu, Vince Gatto, Devanshu Jain, Aviral Kumar, and Ed Chi. 2022 · 2022
Later among the works it cites.
CIRS: Bursting Filter Bubbles by Counterfactual Interactive Recommender System
Chongming Gao, Wenqiang Lei, Jiawei Chen, Shiqi Wang, Xiangnan He, Shijun Li, Biao Li, Yuan Zhang, and Peng Jiang. 2022a · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
What are the statistical limits of offline RL with linear function approximation?
Ruosong Wang, Dean P Foster, and Sham M Kakade. 2020a · 2020
Cited alongside, same era.
Critic Regularized Regression
Ziyu Wang, Alexander Novikov, Konrad Zolna, Josh S Merel, Jost Tobias Springenberg, Scott E Reed, Bobak Shahriari, Noah Siegel, Caglar Gulcehre, Nicolas Heess, et al · 2020
Cited alongside, same era.
Self-Supervised Reinforcement Learning for Recommender Systems. In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20) . 931–940
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M. Jose. 2020 · 2020
Cited alongside, same era.
MOPO: Model-based Offline Policy Optimization. In Advances in Neural Information Processing Systems (NeurIPS ’20, Vol. 33) , H. Larochelle, M. Ranzato, R. Hadsell, M.F. Balcan, and H. Lin (Eds.). 14129–14142
Tianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon, James Y Zou, Sergey Levine, Chelsea Finn, and Tengyu Ma. 2020 · 2020
Cited alongside, same era.
Pseudo Dyna-Q: A Reinforcement Learning Framework for Interactive Recommendation. In WSDM ’20 . 816–824
Lixin Zou, Long Xia, Pan Du, Zhuo Zhang, Ting Bai, Weidong Liu, Jian-Yun Nie, and Dawei Yin. 2020 · 2020
Cited alongside, same era.
Advances and Challenges in Conversational Recommender Systems: A Survey
Chongming Gao, Wenqiang Lei, Xiangnan He, Maarten de Rijke, and Tat-Seng Chua. 2021 · 2021
Cited alongside, same era.
Shifting Consumption towards Diverse Content on Music Streaming Platforms. In Proceedings of the 14th ACM International Conference on Web Search and Data Mining (Virtual Event, Israel) (WSDM ’21) . 238–246
Christian Hansen, Rishabh Mehrotra, Casper Hansen, Brian Brost, Lucas Maystre, and Mounia Lalmas. 2021 · 2021
Cited alongside, same era.
Pessimistic Reward Models for Off-Policy Learning in Recommendation. In Proceedings of the 15th ACM Conference on Recommender Systems (Amsterdam, Netherlands) (RecSys ’21) . 63–74
Olivier Jeunen and Bart Goethals. 2021 · 2021
Cited alongside, same era.
Toward Pareto Efficient Fairness-Utility Trade-off in Recommendation through Reinforcement Learning. In Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining (Virtual Event, AZ, USA) (WSDM ’22) . 316–324
Yingqiang Ge, Xiaoting Zhao, Lucia Yu, Saurabh Paul, Diane Hu, Chu-Cheng Hsieh, and Yongfeng Zhang. 2022 · 2022
Later among the works it cites.
State Encoders in Reinforcement Learning for Recommendation: A Reproducibility Study. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22) . 2738–2748
Jin Huang, Harrie Oosterhuis, Bunyamin Cetinkaya, Thijs Rood, and Maarten de Rijke. 2022 · 2022
Later among the works it cites.
Offline Reinforcement Learning with Implicit Q-Learning. In International Conference on Learning Representations (ICLR ’22)
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2022 · 2022
Later among the works it cites.
Who Are the Best Adopters? User Selection Model for Free Trial Item Promotion
Shiqi Wang, Chongming Gao, Min Gao, Junliang Yu, Zongwei Wang, and Hongzhi Yin. 2022a · 2022
Later among the works it cites.
Dynamics-Aware Adaptation for Reinforcement Learning Based Cross-Domain Interactive Recommendation (SIGIR ’22) . 290–300
Junda Wu, Zhihui Xie, Tong Yu, Handong Zhao, Ruiyi Zhang, and Shuai Li. 2022 · 2022
Later among the works it cites.
Rethinking Reinforcement Learning for Recommendation: A Prompt Perspective. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval (Madrid, Spain) (SIGIR ’22) . 1347–1357
Xin Xin, Tiago Pimentel, Alexandros Karatzoglou, Pengjie Ren, Konstantina Christakopoulou, and Zhaochun Ren. 2022 · 2022
Later among the works it cites.
Dynamic Causal Collaborative Filtering. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management (Atlanta, GA, USA) (CIKM ’22) . 2301–2310
Shuyuan Xu, Juntao Tan, Zuohui Fu, Jianchao Ji, Shelby Heinecke, and Yongfeng Zhang. 2022 · 2022
Later among the works it cites.
Multi-Task Fusion via Reinforcement Learning for Long-Term User Satisfaction in Recommender Systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22) . 4510–4520
Qihua Zhang, Junning Liu, Yuzhuo Dai, Yiyan Qi, Yifan Yuan, Kunlun Zheng, Fan Huang, and Xianfeng Tan. 2022 · 2022
Later among the works it cites.
Reinforcing User Retention in a Billion Scale Short Video Recommender System
Qingpeng Cai, Shuchang Liu, Xueliang Wang, Tianyou Zuo, Wentao Xie, Bin Yang, Dong Zheng, Peng Jiang, and Kun Gai. 2023a · 2023
Closest in time.
Two-Stage Constrained Actor-Critic for Short Video Recommendation
Qingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue, Shuchang Liu, Ruohan Zhan, Xueliang Wang, Tianyou Zuo, Wentao Xie, Dong Zheng, et al · 2023
Closest in time.
Romain Deffayet, Thibaut Thonet, Jean-Michel Renders, and Maarten de Rijke. 2023 · 2023
Closest in time.
Exploration and Regularization of the Latent Action Space in Recommendation
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Kun Gai, Peng Jiang, Xiangyu Zhao, and Yongfeng Zhang. 2023 · 2023
Closest in time.
ResAct: Reinforcing Long-term Engagement in Sequential Recommendation with Residual Actor (ICLR ’23)
Wanqi Xue, Qingpeng Cai, Ruohan Zhan, Dong Zheng, Peng Jiang, and Bo An. 2023 · 2023
Closest in time.
Yuan Zhang, Xue Dong, Weijie Ding, Biao Li, Peng Jiang, and Kun Gai. 2023 · 2023
Closest in time.
Off-Policy Deep Reinforcement Learning without Exploration. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019b · 2062
Closest in time.