Fetching the paper…
Reading the bibliography…
Current advances in recommender systems have been remarkably successful in optimizing immediate engagement.
Returning is believing: Optimizing long-term user engagement in recommender systems. In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management . ACM, 1927–1936
Qingyun Wu, Hongning Wang, Liangjie Hong, and Yue Shi. 2017 · 1936
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. The method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Algorithms for Inverse Reinforcement Learning. In Proceedings of the Seventeenth International Conference on Machine Learning . 663–670
Andrew Y. Ng and Stuart J. Russell. 2000 · 2000
Earlier work this paper cites.
Quantile regression
Roger Koenker and Kevin F Hallock. 2001 · 2001
Earlier work this paper cites.
An MDP-Based Recommender System
Guy Shani, David Heckerman, and Ronen I. Brafman. 2005 · 2005
Earlier work this paper cites.
Solving the apparent diversity-accuracy dilemma of recommender systems
Tao Zhou, Zoltán Kuscsik, Jian-Guo Liu, Matúš Medo, Joseph Rushton Wakeling, and Yi-Cheng Zhang. 2010 · 2010
Earlier work this paper cites.
Maximizing aggregate recommendation diversity: A graph-theoretic approach. In Proc. of the 1st International Workshop on Novelty and Diversity in Recommender Systems . 3–10
Gediminas Adomavicius and YoungOk Kwon. 2011 · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
The self-normalized estimator for counterfactual learning
Adith Swaminathan and Thorsten Joachims. 2015 · 2015
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks. In ICLR
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2016 · 2016
Earlier work this paper cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. 2016 · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2016 · 2016
Earlier work this paper cites.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Exploration: A study of count-based exploration for deep reinforcement learning. In NeurIPS . 2753–2762
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel. 2017 · 2017
Earlier work this paper cites.
Stabilizing reinforcement learning in dynamic environment with application to online recommendation. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining . 1187–1196
Shi-Yong Chen, Yang Yu, Qing Da, Jun Tan, Hai-Kuan Huang, and Hai-Hong Tang. 2018 · 2018
Earlier work this paper cites.
More robust doubly robust off-policy evaluation. In ICML . 1447–1456
Mehrdad Farajtabar, Yinlam Chow, and Mohammad Ghavamzadeh. 2018 · 2018
Earlier work this paper cites.
Addressing function approximation error in actor-critic methods. In ICML . 1587–1596
Scott Fujimoto, Herke Hoof, and David Meger. 2018 · 2018
Cited alongside, same era.
Offline a/b testing for recommender systems. In Proceedings of the 11th ACM International Conference on Web Search and Data Mining . 198–206
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International conference on machine learning . PMLR, 1861–1870
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. 2018 · 2018
Cited alongside, same era.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Cited alongside, same era.
Reinforcement Learning: An Introduction
Richard S Sutton and Andrew G Barto. 2018 · 2018
A minimalist approach to offline reinforcement learning
Scott Fujimoto and Shixiang Shane Gu. 2021 · 2021
Later among the works it cites.
Reinforcement learning to optimize lifetime value in cold-start recommendation. In CIKM . ACM, 782–791
Luo Ji, Qi Qin, Bingqing Han, and Hongxia Yang. 2021 · 2021
Later among the works it cites.
Improving long-term metrics in recommendation systems using short-horizon offline RL
Bogdan Mazoure, Paul Mineiro, Pavithra Srinath, Reza Sharifi Sedeh, Doina Precup, and Adith Swaminathan. 2021 · 2021
Later among the works it cites.
Counterfactual Reward Modification for Streaming Recommendation with Delayed Feedback. In SIGIR ’21: The 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . ACM, 41–50
Xiao Zhang, Haonan Jia, Hanjing Su, Wenhan Wang, Jun Xu, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Rabbit Holes and Taste Distortion: Distribution-Aware Recommendation with Evolving Interests. In WWW ’21: The Web Conference 2021 . 888–899
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
DRN: A deep reinforcement learning framework for news recommendation. In WWW . 167–176
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li. 2018 · 2018
Cited alongside, same era.
A model-based reinforcement learning with adversarial training for online recommendation. In NeurIPS . 10734–10745
Xueying Bai, Jian Guan, and Hongning Wang. 2019 · 2019
Cited alongside, same era.
SlateQ: A Tractable Decomposition for Reinforcement Learning with Recommendation Sets. In IJCAI , Sarit Kraus (Ed.). 2592–2599
Eugene Ie, Vihan Jain, Jing Wang, Sanmit Narvekar, Ritesh Agarwal, Rui Wu, Heng-Tze Cheng, Tushar Chandra, and Craig Boutilier. 2019 · 2019
Cited alongside, same era.
Virtual-Taobao: Virtualizing real-world online retail environment for Reinforcement Learning. In AAAI . 4902–4909
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and Anxiang Zeng. 2019 · 2019
Cited alongside, same era.
BERT4Rec: Sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management . 1441–1450
Fei Sun, Jun Liu, Jian Wu, Changhua Pei, Xiao Lin, Wenwu Ou, and Peng Jiang. 2019 · 2019
Cited alongside, same era.
Grandmaster level in StarCraft II using multi-agent reinforcement learning
Oriol Vinyals, Igor Babuschkin, Wojciech M Czarnecki, Michaël Mathieu, Andrew Dudzik, Junyoung Chung, David H Choi, Richard Powell, Timo Ewalds, Petko Georgiev, et al · 2019
Cited alongside, same era.
Sequential recommender systems: Challenges, progress and prospects. In IJCAI . 6332–6338
Shoujin Wang, Liang Hu, Yan Wang, Longbing Cao, Quan Z. Sheng, and Mehmet Orgun. 2019 · 2019
Cited alongside, same era.
Xing Zhao, Ziwei Zhu, and James Caverlee. 2021 · 2021
Later among the works it cites.
Reward shaping for user satisfaction in a REINFORCE recommender
Konstantina Christakopoulou, Can Xu, Sai Zhang, Sriraj Badam, Trevor Potter, Daniel Li, Hao Wan, Xinyang Yi, Ya Le, Chris Berg, et al · 2022
Closest in time.
Discovering faster matrix multiplication algorithms with reinforcement learning
Alhussein Fawzi, Matej Balog, Aja Huang, Thomas Hubert, Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Francisco J. R. Ruiz, Julian Schrittwieser, Grzegorz Swirszcz, David Silver, Demis Hassabis, and Pushmeet Kohli. 2022 · 2022
Closest in time.
Offline reinforcement learning with in-sample Q-Learning. In ICLR
Ilya Kostrikov, Ashvin Nair, and Sergey Levine. 2022 · 2022
Closest in time.
Training language models to follow instructions with human feedback
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Closest in time.
SURF: Semi-supervised reward learning with data augmentation for feedback-efficient preference-based reinforcement learning
Jongjin Park, Younggyo Seo, Jinwoo Shin, Honglak Lee, Pieter Abbeel, and Kimin Lee. 2022 · 2022
Closest in time.
Surrogate for long-term user experience in recommender systems. In KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . ACM, 4100–4109
Yuyan Wang, Mohit Sharma, Can Xu, Sriraj Badam, Qian Sun, Lee Richardson, Lisa Chung, Ed H. Chi, and Minmin Chen. 2022 · 2022
Closest in time.
Deconfounding duration bias in watch-time prediction for video recommendation. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4472–4481
Ruohan Zhan, Changhua Pei, Qiang Su, Jianfeng Wen, Xueliang Wang, Guanyu Mu, Dong Zheng, Peng Jiang, and Kun Gai. 2022 · 2022
Closest in time.
Multi-Task Fusion via Reinforcement Learning for Long-Term User Satisfaction in Recommender Systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4510–4520
Qihua Zhang, Junning Liu, Yuzhuo Dai, Yiyan Qi, Yifan Yuan, Kunlun Zheng, Fan Huang, and Xianfeng Tan. 2022 · 2022
Closest in time.
Reinforcing User Retention in a Billion Scale Short Video Recommender System. In Companion Proceedings of the ACM Web Conference 2023 . 421–426
Qingpeng Cai, Shuchang Liu, Xueliang Wang, Tianyou Zuo, Wentao Xie, Bin Yang, Dong Zheng, Peng Jiang, and Kun Gai. 2023a · 2023
Closest in time.
Two-Stage Constrained Actor-Critic for Short Video Recommendation. In Proceedings of the ACM Web Conference 2023 . 865–875
Qingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue, Shuchang Liu, Ruohan Zhan, Xueliang Wang, Tianyou Zuo, Wentao Xie, Dong Zheng, Peng Jiang, and Kun Gai. 2023b · 2023
Closest in time.
Off-policy deep reinforcement learning without exploration. In ICML . 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.