Fetching the paper…
Reading the bibliography…
Reinforcement Learning (RL)-Based Recommender Systems (RSs) have gained rising attention for their potential to enhance long-term user engagement.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Actor-Critic Algorithms. In Advances in Neural Information Processing Systems , S. Solla, T. Leen, and K. Müller (Eds.), Vol. 12. MIT Press
Vijay Konda and John Tsitsiklis. 1999 · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. 2002 · 2002
Earlier work this paper cites.
Cumulated gain-based evaluation of IR techniques
Kalervo Järvelin and Jaana Kekäläinen. 2002 · 2002
Earlier work this paper cites.
An MDP-based recommender system
Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier. 2005 · 2005
Earlier work this paper cites.
Collaborative filtering recommender systems
J Ben Schafer, Dan Frankowski, Jon Herlocker, and Shilad Sen. 2007 · 2007
Earlier work this paper cites.
Factorization meets the neighborhood: a multifaceted collaborative filtering model. In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining . 426–434
Yehuda Koren. 2008 · 2008
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky. 2009 · 2009
Earlier work this paper cites.
Collaborative Prediction and Ranking with Non-Random Missing Data. In Proceedings of the Third ACM Conference on Recommender Systems (New York, New York, USA). Association for Computing Machinery, New York, NY, USA, 5–12
Benjamin M. Marlin and Richard S. Zemel. 2009 · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web . 661–670
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Earlier work this paper cites.
Adaptive ε \varepsilon -greedy exploration in reinforcement learning based on value differences. In KI 2010: Advances in Artificial Intelligence: 33rd Annual German Conference on AI, Karlsruhe, Germany, September 21-24, 2010. Proceedings 33 . Springer, 203–210
Michel Tokic. 2010 · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling
Olivier Chapelle and Lihong Li. 2011 · 2011
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Earlier work this paper cites.
Session-based recommendations with recurrent neural networks
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2015 · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. 2015 · 2015
Earlier work this paper cites.
Filter bubbles, echo chambers, and online news consumption
Seth Flaxman, Sharad Goel, and Justin M Rao. 2016 · 2016
Earlier work this paper cites.
Recommendations as Treatments: Debiasing Learning and Evaluation. In Proceedings of the 33rd International Conference on International Conference on Machine Learning (New York, NY, USA). JMLR.org, 1670–1679
Tobias Schnabel, Adith Swaminathan, Ashudeep Singh, Navin Chandak, and Thorsten Joachims. 2016 · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning. In International conference on machine learning . PMLR, 449–458
Marc G Bellemare, Will Dabney, and Rémi Munos. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017 · 2017
Earlier work this paper cites.
DeepFM: A Factorization-Machine Based Neural Network for CTR Prediction. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (Melbourne, Australia) (IJCAI’17) . AAAI Press, 1725–1731
Huifeng Guo, Ruiming Tang, Yunming Ye, Zhenguo Li, and Xiuqiang He. 2017 · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Cited alongside, same era.
BEARS: Towards an evaluation framework for bandit-based interactive recommender systems
Andrea Barraza-Urbina, Georgia Koutrika, Mathieu d’Aquin, and Conor Hayes. 2018 · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods. In International conference on machine learning . PMLR, 1587–1596
Scott Fujimoto, Herke Hoof, and David Meger. 2018 · 2018
Cited alongside, same era.
Self-attentive sequential recommendation. In 2018 IEEE International Conference on Data Mining (ICDM) . IEEE, 197–206
Wang-Cheng Kang and Julian McAuley. 2018 · 2018
Cited alongside, same era.
RLlib: Abstractions for Distributed Reinforcement Learning. In International Conference on Machine Learning (ICML)
Eric Liang, Richard Liaw, Robert Nishihara, Philipp Moritz, Roy Fox, Ken Goldberg, Joseph E. Gonzalez, Michael I. Jordan, and Ion Stoica. 2018 · 2018
Self-Supervised Reinforcement Learning for Recommender Systems. In Proceedings of the 43th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20)
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon Jose. 2020 · 2020
Later among the works it cites.
AutoDebias: Learning to debias for recommendation. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval . 21–30
Jiawei Chen, Hande Dong, Yang Qiu, Xiangnan He, Xin Xin, Liang Chen, Guli Lin, and Keping Yang. 2021 · 2021
Later among the works it cites.
Off-policy actor-critic for recommender systems. In Proceedings of the 16th ACM Conference on Recommender Systems . 338–349
Minmin Chen, Can Xu, Vince Gatto, Devanshu Jain, Aviral Kumar, and Ed Chi. 2022 · 2022
Later among the works it cites.
CIRS: Bursting filter bubbles by counterfactual interactive recommender system
Chongming Gao, Shiqi Wang, Shijun Li, Jiawei Chen, Xiangnan He, Wenqiang Lei, Biao Li, Yuan Zhang, and Peng Jiang. 2022c · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
David Rohde, Stephen Bonner, Travis Dunlop, Flavian Vasile, and Alexandros Karatzoglou. 2018 · 2018
Cited alongside, same era.
Personalized top-n sequential recommendation via convolutional sequence embedding. In Proceedings of the eleventh ACM international conference on web search and data mining . 565–573
Jiaxi Tang and Ke Wang. 2018 · 2018
Cited alongside, same era.
DRN: A deep reinforcement learning framework for news recommendation. In Proceedings of the 2018 world wide web conference . 167–176
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li. 2018 · 2018
Cited alongside, same era.
Top-k off-policy correction for a REINFORCE recommender system. In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining . 456–464
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi. 2019 · 2019
Cited alongside, same era.
Benchmarking batch deep reinforcement learning algorithms
Scott Fujimoto, Edoardo Conti, Mohammad Ghavamzadeh, and Joelle Pineau. 2019a · 2019
Cited alongside, same era.
Recsim: A configurable simulation platform for recommender systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov, Vihan Jain, Sanmit Narvekar, Jing Wang, Rui Wu, and Craig Boutilier. 2019 · 2019
Cited alongside, same era.
Virtual-taobao: Virtualizing real-world online retail environment for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 33. 4902–4909
Jing-Cheng Shi, Yang Yu, Qing Da, Shi-Yong Chen, and An-Xiang Zeng. 2019 · 2019
Cited alongside, same era.
Irec: An interactive recommendation framework. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval . 3165–3175
Thiago Silva, Nícollas Silva, Heitor Werneck, Carlos Mito, Adriano CM Pereira, and Leonardo Rocha. 2022 · 2022
Later among the works it cites.
Tianshou: A Highly Modularized Deep Reinforcement Learning Library
Jiayi Weng, Huayu Chen, Dong Yan, Kaichao You, Alexis Duburcq, Minghao Zhang, Yi Su, Hang Su, and Jun Zhu. 2022 · 2022
Later among the works it cites.
Dynamic causal collaborative filtering. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management . 2301–2310
Shuyuan Xu, Juntao Tan, Zuohui Fu, Jianchao Ji, Shelby Heinecke, and Yongfeng Zhang. 2022 · 2022
Later among the works it cites.
Multi-task fusion via reinforcement learning for long-term user satisfaction in recommender systems. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 4510–4520
Qihua Zhang, Junning Liu, Yuzhuo Dai, Yiyan Qi, Yifan Yuan, Kunlun Zheng, Fan Huang, and Xianfeng Tan. 2022 · 2022
Later among the works it cites.
Reinforcing User Retention in a Billion Scale Short Video Recommender System. In Companion Proceedings of the ACM Web Conference 2023 (WWW ’23 Companion) . Association for Computing Machinery, 421–426
Qingpeng Cai, Shuchang Liu, Xueliang Wang, Tianyou Zuo, Wentao Xie, Bin Yang, Dong Zheng, Peng Jiang, and Kun Gai. 2023a · 2023
Later among the works it cites.
Two-Stage Constrained Actor-Critic for Short Video Recommendation. In Proceedings of the ACM Web Conference 2023 . 865–875
Qingpeng Cai, Zhenghai Xue, Chi Zhang, Wanqi Xue, Shuchang Liu, Ruohan Zhan, Xueliang Wang, Tianyou Zuo, Wentao Xie, Dong Zheng, et al · 2023
Later among the works it cites.
Bias and debias in recommender system: A survey and future directions
Jiawei Chen, Hande Dong, Xiang Wang, Fuli Feng, Meng Wang, and Xiangnan He. 2023 · 2023
Later among the works it cites.
Alleviating Matthew Effect of Offline Reinforcement Learning in Interactive Recommendation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23) . 11 pages
Chongming Gao, Kexin Huang, Jiawei Chen, Yuan Zhang, Biao Li, Peng Jiang, Shiqi Wang, Zhong Zhang, and Xiangnan He. 2023 · 2023
Later among the works it cites.
Exploration and Regularization of the Latent Action Space in Recommendation. In Proceedings of the ACM Web Conference 2023 . 833–844
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, and Yongfeng Zhang. 2023 · 2023
Later among the works it cites.
Contrastive State Augmentations for Reinforcement Learning-Based Recommender Systems. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 922–931
Zhaochun Ren, Na Huang, Yidan Wang, Pengjie Ren, Jun Ma, Jiahuan Lei, Xinlei Shi, Hengliang Luo, Joemon Jose, and Xin Xin. 2023 · 2023
Later among the works it cites.
Gymnasium
Mark Towers, Jordan K. Terry, Ariel Kwiatkowski, John U. Balis, Gianluca de Cola, Tristan Deleu, Manuel Goulão, Andreas Kallinteris, Arjun KG, Markus Krimmel, Rodrigo Perez-Vicente, Andrea Pierré, Sander Schulhoff, Jun Jet Tai, Andrew Tan Jin Shen, and Omar G. Younis. 2023 · 2023
Later among the works it cites.
RL4RS: A Real-World Dataset for Reinforcement Learning based Recommender System. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval . 2935–2944
Kai Wang, Zhene Zou, Minghao Zhao, Qilin Deng, Yue Shang, Yile Liang, Runze Wu, Xudong Shen, Tangjie Lyu, and Changjie Fan. 2023 · 2023
Later among the works it cites.
PrefRec: Recommender Systems with Human Preferences for Reinforcing Long-term User Engagement. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’23) . Association for Computing Machinery, New York, NY, USA, 2874–2884
Wanqi Xue, Qingpeng Cai, Zhenghai Xue, Shuo Sun, Shuchang Liu, Dong Zheng, Peng Jiang, Kun Gai, and Bo An. 2023 · 2023
Later among the works it cites.
KuaiSim: A comprehensive simulator for recommender systems
Kesen Zhao, Shuchang Liu, Qingpeng Cai, Xiangyu Zhao, Ziru Liu, Dong Zheng, Peng Jiang, and Kun Gai. 2023 · 2023
Later among the works it cites.
Off-policy deep reinforcement learning without exploration. In International conference on machine learning . PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019b · 2062
Closest in time.