Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) has gained traction for enhancing user long-term experiences in recommender systems by effectively exploring users' interests.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 1999
Earlier work this paper cites.
Improving collaborative filtering for new-users by smart object selection
A Merialdo · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Getting to know you: learning new user preferences in recommender systems
Al Mamunur Rashid, Istvan Albert, Dan Cosley, Shyong K Lam, Sean M McNee, Joseph A Konstan, and John Riedl · 2002
Earlier work this paper cites.
An mdp-based recommender system
Guy Shani, David Heckerman, Ronen I Brafman, and Craig Boutilier · 2005
Earlier work this paper cites.
Improving recommendation lists through topic diversification
Cai-Nicolas Ziegler, Sean M McNee, Joseph A Konstan, and Georg Lausen · 2005
Earlier work this paper cites.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
Learning to rank for information retrieval
Tie-Yan Liu et al · 2009
Earlier work this paper cites.
Active learning for aspect model in recommender systems
Rasoul Karimi, Christoph Freudenthaler, Alexandros Nanopoulos, and Lars Schmidt-Thieme · 2011
Earlier work this paper cites.
Abandoning objectives: Evolution through the search for novelty alone
Joel Lehman and Kenneth O Stanley · 2011
Earlier work this paper cites.
Analysis of thompson sampling for the multi-armed bandit problem
Shipra Agrawal and Navin Goyal · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Variational information maximisation for intrinsically motivated reinforcement learning
Shakir Mohamed and Danilo Jimenez Rezende · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Vime: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Earlier work this paper cites.
A distributional perspective on reinforcement learning
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Earlier work this paper cites.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Earlier work this paper cites.
Active learning in recommendation systems with multi-level user preferences
Yuheng Bu and Kevin Small · 2018
Earlier work this paper cites.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A Efros · 2018
Cited alongside, same era.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Reinforcement learning to rank in e-commerce search engine: Formalization, analysis, and application
Yujing Hu, Qing Da, Anxiang Zeng, Yang Yu, and Yinghui Xu · 2018
Cited alongside, same era.
Self-attentive sequential recommendation
Wang-Cheng Kang and Julian McAuley · 2018
Cited alongside, same era.
User fairness in recommender systems
Fairness-aware explainable recommendation over knowledge graphs
Zuohui Fu, Yikun Xian, Ruoyuan Gao, Jieyu Zhao, Qiaoying Huang, Yingqiang Ge, Shuyuan Xu, Shijie Geng, Chirag Shah, Yongfeng Zhang, et al · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov · 2020
Later among the works it cites.
Fairmatch: A graph-based approach for improving aggregate diversity in recommender systems
Masoud Mansoury, Himan Abdollahpouri, Mykola Pechenizkiy, Bamshad Mobasher, and Robin Burke · 2020
Later among the works it cites.
Effective diversity in population based reinforcement learning
Jack Parker-Holder, Aldo Pacchiano, Krzysztof M Choromanski, and Stephen J Roberts · 2020
Later among the works it cites.
A semi-personalized system for user cold start recommendation on music streaming apps
Léa Briand, Guillaume Salha-Galvan, Walid Bendada, Mathieu Morlon, and Viet-Anh Tran · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jurek Leonhardt, Avishek Anand, and Megha Khosla · 2018
Cited alongside, same era.
Deep reinforcement learning based recommendation with explicit user-item interactions modeling
Feng Liu, Ruiming Tang, Xutao Li, Weinan Zhang, Yunming Ye, Haokun Chen, Huifeng Guo, and Yuzhou Zhang · 2018
Cited alongside, same era.
Deep learning for matching in search and recommendation
Jun Xu, Xiangnan He, and Hang Li · 2018
Cited alongside, same era.
Towards sample efficient reinforcement learning
Yang Yu · 2018
Cited alongside, same era.
Deep reinforcement learning for page-wise recommendations
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang · 2018
Cited alongside, same era.
Drn: A deep reinforcement learning framework for news recommendation
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li · 2018
Cited alongside, same era.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi · 2019
Cited alongside, same era.
Values of user exploration in recommender systems
Minmin Chen, Yuyan Wang, Can Xu, Ya Le, Mohit Sharma, Lee Richardson, Su-Lin Wu, and Ed Chi · 2021
Later among the works it cites.
Distributional soft actor-critic: Off-policy reinforcement learning for addressing value estimation errors
Jingliang Duan, Yang Guan, Shengbo Eben Li, Yangang Ren, Qi Sun, and Bo Cheng · 2021
Later among the works it cites.
Towards long-term fairness in recommendation
Yingqiang Ge, Shuchang Liu, Ruoyuan Gao, Yikun Xian, Yunqi Li, Xiangyu Zhao, Changhua Pei, Fei Sun, Junfeng Ge, Wenwu Ou, et al · 2021
Later among the works it cites.
Aliexpress learning-to-rank: Maximizing online model performance without going online
Guangda Huzhang, Zhenjia Pang, Yongqing Gao, Yawen Liu, Weijie Shen, Wen-Ji Zhou, Qing Da, Anxiang Zeng, Han Yu, Yang Yu, et al · 2021
Later among the works it cites.
Risk-averse offline reinforcement learning
Núria Armengol Urpí, Sebastian Curi, and Andreas Krause · 2021
Later among the works it cites.
Empowering news recommendation with pre-trained language models
Chuhan Wu, Fangzhao Wu, Tao Qi, and Yongfeng Huang · 2021
Later among the works it cites.
Efficient risk-averse reinforcement learning
Ido Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, and Shie Mannor · 2022
Later among the works it cites.
Multi-objective optimization of notifications using offline reinforcement learning
Prakruthi Prabhakar, Yiping Yuan, Guangyu Yang, Wensheng Sun, and Ajith Muralidharan · 2022
Later among the works it cites.
Counteracting user attention bias in music streaming recommendation via reward modification
Xiao Zhang, Sunhao Dai, Jun Xu, Zhenhua Dong, Quanyu Dai, and Ji-Rong Wen · 2022
Later among the works it cites.
Controllable multi-objective re-ranking with policy hypernetworks
Sirui Chen, Yuan Wang, Zijing Wen, Zhiyu Li, Changshuo Zhang, Xiao Zhang, Quan Lin, Cheng Zhu, and Jun Xu · 2023
Later among the works it cites.
Exploration and regularization of the latent action space in recommendation
Shuchang Liu, Qingpeng Cai, Bowen Sun, Yuhao Wang, Ji Jiang, Dong Zheng, Peng Jiang, Kun Gai, Xiangyu Zhao, and Yongfeng Zhang · 2023
Later among the works it cites.
Riku Togashi, Tatsushi Oka, Naoto Ohsaka, and Tetsuro Morimura · 2023
Later among the works it cites.
A survey on the fairness of recommender systems
Yifan Wang, Weizhi Ma, Min Zhang, Yiqun Liu, and Shaoping Ma · 2023
Later among the works it cites.
Teach and explore: A multiplex information-guided effective and efficient reinforcement learning for sequential recommendation
Surong Yan, Chenglong Shi, Haosen Wang, Lei Chen, Ling Jiang, Ruilin Guo, and Kwei-Jay Lin · 2023
Later among the works it cites.
Cold & warm net: Addressing cold-start users in recommender systems
Xiangyu Zhang, Zongqiang Kuang, Zehao Zhang, Fan Huang, and Xianfeng Tan · 2023
Later among the works it cites.
Deep exploration for recommendation systems
Zheqing Zhu and Benjamin Van Roy · 2023
Later among the works it cites.