Fetching the paper…
Reading the bibliography…
Scaling reinforcement learning (RL) to recommender systems (RS) is promising since maximizing the expected cumulative rewards for RL agents meets the objective of RS, i.e., improving customers' long-term satisfaction.
Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen, Craig Ferguson, Àgata Lapedriza, Noah Jones, Shixiang Gu, and Rosalind W. Picard. 2019 · 1907
Earlier work this paper cites.
Deep Reinforcement Learning for Online Advertising in Recommender Systems
Xiangyu Zhao, Changsheng Gu, Haoshenglun Zhang, Xiaobing Liu, Xiwang Yang, and Jiliang Tang. 2019 · 1909
Earlier work this paper cites.
Behavior Regularized Offline Reinforcement Learning
Yifan Wu, George Tucker, and Ofir Nachum. 2019 · 1911
Earlier work this paper cites.
Statistical Estimates and Transformed Beta Variables
Gunnar Blom. 1958 · 1958
Earlier work this paper cites.
Expected values of normal order statistics
H. Leon Harter. 1961 · 1961
Earlier work this paper cites.
Expected normal order statistics (exact and approximate)
J. P. Royston. 1982 · 1982
Earlier work this paper cites.
Policy Invariance Under Reward Transformations: Theory and Application to Reward Shaping. In (ICML 1999), Bled, Slovenia, June 27 - 30, 1999 . Morgan Kaufmann, 278–287
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell. 1999 · 1999
Earlier work this paper cites.
Actor-critic Algorithms
Vijaymohan Konda. 2002 · 2002
Earlier work this paper cites.
D4RL: Datasets for Deep Data-Driven Reinforcement Learning
Justin Fu, Aviral Kumar, Ofir Nachum, George Tucker, and Sergey Levine. 2020 · 2004
Earlier work this paper cites.
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
Sergey Levine, Aviral Kumar, George Tucker, and Justin Fu. 2020 · 2005
Earlier work this paper cites.
The Importance of Pessimism in Fixed-Dataset Policy Optimization
Jacob Buckman, Carles Gelada, and Marc G. Bellemare. 2020 · 2009
Earlier work this paper cites.
Is Pessimism Provably Efficient for Offline RL?
Ying Jin, Zhuoran Yang, and Zhaoran Wang. 2020 · 2012
Cited alongside, same era.
Self-Attentive Sequential Recommendation. In ICDM 2018, Singapore, November 17-20, 2018 . IEEE Computer Society, 197–206
Wangcheng Kang and Julian J. McAuley. 2018 · 2018
Cited alongside, same era.
Data center cooling using model-predictive control. In NeurIPS 2018, December 3-8, 2018, Montréal, Canada . 3818–3827
Nevena Lazic, Craig Boutilier, Tyler Lu, Eehern Wong, Binz Roy, M. K. Ryu, and Greg Imwalle. 2018 · 2018
Cited alongside, same era.
Personalized Top-N Sequential Recommendation via Convolutional Sequence Embedding. In WSDM 2018, Marina Del Rey, CA, USA, February 5-9, 2018 . ACM, 565–573
Jiaxi Tang and Ke Wang. 2018 · 2018
Cited alongside, same era.
Deep Reinforcement Learning for Page-wise Recommendations. In RecSys 2018, Vancouver, BC, Canada, October 2-7, 2018 . ACM, 95–103
Conservative Q-Learning for Offline Reinforcement Learning. In NeurIPS 2020, December 6-12, 2020, virtual
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine. 2020 · 2020
Later among the works it cites.
Off-Policy Recommendation System Without Exploration. In PAKDD 2020, Singapore, May 11-14, 2020, Proceedings, Part I (Lecture Notes in Computer Science, Vol. 12084) . Springer, 16–27
Chengwei Wang, Tengfei Zhou, Chen Chen, Tianlei Hu, and Gang Chen. 2020 · 2020
Later among the works it cites.
Self-Supervised Reinforcement Learning for Recommender Systems. In SIGIR 2020, Virtual Event, China, July 25-30, 2020 . ACM, 931–940
Xin Xin, Alexandros Karatzoglou, Ioannis Arapakis, and Joemon M. Jose. 2020 · 2020
Later among the works it cites.
Mastering Complex Control in MOBA Games with Deep Reinforcement Learning. In AAAI, New York, NY, USA, February 7-12, 2020 . AAAI Press, 6672–6679
Deheng Ye, Zhao Liu, Mingfei Sun, Bei Shi, Peilin Zhao, Hao Wu, Hongsheng Yu, Shaojie Yang, Xipeng Wu, Qingwei Guo, Qiaobo Chen, Yinyuting Yin, Hao Zhang, Tengfei Shi, Liang Wang, Qiang Fu, Wei Yang, and Lanxiao Huang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018a · 2018
Cited alongside, same era.
Recommendations with Negative Feedback via Pairwise Deep Reinforcement Learning. In KDD 2018, London, UK, August 19-23, 2018 . ACM, 1040–1048
Xiangyu Zhao, Liang Zhang, Zhuoye Ding, Long Xia, Jiliang Tang, and Dawei Yin. 2018b · 2018
Cited alongside, same era.
DRN: A Deep Reinforcement Learning Framework for News Recommendation. In WWW 2018, Lyon, France, April 23-27, 2018 . ACM, 167–176
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li. 2018 · 2018
Cited alongside, same era.
Stabilizing Off-Policy Q-Learning via Bootstrapping Error Reduction. In NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada . 11761–11771
Aviral Kumar, Justin Fu, Matthew Soh, George Tucker, and Sergey Levine. 2019 · 2019
Cited alongside, same era.
A Simple Convolutional Generative Network for Next Item Recommendation. In WSDM 2019, Melbourne, VIC, Australia, February 11-15, 2019 . ACM, 582–590
Fajie Yuan, Alexandros Karatzoglou, Ioannis Arapakis, Joemon M. Jose, and Xiangnan He. 2019 · 2019
Cited alongside, same era.
An Optimistic Perspective on Offline Reinforcement Learning. In ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119) . PMLR, 104–114
Rishabh Agarwal, Dale Schuurmans, and Mohammad Norouzi. 2020 · 2020
Cited alongside, same era.
Cost-Sensitive Portfolio Selection via Deep Reinforcement Learning
Yifan Zhang, Peilin Zhao, Qingyao Wu, Bin Li, Junzhou Huang, and Mingkui Tan. 2022 · 2020
Later among the works it cites.
Reinforcement learning based recommender systems: A survey
M Mehdi Afsar, Trafford Crump, and Behrouz Far. 2021 · 2021
Closest in time.
A Minimalist Approach to Offline Reinforcement Learning
Scott Fujimoto and Shixiang Shane Gu. 2021 · 2021
Closest in time.
Offline Reinforcement Learning with Fisher Divergence Critic Regularization. In ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139) . PMLR, 5774–5783
Ilya Kostrikov, Rob Fergus, Jonathan Tompson, and Ofir Nachum. 2021 · 2021
Closest in time.
Uncertainty Weighted Actor-Critic for Offline Reinforcement Learning. In ICML 2021, 18-24 July 2021, Virtual Event (Proceedings of Machine Learning Research, Vol. 139) . PMLR, 11319–11328
Yue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind, Jian Zhang, Ruslan Salakhutdinov, and Hanlin Goh. 2021 · 2021
Closest in time.
Off-Policy Deep Reinforcement Learning without Exploration. In ICML 2019, 9-15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research, Vol. 97) . PMLR, 2052–2062
Scott Fujimoto, David Meger, and Doina Precup. 2019 · 2062
Closest in time.