Fetching the paper…
Reading the bibliography…
Reinforcement learning is well suited for optimizing policies of recommender systems.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
An introduction, richard s. sutton and andrew g. barto, 1998
Reinforcement Learning · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour · 2000
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta · 2001
Earlier work this paper cites.
Minimax differential dynamic programming: An application to robust biped walking
Jun Morimoto and Christopher G Atkeson · 2003
Earlier work this paper cites.
An mdp-based recommender system
Guy Shani, David Heckerman, and Ronen I Brafman · 2005
Earlier work this paper cites.
Matrix factorization techniques for recommender systems
Yehuda Koren, Robert Bell, and Chris Volinsky · 2009
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Learning to control a low-cost manipulator using data-efficient reinforcement learning
Marc Peter Deisenroth, Carl Edward Rasmussen, and Dieter Fox · 2011
Earlier work this paper cites.
A survey on policy search for robotics
Marc Peter Deisenroth, Gerhard Neumann, Jan Peters, et al · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio · 2014
Cited alongside, same era.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Dj-mc: A reinforcement-learning agent for music playlist recommendation
Elad Liebman, Maytal Saar-Tsechansky, and Peter Stone · 2015
Cited alongside, same era.
Partially observable markov decision process for recommender systems
Zhongqi Lu and Qiang Yang · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare · 2016
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning
Philip Thomas and Emma Brunskill · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel · 2017
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Returning is believing: Optimizing long-term user engagement in recommender systems
Qingyun Wu, Hongning Wang, Liangjie Hong, and Yue Shi · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Learning legged swimming gaits from experience
David Meger, Juan Camilo Gamboa Higuera, Anqi Xu, Philippe Giguere, and Gregory Dudek · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Continuous deep q-learning with model-based acceleration
Shixiang Gu, Timothy Lillicrap, Ilya Sutskever, and Sergey Levine · 2016
Cited alongside, same era.
Fast matrix factorization for online recommendation with implicit feedback
Xiangnan He, Hanwang Zhang, Min-Yen Kan, and Tat-Seng Chua · 2016
Cited alongside, same era.
Top-k off-policy correction for a reinforce recommender system
Minmin Chen, Alex Beutel, Paul Covington, Sagar Jain, Francois Belletti, and Ed H Chi
Cited in the paper.
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
Lantao Yu, Weinan Zhang, Jun Wang, and Yong Yu · 2017
Later among the works it cites.
Offline a/b testing for recommender systems
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé · 2018
Later among the works it cites.
Deep dyna-q: Integrating planning for task-completion dialogue policy learning
Baolin Peng, Xiujun Li, Jianfeng Gao, Jingjing Liu, Kam-Fai Wong, and Shang-Yu Su · 2018
Later among the works it cites.
Drn: A deep reinforcement learning framework for news recommendation
Guanjie Zheng, Fuzheng Zhang, Zihan Zheng, Yang Xiang, Nicholas Jing Yuan, Xing Xie, and Zhenhui Li · 2018
Later among the works it cites.
Model-based reinforcement learning for whole-chain recommendations
Xiangyu Zhao, Long Xia, Yihong Zhao, Dawei Yin, and Jiliang Tang · 2019
Closest in time.