Fetching the paper…
Reading the bibliography…
Industrial recommender systems deal with extremely large action spaces -- many millions of items to recommend.
Asynchronous methods for deep reinforcement learning. In International conference on machine learning . 1928–1937
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu. 2016 · 1937
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams. 1992 · 1992
Earlier work this paper cites.
Temporal difference learning and TD-Gammon
Gerald Tesauro. 1995 · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction
Richard S Sutton, Andrew G Barto, et al · 1998
Earlier work this paper cites.
Eligibility traces for off-policy policy evaluation
Doina Precup. 2000 · 2000
Earlier work this paper cites.
Policy gradient methods for reinforcement learning with function approximation. In Advances in neural information processing systems . 1057–1063
Richard S Sutton, David A McAllester, Satinder P Singh, and Yishay Mansour. 2000 · 2000
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Nikolaus Hansen and Andreas Ostermeier. 2001 · 2001
Earlier work this paper cites.
Off-policy temporal-difference learning with function approximation. In ICML . 417–424
Doina Precup, Richard S Sutton, and Sanjoy Dasgupta. 2001 · 2001
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer. 2002 · 2002
Earlier work this paper cites.
Quick Training of Probabilistic Neural Nets by Importance Sampling.. In AISTATS . 1–9
Yoshua Bengio, Jean-Sébastien Senécal, et al · 2003
Earlier work this paper cites.
Cortical substrates for exploratory decisions in humans
Nathaniel D Daw, John P O’doherty, Peter Dayan, Ben Seymour, and Raymond J Dolan. 2006 · 2006
Earlier work this paper cites.
The netflix prize. In Proceedings of KDD cup and workshop , Vol. 2007. New York, NY, USA, 35
James Bennett, Stan Lanning, et al · 2007
Earlier work this paper cites.
Collaborative filtering for implicit feedback datasets. In ICDM . Ieee, 263–272
Yifan Hu, Yehuda Koren, and Chris Volinsky. 2008 · 2008
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou. 2010 · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th international conference on World wide web . ACM, 661–670
Lihong Li, Wei Chu, John Langford, and Robert E Schapire. 2010 · 2010
Earlier work this paper cites.
Learning from logged implicit exploration data. In Advances in Neural Information Processing Systems . 2217–2225
Alex Strehl, John Langford, Lihong Li, and Sham M Kakade. 2010 · 2010
Earlier work this paper cites.
An empirical evaluation of thompson sampling. In Advances in neural information processing systems . 2249–2257
Olivier Chapelle and Lihong Li. 2011 · 2011
Cited alongside, same era.
Counterfactual reasoning and learning systems: The example of computational advertising
Léon Bottou, Jonas Peters, Joaquin Quiñonero-Candela, Denis X Charles, D Max Chickering, Elon Portugaly, Dipankar Ray, Patrice Simard, and Ed Snelson. 2013 · 2013
Cited alongside, same era.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters. 2013 · 2013
Cited alongside, same era.
Guided policy search. In International Conference on Machine Learning . 1–9
Sergey Levine and Vladlen Koltun. 2013 · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller. 2013 · 2013
Cited alongside, same era.
A recurrent neural network without chaos
Thomas Laurent and James von Brecht. 2016 · 2016
Later among the works it cites.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel. 2016 · 2016
Later among the works it cites.
Safe and efficient off-policy reinforcement learning. In Advances in Neural Information Processing Systems . 1054–1062
Rémi Munos, Tom Stepleton, Anna Harutyunyan, and Marc Bellemare. 2016 · 2016
Later among the works it cites.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Later among the works it cites.
Improved recurrent neural networks for session-based recommendations. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems . ACM, 17–22
Yong Kiam Tan, Xinxing Xu, and Yong Liu. 2016 · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Monte Carlo theory, methods and examples
Art B. Owen. 2013 · 2013
Cited alongside, same era.
Taming the monster: A fast and simple algorithm for contextual bandits. In International Conference on Machine Learning . 1638–1646
Alekh Agarwal, Daniel Hsu, Satyen Kale, John Langford, Lihong Li, and Robert Schapire. 2014 · 2014
Cited alongside, same era.
Empirical evaluation of gated recurrent neural networks on sequence modeling
Junyoung Chung, Caglar Gulcehre, KyungHyun Cho, and Yoshua Bengio. 2014 · 2014
Cited alongside, same era.
Session-based recommendations with recurrent neural networks
Balázs Hidasi, Alexandros Karatzoglou, Linas Baltrunas, and Domonkos Tikk. 2015 · 2015
Cited alongside, same era.
Bandits and recommender systems. In International Workshop on Machine Learning, Optimization and Big Data . Springer, 325–336
Jérémie Mary, Romaric Gaudel, and Philippe Preux. 2015 · 2015
Cited alongside, same era.
Trust region policy optimization. In International Conference on Machine Learning . 1889–1897
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz. 2015 · 2015
Cited alongside, same era.
Autorec: Autoencoders meet collaborative filtering. In Proceedings of the 24th International Conference on World Wide Web . ACM, 111–112
Suvash Sedhain, Aditya Krishna Menon, Scott Sanner, and Lexing Xie. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Data-efficient off-policy policy evaluation for reinforcement learning. In ICML . 2139–2148
Philip Thomas and Emma Brunskill. 2016 · 2016
Later among the works it cites.
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. 2017 · 2017
Later among the works it cites.
Neural collaborative filtering. In Proceedings of the 26th International Conference on World Wide Web . International World Wide Web Conferences Steering Committee, 173–182
Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie, Xia Hu, and Tat-Seng Chua. 2017 · 2017
Later among the works it cites.
Neural survival recommender. In WSDM . ACM, 515–524
How Jing and Alexander J Smola. 2017 · 2017
Later among the works it cites.
Unbiased learning-to-rank with biased feedback. In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining . ACM, 781–789
Thorsten Joachims, Adith Swaminathan, and Tobias Schnabel. 2017 · 2017
Later among the works it cites.
Off-policy evaluation for slate recommendation. In Advances in Neural Information Processing Systems
Adith Swaminathan, Akshay Krishnamurthy, Alekh Agarwal, Miro Dudik, John Langford, Damien Jose, and Imed Zitouni. 2017 · 2017
Later among the works it cites.
Recurrent recommender networks. In Proceedings of the tenth ACM international conference on web search and data mining . ACM, 495–503
Chao-Yuan Wu, Amr Ahmed, Alex Beutel, Alexander J Smola, and How Jing. 2017 · 2017
Later among the works it cites.
Latent Cross: Making Use of Context in Recurrent Recommender Systems. In WSDM . ACM, 46–54
Alex Beutel, Paul Covington, Sagar Jain, Can Xu, Jia Li, Vince Gatto, and Ed H Chi. 2018 · 2018
Closest in time.
Offline A/B testing for Recommender Systems. In WSDM . ACM, 198–206
Alexandre Gilotte, Clément Calauzènes, Thomas Nedelec, Alexandre Abraham, and Simon Dollé. 2018 · 2018
Closest in time.
Short-term satisfaction and long-term coverage: Understanding how users tolerate algorithmic exploration. In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining . ACM, 513–521
Tobias Schnabel, Paul N Bennett, Susan T Dumais, and Thorsten Joachims. 2018 · 2018
Closest in time.
Deep Reinforcement Learning for Page-wise Recommendations
Xiangyu Zhao, Long Xia, Liang Zhang, Zhuoye Ding, Dawei Yin, and Jiliang Tang. 2018 · 2018
Closest in time.