Fetching the paper…
Reading the bibliography…
Many modern commercial sites employ recommender systems to propose relevant content to users.
Better bootstrap confidence intervals
Efron, Bradley · 1987
Earlier work this paper cites.
Updating the inverse of a matrix
Hager, William W · 1989
Earlier work this paper cites.
Dynamic programming and optimal control , volume 1
Bertsekas, Dimitri P · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
Bradtke, Steven J and Barto, Andrew G · 1996
Earlier work this paper cites.
Recommender systems
Resnick, Paul and Varian, Hal R · 1997
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Sutton, Richard S and Barto, Andrew G · 1998
Earlier work this paper cites.
Modeling customer relationships as markov chains
Pfeifer, Phillip E and Carraway, Robert L · 2000
Earlier work this paper cites.
Sequential cost-sensitive decision making with reinforcement learning
Pednault, Edwin, Abe, Naoki, and Zadrozny, Bianca · 2002
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
Peters, Jan, Vijayakumar, Sethu, and Schaal, Stefan · 2003
Earlier work this paper cites.
Joint optimization of customer segmentation and marketing policy to maximize long-term profitability
Jonker, Jedid-Jah, Piersma, Nanda, and Van den Poel, Dirk · 2004
Earlier work this paper cites.
Toward the next generation of recommender systems: A survey of the state-of-the-art and possible extensions
Adomavicius, Gediminas and Tuzhilin, Alexander · 2005
Earlier work this paper cites.
An mdp-based recommender system
Shani, Guy, Heckerman, David, and Brafman, Ronen I · 2005
Cited alongside, same era.
Bayesian probabilistic matrix factorization using markov chain monte carlo
Salakhutdinov, Ruslan and Mnih, Andriy · 2008
Cited alongside, same era.
Matrix factorization techniques for recommender systems
Koren, Yehuda, Bell, Robert, Volinsky, Chris, et al · 2009
Cited alongside, same era.
Bpr: Bayesian personalized ranking from implicit feedback
Rendle, Steffen, Freudenthaler, Christoph, Gantner, Zeno, and Schmidt-Thieme, Lars · 2009
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, Lihong, Chu, Wei, Langford, John, and Schapire, Robert E · 2010
Cited alongside, same era.
Recurrent neural network based language model
Mikolov, Tomas, Karafiát, Martin, Burget, Lukas, Cernockỳ, Jan, and Khudanpur, Sanjeev · 2010
Bayesian nonparametric poisson factorization for recommendation systems
Gopalan, Prem, Ruiz, Francisco J, Ranganath, Rajesh, and Blei, David M · 2014
Later among the works it cites.
Kaggle - coupon purchase prediction challenge by ponpare
Kaggle · 2014
Later among the works it cites.
Weighted importance sampling for off-policy learning with linear function approximation
Mahmood, A Rupam, van Hasselt, Hado P, and Sutton, Richard S · 2014
Later among the works it cites.
Svd free matrix completion with online bias correction for recommender systems
Gogna, Anupriya and Majumdar, Angshul · 2015
Later among the works it cites.
Efficient thompson sampling for online matrix-factorization recommendation
Kawale, Jaya, Bui, Hung H, Kveton, Branislav, Tran-Thanh, Long, and Chawla, Sanjay · 2015
Later among the works it cites.
Ad recommendation systems for life-time value optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nonparametric statistical inference
Gibbons, Jean Dickinson and Chakraborti, Subhabrata · 2011
Cited alongside, same era.
Introduction to recommender systems handbook
Ricci, Francesco, Rokach, Lior, and Shapira, Bracha · 2011
Cited alongside, same era.
Climf: learning to maximize reciprocal rank with collaborative less-is-more filtering
Shi, Yue, Karatzoglou, Alexandros, Baltrunas, Linas, Larson, Martha, Oliver, Nuria, and Hanjalic, Alan · 2012
Cited alongside, same era.
Lifetime value marketing using reinforcement learning
Theocharous, Georgios and Hallak, Assaf · 2013
Cited alongside, same era.
word2vec explained: deriving mikolov et al.’s negative-sampling word-embedding method
Goldberg, Yoav and Levy, Omer · 2014
Cited alongside, same era.
Policy gradient methods for reinforcement learning with function approximation
Sutton, Richard S, McAllester, David A, Singh, Satinder P, Mansour, Yishay, et al
Cited in the paper.
Theocharous, Georgios, Thomas, Philip S, and Ghavamzadeh, Mohammad · 2015
Later among the works it cites.
High-confidence off-policy evaluation
Thomas, Philip S, Theocharous, Georgios, and Ghavamzadeh, Mohammad · 2015
Later among the works it cites.
Benchmarking deep reinforcement learning for continuous control
Duan, Yan, Chen, Xi, Houthooft, Rein, Schulman, John, and Abbeel, Pieter · 2016
Later among the works it cites.
The movielens datasets: History and context
Harper, F Maxwell and Konstan, Joseph A · 2016
Later among the works it cites.
Beyond collaborative filtering: The list recommendation problem
Sar Shalom, Oren, Koenigstein, Noam, Paquet, Ulrich, and Vanchinathan, Hastagiri P · 2016
Later among the works it cites.