Fetching the paper…
Reading the bibliography…
Latent-state environments with long horizons, such as those faced by recommender systems, pose significant challenges for reinforcement learning (RL).
The learning curve equation
L. L. Thurstone · 1919
Earlier work this paper cites.
The optimal control of partially observable Markov processes over a finite horizon
R. D. Smallwood and E. J. Sondik · 1973
Earlier work this paper cites.
Learning without state-estimation in partially observable Markovian decision processes
S. P. Singh, T. Jaakkola, and M. I. Jordan · 1994
Earlier work this paper cites.
On the Lambert W function
R. M. Corless, G. H. Gonnet, D. E. G. Hare, D. J. Jeffrey, and D. E. Knuth · 1996
Earlier work this paper cites.
Tractable inference for complex stochastic processes
X. Boyen and D. Koller · 1998
Earlier work this paper cites.
Hierarchical solution of Markov decision processes using macro-actions
M. Hauskrecht, N. Meuleau, L. P. Kaelbling, T. Dean, and C. Boutilier · 1998
Earlier work this paper cites.
Flexible decomposition algorithms for weakly coupled Markov decision processes
R. Parr · 1998
Earlier work this paper cites.
Reinforcement Learning Through Gradient Descent
L. C. Baird III · 1999
Earlier work this paper cites.
Between MDPs and Semi-MDPs: Learning, planning, and representing knowledge at multiple temporal scales
R. S. Sutton, D. Precup, and S. P. Singh · 1999
Earlier work this paper cites.
Predictive representations of state
M. L. Littman and R. S. Sutton · 2002
Earlier work this paper cites.
Recent advances in hierarchical reinforcement learning
A. G. Barto and S. Mahadevan · 2003
Earlier work this paper cites.
An MDP-based recommender system
G. Shani, D. Heckerman, and R. I. Brafman · 2005
Cited alongside, same era.
Learning and forgetting models and their applications
M. Y Jaber · 2006
Cited alongside, same era.
Usage-based web recommendations: A reinforcement learning approach
N. Taghipour, A. Kardan, and S. S. Ghidary · 2007
Cited alongside, same era.
Temporal aggregation of univariate and multivariate time series models: A survey
A. Silvestrini and D. Veredas · 2008
Cited alongside, same era.
Signal-to-noise ratio analysis of policy gradient algorithms
J. W. Roberts and R. Tedrake · 2009
Cited alongside, same era.
Budget optimization for online campaigns with positive carryover effects
N. Archak, V. Mirrokni, and S. Muthukrishnan · 2012
Cited alongside, same era.
Predictive state recurrent neural networks
C. Downey, A. Hefny, B. Boots, G. J. Gordon, and B. Li · 2017
Later among the works it cites.
Reinforcement learning with a corrupted reward channel
T. Everitt, V. Krakovna, L. Orseau, and S. Legg · 2017
Later among the works it cites.
On overfitting and asymptotic bias in batch reinforcement learning with partial observability, 2017
V. Francois-Lavet, G. Rabusseau, J. Pineau, D. Ernst, and R. Fonteneau · 2017
Later among the works it cites.
Logistic Markov decision processes
M. Mladenov, C. Boutilier, D. Schuurmans, O. Meshi, G. Elidan, and T. Lu · 2017
Later among the works it cites.
Bridging the gap between value and policy based reinforcement learning
O. Nachum, M. Norouzi, K. Xu, and D. Schuurmans · 2017
Later among the works it cites.
Learning to repeat: Fine-grained action repetition for deep reinforcement learning
S. Sharma, A. S. Lakshminarayanan, and B. Ravindran · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Focusing on the long-term: It’s good for users and business
H. Hohnhold, D. O’Brien, and D. Tang · 2015
Cited alongside, same era.
Trust region policy optimization
J. Schulman, S. L., P. Abbeel, M. I. Jordan, and P. Moritz · 2015
Cited alongside, same era.
Hierarchical relative entropy policy search
C. Daniel, G. Neumann, O. Kroemer, and J. Peters · 2016
Cited alongside, same era.
Taming the noise in reinforcement learning via soft updates
R. Fox, A. Pakman, and N. Tishby · 2016
Cited alongside, same era.
Later among the works it cites.
Deep reinforcement learning for list-wise recommendations
X. Zhao, L. Zhang, Z. Ding, D. Yin, Y. Zhao, and J. Tang · 2017
Later among the works it cites.
Reinforcement learning-based recommender system using biclustering technique
S. Choi, H. Ha, U. Hwang, C. Kim, J. Ha, and S. Yoon · 2018
Later among the works it cites.
Temporal regularization for Markov decision process
P. Thodoroff, A. Durand, J. Pineau, and D. Precup · 2018
Later among the works it cites.
Practical diversified recommendations on youtube with determinantal point processes
M. Wilhelm, A. Ramanathan, A. Bonomo, S. Jain, E. H. Chi, and J. Gillenwater · 2018
Later among the works it cites.