Fetching the paper…
Reading the bibliography…
Modern Reinforcement Learning (RL) is commonly applied to practical problems with an enormous number of states, where function approximation must be deployed to approximate either the value function or the policy.
Reinforcement leaning in feature space: Matrix bandit, kernels, and regret bound
L. F. Yang and M. Wang · 1905
Earlier work this paper cites.
Learning to predict by the methods of temporal differences
R. S. Sutton · 1988
Earlier work this paper cites.
Complexity analysis of real-time reinforcement learning
S. Koenig and R. G. Simmons · 1993
Earlier work this paper cites.
Residual algorithms: Reinforcement learning with function approximation
L. Baird · 1995
Earlier work this paper cites.
Generalization in reinforcement learning: Safely approximating the value function
J. A. Boyan and A. W. Moore · 1995
Earlier work this paper cites.
Linear least-squares algorithms for temporal difference learning
S. J. Bradtke and A. G. Barto · 1996
Earlier work this paper cites.
Analysis of temporal-diffference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Using confidence bounds for exploitation-exploration trade-offs
P. Auer · 2002
Earlier work this paper cites.
PAC model-free reinforcement learning
A. L. Strehl, L. Li, E. Wiewiora, J. Langford, and M. L. Littman · 2006
Earlier work this paper cites.
Q-learning with linear function approximation
F. S. Melo and M. I. Ribeiro · 2007
Earlier work this paper cites.
Stochastic linear optimization under bandit feedback
V. Dani, T. P. Hayes, and S. M. Kakade · 2008
Earlier work this paper cites.
Near-optimal regret bounds for reinforcement learning
T. Jaksch, R. Ortner, and P. Auer · 2010
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
L. Li, W. Chu, J. Langford, and R. E. Schapire · 2010
Earlier work this paper cites.
Linearly parameterized bandits
P. Rusmevichientong and J. N. Tsitsiklis · 2010
Earlier work this paper cites.
Algorithms for reinforcement learning
C. Szepesvári · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
R. Vershynin · 2010
Earlier work this paper cites.
Regret bounds for the adaptive control of linear quadratic systems
Y. Abbasi-Yadkori and C. Szepesvári · 2011
Earlier work this paper cites.
Improved algorithms for linear stochastic bandits
Y. Abbasi-Yadkori, D. Pál, and C. Szepesvári · 2011
Cited alongside, same era.
Speedy Q-learning
M. G. Azar, R. Munos, M. Ghavamzadaeh, and H. J. Kappen · 2011
Cited alongside, same era.
Contextual bandits with linear payoff functions
W. Chu, L. Li, L. Reyzin, and R. Schapire · 2011
Cited alongside, same era.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
M. G. Azar, R. Munos, and B. Kappen · 2012
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
S. Bubeck and N. Cesa-Bianchi · 2012
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
M. G. Azar, I. Osband, and R. Munos · 2017
Later among the works it cites.
Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning
C. Dann, T. Lattimore, and E. Brunskill · 2017
Later among the works it cites.
Contextual decision processes with low bellman rank are pac-learnable
N. Jiang, A. Krishnamurthy, A. Agarwal, J. Langford, and R. E. Schapire · 2017
Later among the works it cites.
Efficient reinforcement learning in deterministic systems with value function generalization
Z. Wen and B. Van Roy · 2017
Later among the works it cites.
Improved regret bounds for Thompson sampling in linear quadratic control problems
M. Abeille and A. Lazaric · 2018
Later among the works it cites.
Efficient exploration through bayesian deep q-networks
K. Azizzadenesheli, E. Brunskill, and A. Anandkumar · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Reinforcement learning in robotics: A survey
J. Kober and J. Peters · 2012
Cited alongside, same era.
PAC bounds for discounted MDPs
T. Lattimore and M. Hutter · 2012
Cited alongside, same era.
Playing Atari with deep reinforcement learning
V. Mnih, K. Kavukcuoglu, D. Silver, A. Graves, I. Antonoglou, D. Wierstra, and M. Riedmiller · 2013
Cited alongside, same era.
Efficient exploration and value function generalization in deterministic systems
Z. Wen and B. Van Roy · 2013
Cited alongside, same era.
Generalization and exploration via randomized value functions
I. Osband, B. Van Roy, and Z. Wen · 2014
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
M. L. Puterman · 2014
Cited alongside, same era.
Later among the works it cites.
Regret bounds for robust adaptive control of the linear quadratic regulator
S. Dean, H. Mania, N. Matni, B. Recht, and S. Tu · 2018
Later among the works it cites.
Is Q-learning provably efficient?
C. Jin, Z. Allen-Zhu, S. Bubeck, and M. I. Jordan · 2018
Later among the works it cites.
Bandit algorithms
T. Lattimore and C. Szepesvári · 2018
Later among the works it cites.
Variance reduced value iteration and faster algorithms for solving Markov decision processes
A. Sidford, M. Wang, X. Wu, and Y. Ye · 2018
Later among the works it cites.
Model-free linear quadratic control via reduction to expert prediction
Y. Abbasi-Yadkori, N. Lazic, and C. Szepesvári · 2019
Closest in time.
Learning linear-quadratic regulators efficiently with only T \sqrt{T} regret
A. Cohen, T. Koren, and Y. Mansour · 2019
Closest in time.
S. S. Du, Y. Luo, R. Wang, and H. Zhang · 2019
Closest in time.
Variance-reduced Q-learning is minimax optimal
M. J. Wainwright · 2019
Closest in time.
Towards practical Lipschitz stochastic bandits
T. Wang, W. Ye, D. Geng, and C. Rudin · 2019
Closest in time.
Lipschitz bandit optimization with improved efficiency
X. Zhu and D. B. Dunson · 2019
Closest in time.