Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) with linear function approximation has received increasing attention recently.
Provably efficient exploration in policy optimization
Cai, Q · 1912
Earlier work this paper cites.
Optimism in reinforcement learning with generalized linear function approximation
Wang, Y · 1912
Earlier work this paper cites.
Learning from delayed rewards
Watkins, C. J. C. H · 1989
Earlier work this paper cites.
Du, S. S · 2002
Earlier work this paper cites.
Learning near optimal policies with low inherent bellman error
Zanette, A · 2003
Earlier work this paper cites.
Almost optimal model-free reinforcement learning via reference-advantage decomposition
Zhang, Z · 2004
Earlier work this paper cites.
Local rademacher complexities
Bartlett, P. L · 2005
Earlier work this paper cites.
Model-based reinforcement learning with value-targeted regression
Ayoub, A · 2006
Earlier work this paper cites.
Prediction, learning, and games
Cesa-Bianchi, N · 2006
Earlier work this paper cites.
Pac model-free reinforcement learning
Strehl, A. L · 2006
Earlier work this paper cites.
q q -learning with logarithmic regret
Yang, K · 2006
Earlier work this paper cites.
Provably efficient reinforcement learning for discounted mdps with feature mapping
Zhou, D · 2006
Earlier work this paper cites.
Logarithmic online regret bounds for undiscounted reinforcement learning
Auer, P · 2007
Cited alongside, same era.
Stochastic linear optimization under bandit feedback
Dani, V · 2008
Cited alongside, same era.
On the sample complexity of reinforcement learning with policy space generalization
Mou, W · 2008
Cited alongside, same era.
Optimistic linear programming gives logarithmic regret for irreducible mdps
Tewari, A · 2008
Cited alongside, same era.
Near-optimal regret bounds for reinforcement learning
Jaksch, T · 2010
Cited alongside, same era.
A contextual-bandit approach to personalized news article recommendation
Li, L · 2010
Cited alongside, same era.
Contextual decision processes with low bellman rank are pac-learnable
Jiang, N · 2017
Later among the works it cites.
What doubling tricks can and can’t do for multi-armed bandits
Besson, L · 2018
Later among the works it cites.
Is q-learning provably efficient?
Jin, C · 2018
Later among the works it cites.
Bandit algorithms
Lattimore, T · 2018
Later among the works it cites.
Exploration in structured reinforcement learning
Ok, J · 2018
Later among the works it cites.
Provably efficient q-learning with function approximation via distribution shift error checking oracle
Du, S. S · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved algorithms for linear stochastic bandits
Abbasi-Yadkori, Y · 2011
Cited alongside, same era.
Regret analysis of stochastic and nonstochastic multi-armed bandit problems
Bubeck, S · 2012
Cited alongside, same era.
Nearly minimax optimal reinforcement learning for linear mixture markov decision processes
Zhou, D · 2012
Cited alongside, same era.
Sample complexity of episodic fixed-horizon reinforcement learning
Dann, C · 2015
Cited alongside, same era.
On lower bounds for regret in reinforcement learning
Osband, I · 2016
Cited alongside, same era.
Minimax regret bounds for reinforcement learning
Azar, M. G · 2017
Cited alongside, same era.
Non-asymptotic gap-dependent regret bounds for tabular mdps
Simchowitz, M · 2019
Later among the works it cites.
Introduction to multi-armed bandits
Slivkins, A · 2019
Later among the works it cites.
Sample-optimal parametric q-learning using linearly additive features
Yang, L · 2019
Later among the works it cites.
Tighter problem-dependent regret bounds in reinforcement learning without domain knowledge using value function bounds
Zanette, A · 2019
Later among the works it cites.
Model-based reinforcement learning with value-targeted regression
Jia, Z · 2020
Closest in time.
Provably efficient reinforcement learning with linear function approximation
Jin, C · 2020
Closest in time.