Fetching the paper…
Reading the bibliography…
We propose a new simple and natural algorithm for learning the optimal Q-value function of a discounted-cost Markov Decision Process (MDP) when the transition kernels are unknown.
Merging of opinions with increasing information
Blackwell, D · 1962
Earlier work this paper cites.
Learning from Delayed Rewards
Watkins, C. J. H · 1989
Earlier work this paper cites.
Topics in controlled Markov chains
Borkar, V. S · 1991
Earlier work this paper cites.
Q-learning
Watkins, C. J · 1992
Earlier work this paper cites.
On the convergence of stochastic iterative dynamic programming algorithms
Jaakkola, T · 1994
Earlier work this paper cites.
Asynchronous stochastic approximation and q-learning
Tsitsiklis, J. N · 1994
Earlier work this paper cites.
Probability theory: an advanced course
Borkar, V. S · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
Bertsekas, D. P · 1996
Earlier work this paper cites.
Exact sampling with coupled Markov chains and applications to statistical mechanics
Propp, J. G · 1996
Cited alongside, same era.
Reinforcement learning: An introduction . Vol. 1
Sutton, R. S · 1998
Cited alongside, same era.
Iterated random functions
Diaconis, P · 1999
Cited alongside, same era.
Finite-sample convergence rates for q-learning and indirect algorithms
Kearns, M. J · 1999
Cited alongside, same era.
Actor-critic–type learning algorithms for Markov decision processes
Konda, V. R · 1999
Cited alongside, same era.
Learning algorithms for Markov decision processes with average cost
Abounadi, J · 2001
Cited alongside, same era.
Approximate Dynamic Programming: Solving the curses of dimensionality . Vol. 703
Powell, W. B · 2007
Later among the works it cites.
Stochastic approximation: a dynamical systems viewpoint
Borkar, V. S · 2008
Later among the works it cites.
Markov chains and mixing times
Levin, D. A · 2009
Later among the works it cites.
Algorithms for reinforcement learning
Szepesvári, C · 2010
Later among the works it cites.
Dynamic Programming and Optimal Control vol. 2, 4th ed
Bertsekas, D. P · 2012
Later among the works it cites.
Haskell, W. B · 2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic approximation for nonexpansive maps: Application to q-learning algorithms
Abounadi, J · 2002
Cited alongside, same era.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 2005
Cited alongside, same era.
On boundedness of q-learning iterates for stochastic shortest path problems
Yu, H · 2013
Later among the works it cites.