Fetching the paper…
Reading the bibliography…
We propose empirical dynamic programming algorithms for Markov decision processes (MDPs).
A stochastic approximation method
Robbins, H., S. Monro. 1951 · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Kiefer, J., J. Wolfowitz. 1952 · 1952
Earlier work this paper cites.
Stochastic games
Shapley, Lloyd S. 1953 · 1953
Earlier work this paper cites.
Dynamic Programming
Bellman, R. 1957 · 1957
Earlier work this paper cites.
Functional approximations and dynamic programming
Bellman, R., S. Dreyfus. 1959 · 1959
Earlier work this paper cites.
Steps toward artificial intelligence
Minsky, M. 1961 · 1961
Earlier work this paper cites.
Introduction to dynamic programming
Nemhauser, George L. 1966 · 1966
Earlier work this paper cites.
Dynamic Probabilistic Systems: Vol.: 2.: Semi-Markov and Decision Processes
Howard, R. 1971 · 1971
Earlier work this paper cites.
Beyond regression: New tools for prediction and analysis in the behavioral sciences
Werbos, P. 1974 · 1974
Earlier work this paper cites.
Analysis of recursive stochastic algorithms
Ljung, L. 1977 · 1977
Earlier work this paper cites.
Stochastic approximation methods for constrained and unconstrained systems
Kushner, H. J., D. S. Clark. 1978 · 1978
Earlier work this paper cites.
Approximations of dynamic programs, i
Whitt, W. 1978 · 1978
Earlier work this paper cites.
Approximations of dynamic programs, ii
Whitt, W. 1979 · 1979
Earlier work this paper cites.
Associative search network: A reinforcement learning associative memory
Barto, A., S. Sutton, P. Brouwer. 1981 · 1981
Earlier work this paper cites.
A class of rapidly convergent algorithms for learning automata
Thathachar, M.A. L., P. S. Sastry. 1985 · 1985
Earlier work this paper cites.
The complexity of markov decision processes
Papadimitriou, C. H., J. Tsitsiklis. 1987 · 1987
Earlier work this paper cites.
Q-learning
Watkins, C., P. Dayan. 1992 · 1992
Cited alongside, same era.
Neuro-Dynamic Programming
Bertsekas, D., J. Tsitsiklis. 1996 · 1996
Cited alongside, same era.
Finite time analysis of the pursuit algorithm for learning automata
Rajaraman, K., P. S. Sastry. 1996 · 1996
Cited alongside, same era.
Stochastic processes
Ross, Sheldon M. 1996 · 1996
Cited alongside, same era.
Reinforcement learning: An introduction
Sutton, R. S., A. G. Barto. 1998 · 1998
Cited alongside, same era.
Actor-critic–type learning algorithms for markov decision processes
Konda, V. R., V. S. Borkar. 1999 · 1999
Cited alongside, same era.
The ode method for convergence of stochastic approximation and reinforcement learning
Learning rates for q-learning
Even-Dar, E., Y. Mansour. 2004 · 2004
Later among the works it cites.
Convergence rate of linear two-time-scale stochastic approximation
Konda, V. R., J. Tsitsiklis. 2004 · 2004
Later among the works it cites.
Simulation-Based Algorithms for Markov Decision Processes
Chang, H. S., M. Fu, J. Hu, S. I. Marcus. 2006 · 2006
Later among the works it cites.
Simulation-based uniform value function estimates of markov decision processes
Jain, R., P. Varaiya. 2006 · 2006
Later among the works it cites.
A survey of some simulation-based algorithms for markov decision processes
Chang, H. S., M. Fu, J. Hu, S. I. Marcus. 2007 · 2007
Later among the works it cites.
Approximate Dynamic Programming: Solving the curses of dimensionality
Powell, W. B. 2007 · 2007
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Borkar, V. S., S. P. Meyn. 2000 · 2000
Cited alongside, same era.
Learning algorithms for markov decision processes with average cost
Abounadi, J., D. Bertsekas, V. S. Borkar. 2001 · 2001
Cited alongside, same era.
Q-learning for risk-sensitive control
Borkar, V. S. 2002 · 2002
Cited alongside, same era.
Monotone random systems theory and applications
Chueshov, I. 2002 · 2002
Cited alongside, same era.
Comparison methods for stochastic models and risks
Müller, A., D. Stoyan. 2002 · 2002
Cited alongside, same era.
Convergence of simulation-based policy iteration
Cooper, W., S. Henderson, M. Lewis. 2003 · 2003
Cited alongside, same era.
Stochastic orders
Shaked, M., J. G. Shanthikumar. 2007 · 2007
Later among the works it cites.
Approximate fixed point iteration with an application to infinite horizon markov decision processes
Almudevar, A. 2008 · 2008
Later among the works it cites.
Stochastic approximation: A dynamical systems viewpoint
Borkar, V. S. 2008 · 2008
Later among the works it cites.
Neural network learning: Theoretical foundations
Anthony, M., P. Bartlett. 2009 · 2009
Later among the works it cites.
Markov chains and mixing times
Levin, D., Y. Peres, E. L. Wilmer. 2009 · 2009
Later among the works it cites.
Simulation-based optimization of markov decision processes: An empirical process theory approach
Jain, R., P. Varaiya. 2010 · 2010
Later among the works it cites.
Approximate policy iteration: A survey and some new methods
Bertsekas, D. 2011 · 2011
Later among the works it cites.
Performance guarantees for empirical markov decision processes with applications to multiperiod inventory models
Cooper, W., B. Rangarajan. 2012 · 2012
Later among the works it cites.
Learning automata: an introduction
Narendra, K. S., M.A.L. Thathachar. 2012 · 2012
Later among the works it cites.