Fetching the paper…
Reading the bibliography…
Consider the problem of approximating the optimal policy of a Markov decision process (MDP) by sampling state transitions.
Dynamic Programming
Richard Bellman · 1957
Earlier work this paper cites.
Les problemes de decisions sequentielles
Guy De Ghellinck · 1960
Earlier work this paper cites.
Dynamic programming and Markov processes
Ronald A. Howard · 1960
Earlier work this paper cites.
A probabilistic production and inventory problem
F d’Epenoux · 1963
Earlier work this paper cites.
Solving h-horizon, stationary markov decision problems in time proportional to log (h)
Paul Tseng · 1990
Earlier work this paper cites.
Dynamic programming and optimal control
Dimitri P Bertsekas · 1995
Earlier work this paper cites.
Neuro-dynamic programming: an overview
Dimitri P Bertsekas and John N Tsitsiklis · 1995
Earlier work this paper cites.
On the complexity of solving Markov decision problems
Michael L Littman, Thomas L Dean, and Leslie Pack Kaelbling · 1995
Earlier work this paper cites.
Finite-sample convergence rates for q-learning and indirect algorithms
Michael J Kearns and Satinder P Singh · 1999
Earlier work this paper cites.
On the complexity of policy iteration
Yishay Mansour and Satinder Singh · 1999
Earlier work this paper cites.
A sparse sampling algorithm for near-optimal planning in large markov decision processes
Michael Kearns, Yishay Mansour, and Andrew Y Ng · 2002
Cited alongside, same era.
On the sample complexity of reinforcement learning
Sham M Kakade · 2003
Cited alongside, same era.
An efficient stochastic approximation algorithm for stochastic saddle point problems
Arkadi Nemirovski and Reuven Y Rubinstein · 2005
Cited alongside, same era.
A new complexity result on solving the Markov decision problem
Yinyu Ye · 2005
Cited alongside, same era.
Solving variational inequalities with stochastic mirror-prox algorithm
Anatoli Juditsky, Arkadi Nemirovski, Claire Tauvel, et al · 2011
Cited alongside, same era.
The simplex and policy-iteration methods are strongly polynomial for the Markov decision problem with a fixed discount rate
PAC bounds for discounted MDPs
Tor Lattimore and Marcus Hutter · 2012
Later among the works it cites.
Abstract dynamic programming
Dimitri P Bertsekas · 2013
Later among the works it cites.
Improved and generalized upper bounds on the complexity of policy iteration
Bruno Scherrer · 2013
Later among the works it cites.
Markov decision processes: discrete stochastic dynamic programming
Martin L Puterman · 2014
Later among the works it cites.
Stochastic primal-dual methods and sample complexity of reinforcement learning
Yichen Chen and Mengdi Wang · 2016
Later among the works it cites.
An online primal-dual method for discounted Markov decision processes
Mengdi Wang and Yichen Chen · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yinyu Ye · 2011
Cited alongside, same era.
On the sample complexity of reinforcement learning with a generative model
Mohammad Gheshlaghi Azar, Rémi Munos, and Bert Kappen · 2012
Cited alongside, same era.
Sublinear optimization for machine learning
Kenneth L Clarkson, Elad Hazan, and David P Woodruff · 2012
Cited alongside, same era.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Cited alongside, same era.
Lower bound on the computational complexity of discounted markov decision problems
Yichen Chen and Mengdi Wang · 2017
Closest in time.
Variance reduced value iteration and faster algorithms for solving markov decision processes
A. Sidford, M. Wang, C. Wu, and Y. Ye · 2017
Closest in time.
Mengdi Wang · 2017
Closest in time.