Fetching the paper…
Reading the bibliography…
We consider the problem of controlling a fully specified Markov decision process (MDP), also known as the planning problem, when the state space is very large and calculating the optimal policy is intractable.
Dynamic Programming
R. Bellman · 1957
Earlier work this paper cites.
Dynamic Programming and Markov Processes
R. A. Howard · 1960
Earlier work this paper cites.
Linear programming and sequential decisions
A. S. Manne · 1960
Earlier work this paper cites.
Generalized polynomial approximations in Markovian decision processes
P. Schweitzer and A. Seidmann · 1985
Earlier work this paper cites.
Ergodicity of stochastic processes describing the operation of open queueing networks
A. N. Rybko and A. L. Stolyar · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Martin L. Puterman · 1994
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. Tsitsiklis · 1996
Earlier work this paper cites.
Linear program approximations to factored continuous-state markov decision processes
M. Hauskrecht and B. Kveton · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
On constraint sampling in the linear programming approach to approximate dynamic programming
D. P. de Farias and B. Van Roy · 2004
Cited alongside, same era.
Solving factored mdps with continuous and discrete variables
C. Guestrin, M. Hauskrecht, and B. Kveton · 2004
Cited alongside, same era.
Online convex optimization in the bandit setting: gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Cited alongside, same era.
A cost-shaping linear program for average-cost approximate dynamic programming with performance guarantees
D. P. de Farias and B. Van Roy · 2006
Cited alongside, same era.
Dynamic Programming and Optimal Control
D. P. Bertsekas · 2007
Cited alongside, same era.
Dual representations for dynamic programming
T. Wang, D. Lizotte, M. Bowling, and D. Schuurmans · 2008
Toward off-policy learning control with function approximation
H. R. Maei, Cs. Szepesvári, S. Bhatnagar, and R. S. Sutton · 2010
Later among the works it cites.
Online Learning for Linearly Parametrized Control Problems
Y. Abbasi-Yadkori · 2012
Later among the works it cites.
Approximate dynamic programming via a smoothed linear program
V. V. Desai, V. F. Farias, and C. C. Moallemi · 2012
Later among the works it cites.
Approximate linear programming for average cost mdps
M. H. Veatch · 2013
Later among the works it cites.
A proximal stochastic gradient method with progressive variance reduction
Lin Xiao and Tong Zhang · 2014
Later among the works it cites.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Self-normalized processes: Limit theory and Statistical Applications
V. H. de la Peña, T. L. Lai, and Q-M. Shao · 2009
Cited alongside, same era.
Convergent temporal-difference learning with arbitrary smooth function approximation
H. R. Maei, Cs. Szepesvári, S. Bhatnagar, D. Precup, D. Silver, and R. S. Sutton · 2009
Cited alongside, same era.
Constraint relaxation in approximate linear programs
M. Petrik and S. Zilberstein · 2009
Cited alongside, same era.
The linear programming approach to approximate dynamic programming
D. P. de Farias and B. Van Roy
Cited in the paper.
Approximate linear programming for average-cost dynamic programming
D. P. de Farias and B. Van Roy
Cited in the paper.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, Cs. Szepesvári, and E. Wiewiora
Cited in the paper.
Yichen Chen, Lihong Li, and Mengdi Wang · 2018
Later among the works it cites.
A linearly relaxed approximate linear program for markov decision processes
Chandrashekar Lakshminarayanan, Shalabh Bhatnagar, and Csaba Szepesvári · 2018
Later among the works it cites.
Optimizing over a restricted policy class in Markov decision processes
Ershad Banijamali, Yasin Abbasi-Yadkori, Mohammad Ghavamzadeh, and Nikos Vlassis · 2019
Closest in time.