Fetching the paper…
Reading the bibliography…
We consider large-scale Markov decision processes (MDPs) with parameter uncertainty, under the robust MDP paradigm.
The use of confidence or fiducial limits illustrated in the case of the binomial
C. Clopper and E. S. Pearson · 1934
Earlier work this paper cites.
Option pricing: A simplified approach
J. C. Cox, S. A. Ross, and M. Rubinstein · 1979
Earlier work this paper cites.
Markov decision processes: discrete stochastic dynamic programming
M. L. Puterman · 1994
Earlier work this paper cites.
Stable function approximation in dynamic programming
G. J. Gordon · 1995
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Neuro-Dynamic Programming
D. P. Bertsekas and J. N. Tsitsiklis · 1996
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 1998
Earlier work this paper cites.
Solving uncertain Markov decision problems
A. Bagnell, A. Ng, and J. Schneider · 2001
Cited alongside, same era.
Regression methods for pricing complex american-style options
J. N. Tsitsiklis and B. Van Roy · 2001
Cited alongside, same era.
Technical update: Least-squares temporal difference learning
J. A. Boyan · 2002
Cited alongside, same era.
Least-squares policy iteration
M. G. Lagoudakis and R. Parr · 2003
Cited alongside, same era.
Robust dynamic programming
G. N. Iyengar · 2005
Cited alongside, same era.
Robust control of Markov decision processes with uncertain transition matrices
A. Nilim and L. El Ghaoui · 2005
Cited alongside, same era.
Bias and variance approximation in value function estimates
Projected equation methods for approximate solution of large linear systems
D. P. Bertsekas and H. Yu · 2009
Later among the works it cites.
Learning exercise policies for american options
Y. Li, C. Szepesvari, and D. Schuurmans · 2009
Later among the works it cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, C. Szepesvári, and E. Wiewiora · 2009
Later among the works it cites.
Convergence of least squares temporal difference methods under general conditions
H. Yu · 2010
Later among the works it cites.
Approximate Dynamic Programming
W. B. Powell · 2011
Later among the works it cites.
Dynamic Programming and Optimal Control, Vol II
D. P. Bertsekas · 2012
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Mannor, D. Simester, P. Sun, and J. N. Tsitsiklis · 2007
Cited alongside, same era.
Later among the works it cites.