Fetching the paper…
Reading the bibliography…
Tackling large approximate dynamic programming or reinforcement learning problems requires methods that can exploit regularities, or intrinsic structure, of the problem in hand.
Evolving neural network controllers for unstable systems
A. P. Wieland · 1991
Earlier work this paper cites.
Neuro-Dynamic Programming
D.P. Bertsekas and J.N. Tsitsiklis · 1996
Earlier work this paper cites.
A Probabilistic Theory of Pattern Recognition
L. Devroye, L. Györfi, and G. Lugosi · 1996
Earlier work this paper cites.
On-line policy improvement using Monte-Carlo search
G. Tesauro and G.R. Galperin · 1996
Earlier work this paper cites.
An analysis of temporal difference learning with function approximation
J. N. Tsitsiklis and B. Van Roy · 1997
Earlier work this paper cites.
Reinforcement Learning: An Introduction
R.S. Sutton and A.G. Barto · 1998
Earlier work this paper cites.
Uniform Central Limit Theorems
R. M. Dudley · 1999
Earlier work this paper cites.
Infinite-horizon policy-gradient estimation
J. Baxter and P.L. Bartlett · 2001
Earlier work this paper cites.
A natural policy gradient
S. Kakade · 2001
Earlier work this paper cites.
On actor-critic algorithms
V. R. Konda and J. N. Tsitsiklis · 2001
Earlier work this paper cites.
Simulation-based optimization of Markov reward processes
P. Marbach and J.N. Tsitsiklis · 2001
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
P. L. Bartlett and S. Mendelson · 2002
Earlier work this paper cites.
A Distribution-Free Theory of Nonparametric Regression
L. Györfi, M. Kohler, A. Krzyżak, and H. Walk · 2002
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
S. Kakade and J. Langford · 2002
Earlier work this paper cites.
Error bounds for approximate policy iteration
R. Munos · 2003
Earlier work this paper cites.
Reinforcement learning for humanoid robotics
J. Peters, Vijayakumar S., and S. Schaal · 2003
Earlier work this paper cites.
Dynamic multidrug therapies for HIV: optimal and STI control approaches
B.M. Adams, H.T. Banks, Hee-Dae Kwon, and H.T. Tran · 2004
Earlier work this paper cites.
Policy search by dynamic programming
J. A. Bagnell, S. Kakade, A. Y. Ng, and J. Schneider · 2004
Earlier work this paper cites.
The kernel recursive least-squares algorithm
Y. Engel, S. Mannor, and R. Meir · 2004
Earlier work this paper cites.
Local Rademacher complexities
P. L. Bartlett, O. Bousquet, and S. Mendelson · 2005
Cited alongside, same era.
Dynamic programming and suboptimal control: A survey from ADP to MPC
D.P. Bertsekas · 2005
Cited alongside, same era.
Theory of classification: A survey of some recent advances
S. Boucheron, O. Bousquet, and G. Lugosi · 2005
Cited alongside, same era.
A basic formula for online policy gradient algorithms
X.-R. Cao · 2005
Cited alongside, same era.
Tree-based batch mode reinforcement learning
D. Ernst, P. Geurts, and L. Wehenkel · 2005
Cited alongside, same era.
Neural fitted Q iteration – first experiences with a data efficient neural reinforcement learning method
M. Riedmiller · 2005
Cited alongside, same era.
Regularization and feature selection in least-squares temporal difference learning
J. Z. Kolter and A. Y. Ng · 2009
Later among the works it cites.
Fast gradient-descent methods for temporal-difference learning with linear function approximation
R. S. Sutton, H. R. Maei, D. Precup, S. Bhatnagar, D. Silver, Cs. Szepesvári, and E. Wiewiora · 2009
Later among the works it cites.
Kernelized value function approximation for reinforcement learning
G. Taylor and R. Parr · 2009
Later among the works it cites.
Approximate dynamic programming with a fuzzy parameterization
L. Buşoniu, D. Ernst, B. De Schutter, and R. Babuska · 2010
Later among the works it cites.
Error propagation for approximate policy and value iteration
A-m. Farahmand, R. Munos, and Cs. Szepesvári · 2010
Later among the works it cites.
Analysis of a classification-based policy iteration algorithm
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Convexity, classification, and risk bounds
P. L. Bartlett, M. I. Jordan, and J. D. McAuliffe · 2006
Cited alongside, same era.
Clinical data based optimal STI strategies for HIV: a reinforcement learning approach
D. Ernst, G.-B. Stan, J. Gongalves, and L. Wehenkel · 2006
Cited alongside, same era.
Approximate policy iteration with a policy language bias: Solving relational Markov Decision Processes
A. Fern, S. Yoon, and R. Givan · 2006
Cited alongside, same era.
Extremely randomized trees
P. Geurts, D. Ernst, and L. Wehenkel · 2006
Cited alongside, same era.
Fast learning rates for plug-in classifiers
J.-Y. Audibert and A.B. Tsybakov · 2007
Cited alongside, same era.
Bayesian policy gradient algorithms
M. Ghavamzadeh and Y. Engel · 2007
Cited alongside, same era.
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2010
Later among the works it cites.
Algorithms for Reinforcement Learning
Cs. Szepesvári · 2010
Later among the works it cites.
Action-gap phenomenon in reinforcement learning
A-m. Farahmand · 2011
Later among the works it cites.
Model selection in reinforcement learning
A-m. Farahmand and Cs. Szepesvári · 2011
Later among the works it cites.
Classification-based policy iteration with a critic
V. Gabillon, A. Lazaric, M. Ghavamzadeh, and B. Scherrer · 2011
Later among the works it cites.
Finite-sample analysis of Lasso-TD
M. Ghavamzadeh, A. Lazaric, R. Munos, and M. Hoffman · 2011
Later among the works it cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. Gordon, and J. A. Bagnell · 2011
Later among the works it cites.
Value pursuit iteration
A-m. Farahmand and D. Precup · 2012
Later among the works it cites.
Generalized classification-based approximate policy iteration
A-m. Farahmand, D. Precup, and M. Ghavamzadeh · 2012
Later among the works it cites.
Conservative and greedy approaches to classification-based policy iteration
M. Ghavamzadeh and A. Lazaric · 2012
Later among the works it cites.
Finite-sample analysis of least-squares policy iteration
A. Lazaric, M. Ghavamzadeh, and R. Munos · 2012
Later among the works it cites.
On the use of non-stationary policies for stationary infinite-horizon markov decision processes
B. Scherrer and B. Lesner · 2012
Later among the works it cites.
Approximate modified policy iteration
B. Scherrer, M. Ghavamzadeh, V. Gabillon, and M. Geist · 2012
Later among the works it cites.