Fetching the paper…
Reading the bibliography…
A fundamental problem in control is to learn a model of a system from observations that is useful for controller synthesis.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Williams, R. J · 1992
Earlier work this paper cites.
Markov Decision Processes: Discrete Stochastic Dynamic Programming
Puterman, M. L · 1994
Earlier work this paper cites.
Robot learning from demonstration
Atkeson, C. G. and Schaal, S · 1997
Earlier work this paper cites.
System Identification: Theory for the User
Ljung, L · 1999
Earlier work this paper cites.
Approximately optimal approximate reinforcement learning
Kakade, S. and Langford, J · 2002
Earlier work this paper cites.
Policy search by dynamic programming
Bagnell, J. A., Ng, A. Y., Kakade, S., and Schneider, J · 2003
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Cesa-Bianchi, N., Conconi, A., and Gentile, C · 2004
Cited alongside, same era.
Iterative linear quadratic regulator design for nonlinear biological movement systems
Li, W. and Todorov, E · 2004
Cited alongside, same era.
Exploration and apprenticeship learning in reinforcement learning
Abbeel, P. and Ng, A. Y · 2005
Cited alongside, same era.
Error limiting reductions between classification tasks
Beygelzimer, A., Dani, V., Hayes, T., Langford, J., and Zadrozny, B · 2005
Cited alongside, same era.
Finite time bounds for sampling based fitted value iteration
Szepesvári, C · 2005
Cited alongside, same era.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Kalai, A., Kale, S., and Agarwal, A · 2006
Cited alongside, same era.
Mind the duality gap: Logarithmic regret algorithms for online optimization
Kakade, S. and Shalev-Shwartz, S · 2008
Later among the works it cites.
Reinforcement learning in finite MDPs: PAC analysis
Strehl, A. L., Li, L., and Littman, M. L · 2009
Later among the works it cites.
Model-based reinforcement learning with nearly tight exploration complexity bounds
Szita, I. and Szepesvári, C · 2010
Later among the works it cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Ross, S., Gordon, G., and Bagnell, J. A · 2011
Later among the works it cites.
Agnostic kwik learning and efficient approximate reinforcement learning
Szita, I. and Szepesvári, C · 2011
Later among the works it cites.
Helicopter learning nose-in funnel., 2012
Ross, S · 2012
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…