Differential Dynamic Programming
David Jacobson and David Mayne · 1970
Earlier work this paper cites.
System identification: theory for the user
Lennart Ljung · 1987
Earlier work this paper cites.
Model predictive control: Theory and practice - a survey
Carlos E. Garcia, David M. Prett, and Manfred Morari · 1989
Earlier work this paper cites.
Using local trajectory optimizers to speed up global optimization in dynamic programming
Christopher G. Atkeson · 1993
Earlier work this paper cites.
Learning to act using real-time dynamic programming
Andrew G. Barto, Steven J. Bradtke, and Satinder P. Singh · 1995
Earlier work this paper cites.
Neuro-dynamic Programming
Dimitri Bertsekas and John Tsitsiklis · 1996
Earlier work this paper cites.
Policy invariance under reward transformations: Theory and application to reward shaping
Andrew Y. Ng, Daishi Harada, and Stuart J. Russell · 1999
Earlier work this paper cites.
Finite-time analysis of the multiarmed bandit problem
Peter Auer, Nicolo Cesa-Bianchi, and Paul Fischer · 2002
Earlier work this paper cites.
Near-optimal reinforcement learning in polynomial time
Michael Kearns and Satinder Singh · 2002
Earlier work this paper cites.
A survey of industrial model predictive control technology
S. Joe Qina and Thomas A. Badgwellb · 2003
Earlier work this paper cites.
Policy search by dynamic programming
J. Andrew Bagnell, Sham M. Kakade, Andrew Y. Ng, and Jeff G. Schneider · 2003
Earlier work this paper cites.
Feedback systems an introduction for scientists and engineers
Karl Johan Åström and Richard M. Murray · 2004
Earlier work this paper cites.
A generalized iterative lqg method for locally-optimal feedback control of constrained nonlinear stochastic systems
Emanuel Todorov and Weiwei Li · 2005
Earlier work this paper cites.
Intrinsically motivated reinforcement learning
Nuttapong Chentanez, Andrew G Barto, and Satinder P Singh · 2005
Earlier work this paper cites.
Finite-time bounds for fitted value iteration
Rémi Munos and Csaba Szepesvári · 2008
Earlier work this paper cites.
Compositionality of optimal control laws
Emanuel Todorov · 2009
Earlier work this paper cites.
A contextual-bandit approach to personalized news article recommendation
Lihong Li, Wei Chu, John Langford, and Robert E. Schapire · 2010
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.