Fetching the paper…
Reading the bibliography…
In this paper, we introduce an actor-critic algorithm called Deep Value Model Predictive Control (DMPC), which combines model-based trajectory optimization with value function estimation.
A second-order gradient method for determining optimal trajectories of non-linear discrete-time systems
D. Mayne · 1966
Earlier work this paper cites.
Apprenticeship learning via inverse reinforcement learning
P. Abbeel and A. Y. Ng · 2004
Earlier work this paper cites.
An introduction to stochastic control theory, path integrals and reinforcement learning
H. J. Kappen · 2007
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
E. Theodorou, J. Buchli, and S. Schaal · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
S. Ross, G. J. Gordon, and J. Andrew Bagnell · 2010
Earlier work this paper cites.
Inverse optimal control with linearly-solvable mdps
K. Dvijotham and E. Todorov · 2010
Earlier work this paper cites.
Model predictive quadrotor indoor position control
K. Alexis, C. Papachristos, G. Nikolakopoulos, and A. Tzes · 2011
Earlier work this paper cites.
Guided policy search
S. Levine and V. Koltun · 2013
Earlier work this paper cites.
Whole-body model-predictive control applied to the hrp-2 humanoid
J. Koenemann, A. D. Prete, Y. Tassa, E. Todorov, O. Stasse, M. Bennewitz, and N. Mansard · 2015
Cited alongside, same era.
Path integral control and state-dependent feedback
S. Thijssen and H. Kappen · 2015
Cited alongside, same era.
Universal value function approximators
T. Schaul, D. Horgan, K. Gregor, and D. Silver · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra · 2016
Cited alongside, same era.
Adaptive importance sampling for control and inference
H. J. Kappen and H. C. Ruiz · 2016
Cited alongside, same era.
Aggressive driving with model predictive path integral control
G. Williams, P. Drews, B. Goldfain, J. M. Rehg, and E. A. Theodorou · 2016
An efficient optimal planning and control framework for quadrupedal locomotion
F. Farshidian, M. Neunert, A. W. Winkler, G. Rey, and J. Buchli · 2017
Later among the works it cites.
Reinforcement Learning: An Introduction
R. S. Sutton and A. G. Barto · 2018
Later among the works it cites.
Temporal difference models: Model-free deep rl for model-based control
V. Pong*, S. Gu*, M. Dalal, and S. Levine · 2018
Later among the works it cites.
Learning to search with MCTSnets, 2018
A. Guez, T. Weber, I. Antonoglou, K. Simonyan, O. Vinyals, D. Wierstra, R. Munos, and D. Silver · 2018
Later among the works it cites.
Learning agile and dynamic motor skills for legged robots
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V. Tsounis, V. Koltun, and M. Hutter · 2019
Closest in time.
Plan online, learn offline: Efficient learning and exploration via model-based control
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Real-time motion planning of legged robots: A model predictive control approach
F. Farshidian, E. Jelavic, A. Satapathy, M. Giftthaler, and J. Buchli · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y. Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis · 2017
Cited alongside, same era.
K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch · 2019
Closest in time.
Whole-body mpc for a dynamically stable mobile manipulator
M. V. Minniti, F. Farshidian, R. Grandia, and M. Hutter · 2019
Closest in time.