Fetching the paper…
Reading the bibliography…
We present foundations for using Model Predictive Control (MPC) as a differentiable policy class for reinforcement learning in continuous state and action spaces.
Integrated architectures for learning, planning, and reacting based on approximating dynamic programming
Richard S Sutton · 1990
Earlier work this paper cites.
Exploiting model uncertainty estimates for safe dynamic control learning
Jeff G Schneider · 1997
Earlier work this paper cites.
Iterative linear quadratic regulator design for nonlinear biological movement systems
Weiwei Li and Emanuel Todorov · 2004
Earlier work this paper cites.
Using inaccurate models in reinforcement learning
Pieter Abbeel, Morgan Quigley, and Andrew Y Ng · 2006
Earlier work this paper cites.
Lqr via lagrange multipliers
Stephen Boyd · 2008
Earlier work this paper cites.
A generalized path integral control approach to reinforcement learning
Evangelos Theodorou, Jonas Buchli, and Stefan Schaal · 2010
Earlier work this paper cites.
Model predictive quadrotor indoor position control
Kostas Alexis, Christos Papachristos, George Nikolakopoulos, and Anthony Tzes · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
Marc Deisenroth and Carl E Rasmussen · 2011
Earlier work this paper cites.
Robust tube-based predictive control for mobile robots in off-road conditions
Ramón González, Mirko Fiacchini, José Luis Guzmán, Teodoro Álamo, and Francisco Rodríguez · 2011
Earlier work this paper cites.
Learning-based model predictive control on a quadrotor: Onboard implementation and experimental results
P. Bouffard, A. Aswani, , and C. Tomlin · 2012
Earlier work this paper cites.
Synthesis and stabilization of complex behaviors through online trajectory optimization
T. Erez, Y. Tassa, and E. Todorov · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Approximate real-time optimal control based on sparse gaussian process models
Joschika Boedecker, Jost Tobias Springenberg, Jan Wulfing, and Martin Riedmiller · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Learning neural network policies with guided policy search under unknown dynamics
Sergey Levine and Pieter Abbeel · 2014
Earlier work this paper cites.
Optimization-based autonomous racing of 1:43 scale rc cars
Alexander Liniger, Alexander Domahidi, and Manfred Morari · 2014
Earlier work this paper cites.
Control-limited differential dynamic programming
Yuval Tassa, Nicolas Mansard, and Emo Todorov · 2014
Earlier work this paper cites.
Learning continuous control policies by stochastic value gradients
Nicolas Heess, Gregory Wayne, David Silver, Tim Lillicrap, Tom Erez, and Yuval Tassa · 2015
Cited alongside, same era.
Fast nonlinear model predictive control for multicopter attitude tracking on so (3)
Mina Kamel, Kostas Alexis, Markus Achtelik, and Roland Siegwart · 2015
Cited alongside, same era.
Deepmpc: Learning deep latent features for model predictive control
Ian Lenz, Ross A Knepper, and Ashutosh Saxena · 2015
Cited alongside, same era.
Continuous control with deep reinforcement learning
Timothy P Lillicrap, Jonathan J Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Cited alongside, same era.
Combining model-based and model-free updates for trajectory-centric reinforcement learning
Yevgen Chebotar, Karol Hausman, Marvin Zhang, Gaurav Sukhatme, Stefan Schaal, and Sergey Levine · 2017
Later among the works it cites.
Treeqn and atreec: Differentiable tree planning for deep reinforcement learning
Gregory Farquhar, Tim Rocktäschel, Maximilian Igl, and Shimon Whiteson · 2017
Later among the works it cites.
Qmdp-net: Deep learning for planning under partial observability
Peter Karkus, David Hsu, and Wee Sun Lee · 2017
Later among the works it cites.
Optimal control and planning
Sergey Levine · 2017
Later among the works it cites.
Neural network dynamics for model-based deep reinforcement learning with model-free fine-tuning
Anusha Nagabandi, Gregory Kahn, Ronald S. Fearing, and Sergey Levine · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Embed to control: A locally linear latent dynamics model for control from raw images
Manuel Watter, Jost Springenberg, Joschka Boedecker, and Martin Riedmiller · 2015
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Cited alongside, same era.
Fast Nonlinear Model Predictive Control for Unified Trajectory Optimization and Tracking
Michael Neunert, Cedric de Crousaz, Fardi Furrer, Mina Kamel, Farbod Farshidian, Roland Siegwart, and Jonas Buchli · 2016
Cited alongside, same era.
Control of memory, active perception, and action in minecraft
Junhyuk Oh, Valliappa Chockalingam, Satinder Singh, and Honglak Lee · 2016
Cited alongside, same era.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philpp Moritz, Sergey Levine, Michael I. Jordan, and Pieter Abbeel · 2016
Cited alongside, same era.
Later among the works it cites.
Value prediction network
Junhyuk Oh, Satinder Singh, and Honglak Lee · 2017
Later among the works it cites.
Path integral networks: End-to-end differentiable optimal control
Masashi Okada, Luca Rigazio, and Takenobu Aoshima · 2017
Later among the works it cites.
Learning model-based planning from scratch
Razvan Pascanu, Yujia Li, Oriol Vinyals, Nicolas Heess, Lars Buesing, Sebastien Racanière, David Reichert, Théophane Weber, Daan Wierstra, and Peter Battaglia · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
A fast integrated planning and control framework for autonomous driving via imitation learning
Liting Sun, Cheng Peng, Wei Zhan, and Masayoshi Tomizuka · 2017
Later among the works it cites.
Learning from the hindsight plan—episodic mpc improvement
Aviv Tamar, Garrett Thomas, Tianhao Zhang, Sergey Levine, and Pieter Abbeel · 2017
Later among the works it cites.
Imagination-augmented agents for deep reinforcement learning
Théophane Weber, Sébastien Racanière, David P Reichert, Lars Buesing, Arthur Guez, Danilo Jimenez Rezende, Adria Puigdomènech Badia, Oriol Vinyals, Nicolas Heess, Yujia Li, et al · 2017
Later among the works it cites.
Model predictive path integral control: From theory to parallel computation
Grady Williams, Andrew Aldrich, and Evangelos A Theodorou · 2017
Later among the works it cites.
Differential Dynamic Programming with Nonlinear Constraints
Zhaoming Xie, C. Karen Liu, and Kris Hauser · 2017
Later among the works it cites.
Deepak Pathak, Parsa Mahmoudieh, Guanghao Luo, Pulkit Agrawal, Dian Chen, Yide Shentu, Evan Shelhamer, Jitendra Malik, Alexei A Efros, and Trevor Darrell · 2018
Closest in time.
Mpc-inspired neural network policies for sequential decision making
Marcus Pereira, David D. Fan, Gabriel Nakajima An, and Evangelos Theodorou · 2018
Closest in time.
Temporal difference models: Model-free deep rl for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Closest in time.
Aravind Srinivas, Allan Jabri, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Closest in time.