Fetching the paper…
Reading the bibliography…
Traditionally, model-based reinforcement learning (MBRL) methods exploit neural networks as flexible function approximators to represent $\textit{a priori}$ unknown environment dynamics.
On the adaptive control of robot manipulators
J.-J. E. Slotine and W. Li · 1987
Earlier work this paper cites.
Dyna, an integrated architecture for learning, planning, and reacting
R. S. Sutton · 1991
Earlier work this paper cites.
Planning and acting in partially observable stochastic domains
L. P. Kaelbling, M. L. Littman, and A. R. Cassandra · 1998
Earlier work this paper cites.
Adaptive control design and analysis , volume 37
G. Tao · 2003
Earlier work this paper cites.
Policy gradient reinforcement learning for fast quadrupedal locomotion
N. Kohl and P. Stone · 2004
Earlier work this paper cites.
Stochastic policy gradient reinforcement learning on a simple 3d biped
R. Tedrake, T. W. Zhang, and H. S. Seung · 2004
Earlier work this paper cites.
Learning cpg sensory feedback with policy gradient for biped locomotion for a full-body humanoid
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng · 2005
Earlier work this paper cites.
Multi-step dyna planning for policy evaluation and control
H. Yao, S. Bhatnagar, D. Diao, R. S. Sutton, and C. Szepesvári · 2009
Earlier work this paper cites.
Adaptive control: stability, convergence and robustness
S. Sastry and M. Bodson · 2011
Earlier work this paper cites.
Pilco: A model-based and data-efficient approach to policy search
M. Deisenroth and C. E. Rasmussen · 2011
Earlier work this paper cites.
Dyna-style planning with linear function approximation and prioritized sweeping
R. S. Sutton, C. Szepesvári, A. Geramifard, and M. P. Bowling · 2012
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
E. Todorov, T. Erez, and Y. Tassa · 2012
Earlier work this paper cites.
Localizing external contact using proprioceptive sensors: The contact particle filter
L. Manuelli and R. Tedrake · 2016
Earlier work this paper cites.
Value iteration networks
A. Tamar, Y. Wu, G. Thomas, S. Levine, and P. Abbeel · 2016
Earlier work this paper cites.
Probabilistic foot contact estimation by fusing information from dynamics and differential/forward kinematics
J. Hwangbo, C. D. Bellicoso, P. Fankhauser, and M. Hutter · 2016
Earlier work this paper cites.
Robot collisions: A survey on detection, isolation, and identification
S. Haddadin, A. De Luca, and A. Albu-Schäffer · 2017
Earlier work this paper cites.
Imagination-augmented agents for deep reinforcement learning
S. Racanière, T. Weber, D. Reichert, L. Buesing, A. Guez, D. Jimenez Rezende, A. Puigdomènech Badia, O. Vinyals, N. Heess, Y. Li, et al · 2017
Earlier work this paper cites.
Probabilistic terrain mapping for mobile robots with uncertain localization
P. Fankhauser, M. Bloesch, and M. Hutter · 2018
Earlier work this paper cites.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
K. Chua, R. Calandra, R. McAllister, and S. Levine · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine · 2018
Cited alongside, same era.
JAX: composable transformations of Python+NumPy programs, 2018
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang · 2018
Cited alongside, same era.
Dream to control: Learning behaviors by latent imagination
D. Hafner, T. Lillicrap, J. Ba, and M. Norouzi · 2019
Cited alongside, same era.
When to trust your model: Model-based policy optimization
M. Janner, J. Fu, M. Zhang, and S. Levine · 2019
Cited alongside, same era.
Unitree ros to real
U. Robotics · 2021
Later among the works it cites.
Brax–a differentiable physics engine for large scale rigid body simulation
C. D. Freeman, E. Frey, A. Raichuk, S. Girgin, I. Mordatch, and O. Bachem · 2021
Later among the works it cites.
Reinforcement learning for robust parameterized locomotion control of bipedal robots
Z. Li, X. Cheng, X. B. Peng, P. Abbeel, S. Levine, G. Berseth, and K. Sreenath · 2021
Later among the works it cites.
Neural networks with physics-informed architectures and constraints for dynamical systems modeling
F. Djeumou, C. Neary, E. Goubault, S. Putot, and U. Topcu · 2022
Later among the works it cites.
Rloc: Terrain-aware legged locomotion using reinforcement learning and optimal control
S. Gangapurwala, M. Geisert, R. Orsolino, M. Fallon, and I. Havoutis · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep lagrangian networks: Using physics as model prior for deep learning
M. Lutter, C. Ritter, and J. Peters · 2019
Cited alongside, same era.
Trajectory-based probabilistic policy gradient for learning locomotion behaviors
S. Choi and J. Kim · 2019
Cited alongside, same era.
Mopo: Model-based offline policy optimization
T. Yu, G. Thomas, L. Yu, S. Ermon, J. Y. Zou, S. Levine, C. Finn, and T. Ma · 2020
Cited alongside, same era.
Feedback linearization for uncertain systems via reinforcement learning
T. Westenbroek, D. Fridovich-Keil, E. Mazumdar, S. Arora, V. Prabhu, S. S. Sastry, and C. J. Tomlin · 2020
Cited alongside, same era.
M. Cranmer, S. Greydanus, S. Hoyer, P. Battaglia, D. Spergel, and S. Ho · 2020
Cited alongside, same era.
Learning quadrupedal locomotion over challenging terrain
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2020
Cited alongside, same era.
Data efficient reinforcement learning for legged robots
Y. Yang, K. Caluwaerts, A. Iscen, T. Zhang, J. Tan, and V. Sindhwani · 2020
Cited alongside, same era.
Legged robots that keep on learning: Fine-tuning locomotion policies in the real world
L. Smith, J. C. Kew, X. B. Peng, S. Ha, J. Tan, and S. Levine · 2022
Later among the works it cites.
Cpg-rl: Learning central pattern generators for quadruped locomotion
G. Bellegarda and A. Ijspeert · 2022
Later among the works it cites.
Safe reinforcement learning for legged locomotion
T.-Y. Yang, T. Zhang, L. Luu, S. Ha, J. Tan, and W. Yu · 2022
Later among the works it cites.
A walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning
L. Smith, I. Kostrikov, and S. Levine · 2022
Later among the works it cites.
The primacy bias in deep reinforcement learning
E. Nikishin, M. Schwarzer, P. D’Oro, P.-L. Bacon, and A. Courville · 2022
Later among the works it cites.
Bundled gradients through contact via randomized smoothing
H. J. T. Suh, T. Pang, and R. Tedrake · 2022
Later among the works it cites.
Lyapunov design for robust and efficient robotic reinforcement learning
T. Westenbroek, F. Castaneda, A. Agrawal, S. Sastry, and K. Sreenath · 2022
Later among the works it cites.
Learning robust perceptive locomotion for quadrupedal robots in the wild
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter · 2022
Later among the works it cites.
Model-based reinforcement learning: A survey
T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, et al · 2023
Later among the works it cites.
Neural volumetric memory for visual locomotion control
R. Yang, G. Yang, and X. Wang · 2023
Later among the works it cites.
Global planning for contact-rich manipulation via local smoothing of quasi-dynamic contact models
T. Pang, H. T. Suh, L. Yang, and R. Tedrake · 2023
Later among the works it cites.
TD-MPC2: Scalable, robust world models for continuous control
N. Hansen, H. Su, and X. Wang · 2024
Closest in time.
Smoothed online learning for prediction in piecewise affine systems
A. Block, M. Simchowitz, and R. Tedrake · 2024
Closest in time.