Fetching the paper…
Reading the bibliography…
Deep reinforcement learning is a promising approach to learning policies in uncontrolled environments that do not require domain knowledge.
N. Kohl and P. Stone, “Policy gradient reinforcement learning for fast quadrupedal locomotion,” IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA ’04. 2004 , vol. 3, pp. 2619–2624 Vol.3, 2004
2004
Earlier work this paper cites.
R. Tedrake, T. Zhang, and H. Seung, “Stochastic policy gradient reinforcement learning on a simple 3d biped,” 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566) , vol. 3, pp. 2849–2854 vol.3, 2004
2004
Earlier work this paper cites.
G. Endo, J. Morimoto, T. Matsubara, J. Nakanishi, and G. Cheng, “Learning cpg sensory feedback with policy gradient for biped locomotion for a full-body humanoid,” in AAAI , 2005
2005
Earlier work this paper cites.
R. Tedrake and H. Seung, “Learning to walk in 20 minutes,” 2005
2005
Earlier work this paper cites.
E. Todorov, T. Erez, and Y. Tassa, “Mujoco: A physics engine for model-based control,” in 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems . IEEE, 2012, pp. 5026–5033
2012
Earlier work this paper cites.
M. Cutler, T. J. Walsh, and J. P. How, “Reinforcement learning with multi-fidelity simulators,” 2014 IEEE International Conference on Robotics and Automation (ICRA) , pp. 3888–3895, 2014
2014
Earlier work this paper cites.
N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: a simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res. , vol. 15, pp. 1929–1958, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
G. D. Kenneally, A. De, and D. E. Koditschek, “Design principles for a family of direct-drive legged robots,” IEEE Robotics and Automation Letters , vol. 1, pp. 900–907, 2016
2016
Earlier work this paper cites.
M. Hutter, C. Gehring, D. Jud, A. Lauber, D. Bellicoso, V. Tsounis, J. Hwangbo, K. Bodie, P. Fankhauser, M. Bloesch, R. Diethelm, S. Bachmann, A. Melzer, and M. Höpflinger, “Anymal - a highly mobile and dynamic quadrupedal robot,” IEEE International Conference on Intelligent Robots and Systems (IROS) , pp. 38–44, 2016
2016
Earlier work this paper cites.
J. Tan, Z. Xie, B. Boots, and C. Liu, “Simulation-based design of dynamic controllers for humanoid balancing,” 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 2729–2736, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
J. Ba, J. R. Kiros, and G. E. Hinton, “Layer normalization,” ArXiv , vol. abs/1607.06450, 2016
2016
Earlier work this paper cites.
F. Sadeghi and S. Levine, “ CAD 2 \text{CAD}^{2} RL: Real single-image flight without a single real image,” Robotics: Science and Systems (RSS) , 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel, “Domain randomization for transferring deep neural networks from simulation to the real world,” 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 23–30, 2017
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International Conference on Machine Learning (ICML) , 2017
2017
Earlier work this paper cites.
H.-W. Park, P. M. Wensing, and S. Kim, “High-speed bounding with the mit cheetah 2: Control design and experiments,” The International Journal of Robotics Research , vol. 36, no. 2, pp. 167–192, 2017. [Online]. Available: https://doi.org/10.1177/0278364917694244
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Babaeizadeh, I. Frosio, S. Tyree, J. Clemons, and J. Kautz, “Reinforcement learning through asynchronous advantage actor-critic on a gpu,” in ICLR , 2017
2017
Earlier work this paper cites.
S. S. Gu, E. Holly, T. P. Lillicrap, and S. Levine, “Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates,” 2017 IEEE International Conference on Robotics and Automation (ICRA) , pp. 3389–3396, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
S. Levine, P. Pastor, A. Krizhevsky, and D. Quillen, “Learning hand-eye coordination for robotic grasping with deep learning and large-scale data collection,” The International Journal of Robotics Research , vol. 37, pp. 421 – 436, 2018
2018
Cited alongside, same era.
X. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel, “Sim-to-real transfer of robotic control with dynamics randomization,” 2018 IEEE International Conference on Robotics and Automation (ICRA) , pp. 1–8, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
L. Liu and J. Hodgins, “Learning basketball dribbling skills using trajectory optimization and deep reinforcement learning,” ACM Transactions on Graphics (TOG) , vol. 37, pp. 1 – 14, 2018
2018
Cited alongside, same era.
2019
Later among the works it cites.
2019
Later among the works it cites.
W. Yu, J. Tan, Y. Bai, E. Coumans, and S. Ha, “Learning fast adaptation with meta strategy optimization,” IEEE Robotics and Automation Letters , vol. 5, pp. 2950–2957, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
X. Peng, P. Abbeel, S. Levine, and M. V. D. Panne, “Deepmimic: Example-guided deep reinforcement learning of physics-based character skills,” ACM Trans. Graph. , vol. 37, pp. 143:1–143:14, 2018
2018
Cited alongside, same era.
G. Bledt, M. J. Powell, B. Katz, J. Carlo, P. Wensing, and S. Kim, “Mit cheetah 3: Design and control of a robust, dynamic quadruped robot,” IEEE International Conference on Intelligent Robots and Systems (IROS) , pp. 2245–2252, 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” in ICML , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
R. S. Sutton and A. G. Barto, Reinforcement learning: An introduction . MIT press, 2018
2018
Cited alongside, same era.
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, and Q. Zhang, “JAX: composable transformations of Python+NumPy programs,” 2018. [Online]. Available: http://github.com/google/jax
2018
Cited alongside, same era.
X. Song, Y. Yang, K. Choromanski, K. Caluwaerts, W. Gao, C. Finn, and J. Tan, “Rapidly adaptable legged robots via evolutionary meta-learning,” 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pp. 3769–3776, 2020
2020
Later among the works it cites.
J. Lee, J. Hwangbo, L. Wellhausen, V. Koltun, and M. Hutter, “Learning quadrupedal locomotion over challenging terrain,” Science Robotics , vol. 5, 2020
2020
Later among the works it cites.
2020
Later among the works it cites.
T. Haarnoja, S. Ha, A. Zhou, J. Tan, G. Tucker, and S. Levine, “Learning to walk via deep reinforcement learning,” Robotics: Science and Systems (RSS) , 2020
2020
Later among the works it cites.
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi, “Dream to control: Learning behaviors by latent imagination,” International Conference on Learning Representations (ICLR) , 2020
2020
Later among the works it cites.
Y.-H. Xu, C.-C. Yang, M. Hua, and W. Zhou, “Deep deterministic policy gradient (ddpg)-based resource allocation scheme for noma vehicular communications,” IEEE Access , vol. 8, pp. 18 797–18 807, 2020
2020
Later among the works it cites.
S. Tunyasuvunakool, A. Muldal, Y. Doron, S. Liu, S. Bohez, J. Merel, T. Erez, T. Lillicrap, N. Heess, and Y. Tassa, “dm_control: Software and tasks for continuous control,” Software Impacts , vol. 6, p. 100022, 2020. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S2665963820300099
2020
Later among the works it cites.
2021
Later among the works it cites.
A. Kumar, Z. Fu, D. Pathak, and J. Malik, “Rma: Rapid motor adaptation for legged robots,” Robotics: Science and Systems (RSS) , 2021
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Fu, A. Kumar, J. Malik, and D. Pathak, “Minimizing energy consumption leads to the emergence of gaits in legged robots,” Conference on Robot Learning (CoRL) , 2021
2021
Later among the works it cites.
2022
Closest in time.
2022
Closest in time.
L. Smith, J. C. Kew, X. B. Peng, S. Ha, J. Tan, and S. Levine, “Legged robots that keep on learning: Fine-tuning locomotion policies in the real world,” IEEE International Conference on Robotics and Automation (ICRA) , 2022
2022
Closest in time.
2022
Closest in time.
T. Hiraoka, T. Imagawa, T. Hashimoto, T. Onishi, and Y. Tsuruoka, “Dropout q-functions for doubly efficient reinforcement learning,” International Conference on Learning Representations (ICLR) , 2022
2022
Closest in time.