Fetching the paper…
Reading the bibliography…
Learning adaptable policies is crucial for robots to operate autonomously in our complex and quickly changing world.
Policy gradient methods for reinforcement learning with function approximation
Richard S. Sutton, David McAllester, Satinder Singh, and Yishay Mansour · 1999
Earlier work this paper cites.
Deterministic policy gradient algorithms
David Silver, Guy Lever, Nicolas Heess, Thomas Degris, Daan Wierstra, and Martin A. Riedmiller · 2014
Earlier work this paper cites.
Robots that can adapt like animals
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret · 2015
Earlier work this paper cites.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael I. Jordan, and Philipp Moritz · 2015
Earlier work this paper cites.
Transfer from simulation to real world through learning deep inverse dynamics model
Paul F. Christiano, Zain Shah, Igor Mordatch, Jonas Schneider, Trevor Blackwell, Joshua Tobin, Pieter Abbeel, and Wojciech Zaremba · 2016
Earlier work this paper cites.
Design principles for a family of direct-drive legged robots
Gavin Kenneally, Avik De, and Daniel E Koditschek · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Earlier work this paper cites.
Grounded action transformation for robot learning in simulation
Josiah P Hanna and Peter Stone · 2017
Earlier work this paper cites.
Towards generalization and simplicity in continuous control
Aravind Rajeswaran, Kendall Lowrey, Emanuel Todorov, and Sham M. Kakade · 2017
Earlier work this paper cites.
CAD2RL: real single-image flight without a single real image
Fereshteh Sadeghi and Sergey Levine · 2017
Earlier work this paper cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Earlier work this paper cites.
Learning to reinforcement learn
Jane X. Wang, Zeb Kurth-Nelson, Hubert Soyer, Joel Z. Leibo, Dhruva Tirumala, Rémi Munos, Charles Blundell, Dharshan Kumaran, and Matt M. Botvinick · 2017
Earlier work this paper cites.
Preparing for the unknown: Learning a universal policy with online system identification
Wenhao Yu, Jie Tan, C. Karen Liu, and Greg Turk · 2017
Earlier work this paper cites.
Meta-learning by the baldwin effect
Chrisantha Fernando, Jakub Sygnowski, Simon Osindero, Jane Wang, Tom Schaul, Denis Teplyashin, Pablo Sprechmann, Alexander Pritzel, and Andrei A. Rusu · 2018
Earlier work this paper cites.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Chelsea Finn and Sergey Levine · 2018
Cited alongside, same era.
Probabilistic model-agnostic meta-learning
Chelsea Finn, Kelvin Xu, and Sergey Levine · 2018
Cited alongside, same era.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, Yuxuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Automated deep reinforcement learning environment for hardware of a modular legged robot
Sehoon Ha, Joohyung Kim, and Katsu Yamane · 2018
Cited alongside, same era.
Evolved policy gradients
Rein Houthooft, Yuhua Chen, Phillip Isola, Bradly C. Stadie, Filip Wolski, Jonathan Ho, and Pieter Abbeel · 2018
Cited alongside, same era.
Simple random search provides a competitive approach to reinforcement learning
Bhairav Mehta, Manfred Diaz, Florian Golemo, Christopher J. Pal, and Liam Paull · 2019
Later among the works it cites.
Multi-agent manipulation via locomotion using hierarchical sim2real
Ofir Nachum, Michael Ahn, Hugo Ponte, Shixiang Gu, and Vikash Kumar · 2019
Later among the works it cites.
Learning to adapt in dynamic, real-world environments through meta-reinforcement learning
Anusha Nagabandi, Ignasi Clavera, Simin Liu, Ronald S. Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2019
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Later among the works it cites.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Cited alongside, same era.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Cited alongside, same era.
On first-order meta-learning algorithms
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Cited alongside, same era.
Sim-to-real transfer of robotic control with dynamics randomization
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Sim-to-real: Learning agile locomotion for quadruped robots
Jie Tan, Tingnan Zhang, Erwin Coumans, Atil Iscen, Yunfei Bai, Danijar Hafner, Steven Bohez, and Vincent Vanhoucke · 2018
Cited alongside, same era.
Meta reinforcement learning for sim-to-real domain adaptation
Karol Arndt, Murtaza Hazara, Ali Ghadirzadeh, and Ville Kyrki · 2019
Cited alongside, same era.
When random search is not enough: Sample-efficient and noise-robust blackbox optimization of RL policies
Krzysztof Choromanski, Aldo Pacchiano, Jack Parker-Holder, Yunhao Tang, Deepali Jain, Yuxiang Yang, Atil Iscen, Jasmine Hsu, and Vikas Sindhwani · 2019
Cited alongside, same era.
Realizing learned quadruped locomotion behaviors through kinematic motion primitives
Abhik Singla, Shounak Bhattacharya, Dhaivat Dholakiya, Shalabh Bhatnagar, Ashitava Ghosal, Bharadwaj Amrutur, and Shishir Kolathaya · 2019
Later among the works it cites.
Learning locomotion skills for cassie: Iterative design and sim-to-real
Zhaoming Xie, Patrick Clary, Jeremy Dao, Pedro Morais, Jonathan Hurst, and Michiel van de Panne · 2019
Later among the works it cites.
Norml: No-reward meta learning
Yuxiang Yang, Ken Caluwaerts, Atil Iscen, Jie Tan, and Chelsea Finn · 2019
Later among the works it cites.
Policy transfer with strategy optimization
Wenhao Yu, C Karen Liu, and Greg Turk · 2019
Later among the works it cites.
Varibad: A very good method for bayes-adaptive deep RL via meta-learning
Luisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze, Yarin Gal, Katja Hofmann, and Shimon Whiteson · 2019
Later among the works it cites.
Pybullet, a python module for physics simulation for games, robotics and machine learning
Erwin Coumans and Yunfei Bai · 2020
Closest in time.
Gradientless descent: High-dimensional zeroth-order optimization
Daniel Golovin, John Karro, Greg Kochanski, Chansoo Lee, Xingyou Song, and Qiuyi Zhang · 2020
Closest in time.
Learning generalizable locomotion skills with hierarchical reinforcement learning
Tianyu Li, Nathan Lambert, Roberto Calandra, Franziska Meier, and Akshara Rai · 2020
Closest in time.
Es-maml: Simple hessian-free meta learning
Xingyou Song, Wenbo Gao, Yuxiang Yang, Krzysztof Choromanski, Aldo Pacchiano, and Yunhao Tang · 2020
Closest in time.
Learning fast adaptation with meta strategy optimization
Wenhao Yu, Jie Tan, Yunfei Bai, Erwin Coumans, and Sehoon Ha · 2020
Closest in time.