Fetching the paper…
Reading the bibliography…
Reinforcement learning (RL) algorithms have demonstrated promising results on complex tasks, yet often require impractical numbers of samples since they learn from scratch.
Evolutionary principles in self-referential learning, or on learning how to learn: the meta-meta-… hook
Jürgen Schmidhuber · 1987
Earlier work this paper cites.
Learning a synaptic learning rule
Yoshua Bengio, Samy Bengio, and Jocelyn Cloutier · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Asymptopia: an exposition of statistical asymptotic theory
D. Pollard · 2000
Earlier work this paper cites.
Divergences, surrogate loss functions and experimental design
X. Nguyen, M. J. Wainwright, and M. I. Jordan · 2005
Earlier work this paper cites.
Policy gradient methods for robotics
Jan Peters and Stefan Schaal · 2006
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan R Peters · 2009
Earlier work this paper cites.
Robot motor skill coordination with em-based reinforcement learning
Petar Kormushev, Sylvain Calinon, and Darwin G Caldwell · 2010
Earlier work this paper cites.
A reduction of imitation learning and structured prediction to no-regret online learning
Stéphane Ross, Geoffrey Gordon, and Drew Bagnell · 2011
Earlier work this paper cites.
Integrating reinforcement learning with human demonstrations of varying ability
Matthew E Taylor, Halit Bener Suay, and Sonia Chernova · 2011
Earlier work this paper cites.
Autonomous reinforcement learning on raw visual input data in a real world application
Sascha Lange, Martin Riedmiller, and Arne Voigtländer · 2012
Earlier work this paper cites.
Learning to learn
Sebastian Thrun and Lorien Pratt · 2012
Earlier work this paper cites.
Reinforcement learning in robotics: A survey
Jens Kober, J Andrew Bagnell, and Jan Peters · 2013
Earlier work this paper cites.
Reinforcement learning from demonstration through shaping
Tim Brys, Anna Harutyunyan, Halit Bener Suay, Sonia Chernova, Matthew E Taylor, and Ann Nowé · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
End to end learning for self-driving cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba · 2016
Earlier work this paper cites.
Rl 2 : Fast reinforcement learning via slow reinforcement learning
Yan Duan, John Schulman, Xi Chen, Peter L Bartlett, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
A machine learning approach to visual perception of forest trails for mobile robots
Alessandro Giusti, Jérôme Guzzi, Dan C Cireşan, Fang-Lin He, Juan P Rodríguez, Flavio Fontana, Matthias Faessler, Christian Forster, Jürgen Schmidhuber, Gianni Di Caro, et al · 2016
Earlier work this paper cites.
PLATO: policy learning using adaptive trajectory optimization
Gregory Kahn, Tianhao Zhang, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Reward augmented maximum likelihood for neural structured prediction
Mohammad Norouzi, Samy Bengio, Zhifeng Chen, Navdeep Jaitly, Mike Schuster, Yonghui Wu, and Dale Schuurmans · 2016
Cited alongside, same era.
Actor-mimic: Deep multitask and transfer reinforcement learning
Emilio Parisotto, Jimmy Lei Ba, and Ruslan Salakhutdinov · 2016
Cited alongside, same era.
Policy distillation
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2016
Cited alongside, same era.
Query-efficient imitation learning for end-to-end simulated driving
Jiakai Zhang and Kyunghyun Cho · 2017
Later among the works it cites.
Learning to adapt: Meta-learning for model-based control
Ignasi Clavera, Anusha Nagabandi, Ronald S Fearing, Pieter Abbeel, Sergey Levine, and Chelsea Finn · 2018
Later among the works it cites.
Meta-learning and universality: Deep representations and gradient descent can approximate any learning algorithm
Chelsea Finn and Sergey Levine · 2018
Later among the works it cites.
Divide-and-conquer reinforcement learning
Dibya Ghosh, Avi Singh, Aravind Rajeswaran, Vikash Kumar, and Sergey Levine · 2018
Later among the works it cites.
Meta-reinforcement learning of structured exploration strategies
Abhishek Gupta, Russell Mendonca, YuXuan Liu, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Exploration from demonstration for interactive reinforcement learning
Kaushik Subramanian, Charles L Isbell Jr, and Andrea L Thomaz · 2016
Cited alongside, same era.
Learning to reinforcement learn
Jane X Wang, Zeb Kurth-Nelson, Dhruva Tirumala, Hubert Soyer, Joel Z Leibo, Remi Munos, Charles Blundell, Dharshan Kumaran, and Matt Botvinick · 2016
Cited alongside, same era.
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2017
Cited alongside, same era.
One-shot imitation learning
Yan Duan, Marcin Andrychowicz, Bradly C. Stadie, Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Model-agnostic meta-learning for fast adaptation of deep networks
Chelsea Finn, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Deep visual foresight for planning robot motion
Chelsea Finn and Sergey Levine · 2017
Cited alongside, same era.
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Later among the works it cites.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Gabriel Dulac-Arnold, et al · 2018
Later among the works it cites.
Rein Houthooft, Richard Y Chen, Phillip Isola, Bradly C Stadie, Filip Wolski, Jonathan Ho, and Pieter Abbeel · 2018
Later among the works it cites.
A simple neural attentive meta-learner
Nikhil Mishra, Mostafa Rohaninejad, Xi Chen, and Pieter Abbeel · 2018
Later among the works it cites.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Later among the works it cites.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Later among the works it cites.
Promp: Proximal meta-policy search
Jonas Rothfuss, Dennis Lee, Ignasi Clavera, Tamim Asfour, and Pieter Abbeel · 2018
Later among the works it cites.
Meta reinforcement learning with latent variable gaussian processes
Steindór Sæmundsson, Katja Hofmann, and Marc Peter Deisenroth · 2018
Later among the works it cites.
Some considerations on learning to explore via meta-reinforcement learning
Bradly C Stadie, Ge Yang, Rein Houthooft, Xi Chen, Yan Duan, Yuhuai Wu, Pieter Abbeel, and Ilya Sutskever · 2018
Later among the works it cites.
Truncated horizon policy search: Combining reinforcement learning & imitation learning
Wen Sun, J Andrew Bagnell, and Byron Boots · 2018
Later among the works it cites.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.
Efficient off-policy meta-reinforcement learning via probabilistic context variables
Kate Rakelly, Aurick Zhou, Deirdre Quillen, Chelsea Finn, and Sergey Levine · 2019
Closest in time.