Fetching the paper…
Reading the bibliography…
Designing rewards for Reinforcement Learning (RL) is challenging because it needs to convey the desired task, be efficient to optimize, and be easy to compute.
Self-supervised Learning of Image Embedding for Continuous Control
Carlos Florensa, Jonas Degrave, Nicolas Heess, Jost Tobias Springenberg, and Martin Riedmiller · 1901
Earlier work this paper cites.
Adaptive Variance for Changing Sparse-Reward Environments
Xingyu Lin, Pengsheng Guo, Carlos Florensa, and David Held · 1903
Earlier work this paper cites.
Reverse Curriculum Generation for Reinforcement Learning
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang, and Pieter Abbeel · 1938
Earlier work this paper cites.
ALVINN: an autonomous land vehicle in a neural network
Dean A Pomerleau · 1989
Earlier work this paper cites.
Learning to Achieve Goals
Leslie P. Kaelbling · 1993
Earlier work this paper cites.
Maximum Entropy Inverse Reinforcement Learning
Brian D Ziebart, Andrew Maas, J Andrew Bagnell, and Anind K Dey · 2008
Earlier work this paper cites.
Learning locomotion over rough terrain using terrain templates
Mrinal Kalakrishnan, Jonas Buchli, Peter Pastor, and Stefan Schaal · 2009
Earlier work this paper cites.
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning
Stéphane Ross, Geoffrey J Gordon, and J Andrew Bagnell · 2011
Earlier work this paper cites.
MuJoCo : A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Overcoming Exploration in Reinforcement Learning with Demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2013
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei a Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Universal Value Function Approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Cited alongside, same era.
Learning movement primitive attractor goals and sequential skills from kinesthetic demonstrations
Simon Manschitz, Jens Kober, Michael Gienger, and Jan Peters · 2015
Cited alongside, same era.
Trust Region Policy Optimization
John Schulman, Philipp Moritz, Michael Jordan, and Pieter Abbeel · 2015
Cited alongside, same era.
End to End Learning for Self-Driving Cars
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba · 2016
Cited alongside, same era.
Guided Cost Learning: Deep Inverse Optimal Control via Policy Optimization
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2018
Later among the works it cites.
Learning Robust Rewards with Adversarial Inverse Reinforcement Learning
Justin Fu, Katie Luo, and Sergey Levine · 2018
Later among the works it cites.
Reinforcement Learning from Imperfect Demonstrations
Yang Gao, Huazhe Harry, Xu Ji, Lin Fisher, Yu Sergey, and Levine Trevor · 2018
Later among the works it cites.
Truncated Horizon Policy Search: Combining Reinforcement Learning & Imitation Learning
Wen Sun, J. Andrew Bagnell, and Byron Boots · 2018
Later among the works it cites.
DeepMimic: Example-Guided Deep Reinforcement Learning of Physics-Based Character Skills
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chelsea Finn, Sergey Levine, and Pieter Abbeel · 2016
Cited alongside, same era.
Generative Adversarial Imitation Learning
Jonathan Ho and Stefano Ermon · 2016
Cited alongside, same era.
Mastering the game of Go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, Yutian Chen, Timothy Lillicrap, Fan Hui, Laurent Sifre, George Van Den Driessche, Thore Graepel, and Demis Hassabis · 2017
Cited alongside, same era.
Emergence of Locomotion Behaviours in Rich Environments
Nicolas Heess, Dhruva TB, Srinivasan Sriram, Jay Lemmon, Josh Merel, Greg Wayne, Yuval Tassa, Tom Erez, Ziyu Wang, S. M. Ali Eslami, Martin Riedmiller, and David Silver · 2017
Cited alongside, same era.
Third-Person Imitation Learning
Bradly C. Stadie, Pieter Abbeel, and Ilya Sutskever · 2017
Cited alongside, same era.
Leveraging Demonstrations for Deep Reinforcement Learning on Robotics Problems with Sparse Rewards
Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2017
Cited alongside, same era.
Learning by Playing – Solving Sparse Reward Tasks from Scratch
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van De Wiele, Volodymyr Mnih, Nicolas Heess, and Tobias Springenberg · 2018
Cited alongside, same era.
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, Sergey Levine, and Computer Sciences · 2018
Cited alongside, same era.
Xue Bin Peng, Pieter Abbeel, Sergey Levine, and Michiel van de Panne · 2018
Later among the works it cites.
Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, John Schulman, and L G Sep · 2018
Later among the works it cites.
AlphaStar: Mastering the Real-Time Strategy Game StarCraft II
Oriol Vinyals, Igor Babuschkin, Junyoung Chung, Michael Mathieu, Max Jaderberg, Wojciech M Czarnecki, Andrew Dudzik, Aja Huang, Petko Georgiev, Richard Powell, Timo Ewalds, Dan Horgan, Manuel Kroiss, Ivo Danihelka, John Agapiou, Junhyuk Oh, Valentin Dalibard, David Choi, Laurent Sifre, Yury Sulsky, Sasha Vezhnevets, James Molloy, Trevor Cai, David Budden, Tom Paine, Caglar Gulcehre, Ziyu Wang, Tobias Pfaff, Toby Pohlen, Yuhuai Wu, Dani Yogatama, Julia Cohen, Katrina McKinney, Oliver Smith, Tom Schaul, Timothy Lillicrap, Chris Apps, Koray Kavukcuoglu, Demis Hassabis, and David Silver · 2019
Closest in time.
Sample-Efficient Imitation Learning via Generative Adversarial Nets
Lionel Blondé and Alexandros Kalousis · 2019
Closest in time.
Sample Efficient Imitation Learning for Continuous Control
Fumihiro Sasaki, Tetsuya Yohira, and Atsuo Kawaguchi · 2019
Closest in time.
Generative predecessor models for sample-efficient imitation learning
Yannick Schroecker, Mel Vecerik, and Jonathan Scholz · 2019
Closest in time.
DISCRIMINATOR-ACTOR-CRITIC: ADDRESSING SAMPLE INEFFICIENCY AND REWARD BIAS IN ADVERSARIAL IMITATION LEARNING
Ilya Kostrikov, Kumar Krishna Agrawal2, Debidatta Dwibedi, Sergey Levine, and Jonathan Tompson · 2019
Closest in time.