Fetching the paper…
Reading the bibliography…
Providing densely shaped reward functions for RL algorithms is often exceedingly challenging, motivating the development of RL algorithms that can learn from easier-to-specify sparse reward functions.
Learning from demonstration
Stefan Schaal et al · 1997
Earlier work this paper cites.
A survey of robot learning from demonstration
Brenna D Argall, Sonia Chernova, Manuela Veloso, and Brett Browning · 2009
Earlier work this paper cites.
Td_gamma: Re-evaluating complex backups in temporal difference learning
George Konidaris, Scott Niekum, and Philip S Thomas · 2011
Earlier work this paper cites.
Mujoco: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Learning from limited demonstrations
Beomjoon Kim, Amir-massoud Farahmand, Joelle Pineau, and Doina Precup · 2013
Earlier work this paper cites.
Exploiting multi-step sample trajectories for approximate value iteration
Robert Wright, Steven Loscalzo, Philip Dexter, and Lei Yu · 2013
Earlier work this paper cites.
Policy search for motor primitives in robotics
Jens Kober and Jan Peters · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Cfqi: Fitted q-iteration with complex returns
Robert William Wright, Xingye Qiao, Lei Yu, and Steven Loscalzo · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
High-dimensional continuous control using generalized advantage estimation
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel · 2016
Earlier work this paper cites.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Earlier work this paper cites.
Deep reinforcement learning with double q-learning
Hado Van Hasselt, Arthur Guez, and David Silver · 2016
Cited alongside, same era.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aäron Oord, and Rémi Munos · 2017
Cited alongside, same era.
Proximal policy optimization algorithms, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
Addressing function approximation error in actor-critic methods
Scott Fujimoto, Herke Hoof, and David Meger · 2018
Cited alongside, same era.
Reinforcement learning from imperfect demonstrations
Yang Gao, Huazhe Xu, Ji Lin, Fisher Yu, Sergey Levine, and Trevor Darrell · 2018
Cited alongside, same era.
Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards, 2018
Mel Vecerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe, and Martin Riedmiller · 2018
Later among the works it cites.
Advantage-weighted regression: Simple and scalable off-policy reinforcement learning, 2019
Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine · 2019
Later among the works it cites.
Deep reinforcement learning with optimized reward functions for robotic trajectory planning
Jiexin Xie, Zhenzhou Shao, Yue Li, Yong Guan, and Jindong Tan · 2019
Later among the works it cites.
Reinforcement learning from imperfect demonstrations under soft expert guidance
Mingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun, Chao Yang, Bin Fang, and Huaping Liu · 2020
Later among the works it cites.
Specification gaming: the flip side of ai ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine · 2018
Cited alongside, same era.
Deep q-learning from demonstrations
Todd Hester, Matej Vecerik, Olivier Pietquin, Marc Lanctot, Tom Schaul, Bilal Piot, Dan Horgan, John Quan, Andrew Sendonaris, Ian Osband, et al · 2018
Cited alongside, same era.
Policy optimization with demonstrations
Bingyi Kang, Zequn Jie, and Jiashi Feng · 2018
Cited alongside, same era.
Overcoming exploration in reinforcement learning with demonstrations
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel · 2018
Cited alongside, same era.
Learning complex dexterous manipulation with deep reinforcement learning and demonstrations, 2018
Aravind Rajeswaran, Vikash Kumar, Abhishek Gupta, Giulia Vezzani, John Schulman, Emanuel Todorov, and Sergey Levine · 2018
Cited alongside, same era.
Learning to mix n-step returns: Generalizing lambda-returns for deep reinforcement learning
Sahil Sharma, Girish Raguvir J, Srivatsan Ramesh, and Balaraman Ravindran · 2018
Cited alongside, same era.
Reinforcement learning: An introduction
Richard S Sutton and Andrew G Barto · 2018
Cited alongside, same era.
Conservative q-learning for offline reinforcement learning
Aviral Kumar, Aurick Zhou, George Tucker, and Sergey Levine · 2020
Later among the works it cites.
Controlling overestimation bias with truncated mixture of continuous distributional quantile critics
Arsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, and Dmitry Vetrov · 2020
Later among the works it cites.
Safety augmented value estimation from demonstrations (saved): Safe deep model-based rl for sparse cost robotic tasks, 2020
Brijen Thananjeyan, Ashwin Balakrishna, Ugo Rosolia, Felix Li, Rowan McAllister, Joseph E. Gonzalez, Sergey Levine, Francesco Borrelli, and Ken Goldberg · 2020
Later among the works it cites.
Learning dense rewards for contact-rich manipulation tasks
Zheng Wu, Wenzhao Lian, Vaibhav Unhelkar, Masayoshi Tomizuka, and Stefan Schaal · 2020
Later among the works it cites.
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, and Roberto Martín-Martín · 2020
Later among the works it cites.
Awac: Accelerating online reinforcement learning with offline datasets, 2021
Ashvin Nair, Abhishek Gupta, Murtaza Dalal, and Sergey Levine · 2021
Later among the works it cites.
Recovery rl: Safe reinforcement learning with learned recovery zones
Brijen Thananjeyan, Ashwin Balakrishna, Suraj Nair, Michael Luo, Krishnan Srinivasan, Minho Hwang, Joseph E Gonzalez, Julian Ibarz, Chelsea Finn, and Ken Goldberg · 2021
Later among the works it cites.
Ls3: Latent space safe sets for long-horizon visuomotor control of sparse reward iterative tasks
Albert Wilcox, Ashwin Balakrishna, Brijen Thananjeyan, Joseph E. Gonzalez, and Ken Goldberg · 2021
Later among the works it cites.