Fetching the paper…
Reading the bibliography…
While using shaped rewards can be beneficial when solving sparse reward tasks, their successful application often requires careful engineering and is problem specific.
Learning to Achieve Goals
Leslie Pack Kaelbling · 1993
Earlier work this paper cites.
Reward Functions for Accelerated Learning
Maja J Mataric · 1994
Earlier work this paper cites.
Policy invariance under reward transformations: theory and application to reward shaping
Andrew Y Ng, Daishi Harada, and Stuart Russel · 1999
Earlier work this paper cites.
Curriculum Learning
Yoshua Bengio, Jerome Louradour, Ronan Collobert, and Jason Weston · 2009
Earlier work this paper cites.
Horde : A Scalable Real-time Architecture for Learning Knowledge from Unsupervised Sensorimotor Interaction Categories and Subject Descriptors
Richard S Sutton, Joseph Modayil, Michael Delp, Thomas Degris, Patrick M Pilarski, and Adam White · 2011
Earlier work this paper cites.
MuJoCo: A physics engine for model-based control
Emanuel Todorov, Tom Erez, and Yuval Tassa · 2012
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Universal Value Function Approximators
Tom Schaul, Dan Horgan, Karol Gregor, and David Silver · 2015
Earlier work this paper cites.
Unifying Count-Based Exploration and Intrinsic Motivation
Marc G. Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Faulty reward functions in the wild
Jack Clark and Dario Amodei · 2016
Earlier work this paper cites.
The Malmo Platform for Artificial Intelligence Experimentation
Matthew Johnson, Katja Hofmann, Tim Hutton, and David Bignell · 2016
Earlier work this paper cites.
Continuous control with deep reinforcement learning
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra · 2016
Cited alongside, same era.
Hindsight Experience Replay
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, Pieter Abbeel, and Wojciech Zaremba · 2017
Cited alongside, same era.
Improving Stochastic Policy Gradients in Continuous Control with Deep Reinforcement Learning using the Beta Distribution
Po-Wei Chou, Maturana Daniel, and Sebastian Scherer · 2017
Cited alongside, same era.
Reinforcement Learning with Deep Energy-Based Policies
Tuomas Haarnoja, Haoran Tang, Pieter Abbeel, and Sergey Levine · 2017
Cited alongside, same era.
Curiosity-driven Exploration by Self-supervised Prediction
Deepak Pathak, Pulkit Agrawal, Alexei A. Efros, and Trevor Darrell · 2017
Cited alongside, same era.
Proximal Policy Optimization Algorithms
Region Growing Curriculum Generation for Reinforcement Learning
Artem Molchanov, Karol Hausman, Stan Birchfield, and Gaurav Sukhatme · 2018
Later among the works it cites.
Data-Efficient Hierarchical Reinforcement Learning
Ofir Nachum, Shixiang Gu, Honglak Lee, and Sergey Levine · 2018
Later among the works it cites.
Visual Reinforcement Learning with Imagined Goals
Ashvin Nair, Vitchyr Pong, Murtaza Dalal, Shikhar Bahl, Steven Lin, and Sergey Levine · 2018
Later among the works it cites.
Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration
Alexandre Péré, Sébastien Forestier, Olivier Sigaud, and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, Timothy Lillicrap, Karen Simonyan, and Demis Hassabis · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov · 2017
Cited alongside, same era.
#Exploration: A Study of Count-Based Exploration for Deep Reinforcement Learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2017
Cited alongside, same era.
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, Shane Legg, and Koray Kavukcuoglu · 2018
Cited alongside, same era.
Automatic Goal Generation for Reinforcement Learning Agents
Carlos Florensa, David Held, Xinyang Geng, and Pieter Abbeel · 2018
Cited alongside, same era.
Learning Actionable Representations with Goal-Conditioned Policies
Dibya Ghosh, Abhishek Gupta, and Sergey Levine · 2018
Cited alongside, same era.
Large-Scale Study of Curiosity-Driven Learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros
Cited in the paper.
Exploration by Random Network Distillation
Yuri Burda, Harrison Edwards, Amos Storkey, and Oleg Klimov
Cited in the paper.
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
Diversity Is All You Need: Learning Skills Without a Reward Function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2019
Closest in time.
Competitive Experience Replay
Hao Liu, Alexander Trott, Richard Socher, and Caiming Xiong · 2019
Closest in time.
Episodic Curiosity through Reachability
Nikolay Savinov, Anton Raichuk, Raphaël Marinier, Damien Vincent, Marc Pollefeys, Timothy Lillicrap, and Sylvain Gelly · 2019
Closest in time.
Unsupervised Control Through Non-Parametric Discriminative Rewards
David Warde-Farley, Tom Van de Wiele, Tejas Kulkarni, Catalin Ionescu, Steven Hansen, and Volodymyr Mnih · 2019
Closest in time.