Fetching the paper…
Reading the bibliography…
Dealing with sparse rewards is a longstanding challenge in reinforcement learning.
“A possibility for implementing curiosity and boredom in model-building neural controllers”
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
“Empowerment: A universal agent-centric measure of control”
Alexander Klyubin, Daniel Polani and Chrystopher Nehaniv · 2005
Earlier work this paper cites.
“Curriculum learning”
Yoshua Bengio, Jérôme Louradour, Ronan Collobert and Jason Weston · 2009
Earlier work this paper cites.
“Formal theory of creativity, fun, and intrinsic motivation (1990–2010)”
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
“PILCO: A model-based and data-efficient approach to policy search”
Marc Deisenroth and Carl Rasmussen · 2011
Earlier work this paper cites.
“Learning to control a low-cost manipulator using data-efficient reinforcement learning”
Marc Deisenroth, Carl Rasmussen and Dieter Fox · 2011
Earlier work this paper cites.
“Hierarchical curiosity loops and active sensing”
Goren Gordon and Ehud Ahissar · 2012
Earlier work this paper cites.
“Learning and exploration in action-perception loops”
Daniel-Jeh Little and Friedrich Sommer · 2013
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“Continuous control with deep reinforcement learning”
Timothy Lillicrap, Jonathan Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver and Daan Wierstra · 2015
Earlier work this paper cites.
“Variational information maximisation for intrinsically motivated reinforcement learning”
Shakir Mohamed and Danilo Rezende · 2015
Earlier work this paper cites.
“Unifying count-based exploration and intrinsic motivation”
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton and Remi Munos · 2016
Earlier work this paper cites.
“Variational intrinsic control”
Karol Gregor, Danilo Rezende and Daan Wierstra · 2016
Earlier work this paper cites.
“Learning values across many orders of magnitude”
Hado van Hasselt, Arthur Guez, Matteo Hessel, Volodymyr Mnih and David Silver · 2016
Earlier work this paper cites.
“Vime: Variational information maximizing exploration”
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De and Pieter Abbeel · 2016
Earlier work this paper cites.
“Hindsight experience replay”
Marcin Andrychowicz, Filip Wolski, Alex Ray, Jonas Schneider, Rachel Fong, Peter Welinder, Bob McGrew, Josh Tobin, OpenAI Abbeel and Wojciech Zaremba · 2017
Cited alongside, same era.
“One-shot imitation learning”
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel and Wojciech Zaremba · 2017
Cited alongside, same era.
“Reverse curriculum generation for reinforcement learning”
Carlos Florensa, David Held, Markus Wulfmeier, Michael Zhang and Pieter Abbeel · 2017
Cited alongside, same era.
“Intrinsically motivated goal exploration processes with automatic curriculum learning”
Sébastien Forestier, Yoan Mollard and Pierre-Yves Oudeyer · 2017
Cited alongside, same era.
“Count-based exploration with neural density models”
Georg Ostrovski, Marc Bellemare, Aaron Oord and Rémi Munos · 2017
“Openai baselines” https://github.com/openai/baselines , 2018
Prafulla Dhariwal, Christopher Hesse, Oleg Klimov, Alex Nichol, Matthias Plappert, Alec Radford, John Schulman, Szymon Sidor, Yuhuai Wu and Peter Zhokhov · 2018
Later among the works it cites.
“Curriculum goal masking for continuous deep reinforcement learning”
Manfred Eppe, Sven Magg and Stefan Wermter · 2018
Later among the works it cites.
“Accuracy-based Curriculum Learning in Deep Reinforcement Learning”
Pierre Fournier, Olivier Sigaud, Mohamed Chetouani and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
“Learning to Play with Intrinsically-Motivated Self-Aware Agents”
Nick Haber, Damian Mrowca, Li Fei-Fei and Daniel Yamins · 2018
Later among the works it cites.
“Multi-objective Model-based Policy Search for Data-efficient Learning with Sparse Rewards”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Curiosity-driven exploration by self-supervised prediction”
Deepak Pathak, Pulkit Agrawal, Alexei Efros and Trevor Darrell · 2017
Cited alongside, same era.
“Parameter space noise for exploration”
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Chen, Xi Chen, Tamim Asfour, Pieter Abbeel and Marcin Andrychowicz · 2017
Cited alongside, same era.
“Data-efficient deep reinforcement learning for dexterous manipulation”
Ivaylo Popov, Nicolas Heess, Timothy Lillicrap, Roland Hafner, Gabriel Barth-Maron, Matej Vecerik, Thomas Lampe, Yuval Tassa, Tom Erez and Martin Riedmiller · 2017
Cited alongside, same era.
Paulo Rauber, Avinash Ummadisingu, Filipe Mutz and Juergen Schmidhuber · 2017
Cited alongside, same era.
“Leveraging demonstrations for deep reinforcement learning on robotics problems with sparse rewards”
Matej Večerik, Todd Hester, Jonathan Scholz, Fumin Wang, Olivier Pietquin, Bilal Piot, Nicolas Heess, Thomas Rothörl, Thomas Lampe and Martin Riedmiller · 2017
Cited alongside, same era.
“Large-scale study of curiosity-driven learning”
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell and Alexei Efros · 2018
Cited alongside, same era.
“Exploration by random network distillation”
Yuri Burda, Harri Edwards, Amos Storkey and Oleg Klimov · 2018
Cited alongside, same era.
Rituraj Kaushik, Konstantinos Chatzilygeroudis and Jean-Baptiste Mouret · 2018
Later among the works it cites.
“Curiosity driven exploration of learned disentangled goal spaces”
Adrien Laversanne-Finot, Alexandre Péré and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
“Overcoming exploration in reinforcement learning with demonstrations”
Ashvin Nair, Bob McGrew, Marcin Andrychowicz, Wojciech Zaremba and Pieter Abbeel · 2018
Later among the works it cites.
“Unsupervised Learning of Goal Spaces for Intrinsically Motivated Goal Exploration”
Alexandre Péré, Sébastien Forestier, Olivier Sigaud and Pierre-Yves Oudeyer · 2018
Later among the works it cites.
“Multi-goal reinforcement learning: Challenging robotics environments and request for research”
Matthias Plappert, Marcin Andrychowicz, Alex Ray, Bob McGrew, Bowen Baker, Glenn Powell, Jonas Schneider, Josh Tobin, Maciek Chociej and Peter Welinder · 2018
Later among the works it cites.
“Learning by Playing-Solving Sparse Reward Tasks from Scratch”
Martin Riedmiller, Roland Hafner, Thomas Lampe, Michael Neunert, Jonas Degrave, Tom Van, Volodymyr Mnih, Nicolas Heess and Jost Springenberg · 2018
Later among the works it cites.
“Curiosity-Driven Experience Prioritization via Density Estimation”, 2018
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
“Energy-Based Hindsight Experience Prioritization”
Rui Zhao and Volker Tresp · 2018
Later among the works it cites.
“Go-Explore: a New Approach for Hard-Exploration Problems”
Adrien Ecoffet, Joost Huizinga, Joel Lehman, Kenneth Stanley and Jeff Clune · 2019
Closest in time.
“Solving the Rubik’s Cube with Approximate Policy Iteration”
Stephen McAleer, Forest Agostinelli, Alexander Shmakov and Pierre Baldi · 2019
Closest in time.