Fetching the paper…
Reading the bibliography…
We introduce an exploration bonus for deep reinforcement learning methods that is easy to implement and adds minimal overhead to the computation performed.
Curious model-building control systems
Jürgen Schmidhuber · 1991
Earlier work this paper cites.
Intrinsic motivation systems for autonomous mental development
Pierre-Yves Oudeyer, Frdric Kaplan, and Verena V Hafner · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
Ali Rahimi and Benjamin Recht · 2008
Earlier work this paper cites.
An analysis of model-based interval estimation for markov decision processes
Alexander L Strehl and Michael L Littman · 2008
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Yann LeCun, et al · 2009
Earlier work this paper cites.
On random weights and unsupervised feature learning
Andrew M Saxe, Pang Wei Koh, Zhenghao Chen, Maneesh Bhand, Bipin Suresh, and Andrew Y Ng · 2011
Earlier work this paper cites.
Exploration in model-based reinforcement learning by empirically estimating learning progress
Manuel Lopes, Tobias Lang, Marc Toussaint, and Pierre-Yves Oudeyer · 2012
Earlier work this paper cites.
An information-theoretic approach to curiosity-driven reinforcement learning
Susanne Still and Doina Precup · 2012
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A. Rusu, Joel Veness, Marc G. Bellemare, Alex Graves, Martin Riedmiller, Andreas K. Fidjeland, Georg Ostrovski, Stig Petersen, Charles Beattie, Amir Sadik, Ioannis Antonoglou, Helen King, Dharshan Kumaran, Daan Wierstra, Shane Legg, and Demis Hassabis · 2015
Earlier work this paper cites.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Earlier work this paper cites.
Deep fried convnets
Zichao Yang, Marcin Moczulski, Misha Denil, Nando de Freitas, Alex Smola, Le Song, and Ziyu Wang · 2015
Earlier work this paper cites.
Unifying count-based exploration and intrinsic motivation
Marc Bellemare, Sriram Srinivasan, Georg Ostrovski, Tom Schaul, David Saxton, and Remi Munos · 2016
Earlier work this paper cites.
Learning to perform physics experiments via deep reinforcement learning
Misha Denil, Pulkit Agrawal, Tejas D Kulkarni, Tom Erez, Peter Battaglia, and Nando de Freitas · 2016
Earlier work this paper cites.
VIME: Variational information maximizing exploration
Rein Houthooft, Xi Chen, Yan Duan, John Schulman, Filip De Turck, and Pieter Abbeel · 2016
Earlier work this paper cites.
Hierarchical deep reinforcement learning: Integrating temporal abstraction and intrinsic motivation
Tejas D Kulkarni, Karthik Narasimhan, Ardavan Saeedi, and Josh Tenenbaum · 2016
Earlier work this paper cites.
Asynchronous methods for deep reinforcement learning
Volodymyr Mnih, Adria Puigdomenech Badia, Mehdi Mirza, Alex Graves, Timothy Lillicrap, Tim Harley, David Silver, and Koray Kavukcuoglu · 2016
Earlier work this paper cites.
Deep exploration via bootstrapped DQN
Ian Osband, Charles Blundell, Alexander Pritzel, and Benjamin Van Roy · 2016
Cited alongside, same era.
Mastering the game of Go with deep neural networks and tree search
David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, George van den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, Sander Dieleman, Dominik Grewe, John Nham, Nal Kalchbrenner, Ilya Sutskever, Timothy Lillicrap, Madeleine Leach, Koray Kavukcuoglu, Thore Graepel, and Demis Hassabis · 2016
Cited alongside, same era.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Cited alongside, same era.
Surprise-based intrinsic motivation for deep reinforcement learning
Joshua Achiam and Shankar Sastry · 2017
Cited alongside, same era.
Large-scale study of curiosity-driven learning
Yuri Burda, Harri Edwards, Deepak Pathak, Amos Storkey, Trevor Darrell, and Alexei A. Efros · 2018
Closest in time.
IMPALA: Scalable distributed Deep-RL with importance weighted actor-learner architectures
Lasse Espeholt, Hubert Soyer, Remi Munos, Karen Simonyan, Volodymir Mnih, Tom Ward, Yotam Doron, Vlad Firoiu, Tim Harley, Iain Dunning, et al · 2018
Closest in time.
Diversity is all you need: Learning skills without a reward function
Benjamin Eysenbach, Abhishek Gupta, Julian Ibarz, and Sergey Levine · 2018
Closest in time.
Dora the explorer: Directed outreaching reinforcement action-selection
Lior Fox, Leshem Choshen, and Yonatan Loewenstein · 2018
Closest in time.
Expert-augmented actor-critic for vizdoom and montezumas revenge
Michał Garmulewicz, Henryk Michalewski, and Piotr Miłoś · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marc G Bellemare, Will Dabney, and Rémi Munos · 2017
Cited alongside, same era.
UCB and infogain exploration via q q -ensembles
Richard Y Chen, John Schulman, Pieter Abbeel, and Szymon Sidor · 2017
Cited alongside, same era.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, Charles Blundell, and Shane Legg · 2017
Cited alongside, same era.
EX2: Exploration with exemplar models for deep reinforcement learning
Justin Fu, John D Co-Reyes, and Sergey Levine · 2017
Cited alongside, same era.
Variational intrinsic control
Karol Gregor, Danilo Jimenez Rezende, and Daan Wierstra · 2017
Cited alongside, same era.
Rainbow: Combining improvements in deep reinforcement learning
Matteo Hessel, Joseph Modayil, Hado Van Hasselt, Tom Schaul, Georg Ostrovski, Will Dabney, Dan Horgan, Bilal Piot, Mohammad Azar, and David Silver · 2017
Cited alongside, same era.
Marlos C Machado, Marc G Bellemare, Erik Talvitie, Joel Veness, Matthew Hausknecht, and Michael Bowling · 2017
Cited alongside, same era.
The uncertainty Bellman equation and exploration
Brendan O’Donoghue, Ian Osband, Remi Munos, and Volodymyr Mnih · 2017
Cited alongside, same era.
Closest in time.
Learning to play with intrinsically-motivated self-aware agents
Nick Haber, Damian Mrowca, Li Fei-Fei, and Daniel LK Yamins · 2018
Closest in time.
Distributed prioritized experience replay
Dan Horgan, John Quan, David Budden, Gabriel Barth-Maron, Matteo Hessel, Hado Van Hasselt, and David Silver · 2018
Closest in time.
Count-based exploration with the successor representation
Marlos C Machado, Marc G Bellemare, and Michael Bowling · 2018
Closest in time.
Junhyuk Oh, Yijie Guo, Satinder Singh, and Honglak Lee · 2018
Closest in time.
OpenAI Five
OpenAI · 2018
Closest in time.
Learning Dexterous In-Hand Manipulation
OpenAI, :, M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba · 2018
Closest in time.
Randomized prior functions for deep reinforcement learning
Ian Osband, John Aslanides, and Albin Cassirer · 2018
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos · 2018
Closest in time.
Observe and look further: Achieving consistent performance on Atari
Tobias Pohlen, Bilal Piot, Todd Hester, Mohammad Gheshlaghi Azar, Dan Horgan, David Budden, Gabriel Barth-Maron, Hado van Hasselt, John Quan, Mel Večerík, et al · 2018
Closest in time.
Temporal difference models: Model-free deep RL for model-based control
Vitchyr Pong, Shixiang Gu, Murtaza Dalal, and Sergey Levine · 2018
Closest in time.
Learning Montezuma’s Revenge from a single demonstration
Tim Salimans and Richard Chen · 2018
Closest in time.
Christopher Stanton and Jeff Clune · 2018
Closest in time.
Intrinsic motivation and automatic curricula via asymmetric self-play
Sainbayar Sukhbaatar, Ilya Kostrikov, Arthur Szlam, and Rob Fergus · 2018
Closest in time.