Fetching the paper…
Reading the bibliography…
Evolution strategies (ES) are a family of black-box optimization algorithms able to train deep neural networks roughly as well as Q-learning and policy gradient methods on challenging deep reinforcement learning (RL) problems, but are much faster (e.g.
Evolutionsstrategien
Ingo Rechenberg · 1978
Earlier work this paper cites.
Deceptiveness and genetic algorithm dynamics
Gunar E. Liepins and Michael D. Vose · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J Williams · 1992
Earlier work this paper cites.
Reinforcement learning: An introduction , volume 1
Richard S Sutton and Andrew G Barto · 1998
Earlier work this paper cites.
Catastrophic forgetting in connectionist networks
Robert M French · 1999
Earlier work this paper cites.
Reducing the time complexity of the derandomized evolution strategy with covariance matrix adaptation (cma-es)
Nikolaus Hansen, Sibylle D Müller, and Petros Koumoutsakos · 2003
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Jan Peters, and Juergen Schmidhuber · 2008
Earlier work this paper cites.
What is intrinsic motivation? a typology of computational approaches
Pierre-Yves Oudeyer and Frederic Kaplan · 2009
Earlier work this paper cites.
Formal theory of creativity, fun, and intrinsic motivation (1990–2010)
Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Parameter-exploring policy gradients
Frank Sehnke, Christian Osendorfer, Thomas Rückstieß, Alex Graves, Jan Peters, and Jürgen Schmidhuber · 2010
Earlier work this paper cites.
Game-independent ai agents for playing atari 2600 console games
Yavar Naddaf · 2010
Earlier work this paper cites.
Deep auto-encoder neural networks in reinforcement learning
Sascha Lange and Martin Riedmiller · 2010
Earlier work this paper cites.
Novelty search and the problem with objectives
Joel Lehman and Kenneth O. Stanley · 2011
Earlier work this paper cites.
When novelty is not enough
Giuseppe Cuccu and Faustino Gomez · 2011
Earlier work this paper cites.
Novelty-based restarts for evolution strategies
Giuseppe Cuccu, Faustino Gomez, and Tobias Glasmachers · 2011
Earlier work this paper cites.
The arcade learning environment: An evaluation platform for general agents
Marc G Bellemare, Yavar Naddaf, Joel Veness, and Michael Bowling · 2013
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling · 2013
Earlier work this paper cites.
Behavioral repertoire learning in robotics
Antoine Cully and Jean-Baptiste Mouret · 2013
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann Dauphin, Razvan Pascanu, Çaglar Gülçehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
Novelty search creates robots with general skills for exploration
Roby Velez and Jeff Clune · 2014
Cited alongside, same era.
Systematic derivation of behaviour characterisations in evolutionary robotics
Jorge Gomes, Pedro Mariano, and Anders Lyhne Christensen · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Openai gym, 2016
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba · 2016
Later among the works it cites.
Conditional image generation with pixelcnn decoders
Aaron van den Oord, Nal Kalchbrenner, Lasse Espeholt, Oriol Vinyals, Alex Graves, et al · 2016
Later among the works it cites.
Curiosity search: producing generalists by encouraging individuals to continually explore and acquire skills throughout their lifetime
Christopher Stanton and Jeff Clune · 2016
Later among the works it cites.
Super mario bros. in openai gym, 2016
Phillip Paquette · 2016
Later among the works it cites.
Learning behavior characterizations for novelty search
Elliot Meyerson, Joel Lehman, and Risto Miikkulainen · 2016
Later among the works it cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen · 2016
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Trust region policy optimization
John Schulman, Sergey Levine, Pieter Abbeel, Michael Jordan, and Philipp Moritz · 2015
Cited alongside, same era.
Incentivizing exploration in reinforcement learning with deep predictive models
Bradly C Stadie, Sergey Levine, and Pieter Abbeel · 2015
Cited alongside, same era.
Robots that can adapt like animals
A. Cully, J. Clune, D. Tarapore, and J.-B. Mouret · 2015
Cited alongside, same era.
Illuminating search spaces by mapping elites
Jean-Baptiste Mouret and Jeff Clune · 2015
Cited alongside, same era.
Confronting the challenge of quality diversity
Justin K Pugh, Lisa B Soros, Paul A Szerlip, and Kenneth O Stanley · 2015
Cited alongside, same era.
Andrei A Rusu, Sergio Gomez Colmenarejo, Caglar Gulcehre, Guillaume Desjardins, James Kirkpatrick, Razvan Pascanu, Volodymyr Mnih, Koray Kavukcuoglu, and Raia Hadsell · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Later among the works it cites.
# exploration: A study of count-based exploration for deep reinforcement learning
Haoran Tang, Rein Houthooft, Davis Foote, Adam Stooke, OpenAI Xi Chen, Yan Duan, John Schulman, Filip DeTurck, and Pieter Abbeel · 2017
Closest in time.
Count-based exploration with neural density models
Georg Ostrovski, Marc G Bellemare, Aaron van den Oord, and Rémi Munos · 2017
Closest in time.
Curiosity-driven exploration by self-supervised prediction
Deepak Pathak, Pulkit Agrawal, Alexei A Efros, and Trevor Darrell · 2017
Closest in time.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, and Ilya Sutskever · 2017
Closest in time.
Parameter space noise for exploration
Matthias Plappert, Rein Houthooft, Prafulla Dhariwal, Szymon Sidor, Richard Y Chen, Xi Chen, Tamim Asfour, Pieter Abbeel, and Marcin Andrychowicz · 2017
Closest in time.
Noisy networks for exploration
Meire Fortunato, Mohammad Gheshlaghi Azar, Bilal Piot, Jacob Menick, Ian Osband, Alex Graves, Vlad Mnih, Remi Munos, Demis Hassabis, Olivier Pietquin, et al · 2017
Closest in time.
Stein variational policy gradient
Yang Liu, Prajit Ramachandran, Qiang Liu, and Jian Peng · 2017
Closest in time.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Closest in time.
Diffusion-based neuromodulation can eliminate catastrophic forgetting in simple neural networks
Roby Velez and Jeff Clune · 2017
Closest in time.
Population based training of neural networks
Max Jaderberg, Valentin Dalibard, Simon Osindero, Wojciech M Czarnecki, Jeff Donahue, Ali Razavi, Oriol Vinyals, Tim Green, Iain Dunning, Karen Simonyan, et al · 2017
Closest in time.
Risto Miikkulainen, Jason Liang, Elliot Meyerson, Aditya Rawal, Dan Fink, Olivier Francon, Bala Raju, Arshak Navruzyan, Nigel Duffy, and Babak Hodjat · 2017
Closest in time.