Fetching the paper…
Reading the bibliography…
In this work we show that Evolution Strategies (ES) are a viable method for learning non-differentiable parameters of large supervised models.
The state of sparsity in deep neural networks
Trevor Gale, Erich Elsen, and Sara Hooker · 1902
Earlier work this paper cites.
Evolutionsstrategie–optimierung technisher systeme nach prinzipien der biologischen evolution
Ingo Rechenberg · 1973
Earlier work this paper cites.
Prefix sums and their applications
Guy E Blelloch · 1990
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S. Denker, and Sara A. Solla · 1990
Earlier work this paper cites.
Simple statistical gradient-following algorithms for connectionist reinforcement learning
Ronald J. Williams · 1992
Earlier work this paper cites.
Sparse connection and pruning in large dynamic artificial neural networks
Nikko Ström · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Evolving artificial neural networks
Xin Yao · 1999
Earlier work this paper cites.
Completely derandomized self-adaptation in evolution strategies
Nikolaus Hansen and Andreas Ostermeier · 2001
Earlier work this paper cites.
A comparison of evolution strategies and backpropagation for neural network training
Martin Mandischer · 2002
Earlier work this paper cites.
Neuroevolution for reinforcement learning using evolution strategies
Christian Igel · 2003
Earlier work this paper cites.
Parallel prefix sum (scan) with cuda
Mark Harris, Shubhabrata Sengupta, and John D Owens · 2007
Earlier work this paper cites.
A simple modification in cma-es achieving linear time and space complexity
Raymond Ros and Nikolaus Hansen · 2008
Earlier work this paper cites.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Jan Peters, and Juergen Schmidhuber · 2008
Earlier work this paper cites.
Genesis of organic computing systems: Coupling evolution and learning
Christian Igel and Bernhard Sendhoff · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
High dimensions and heavy tails for natural evolution strategies
Tom Schaul, Tobias Glasmachers, and Jürgen Schmidhuber · 2011
Cited alongside, same era.
Thrust: A productivity-oriented library for cuda
Nathan Bell and Jared Hoberock · 2012
Cited alongside, same era.
Regularization of neural networks using dropconnect
Li Wan, Matthew Zeiler, Sixin Zhang, Yann Le Cun, and Rob Fergus · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Sharan Narang, Erich Elsen Gregory F. Diamos, and Shubho Sengupta · 2017
Later among the works it cites.
Open sourcing Sonnet - a new library for constructing neural networks
Malcolm Reynolds, Gabriel Barth-Maron, Frederic Besse, Diego de Las Casas, Andreas Fidjeland, Tim Green, Adrià Puigdomènech, Sébastien Racanière, Jack Rae, and Fabio Viola · 2017
Later among the works it cites.
Evolution strategies as a scalable alternative to reinforcement learning
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever · 2017
Later among the works it cites.
Don’t Decay the Learning Rate, Increase the Batch Size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Later among the works it cites.
On the relationship between the openai evolution strategy and stochastic gradient descent
Xingwen Zhang, Jeff Clune, and Kenneth O. Stanley · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Cited alongside, same era.
Natural evolution strategies
Daan Wierstra, Tom Schaul, Tobias Glasmachers, Yi Sun, Jan Peters, and Jürgen Schmidhuber · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
François Chollet · 2016
Cited alongside, same era.
EIE: efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally · 2016
Cited alongside, same era.
Deep rewiring: Training very sparse deep networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2017
Cited alongside, same era.
Michael Zhu and Suyog Gupta · 2017
Later among the works it cites.
Efficient neural audio synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron van den Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Later among the works it cites.
Es is more than just a traditional finite-difference approximator
Joel Lehman, Jay Chen, Jeff Clune, and Kenneth O. Stanley · 2018
Later among the works it cites.
Learning sparse neural networks through l0 regularization
Christos Louizos, Max Welling, and Diederik P. Kingma · 2018
Later among the works it cites.
Guided evolutionary strategies: escaping the curse of dimensionality in random search
Niru Maheswaranathan, Luke Metz, George Tucker, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
Simple random search provides a competitive approach to reinforcement learning
Horia Mania, Aurelia Guy, and Benjamin Recht · 2018
Later among the works it cites.
Regularized evolution for image classifier architecture search
Esteban Real, Alok Aggarwal, Yanping Huang, and Quoc V Le · 2018
Later among the works it cites.
Faster gaze prediction with dense networks and fisher pruning
Lucas Theis, Iryna Korshunova, Alykhan Tejani, and Ferenc Huszár · 2018
Later among the works it cites.
A comparative study of large-scale variants of cma-es
Konstantinos Varelas, Anne Auger, Dimo Brockhoff, Nikolaus Hansen, Ouassim Ait ElHara, Yann Semet, Rami Kassab, and Frédéric Barbaresco · 2018
Later among the works it cites.
TF-Replicator: Distributed machine learning for researchers
Peter Buchlovsky, David Budden, Dominik Grewe, Chris Jones, John Aslanides, Frederic Besse, Andy Brock, Aidan Clark, Sergio Gómez Colmenarejo, Aedan Pope, Fabio Viola, and Dan Belov · 2019
Closest in time.