Fetching the paper…
Reading the bibliography…
Efforts to improve the learning abilities of neural networks have focused mostly on the role of optimization methods rather than on weight initializations.
Daniel. Park, Jascha Sohl-Dickstein, Quoc. Le and Samuel. Smith · 1905
Earlier work this paper cites.
“Deconstructing Lottery Tickets: Zeros, Signs, and the Supermask” arXiv: 1905.01067
Hattie Zhou, Janice Lan, Rosanne Liu and Jason Yosinski · 1905
Earlier work this paper cites.
“Optimal brain damage”
Yann LeCun, John Denker and Sara Solla · 1990
Earlier work this paper cites.
“Second order derivatives for network pruning: Optimal brain surgeon”
Babak Hassibi and David Stork · 1993
Earlier work this paper cites.
“An ensemble of neural networks for weather forecasting”
Imran Maqsood, Muhammad Khan and Ajith Abraham · 2004
Earlier work this paper cites.
“An efficient weather forecasting system using artificial neural network”
S Baboo and I Shereef · 2010
Earlier work this paper cites.
“Understanding the difficulty of training deep feedforward neural networks”
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
“Application of artificial neural networks in weather forecasting: a comprehensive literature review”
Gyanesh Shrivastava, Sanjeev Karmakar, Manoj Kowar and Pulak Guhathakurta · 2012
Earlier work this paper cites.
Alex Graves, Greg Wayne and Ivo Danihelka · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization”
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“The loss surfaces of multilayer networks”
Anna Choromanska et al · 2015
Earlier work this paper cites.
“Learning both Weights and Connections for Efficient Neural Network”
Song Han, Jeff Pool, John Tran and William Dally · 2015
Cited alongside, same era.
“Distilling the knowledge in a neural network”
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Cited alongside, same era.
“Inferring algorithmic patterns with stack-augmented recurrent nets”
Armand Joulin and Tomas Mikolov · 2015
Cited alongside, same era.
“Neural gpus learn algorithms”
Łukasz Kaiser and Ilya Sutskever · 2015
Cited alongside, same era.
“TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems” Software available from tensorflow.org, 2015
Martín et al · 2015
Cited alongside, same era.
“Gradient descent finds global minima of deep neural networks”
Simon Du et al · 2018
Later among the works it cites.
“The lottery ticket hypothesis: Finding sparse, trainable neural networks”
Jonathan Frankle and Michael Carbin · 2018
Later among the works it cites.
“Measuring the Intrinsic Dimension of Objective Landscapes” arXiv: 1804.08838
Chunyuan Li, Heerad Farkhoor, Rosanne Liu and Jason Yosinski · 2018
Later among the works it cites.
Arvind Mohan and Datta Gaitonde · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Deep learning without poor local minima”
Kenji Kawaguchi · 2016
Cited alongside, same era.
“Pruning filters for efficient convnets”
Hao Li et al · 2016
Cited alongside, same era.
“All you need is a good init” arXiv: 1511.06422
Dmytro Mishkin and Jiri Matas · 2016
Cited alongside, same era.
“The loss surface of deep and wide neural networks”
Quynh Nguyen and Matthias Hein · 2017
Cited alongside, same era.
Sanjeev Arora, Nadav Cohen and Elad Hazan · 2018
Cited alongside, same era.
“Winning Ways for Your Mathematical Plays, Volume 2”
Elwyn Berlekamp, John Conway and Richard Guy · 2018
Cited alongside, same era.
Behnam Neyshabur et al · 2018
Later among the works it cites.
“Are efficient deep representations learnable?”
Maxwell Nye and Andrew Saxe · 2018
Later among the works it cites.
“Neural arithmetic logic units”
Andrew Trask et al · 2018
Later among the works it cites.
“MetaInit: Initializing learning by learning to initialize”
Yann Dauphin and Samuel Schoenholz · 2019
Later among the works it cites.
“Elimination of all bad local minima in deep learning”
Kenji Kawaguchi and Leslie Kaelbling · 2019
Later among the works it cites.
“Towards moderate overparameterization: global convergence guarantees for training shallow neural networks”
S. Oymak and M. Soltanolkotabi · 2020
Closest in time.