Fetching the paper…
Reading the bibliography…
The selection of initial parameter values for gradient-based optimization of deep neural networks is one of the most impactful hyperparameter choices in deep learning systems, affecting both convergence times and model performance.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
Geoffrey Hinton, Li Deng, Dong Yu, George E. Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, et al · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Bounds for the ratio of two gamma functions—from wendel’s and related inequalities to logarithmically completely monotonic functions
Feng Qi and Qiu-Ming Luo · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Earlier work this paper cites.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Earlier work this paper cites.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
A recurrent neural network without chaos
Thomas Laurent and James von Brecht · 2016
Earlier work this paper cites.
Full-capacity unitary recurrent neural networks
Scott Wisdom, Thomas Powers, John Hershey, Jonathan Le Roux, and Les Atlas · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Stable architectures for deep neural networks
Eldad Haber and Lars Ruthotto · 2017
Cited alongside, same era.
Depth creates no bad local minima
Haihao Lu and Kenji Kawaguchi · 2017
Cited alongside, same era.
Efficient orthogonal parametrisation of recurrent neural networks using householder reflections
Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman, and James Bailey · 2017
Minmin Chen, Jeffrey Pennington, and Samuel S Schoenholz · 2018
Later among the works it cites.
Deep linear networks with arbitrary loss: All local minima are global
Thomas Laurent and James von Brecht · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2018
Later among the works it cites.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2017
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
Gradient descent with identity initialization efficiently learns positive definite linear transformations
Peter Bartlett, Dave Helmbold, and Phil Long · 2018
Cited alongside, same era.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Later among the works it cites.
Critical points of linear neural networks: Analytical forms and landscape properties
Yi Zhou and Yingbin Liang · 2018
Later among the works it cites.
Width provably matters in optimization for deep linear neural networks
Simon Du and Wei Hu · 2019
Later among the works it cites.
Dynamical isometry and a mean field theory of lstms and grus
Dar Gilboa, Bo Chang, Minmin Chen, Greg Yang, Samuel S Schoenholz, Ed H Chi, and Jeffrey Pennington · 2019
Later among the works it cites.
Spectrum concentration in deep residual learning: a free probability approach
Zenan Ling and Robert C Qiu · 2019
Later among the works it cites.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzȩbski, Jacek Tabor, and Maciej Nowak · 2019
Later among the works it cites.