Fetching the paper…
Reading the bibliography…
Randomly initialized neural networks are known to become harder to train with increasing depth, unless architectural enhancements like residual connections and batch normalization are used.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Products of Random Matrices with Applications to Schrödinger Operators
Philippe Bougerol · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Earlier work this paper cites.
Lyapunov exponents for products of complex Gaussian random matrices
Peter J. Forrester · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A. Tropp · 2015
Earlier work this paper cites.
Proofs, beliefs, and algorithms through the lens of sum-of-squares
Boaz Barak and David Steurer · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Bulk and soft-edge universality for singular values of products of ginibre random matrices
Dang-Zheng Liu, Dong Wang, and Lun Zhang · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Cited alongside, same era.
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Cited alongside, same era.
Benefits of depth in neural networks
Matus Telgarsky · 2016
Cited alongside, same era.
Bridging the gap between constant step size stochastic gradient descent and Markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Understanding batch normalization, 2018
Nils Bjorck, Carla P. Gomes, Bart Selman, and Kilian Q. Weinberger · 2018
Later among the works it cites.
Markov Chains
Randal Douc, Eric Moulines, Pierre Priouret, and Philippe Soulier · 2018
Later among the works it cites.
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Ming Zhou, Klaus Neymeyr, and Thomas Hofmann · 2018
Later among the works it cites.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2018
Later among the works it cites.
How does batch normalization help optimization?(no, it is not about internal covariate shift)
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Cited alongside, same era.
Fashion-mnist: a novel image dataset for benchmarking machine learning algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Cited alongside, same era.
Later among the works it cites.
Group normalization
Yuxin Wu and Kaiming He · 2018
Later among the works it cites.
Gradient descent with identity initialization efficiently learns positive-definite linear transformations by deep residual networks
Peter L. Bartlett, David P. Helmbold, and Philip M. Long · 2019
Later among the works it cites.
The normalization method for alleviating pathological sharpness in wide neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Later among the works it cites.
PyHessian: Neural networks through the lens of the Hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael Mahoney · 2019
Later among the works it cites.