Fetching the paper…
Reading the bibliography…
Training very deep neural networks is still an extremely challenging task.
An efficient method for finding the minimum of a function of several variables without calculating derivatives
Michael JD Powell · 1964
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1996
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 1998
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
Sepp Hochreiter, Yoshua Bengio, Paolo Frasconi, Jürgen Schmidhuber, et al · 2001
Earlier work this paper cites.
Nonlinear systems third edition
Hassan K. Khalil · 2008
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Avoiding pathologies in very deep networks
David Duvenaud, Oren Rippel, Ryan Adams, and Zoubin Ghahramani · 2014
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Rupesh Kumar Srivastava, Klaus Greff, and Jürgen Schmidhuber · 2015
Earlier work this paper cites.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Cited alongside, same era.
Residual networks behave like ensembles of relatively shallow networks
Andreas Veit, Michael J Wilber, and Serge Belongie · 2016
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Later among the works it cites.
Scalable second order optimization for deep learning
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
Thomas Bachlechner, Bodhisattwa Prasad Majumder, Huanru Henry Mao, Garrison W Cottrell, and Julian McAuley · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Distributed second-order optimization using kronecker-factored approximations
Jimmy Ba, Roger Grosse, and James Martens · 2017
Cited alongside, same era.
The shattered gradients problem: If resnets are the answer, then what is the question?
David Balduzzi, Marcus Frean, Lennox Leary, JP Lewis, Kurt Wan-Duo Ma, and Brian McWilliams · 2017
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Cited alongside, same era.
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2020
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Sam Smith · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Tom Hennigan, Trevor Cai, Tamara Norman, and Igor Babuschkin · 2020
Later among the works it cites.
Optax: composable gradient transformation and optimisation, in jax!, 2020
Matteo Hessel, David Budden, Fabio Viola, Mihaela Rosca, Eren Sezener, and Tom Hennigan · 2020
Later among the works it cites.
Bidirectional self-normalizing neural networks
Yao Lu, Stephen Gould, and Thalaiyasingam Ajanthan · 2020
Later among the works it cites.
Going deeper with neural networks without skip connections
Oyebade K Oyedotun, Djamila Aouada, Björn Ottersten, et al · 2020
Later among the works it cites.
Deep isometric learning for visual recognition
Haozhi Qi, Chong You, Xiaolong Wang, Yi Ma, and Jitendra Malik · 2020
Later among the works it cites.
Is normalization indispensable for training deep neural network?
Jie Shao, Kai Hu, Changhu Wang, Xiangyang Xue, and Bhiksha Raj · 2020
Later among the works it cites.
Disentangling trainability and generalization in deep neural networks
Lechao Xiao, Jeffrey Pennington, and Samuel Schoenholz · 2020
Later among the works it cites.
Repvgg: Making vgg-style convnets great again
Xiaohan Ding, Xiangyu Zhang, Ningning Ma, Jungong Han, Guiguang Ding, and Jian Sun · 2021
Later among the works it cites.
Stable resnet
Soufiane Hayou, Eugenio Clerico, Bobby He, George Deligiannidis, Arnaud Doucet, and Judith Rousseau · 2021
Later among the works it cites.
On the validity of kernel approximations for orthogonally-initialized neural networks
James Martens · 2021
Later among the works it cites.
James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2021
Later among the works it cites.