Fetching the paper…
Reading the bibliography…
Normalization layers are one of the key building blocks for deep neural networks.
Asymptotic behavior of group integrals in the limit of infinite rank
Don Weingarten · 1978
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Learning in linear neural networks: A survey
Pierre F Baldi and Kurt Hornik · 1995
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
Effect of batch learning in multilayer neural networks
Kenji Fukumizu · 1998
Earlier work this paper cites.
Moments and cumulants of polynomial random variables on unitary groups, the Itzykson-Zuber integral, and free probability
Benoît Collins · 2003
Earlier work this paper cites.
Integration with respect to the haar measure on unitary, orthogonal and symplectic group
Benoît Collins and Piotr Śniady · 2006
Earlier work this paper cites.
On some properties of orthogonal Weingarten functions
Benoît Collins and Sho Matsumoto · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Earlier work this paper cites.
On polynomial integrals over the orthogonal group
Teodor Banica, Benoit Collins, and Jean-Marc Schlenker · 2011
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
Andrew L Maas, Awni Y Hannun, and Andrew Y Ng · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on ImageNet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Dmytro Mishkin and Jiri Matas · 2015
Earlier work this paper cites.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Unitary evolution recurrent neural networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2016
Earlier work this paper cites.
Recurrent orthogonal networks and long-memory tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Cited alongside, same era.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Global optimality conditions for deep neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2017
Cited alongside, same era.
Nonlinear random matrix theory for deep learning
Jeffrey Pennington and Pratik Worah · 2017
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Later among the works it cites.
Batch normalization provably avoids ranks collapse for randomly initialised deep networks
Hadi Daneshmand, Jonas Kohler, Francis Bach, Thomas Hofmann, and Aurelien Lucchi · 2020
Later among the works it cites.
Larger-scale transformers for multilingual masked language modeling
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Cited alongside, same era.
Efficient orthogonal parametrisation of recurrent neural networks using Householder reflections
Zakaria Mhammedi, Andrew Hellicar, Ashfaqur Rahman, and James Bailey · 2017
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Cited alongside, same era.
Kronecker recurrent units
Cijo Jose, Moustapha Cissé, and Francois Fleuret · 2018
Cited alongside, same era.
Haiping Wu, Bin Xiao, Noel Codella, Mengchen Liu, Xiyang Dai, Lu Yuan, and Lei Zhang · 2021
Later among the works it cites.
Attention is not all you need: Pure attention loses rank doubly exponentially with depth
Yihe Dong, Jean-Baptiste Cordonnier, and Andreas Loukas · 2021
Later among the works it cites.
Batch normalization orthogonalizes representations in deep random networks
Hadi Daneshmand, Amir Joudaki, and Francis Bach · 2021
Later among the works it cites.
James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2021
Later among the works it cites.
Beyond BatchNorm: towards a unified understanding of normalization in deep learning
Ekdeep S Lubana, Robert Dick, and Hidenori Tanaka · 2021
Later among the works it cites.
A ConvNet for the 2020s
Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie · 2022
Later among the works it cites.
Rank diminishing in deep neural networks
Ruili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao, Michael Jordan, and Zheng-Jun Zha · 2022
Later among the works it cites.
Signal propagation in Transformers: Theoretical perspectives and the role of rank collapse
Lorenzo Noci, Sotiris Anagnostidis, Luca Biggio, Antonio Orvieto, Sidak Pal Singh, and Aurelien Lucchi · 2022
Later among the works it cites.
The neural covariance SDE: Shaped infinite depth-and-width networks at initialization
Mufan Li, Mihai Nica, and Dan Roy · 2022
Later among the works it cites.
Benoit Collins, Sho Matsumoto, and Jonathan Novak · 2022
Later among the works it cites.
Deep learning without shortcuts: Shaping the kernel with tailored rectifiers
Guodong Zhang, Aleksandar Botev, and James Martens · 2022
Later among the works it cites.
ConvNeXt V2: Co-designing and scaling convnets with masked autoencoders
Sanghyun Woo, Shoubhik Debnath, Ronghang Hu, Xinlei Chen, Zhuang Liu, In So Kweon, and Saining Xie · 2023
Closest in time.
Deep transformers without shortcuts: Modifying self-attention for faithful signal propagation
Bobby He, James Martens, Guodong Zhang, Aleksandar Botev, Andrew Brock, Samuel L Smith, and Yee Whye Teh · 2023
Closest in time.
The shaped transformer: Attention models in the infinite depth-and-width limit
Lorenzo Noci, Chuning Li, Mufan Bill Li, Bobby He, Thomas Hofmann, Chris Maddison, and Daniel M Roy · 2023
Closest in time.