Fetching the paper…
Reading the bibliography…
A critically important, ubiquitous, and yet poorly understood ingredient in modern deep networks (DNs) is batch normalization (BN), which centers and normalizes the feature maps.
A note on alternative regressions
Paul A Samuelson · 1942
Earlier work this paper cites.
Facing up to arrangements: Face-count formulas for partitions of space by hyperplanes: Face-count formulas for partitions of space by hyperplanes , volume 154
Thomas Zaslavsky · 1975
Earlier work this paper cites.
An analysis of the total least squares problem
Gene H Golub and Charles F Van Loan · 1980
Earlier work this paper cites.
Constructive approximation , volume 303
Ronald A DeVore and George G Lorentz · 1993
Earlier work this paper cites.
Efficient backprop, neural networks: Tricks of the trade
Y LeCun, L Bottou, G Orr, and K Muller · 1998
Earlier work this paper cites.
An introduction to hyperplane arrangements
Richard P Stanley et al · 2004
Earlier work this paper cites.
On decompositional algorithms for uniform sampling from n-spheres and n-balls
Radoslav Harman and Vladimír Lacko · 2010
Earlier work this paper cites.
A distance-based point-reassignment heuristic for the k-hyperplane clustering problem
Edoardo Amaldi and Stefano Coniglio · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Improving neural networks with dropout
Nitish Srivastava · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
Guido F Montufar, Razvan Pascanu, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Dropout improves recurrent neural networks for handwriting recognition
Vu Pham, Théodore Bluche, Christopher Kermorvant, and Jérôme Louradour · 2014
Earlier work this paper cites.
Mathematical theory of probability and statistics
Richard Von Mises · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Dmytro Mishkin and Jiri Matas · 2015
Cited alongside, same era.
Deep Learning , volume 1
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Cited alongside, same era.
Caglar Gulcehre, Marcin Moczulski, Francesco Visin, and Yoshua Bengio · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Durk P Kingma · 2016
Cited alongside, same era.
Can we gain more from orthogonality regularizations in training deep networks?
Nitin Bansal, Xiaohan Chen, and Zhangyang Wang · 2018
Later among the works it cites.
Understanding batch normalization
Nils Bjorck, Carla P Gomes, Bart Selman, and Kilian Q Weinberger · 2018
Later among the works it cites.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Later among the works it cites.
Bounding and counting linear regions of deep neural networks
Thiago Serra, Christian Tjandraatmadja, and Srikumar Ramalingam · 2018
Later among the works it cites.
A max-affine spline perspective of recurrent neural networks
Zichao Wang, Randall Balestriero, and Richard Baraniuk · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Sergey Ioffe · 2017
Cited alongside, same era.
Improving training of deep neural networks via singular value bounding
Kui Jia, Dacheng Tao, Shenghua Gao, and Xiangmin Xu · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Dmitry Molchanov, Arsenii Ashukha, and Dmitry Vetrov · 2017
Cited alongside, same era.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl Dickstein · 2017
Cited alongside, same era.
Efficiently sampling vectors and coordinates from the n-sphere and n-ball
Aaron R Voelker, Jan Gosmann, and Terrence C Stewart · 2017
Cited alongside, same era.
All you need is beyond a good init: Exploring better solution for training extremely deep convolutional neural networks with orthonormality and modulation
Di Xie, Jiang Xiong, and Shiliang Pu · 2017
Cited alongside, same era.
Later among the works it cites.
From hard to soft: Understanding deep network nonlinearities via vector quantization and statistical inference
Randall Balestriero and Richard Baraniuk · 2019
Later among the works it cites.
The geometry of deep networks: Power diagram subdivision
Randall Balestriero, Romain Cosentino, Behnaam Aazhang, and Richard Baraniuk · 2019
Later among the works it cites.
Exponential convergence rates for batch normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, Thomas Hofmann, Ming Zhou, and Klaus Neymeyr · 2019
Later among the works it cites.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2019
Later among the works it cites.
Mad max: Affine spline insights into deep learning
Randall Balestriero and Richard G. Baraniuk · 2020
Later among the works it cites.
Neural architecture search on imagenet in four GPU hours: A theoretically inspired perspective
Wuyang Chen, Xinyu Gong, and Zhangyang Wang · 2021
Later among the works it cites.
Singular value perturbation and deep network optimization
Rudolf H Riedi, Randall Balestriero, and Richard G Baraniuk · 2022
Closest in time.