Fetching the paper…
Reading the bibliography…
Batch Normalization (BN) improves both convergence and generalization in training neural networks.
On the ability of the optimal perceptron to generalise
M. Opper, W. Kinzel, J. Kleinz, and R. Nehl · 1990
Earlier work this paper cites.
Generalization in a linear perceptron in the presence of noise
Anders Krogh and John A. Hertz · 1992
Earlier work this paper cites.
Statistical mechanics of learning from examples
H. S. Seung, Haim Sompolinsky, and N. Tishby · 1992
Earlier work this paper cites.
Training with Noise is Equivalent to Tikhonov Regularization
Chris M. Bishop · 1995
Earlier work this paper cites.
Dynamics of on-line gradient descent learning for multilayer neural networks
David Saad and Sara A. Solla · 1996
Earlier work this paper cites.
Statistical mechanics approach to early stopping and weight decay
Siegfried Bös · 1998
Earlier work this paper cites.
Dynamics of batch training in a perceptron
Siegfried Bs and Manfred Opper · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Adding noise to the input of a model trained with a regularized objective
Salah Rifai, Xavier Glorot, Yoshua Bengio, and Pascal Vincent · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton · 2012
Earlier work this paper cites.
Dropout Training as Adaptive Regularization
Stefan Wager, Sida Wang, and Percy Liang · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Rethinking the Inception Architecture for Computer Vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jonathon Shlens, and Zbigniew Wojna · 2015
Cited alongside, same era.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
On the expressive power of deep neural networks
Maithra Raghu, Ben Poole, Jon Kleinberg, Surya Ganguli, and Jascha Sohl Dickstein · 2017
Later among the works it cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Later among the works it cites.
L2 regularization versus batch and weight normalization
Twan. van Laarhoven · 2017
Later among the works it cites.
Statistical mechanical analysis of online learning with weight normalization in single layer perceptron
Yuki Yoshida, Ryo Karakida, Masato Okada, and Shun ichi Amari · 2017
Later among the works it cites.
Robust Implicit Backpropagation
Francois Fagan and Garud Iyengar · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Song Mei, Yu Bai, and Andrea Montanari · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S. Advani and Andrew M. Saxe · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Igor Gitman and Boris Ginsburg · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Understanding the disharmony between dropout and batch normalization by variance shift
Xiang Li, Shuo Chen, Xiaolin Hu, and Jian Yang · 2018
Closest in time.
Do normalization layers in a deep convnet really need to be distinct?
Ping Luo, Zhanglin Peng, Jiamin Ren, and Ruimao Zhang · 2018
Closest in time.
On the importance of single directions for generalization
Ari S. Morcos, David G.T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Closest in time.
How Does Batch Normalization Help Optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Closest in time.
Bayesian uncertainty estimation for batch normalized deep networks
Mattias Teye, Hossein Azizpour, and Kevin Smith · 2018
Closest in time.
Yuxin Wu and Kaiming He · 2018
Closest in time.
Differentiable learning-to-normalize via switchable normalization
Ping Luo, Jiamin Ren, Zhanglin Peng, Ruimao Zhang, and Jingyu Li · 2019
Closest in time.
Switchable whitening for deep representation learning
Xingang Pan, Xiaohang Zhan, Jianping Shi, Xiaoou Tang, and Ping Luo · 2019
Closest in time.
Ssn: Learning sparse switchable normalization via sparsestmax
Wenqi Shao, Tianjian Meng, Jingyu Li, Ruimao Zhang, Yudian Li, Xiaogang Wang, and Ping Luo · 2019
Closest in time.