Fetching the paper…
Reading the bibliography…
Batch Normalization (BN) is capable of accelerating the training of deep models by centering and scaling activations within mini-batches.
Possible principles underlying the transformations of sensory messages
H. B. Barlow · 1961
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Learning factorial codes by predictability minimization
J. Schmidhuber · 1992
Earlier work this paper cites.
The ”independent components” of natural scenes are edge filters
A. J. Bell and T. J. Sejnowski · 1997
Earlier work this paper cites.
Effiicient backprop
Y. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Earlier work this paper cites.
Accelerated gradient descent by factor-centering decomposition
N. N. Schraudolph · 1998
Earlier work this paper cites.
From few to many: Illumination cone models for face recognition under variable lighting and pose
A. Georghiades, P. Belhumeur, and D. Kriegman · 2001
Earlier work this paper cites.
The cmu pose, illumination, and expression (pie) database
T. Sim, S. Baker, and M. Bsat · 2002
Earlier work this paper cites.
Learning a spatially smooth subspace for face recognition
D. Cai, X. He, Y. Hu, J. Han, and T. Huang · 2007
Earlier work this paper cites.
Collected Matrix Derivative Results for Forward and Reverse Mode Algorithmic Differentiation
M. B. Giles · 2008
Earlier work this paper cites.
Slow, decorrelated features for pretraining complex cell-like networks
Y. Bengio and J. S. Bergstra · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
A convergence analysis of log-linear training
S. Wiesler and H. Ney · 2011
Earlier work this paper cites.
Deep Boltzmann Machines and the Centering Trick
G. Montavon and K.-R. Müller · 2012
Earlier work this paper cites.
Deep learning made easier by linear transformations in perceptrons
T. Raiko, H. Valpola, and Y. LeCun · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
A. M. Saxe, J. L. McClelland, and S. Ganguli · 2013
Earlier work this paper cites.
cudnn: Efficient primitives for deep learning
S. Chetlur, C. Woolley, P. Vandermersch, J. Cohen, J. Tran, B. Catanzaro, and E. Shelhamer · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Mean-normalized stochastic gradient for large-scale deep learning
S. Wiesler, A. Richard, R. Schlüter, and H. Ney · 2014
Cited alongside, same era.
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, and k. kavukcuoglu · 2015
Cited alongside, same era.
Batch normalized recurrent neural networks
C. Laurent, G. Pereyra, P. Brakel, Y. Zhang, and Y. Bengio · 2016
Later among the works it cites.
Q. Liao, K. Kawaguchi, and T. Poggio · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. P. Kingma · 2016
Later among the works it cites.
Relative natural gradient for learning large complex models
K. Sun and F. Nielsen · 2016
Later among the works it cites.
Inception-v4, inception-resnet and the impact of residual connections on learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Natural neural networks
G. Desjardins, K. Simonyan, R. Pascanu, and K. Kavukcuoglu · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Training deep networks with structured layers by matrix backpropagation
C. Ionescu, O. Vantzos, and C. Sminchisescu · 2015
Cited alongside, same era.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. Salakhutdinov, and N. Srebro · 2015
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2015
Cited alongside, same era.
C. Szegedy, S. Ioffe, and V. Vanhoucke · 2016
Later among the works it cites.
Regularizing deep convolutional neural networks with a structured decorrelation constraint
W. Xiong, B. Du, L. Zhang, R. Hu, and D. Tao · 2016
Later among the works it cites.
S. Zagoruyko and N. Komodakis · 2016
Later among the works it cites.
Riemannian approach to batch normalization
M. Cho and J. Lee · 2017
Later among the works it cites.
Recurrent batch normalization
T. Cooijmans, N. Ballas, C. Laurent, and A. C. Courville · 2017
Later among the works it cites.
Projection based weight normalization for deep neural networks
L. Huang, X. Liu, B. Lang, and B. Li · 2017
Later among the works it cites.
Centered weight normalization in accelerating training of deep neural networks
L. Huang, X. Liu, Y. Liu, B. Lang, and D. Tao · 2017
Later among the works it cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
S. Ioffe · 2017
Later among the works it cites.
Optimal whitening and decorrelation
A. Kessy, A. Lewin, and K. Strimmer · 2017
Later among the works it cites.
Learning deep architectures via generalized whitened neural networks
P. Luo · 2017
Later among the works it cites.
Normalizing the normalizers: Comparing and extending network normalization schemes
M. Ren, R. Liao, R. Urtasun, F. H. Sinz, and R. S. Zemel · 2017
Later among the works it cites.
Regularizing cnns with locally constrained decorrelations
P. Rodríguez, J. Gonzàlez, G. Cucurull, J. M. Gonfaus, and F. X. Roca · 2017
Later among the works it cites.
On the effects of batch and weight normalization in generative adversarial networks
S. Xiang and H. Li · 2017
Later among the works it cites.
Orthogonal weight normalization: Solution to optimization over multiple dependent stiefel manifolds in deep neural networks
L. Huang, X. Liu, B. Lang, A. W. Yu, Y. Wang, and B. Li · 2018
Closest in time.