Fetching the paper…
Reading the bibliography…
Batch Normalization (BatchNorm) is commonly used in Convolutional Neural Networks (CNNs) to improve training speed and stability.
B. Widrow, “An adaptive adaline neuron using chemical memistors,” Tech. Rep. TR-1553-2, Stanford Electronics Labs, Stanford, CA, Oct. 1960
1960
Earlier work this paper cites.
B. Widrow and M. Hoff, “Adaptive switching circuits,” Tech. Rep. TR-1553-1, Stanford Electronics Labs, Stanford, CA, June 1960
1960
Earlier work this paper cites.
B. Widrow, “Adaptive filters,” in Aspects of Network and System Theory
1971
Earlier work this paper cites.
Y. LeCun, L. Jackel, B. Boser, J. Denker, H. Graf, I. Guyon, D. Henderson, R. Howard, and W. Hubbard, “Handwritten digit recognition: Applications of neural network chips and automatic learning,” IEEE Communications Magazine
1989
Earlier work this paper cites.
Y. LeCun, P. Simard, and B. Pearlmutter, “Automatic learning rate maximization by on-line estimation of the hessian’s eigenvectors,” in Advances in Neural Information Processing Systems 7
1993
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE
1998
Earlier work this paper cites.
Upper Saddle River, NJ: Prentice Hall, 4th ed., 2002
S. Haykin, Adaptive Filter Theory · 2002
Earlier work this paper cites.
2002
Earlier work this paper cites.
A. Krizhevsky and G. Hinton, “Learning multiple layers of features from tiny images,” 2009
2009
Cited alongside, same era.
X. Glorot and Y. Bengio, “Understanding the difficulty of training deep feedforward neural networks,” in Proceedings of the Thirteenth International Conference on Artificial Intelligence and Statistics
2010
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 25
2012
Cited alongside, same era.
2014
Cited alongside, same era.
S. Ioffe and C. Szegedy, “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” in Proceedings of the 32nd International Conference on Machine Learning
D. Balduzzi, M. Frean, L. Leary, J. Lewis, K. Ma, and B. McWilliams, “The shattered gradients problem: If resnets are the answer, then what is the question?,” in Proceedings of the 34th International Conference on Machine Learning
2017
Later among the works it cites.
B. Hanin, “Which neural net architectures give rise to exploding and vanishing gradients?,” in Advances in Neural Information Processing Systems 32
2018
Later among the works it cites.
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry, “How does batch normalization help optimization?,” in Advances in Neural Information Processing Systems 32
2018
Later among the works it cites.
H. Zhang, W. Chen, and T. Liu, “On the local hessian in back-propagation,” in Advances in Neural Information Processing Systems 32
2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” in Proceedings of the IEEE International Conference on Computer Vision
2015
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
2016
Cited alongside, same era.
D. Mishkin and J. Matas, “All you need is a good init,” in 4th International Conference on Learning Representations
2016
Cited alongside, same era.
2019
Later among the works it cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 33
2019
Later among the works it cites.
H. Zhang, Y. Dauphin, and T. Ma, “Fixup initialization: Residual learning without normalization,” in 7th International Conference on Learning Representations
2019
Later among the works it cites.