Fetching the paper…
Reading the bibliography…
Normalization Layers (NLs) are widely used in modern deep-learning architectures.
A. Blum, R. L. Rivest, Training a 3-node neural network is np-complete, in: Advances in Neural Information Processing Systems (NeurIPS), 1989
1989
Earlier work this paper cites.
A. Krizhevsky, G. Hinton, et al., Learning multiple layers of features from tiny images, Tech. rep., University of Toronto (2009)
2009
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng, Reading digits in natural images with unsupervised feature learning, in: NIPS Workshop on Deep Learning and Unsupervised Feature Learning, 2011
2011
Earlier work this paper cites.
I. Sutskever, J. Martens, G. Dahl, G. Hinton, On the importance of initialization and momentum in deep learning, in: 30th International Conference on Machine Learning (ICML), Atlanta, Georgia, USA, 2013
2013
Earlier work this paper cites.
S. Ioffe, C. Szegedy, Batch normalization: Accelerating deep network training by reducing internal covariate shift, in: 32nd International Conference on Machine Learning (ICML), 2015
2015
Earlier work this paper cites.
D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: International Conference on Learning Representations (ICLR), 2015
2015
Earlier work this paper cites.
K. Simonyan, A. Zisserman, Very deep convolutional networks for large-scale image recognition, in: International Conference on Learning Representations (ICLR), 2015
2015
Earlier work this paper cites.
T. Salimans, D. P. Kingma, Weight normalization: A simple reparameterization to accelerate training of deep neural networks, in: Advances in neural information processing systems (NeurIPS), 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016
2016
Earlier work this paper cites.
M. Cho, J. Lee, Riemannian approach to batch normalization, in: Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
Earlier work this paper cites.
2017
Cited alongside, same era.
W. Liu, Y.-M. Zhang, X. Li, Z. Yu, B. Dai, T. Zhao, L. Song, Deep hyperspherical learning, in: Advances in Neural Information Processing Systems (NeurIPS), 2017
2017
Cited alongside, same era.
Y. Wu, K. He, Group normalization, in: Proceedings of the European Conference on Computer Vision (ECCV), 2018
2018
Cited alongside, same era.
S. Santurkar, D. Tsipras, A. Ilyas, A. Madry, How does batch normalization help optimization?, in: Advances in Neural Information Processing Systems (NeurIPS), 2018
2018
Cited alongside, same era.
N. Bjorck, C. P. Gomes, B. Selman, K. Q. Weinberger, Understanding batch normalization, in: Advances in Neural Information Processing Systems (NeurIPS), 2018
I. Loshchilov, F. Hutter, Decoupled weight decay regularization, in: International Conference on Learning Representations (ICLR), 2019
2019
Later among the works it cites.
B. Ghorbani, S. Krishnan, Y. Xiao, An investigation into neural net optimization via hessian eigenvalue density, in: 36th International Conference on Machine Learning (ICML), 2019
2019
Later among the works it cites.
X. Lian, J. Liu, Revisit batch normalization: New understanding and refinement via composition optimization, in: The 22nd International Conference on Artificial Intelligence and Statistics (AISTATS), 2019
2019
Later among the works it cites.
R. Karakida, S. Akaho, S.-i. Amari, The normalization method for alleviating pathological sharpness in wide neural networks, in: NeurIPS, 2019
2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
E. Hoffer, I. Hubara, D. Soudry, Fix your classifier: the marginal value of training the last weight layer, in: International Conference on Learning Representations (ICLR), 2018
2018
Cited alongside, same era.
E. Hoffer, R. Banner, I. Golan, D. Soudry, Norm matters: efficient and accurate normalization schemes in deep networks, in: Advances in Neural Information Processing Systems (NeurIPS), 2018
2018
Cited alongside, same era.
W. Liu, Z. Liu, Z. Yu, B. Dai, R. Lin, Y. Wang, J. M. Rehg, L. Song, Decoupled networks, in: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018
2018
Cited alongside, same era.
S. Arora, Z. Li, K. Lyu, Theoretical analysis of auto rate-tuning by batch normalization, in: International Conference on Learning Representations (ICLR), 2019
2019
Cited alongside, same era.
J. L. Ba, J. R. Kiros, G. E. Hinton, Layer normalization, arXiv preprint arXiv:1607.06450
Cited in the paper.
Cited in the paper.
J. Duchi, E. Hazan, Y. Singer, Adaptive subgradient methods for online learning and stochastic optimization, Journal of Machine Learning Research (JMLR)
Cited in the paper.
2019
Later among the works it cites.
Y. Cai, Q. Li, Z. Shen, A quantitative analysis of the effect of batch normalization on gradient descent, in: 36th International Conference on Machine Learning (ICML), 2019
2019
Later among the works it cites.
G. Zhang, C. Wang, B. Xu, R. Grosse, Three mechanisms of weight decay regularization, in: International Conference on Learning Representations (ICLR), 2019
2019
Later among the works it cites.
H. Daneshmand, J. Kohler, F. Bach, T. Hofmann, A. Lucchi, Batch normalization provably avoids rank collapse for randomly initialised deep networks, in: NeurIPS, 2020
2020
Closest in time.
Z. Li, S. Arora, An exponential learning rate schedule for deep learning, in: International Conference on Learning Representations (ICLR), 2020
2020
Closest in time.