Fetching the paper…
Reading the bibliography…
Batch Normalization (BN) is widely used to stabilize the optimization process and improve the test performance of deep neural networks.
Micro-Batch Training with Batch-Channel Normalization and Weight Standardization
Qiao, S.; Wang, H.; Liu, C.; Shen, W.; and Yuille, A. 2020 · 1903
Earlier work this paper cites.
Online normalization for training neural networks
Chiley, V.; Sharapov, I.; Kosson, A.; Koster, U.; Reece, R.; Samaniego de la Fuente, S.; Subbiah, V.; and James, M. 2019 · 1905
Earlier work this paper cites.
Four Things Everyone Should Know to Improve Batch Normalization
Summers, C.; and Dinneen, M. J. 2020 · 1906
Earlier work this paper cites.
Dropout: A Simple Way to Prevent Neural Networks from Overfitting
Srivastava, N.; Hinton, G.; Krizhevsky, A.; Sutskever, I.; and Salakhutdinov, R. 2014 · 1958
Earlier work this paper cites.
CIFAR-100 (Canadian Institute For Advanced Research)
Krizhevsky, A.; and Hinton, G. 2009 · 2009
Earlier work this paper cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N. 2021 · 2010
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research
Deng, L. 2012 · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
Ba, J. L.; Kiros, J. R.; and Hinton, G. E. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016 · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T.; and Kingma, D. P. 2016 · 2016
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E.; Hubara, I.; and Soudry, D. 2017 · 2017
Cited alongside, same era.
Arbitrary style transfer in real-time with adaptive instance normalization
Huang, X.; and Belongie, S. 2017 · 2017
Cited alongside, same era.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Ioffe, S. 2017 · 2017
Cited alongside, same era.
On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
Keskar, N. S.; Mudigere, D.; Nocedal, J.; Smelyanskiy, M.; and Tang, P. T. P. 2017 · 2017
Cited alongside, same era.
Stochastic normalizations as bayesian learning
Shekhovtsov, A.; and Flach, B. 2019 · 2018
Cited alongside, same era.
Group normalization
Explicit regularisation in gaussian noise injections
Camuto, A.; Willetts, M.; Simsekli, U.; Roberts, S. J.; and Holmes, C. C. 2020 · 2020
Later among the works it cites.
On the Relationship between Self-Attention and Convolutional Layers
Cordonnier, J.-B.; Loukas, A.; and Jaggi, M. 2020 · 2020
Later among the works it cites.
The implicit and explicit regularization effects of dropout
Wei, C.; Kakade, S.; and Ma, T. 2020 · 2020
Later among the works it cites.
High-performance large-scale image recognition without normalization
Brock, A.; De, S.; Smith, S. L.; and Simonyan, K. 2021 · 2021
Later among the works it cites.
Rethinking" batch" in batchnorm
Wu, Y.; and Johnson, J. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wu, Y.; and He, K. 2018 · 2018
Cited alongside, same era.
Weighted Channel Dropout for Regularization of Deep Convolutional Neural Network
Hou, S.; and Wang, Z. 2019 · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Cited alongside, same era.
Evalnorm: Estimating batch normalization statistics for evaluation
Singh, S.; and Shrivastava, A. 2019 · 2019
Cited alongside, same era.
PyTorch Image Models
Wightman, R. 2019 · 2019
Cited alongside, same era.
Zagoruyko, S.; and Komodakis, N. 2017 · 2019
Cited alongside, same era.
Beyer, L.; Zhai, X.; and Kolesnikov, A. 2022 · 2022
Later among the works it cites.
Locality-Aware Channel-Wise Dropout for Occluded Face Recognition
He, M.; Zhang, J.; Shan, S.; Liu, X.; Wu, Z.; and Chen, X. 2022 · 2022
Later among the works it cites.
NoMorelization: Building Normalizer-Free Models from a Sample’s Perspective
Liu, C.; Yang, Y.; Ding, Y.; and Lu, H. 2022 · 2022
Later among the works it cites.
Patches Are All You Need?
Trockman, A.; and Kolter, J. Z. 2023 · 2023
Closest in time.
vit-pytorch
Wang, P. 2023 · 2023
Closest in time.