Fetching the paper…
Reading the bibliography…
Batch normalization is a key component of most image classification models, but it has many undesirable properties stemming from its dependence on the batch size and interactions between examples.
Fixup initialization: Residual learning without normalization
Zhang, H., Dauphin, Y. N., and Ma, T · 1901
Earlier work this paper cites.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B · 1964
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Efficient backprop
LeCun, Y. A., Bottou, L., Orr, G. B., and Müller, K.-R · 2012
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Pascanu, R., Mikolov, T., and Bengio, Y · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R · 2014
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
ImageNet large scale visual recognition challenge
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2015
Earlier work this paper cites.
Srivastava, R. K., Greff, K., and Schmidhuber, J · 2015
Earlier work this paper cites.
Normalization propagation: A parametric technique for removing internal covariate shift in deep networks
Arpit, D., Zhou, Y., Kota, B., and Govindaraju, V · 2016
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Hendrycks, D. and Gimpel, K · 2016
Earlier work this paper cites.
Deep networks with stochastic depth
Huang, G., Sun, Y., Liu, Z., Sedra, D., and Weinberger, K. Q · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Loshchilov, I. and Hutter, F · 2016
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks
Radford, A., Metz, L., and Chintala, S · 2016
Earlier work this paper cites.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
Balduzzi, D., Frean, M., Leary, L., Lewis, J., Ma, K. W.-D., and McWilliams, B · 2017
Earlier work this paper cites.
Gitman, I. and Ginsburg, B · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Goyal, P., Dollár, P., Girshick, R., Noordhuis, P., Wesolowski, L., Kyrola, A., Tulloch, A., Jia, Y., and He, K · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Hoffer, E., Hubara, I., and Soudry, D · 2017
Earlier work this paper cites.
Centered weight normalization in accelerating training of deep neural networks
Huang, L., Liu, X., Liu, Y., Lang, B., and Tao, D · 2017
Earlier work this paper cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Ioffe, S · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Revisiting unreasonable effectiveness of data in deep learning era
Sun, C., Shrivastava, A., Singh, S., and Gupta, A · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Cited alongside, same era.
Aggregated residual transformations for deep neural networks
Xie, S., Girshick, R., Dollár, P., Tu, Z., and He, K · 2017
Cited alongside, same era.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
mixup: Beyond empirical risk minimization
Zhang, H., Cisse, M., Dauphin, Y. N., and Lopez-Paz, D · 2017
Cited alongside, same era.
Understanding batch normalization
Bjorck, N., Gomes, C. P., Selman, B., and Weinberger, K. Q · 2018
Cited alongside, same era.
The DeepMind JAX Ecosystem, 2020
Babuschkin, I., Baumli, K., Bell, A., Bhupatiraju, S., Bruce, J., Buchlovsky, P., Budden, D., Cai, T., Clark, A., Danihelka, I., Fantacci, C., Godwin, J., Jones, C., Hennigan, T., Hessel, M., Kapturowski, S., Keck, T., Kemaev, I., King, M., Martens, L., Mikulik, V., Norman, T., Quan, J., Papamakarios, G., Ring, R., Ruiz, F., Sanchez, A., Schneider, R., Sezener, E., Spencer, S., Srinivasan, S., Stokowiec, W., and Viola, F · 2020
Later among the works it cites.
Rezero is all you need: Fast convergence at large depth
Bachlechner, T., Majumder, B. P., Mao, H. H., Cottrell, G. W., and McAuley, J · 2020
Later among the works it cites.
On the distance between two neural networks and the stability of learning
Bernstein, J., Vahdat, A., Yue, Y., and Liu, M.-Y · 2020
Later among the works it cites.
A simple framework for contrastive learning of visual representations
Chen, T., Kornblith, S., Norouzi, M., and Hinton, G · 2020
Later among the works it cites.
Randaugment: Practical automated data augmentation with a reduced search space
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
JAX: composable transformations of Python+NumPy programs, 2018
Bradbury, J., Frostig, R., Hawkins, P., Johnson, M. J., Leary, C., Maclaurin, D., and Wanderman-Milne, S · 2018
Cited alongside, same era.
Faster neural networks straight from jpeg
Gueguen, L., Sergeev, A., Kadlec, B., Liu, R., and Yosinski, J · 2018
Cited alongside, same era.
How to start training: The effect of initialization and architecture
Hanin, B. and Rolnick, D · 2018
Cited alongside, same era.
Squeeze-and-excitation networks
Hu, J., Shen, L., and Sun, G · 2018
Cited alongside, same era.
Towards understanding regularization in batch normalization
Luo, P., Wang, X., Shao, W., and Peng, Z · 2018
Cited alongside, same era.
Exploring the limits of weakly supervised pretraining
Mahajan, D., Girshick, R., Ramanathan, V., He, K., Paluri, M., Li, Y., Bharambe, A., and Van Der Maaten, L · 2018
Cited alongside, same era.
Regularizing and optimizing LSTM language models
Merity, S., Keskar, N. S., and Socher, R · 2018
Cited alongside, same era.
Cubuk, E. D., Zoph, B., Shlens, J., and Le, Q. V · 2020
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
De, S. and Smith, S · 2020
Later among the works it cites.
Maxup: A simple way to improve generalization of neural network training
Gong, C., Ren, T., Ye, M., and Liu, Q · 2020
Later among the works it cites.
Array programming with numpy
Harris, C. R., Millman, K. J., van der Walt, S. J., Gommers, R., Virtanen, P., Cournapeau, D., Wieser, E., Taylor, J., Berg, S., Smith, N. J., Kern, R., Picus, M., Hoyer, S., van Kerkwijk, M. H., Brett, M., Haldane, A., del Río, J. F., Wiebe, M., Peterson, P., Gérard-Marchant, P., Sheppard, K., Reddy, T., Weckesser, W., Abbasi, H., Gohlke, C., and Oliphant, T. E · 2020
Later among the works it cites.
Momentum contrast for unsupervised visual representation learning
He, K., Fan, H., Wu, Y., Xie, S., and Girshick, R · 2020
Later among the works it cites.
Haiku: Sonnet for JAX, 2020
Hennigan, T., Cai, T., Norman, T., and Babuschkin, I · 2020
Later among the works it cites.
Hooker, S · 2020
Later among the works it cites.
Normalization techniques in training dnns: Methodology, analysis and application
Huang, L., Qin, J., Zhou, Y., Zhu, F., Liu, L., and Shao, L · 2020
Later among the works it cites.
Scaling laws for neural language models
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2020
Later among the works it cites.
Pham, H., Xie, Q., Dai, Z., and Le, Q. V · 2020
Later among the works it cites.
Resizemix: Mixing data with preserved object information and true labels
Qin, J., Fang, J., Zhang, Q., Liu, W., Wang, X., and Wang, X · 2020
Later among the works it cites.
Designing network design spaces
Radosavovic, I., Kosaraju, R. P., Girshick, R., He, K., and Dollár, P · 2020
Later among the works it cites.
Is normalization indispensable for training deep neural network?
Shao, J., Hu, K., Wang, C., Xue, X., and Raj, B · 2020
Later among the works it cites.
Powernorm: Rethinking batch normalization in transformers
Shen, S., Yao, Z., Gholami, A., Mahoney, M., and Keutzer, K · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent
Smith, S., Elsen, E., and De, S · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2020
Later among the works it cites.
Self-training with noisy student improves imagenet classification
Xie, Q., Luong, M.-T., Hovy, E., and Le, Q. V · 2020
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Zhang, J., He, T., Sra, S., and Jadbabaie, A · 2020
Later among the works it cites.
Lambdanetworks: Modeling long-range interactions without attention
Bello, I · 2021
Closest in time.
Characterizing signal propagation to close the performance gap in unnormalized resnets
Brock, A., De, S., and Smith, S. L · 2021
Closest in time.
An image is worth 16x16 words: Transformers for image recognition at scale
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2021
Closest in time.
Sharpness-aware minimization for efficiently improving generalization
Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B · 2021
Closest in time.
Cloud TPU Performance Guide
Google · 2021
Closest in time.
Bottleneck transformers for visual recognition
Srinivas, A., Lin, T.-Y., Parmar, N., Shlens, J., Abbeel, P., and Vaswani, A · 2021
Closest in time.