Fetching the paper…
Reading the bibliography…
The standard normalization method for neural network (NN) models used in Natural Language Processing (NLP) is layer normalization (LN).
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J · 2002
Earlier work this paper cites.
Neighbourhood components analysis
Goldberger, J., Hinton, G. E., Roweis, S. T., and Salakhutdinov, R. R · 2005
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Jarrett, K., Kavukcuoglu, K., Ranzato, M., and LeCun, Y · 2009
Earlier work this paper cites.
Empirical evaluation and combination of advanced language modeling techniques
Mikolov, T., Deoras, A., Kombrink, S., Burget, L., and Černockỳ, J · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E · 2012
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Ba, J. L., Kiros, J. R., and Hinton, G. E · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Salimans, T. and Kingma, D. P · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Instance normalization: The missing ingredient for fast stylization
Ulyanov, D., Vedaldi, A., and Lempitsky, V · 2016
Earlier work this paper cites.
Zagoruyko, S. and Komodakis, N · 2016
Earlier work this paper cites.
Recurrent batch normalization
Cooijmans, T., Ballas, N., Laurent, C., Gülçehre, Ç., and Courville, A · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Huang, G., Liu, Z., Van Der Maaten, L., and Weinberger, K. Q · 2017
Earlier work this paper cites.
Tying word vectors and word classifiers: A loss framework for language modeling
Inan, H., Khosravi, K., and Socher, R · 2017
Earlier work this paper cites.
Batch renormalization: Towards reducing minibatch dependence in batch-normalized models
Ioffe, S · 2017
Cited alongside, same era.
Focal loss for dense object detection
Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Dollár, P · 2017
Cited alongside, same era.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2017
Cited alongside, same era.
Nuerips 2017 test-of-time award presentation, December 2017
Rahimi, A · 2017
Cited alongside, same era.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Positional normalization
Li, B., Wu, F., Weinberger, K. Q., and Belongie, S · 2019
Later among the works it cites.
Roberta: A robustly optimized bert pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V · 2019
Later among the works it cites.
A tensorized transformer for language modeling
Ma, X., Zhang, P., Zhang, S., Duan, N., Hou, Y., Zhou, M., and Song, D · 2019
Later among the works it cites.
Transformers without tears: Improving the normalization of self-attention
Nguyen, T. Q. and Salazar, J · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Ott, M., Edunov, S., Baevski, A., Fan, A., Gross, S., Ng, N., Grangier, D., and Auli, M · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scaling neural machine translation
Ott, M., Edunov, S., Grangier, D., and Auli, M · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Santurkar, S., Tsipras, D., Ilyas, A., and Madry, A · 2018
Cited alongside, same era.
Group normalization
Wu, Y. and He, K · 2018
Cited alongside, same era.
Breaking the softmax bottleneck: A high-rank rnn language model
Yang, Z., Dai, Z., Salakhutdinov, R., and Cohen, W. W · 2018
Cited alongside, same era.
Adaptive input representations for neural language modeling
Baevski, A. and Auli, M · 2019
Cited alongside, same era.
Qiao, S., Wang, H., Liu, C., Shen, W., and Yuille, A · 2019
Later among the works it cites.
Evalnorm: Estimating batch normalization statistics for evaluation
Singh, S. and Shrivastava, A · 2019
Later among the works it cites.
Learning deep transformer models for machine translation
Wang, Q., Li, B., Xiao, T., Zhu, J., Li, C., Wong, D. F., and Chao, L. S · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
Wu, F., Fan, A., Baevski, A., Dauphin, Y., and Auli, M · 2019
Later among the works it cites.
Understanding and improving layer normalization
Xu, J., Sun, X., Zhang, Z., Zhao, G., and Lin, J · 2019
Later among the works it cites.
PyHessian: Neural networks through the lens of the Hessian
Yao, Z., Gholami, A., Keutzer, K., and Mahoney, M. W · 2019
Later among the works it cites.
Root mean square layer normalization
Zhang, B. and Sennrich, R · 2019
Later among the works it cites.
Reducing transformer depth on demand with structured dropout
Fan, A., Grave, E., and Joulin, A · 2020
Closest in time.
Improving neural language generation with spectrum control
Wang, L., Huang, J., Huang, K., Hu, Z., Wang, G., and Gu, Q · 2020
Closest in time.
Towards stabilizing batch statistics in backward propagation of batch normalization
Yan, J., Wan, R., Zhang, X., Zhang, W., Wei, Y., and Sun, J · 2020
Closest in time.