Fetching the paper…
Reading the bibliography…
Layer normalization (LayerNorm) is a technique to normalize the distributions of intermediate layers.
Building a large annotated corpus of english: The penn treebank
M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz · 1993
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W. Zhu · 2002
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
B. Pang and L. Lee · 2005
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Earlier work this paper cites.
The iwslt 2015 evaluation campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, and M. Federico · 2014
Earlier work this paper cites.
A fast and accurate dependency parser using neural networks
D. Chen and C. D. Manning · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
D. Bahdanau, K. Cho, and Y. Bengio · 2015
Earlier work this paper cites.
The iwslt 2015 evaluation campaign
M. Cettolo, J. Niehues, S. Stüker, L. Bentivogli, R. Cattoni, and M. Federico · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
J. Lei Ba, J. R. Kiros, and G. E. Hinton · 2016
Cited alongside, same era.
Sequence level training with recurrent neural networks
M. Ranzato, S. Chopra, M. Auli, and W. Zaremba · 2016
Cited alongside, same era.
Instance normalization: The missing ingredient for fast stylization
D. Ulyanov, A. Vedaldi, and V. S. Lempitsky · 2016
Cited alongside, same era.
Sequence-to-sequence learning as beam-search optimization
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Later among the works it cites.
Character-level language modeling with deeper self-attention
R. Al-Rfou, D. Choe, N. Constant, M. Guo, and L. Jones · 2018
Later among the works it cites.
Understanding batch normalization
N. Bjorck, C. P. Gomes, B. Selman, and K. Q. Weinberger · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Wiseman and A. M. Rush · 2016
Cited alongside, same era.
Hierarchical multiscale recurrent neural networks
J. Chung, S. Ahn, and Y. Bengio · 2017
Cited alongside, same era.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Cited alongside, same era.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A. Mayr, and S. Hochreiter · 2017
Cited alongside, same era.
Online and linear-time attention by enforcing monotonic alignments
C. Raffel, M. Luong, P. J. Liu, R. J. Weiss, and D. Eck · 2017
Cited alongside, same era.
S. Santurkar, D. Tsipras, A. Ilyas, and A. Madry · 2018
Later among the works it cites.
Group normalization
Y. Wu and K. He · 2018
Later among the works it cites.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, W. W. Cohen, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Closest in time.
Fixup initialization: Residual learning without normalization
H. Zhang, Y. N. Dauphin, and T. Ma · 2019
Closest in time.