Fetching the paper…
Reading the bibliography…
Attention based Transformer architecture has enabled significant advances in the field of natural language processing.
Long short-term memory
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
One billion word benchmark for measuring progress in statistical language modeling
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, and P. Koehn · 2013
Earlier work this paper cites.
On the properties of neural machine translation: Encoder–decoder approaches
K. Cho, B. van Merrienboer, D. Bahdanau, and Y. Bengio · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
I. Sutskever, O. Vinyals, and Q. V. Le · 2014
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
P. Rajpurkar, J. Zhang, K. Lopyrev, and P. Liang · 2016
Earlier work this paper cites.
Convolutional sequence to sequence learning
J. Gehring, M. Auli, D. Grangier, D. Yarats, and Y. N. Dauphin · 2017
Earlier work this paper cites.
Deep reinforcement learning: An overview
Y. Li · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Breaking the softmax bottleneck: A high-rank rnn language model
Z. Yang, Z. Dai, R. Salakhutdinov, and W. W. Cohen · 2017
Earlier work this paper cites.
State-of-the-art speech recognition with sequence-to-sequence models
C.-C. Chiu, T. N. Sainath, Y. Wu, R. Prabhavalkar, P. Nguyen, Z. Chen, A. Kannan, R. J. Weiss, K. Rao, E. Gonina, et al · 2018
Cited alongside, same era.
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Generating wikipedia by summarizing long sequences
P. J. Liu, M. Saleh, E. Pot, B. Goodrich, R. Sepassi, L. Kaiser, and N. Shazeer · 2018
Cited alongside, same era.
Scaling neural machine translation
M. Ott, S. Edunov, D. Grangier, and M. Auli · 2018
Cited alongside, same era.
Generating long sequences with sparse transformers
R. Child, S. Gray, A. Radford, and I. Sutskever · 2019
Later among the works it cites.
Adaptively sparse transformers
G. M. Correia, V. Niculae, and A. F. Martins · 2019
Later among the works it cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Later among the works it cites.
Token-level ensemble distillation for grapheme-to-phoneme conversion
H. Sun, X. Tan, J.-W. Gan, H. Liu, S. Zhao, T. Qin, and T.-Y. Liu · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Cited alongside, same era.
Tensor2tensor for neural machine translation
A. Vaswani, S. Bengio, E. Brevdo, F. Chollet, A. N. Gomez, S. Gouws, L. Jones, L. Kaiser, N. Kalchbrenner, N. Parmar, R. Sepassi, N. Shazeer, and J. Uszkoreit · 2018
Cited alongside, same era.
Non-local neural networks
X. Wang, R. Girshick, A. Gupta, and K. He · 2018
Cited alongside, same era.
A broad-coverage challenge corpus for sentence understanding through inference
A. Williams, N. Nangia, and S. Bowman · 2018
Cited alongside, same era.
Relational deep reinforcement learning
V. Zambaldi, D. Raposo, A. Santoro, V. Bapst, Y. Li, I. Babuschkin, K. Tuyls, D. Reichert, T. Lillicrap, E. Lockhart, et al · 2018
Cited alongside, same era.
Self-attention generative adversarial networks
H. Zhang, I. Goodfellow, D. Metaxas, and A. Odena · 2018
Cited alongside, same era.
E. Voita, D. Talbot, F. Moiseev, R. Sennrich, and I. Titov · 2019
Later among the works it cites.
Pay less attention with lightweight and dynamic convolutions
F. Wu, A. Fan, A. Baevski, Y. N. Dauphin, and M. Auli · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. Salakhutdinov, and Q. V. Le · 2019
Later among the works it cites.
Reducing bert pre-training time from 3 days to 76 minutes
Y. You, J. Li, J. Hseu, X. Song, J. Demmel, and C.-J. Hsieh · 2019
Later among the works it cites.
Are transformers universal approximators of sequence-to-sequence functions?
C. Yun, S. Bhojanapalli, A. S. Rawat, S. J. Reddi, and S. Kumar · 2019
Later among the works it cites.