Fetching the paper…
Reading the bibliography…
Transformers are responsible for the vast majority of recent advances in natural language processing.
1907
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
1910
Earlier work this paper cites.
P. Gage, A new algorithm for data compression, C Users Journal 12 (2) (1994) 23–38
1994
Earlier work this paper cites.
2010
Earlier work this paper cites.
M. Schuster, K. Nakajima, Japanese and korean voice search, in: 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2012, pp. 5149–5152
2012
Earlier work this paper cites.
doi:10.18653/v1/P16-1162
R. Sennrich, B. Haddow, A. Birch, Neural machine translation of rare words with subword units , in: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, Berlin, Germany, 2016, pp. 1715–1725 · 2016
Cited alongside, same era.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Attention is all you need, in: Advances in neural information processing systems, 2017, pp. 5998–6008
2017
Cited alongside, same era.
T. Kudo, Subword regularization: Improving neural network translation models with multiple subword candidates, in: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2018, pp. 66–75
2018
Cited alongside, same era.
T. Kudo, J. Richardson, Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing, in: Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, 2018, pp. 66–71
2018
K. Bostrom, G. Durrett, Byte pair encoding is suboptimal for language model pretraining, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, 2020, pp. 4617–4624
2020
Later among the works it cites.
I. Provilkov, D. Emelianenko, E. Voita, Bpe-dropout: Simple and effective subword regularization, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 1882–1892
2020
Later among the works it cites.
A. F. Aji, N. Bogoychev, K. Heafield, R. Sennrich, In neural machine translation, what does transfer learning transfer?, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 7701–7710
2020
Later among the works it cites.
S. Sato, J. Sakuma, N. Yoshinaga, M. Toyoda, M. Kitsuregawa, Vocabulary adaptation for domain adaptation in neural machine translation, in: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: Findings, 2020, pp. 4269–4279
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers for language understanding, in: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 2019, pp. 4171–4186
2019
Cited alongside, same era.
Y. Arase, J. Tsujii, Transfer fine-tuning: A bert case study, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), 2019, pp. 5396–5407
2019
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, Improving language understanding by generative pre-training, URL https://s3-us-west-2. amazonaws. com/openai-assets/researchcovers/languageunsupervised/language understanding paper. pdf
Cited in the paper.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, OpenAI Blog 1 (8)
Cited in the paper.
2020
Later among the works it cites.
Z. Wen, X. H. Lu, S. Reddy, Medal: Medical abbreviation disambiguation dataset for natural language understanding pretraining, in: CLINICALNLP, 2020
2020
Later among the works it cites.
X. Wang, S. Ruder, G. Neubig, Multi-view subword regularization, in: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 473–482
2021
Closest in time.