Fetching the paper…
Reading the bibliography…
Pretrained masked language models (MLMs) require finetuning for most NLP tasks.
Rodrigo Nogueira and Kyunghyun Cho. 2019 · 1901
Earlier work this paper cites.
Language models with transformers
Chenguang Wang, Mu Li, and Alexander J Smola. 2019 · 1904
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
HuggingFace’s transformers: State-of-the-art natural language processing
Thomas Wolf, L Debut, V Sanh, J Chaumond, C Delangue, A Moi, P Cistac, T Rault, R Louf, M Funtowicz, et al. 2019 · 1910
Earlier work this paper cites.
Syntactic structures
Noam Chomsky. 1957 · 1957
Earlier work this paper cites.
Statistical analysis of non-lattice data
Julian Besag. 1975 · 1975
Earlier work this paper cites.
Design of a linguistic statistical decoder for the recognition of continuous speech
Frederick Jelinek, Lalit Bahl, and Robert Mercer. 1975 · 1975
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F Brown, Vincent J Della Pietra, Stephen A Della Pietra, and Robert L Mercer. 1993 · 1993
Earlier work this paper cites.
The empirical base of linguistics: Grammaticality judgments and linguistic methodology
Carson T Schütze. 1996 · 1996
Earlier work this paper cites.
Discriminative language modeling with conditional random fields and the perceptron algorithm
Brian Roark, Murat Saraclar, Michael Collins, and Mark Johnson. 2004 · 2004
Earlier work this paper cites.
Bidirectional phrase-based statistical machine translation
Andrew Finch and Eiichiro Sumita. 2009 · 2009
Earlier work this paper cites.
Enhancing language models in statistical machine translation with backward n-grams and mutual information triggers
Deyi Xiong, Min Zhang, and Haizhou Li. 2011 · 2011
Earlier work this paper cites.
Exploiting the succeeding words in recurrent neural network language models
Yangyang Shi, Martha Larson, Pascal Wiggers, and Catholijn M Jonker. 2013 · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Bidirectional recurrent neural network language models for automatic speech recognition
Ebru Arisoy, Abhinav Sethy, Bhuvana Ramabhadran, and Stanley Chen. 2015 · 2015
Earlier work this paper cites.
The IWSLT 2015 evaluation campaign
Mauro Cettolo, Jan Niehues, Sebastian Stüker, Luisa Bentivogli, R Cattoni, and Marcello Federico. 2015 · 2015
Earlier work this paper cites.
On using monolingual corpora in neural machine translation
Caglar Gulcehre, Orhan Firat, Kelvin Xu, Kyunghyun Cho, Loic Barrault, Huei-Chi Lin, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
LibriSpeech: An ASR corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur. 2015 · 2015
Earlier work this paper cites.
Listen, attend and spell: A neural network for large vocabulary conversational speech recognition
William Chan, Navdeep Jaitly, Quoc Le, and Oriol Vinyals. 2016 · 2016
Cited alongside, same era.
Sequence-level knowledge distillation
Yoon Kim and Alexander M Rush. 2016 · 2016
Cited alongside, same era.
Multi-language neural network language models
Anton Ragni, Edgar Dakin, Xie Chen, Mark JF Gales, and Kate M Knill. 2016 · 2016
Cited alongside, same era.
Modeling coverage for neural machine translation
Zhaopeng Tu, Zhengdong Lu, Yang Liu, Xiaohua Liu, and Hang Li. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, et al. 2016 · 2016
Cited alongside, same era.
SwitchOut: An efficient data augmentation algorithm for neural machine translation
Xinyi Wang, Hieu Pham, Zihang Dai, and Graham Neubig. 2018 · 2018
Later among the works it cites.
ESPnet: End-to-end speech processing toolkit
Shinji Watanabe, Takaaki Hori, Shigeki Karita, Tomoki Hayashi, Jiro Nishitoba, Yuya Unno, Nelson Enrique Yalta Soplin, Jahn Heymann, Matthew Wiesner, Nanxin Chen, et al. 2018 · 2018
Later among the works it cites.
Breaking the beam search curse: A study of (re-) scoring methods and stopping criteria for neural machine translation
Yilin Yang, Liang Huang, and Mingbo Ma. 2018 · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Closest in time.
On the use of BERT for neural machine translation
Stephane Clinchant, Kweon Woo Jung, and Vassilina Nikoulina. 2019 · 2019
Closest in time.
Cross-lingual language model pretraining
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Investigating bidirectional recurrent neural network language models for speech recognition
Xie Chen, Anton Ragni, Xunying Liu, and Mark JF Gales. 2017 · 2017
Cited alongside, same era.
Grammaticality, acceptability, and probability: A probabilistic view of linguistic knowledge
Jey Han Lau, Alexander Clark, and Shalom Lappin. 2017 · 2017
Cited alongside, same era.
Unsupervised pretraining for sequence to sequence learning
Prajit Ramachandran, Peter J Liu, and Quoc V Le. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Large margin neural language model
Jiaji Huang, Yi Li, Wei Ping, and Liang Huang. 2018 · 2018
Cited alongside, same era.
Correcting length bias in neural machine translation
Kenton Murray and David Chiang. 2018 · 2018
Cited alongside, same era.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Cited alongside, same era.
Alexis Conneau and Guillaume Lample. 2019 · 2019
Closest in time.
Transformer-XL: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
Jasper: An end-to-end convolutional neural acoustic model
Jason Li, Vitaly Lavrukhin, Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen, Huyen Nguyen, and Ravi Teja Gadde. 2019 · 2019
Closest in time.
Transformers without tears: Improving the normalization of self-attention
Toan Q Nguyen and Julian Salazar. 2019 · 2019
Closest in time.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Closest in time.
Effective sentence scoring method using BERT for speech recognition
Joongbo Shin, Yoonhyung Lee, and Kyomin Jung. 2019 · 2019
Closest in time.
BERT has a mouth, and it must speak: BERT as a Markov random field language model
Alex Wang and Kyunghyun Cho. 2019 · 2019
Closest in time.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman. 2019 · 2019
Closest in time.
XLNet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Closest in time.
GluonCV and GluonNLP: Deep learning in computer vision and natural language processing
Jian Guo, He He, Tong He, Leonard Lausen, Mu Li, Haibin Lin, Xingjian Shi, Chenguang Wang, Junyuan Xie, Sheng Zha, et al. 2020 · 2020
Closest in time.
BLiMP: The benchmark of linguistic minimal pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R Bowman. 2020 · 2020
Closest in time.
BERTScore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.
Incorporating BERT into neural machine translation
Jinhua Zhu, Yingce Xia, Lijun Wu, Di He, Tao Qin, Wengang Zhou, Houqiang Li, and Tie-Yan Liu. 2020 · 2020
Closest in time.