Roberta: A robustly optimized BERT pretraining approach
Original
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Transformation Invariance in Pattern Recognition — Tangent Distance and Tangent Propagation , pages 239–274. Springer Berlin Heidelberg, Berlin, Heidelberg
Patrice Y. Simard, Yann A. LeCun, John S. Denker, and Bernard Victorri. 1998 · 1998
Earlier work this paper cites.
Vicinal risk minimization
Olivier Chapelle, Jason Weston, Léon Bottou, and Vladimir Vapnik. 2001 · 2001
Earlier work this paper cites.
Introduction to the conll-2002 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang. 2002 · 2002
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Original
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
Introduction to the conll-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
Tri-training: exploiting unlabeled data using three classifiers
Zhi-Hua Zhou and Ming Li. 2005 · 2005
Earlier work this paper cites.
Cross-lingual word clusters for direct transfer of linguistic structure
Oscar Täckström, Ryan McDonald, and Jakob Uszkoreit. 2012 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Learned in translation: Contextualized word vectors
Bryan McCann, James Bradbury, Caiming Xiong, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
Regularizing neural networks by penalizing confident output distributions
Original
Gabriel Pereyra, George Tucker, Jan Chorowski, Lukasz Kaiser, and Geoffrey E. Hinton. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.