Fetching the paper…
Reading the bibliography…
We re-evaluate the standard practice of sharing weights between input and output embeddings in state-of-the-art pre-trained language models.
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Large-scale learning of word relatedness with constraints
Guy Halawi, Gideon Dror, Evgeniy Gabrilovich, and Yehuda Koren · 2012
Earlier work this paper cites.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson · 2012
Earlier work this paper cites.
Better word representations with recursive neural networks for morphology
Minh-Thang Luong, Richard Socher, and Christopher D Manning · 2013
Earlier work this paper cites.
An unsupervised model for instance level subcategorization acquisition
Simon Baker, Roi Reichart, and Anna Korhonen · 2014
Earlier work this paper cites.
Multimodal distributional semantics
Elia Bruni, Nam-Khanh Tran, and Marco Baroni · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson · 2014
Earlier work this paper cites.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Felix Hill, Roi Reichart, and Anna Korhonen · 2015
Earlier work this paper cites.
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Results of the wmt16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević · 2016
Earlier work this paper cites.
Zero-Resource Translation with Multi-Lingual Neural Machine Translation
Orhan Firat, Baskaran Sankaran, Yaser Al-onaizan, Fatos T. Yarman Vural, and Kyunghyun Cho · 2016
Earlier work this paper cites.
SQuAD: 100,000+ Questions for Machine Comprehension of Text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Semantics derived automatically from language corpora contain human-like biases
Aylin Caliskan, Joanna J Bryson, and Arvind Narayanan · 2017
Earlier work this paper cites.
Tying Word Vectors and Word Classifiers: A Loss Framework for Language Modeling
Hakan Inan, Khashayar Khosravi, and Richard Socher · 2017
Earlier work this paper cites.
Google’s Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation
Melvin Johnson, Mike Schuster, Quoc V Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean · 2017
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov · 2018
Earlier work this paper cites.
Universal Language Model Fine-tuning for Text Classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Universal dependencies 2.2
Joakim Nivre, Mitchell Abrams, Željko Agić, Lars Ahrenberg, Lene Antonsen, Maria Jesus Aranzabe, Gashaw Arutie, Masayuki Asahara, Luma Ateyah, Mohammed Attia, et al · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Noam Shazeer, Youlong Cheng, Niki Parmar, Dustin Tran, Ashish Vaswani, Penporn Koanantakool, Peter Hawkins, HyoukJoong Lee, Mingsheng Hong, Cliff Young, Ryan Sepassi, and Blake Hechtman · 2018
Earlier work this paper cites.
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Adina Williams, Nikita Nangia, and Samuel R. Bowman · 2018
Earlier work this paper cites.
Overview of the third bucc shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp · 2018
Cited alongside, same era.
Massively Multilingual Neural Machine Translation
Roee Aharoni, Melvin Johnson, and Orhan Firat · 2019
Cited alongside, same era.
Towards better substitution-based word sense induction
Asaf Amrami and Yoav Goldberg · 2019
Cited alongside, same era.
Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond
Mikel Artetxe and Holger Schwenk · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Closest in time.
Emerging cross-lingual structure in pretrained language models
Alexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer, and Veselin Stoyanov · 2020
Closest in time.
FILTER: An Enhanced Fusion Method for Cross-lingual Language Understanding
Yuwei Fang, Shuohang Wang, Zhe Gan, Siqi Sun, and Jingjing Liu · 2020
Closest in time.
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith · 2020
Closest in time.
XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yanai Elazar and Yoav Goldberg · 2019
Cited alongside, same era.
Linguistic Knowledge and Transferability of Contextual Representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Cited alongside, same era.
Are Sixteen Heads Really Better than One?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette · 2019
Cited alongside, same era.
Transfer learning in natural language processing
Sebastian Ruder, Matthew E Peters, Swabha Swayamdipta, and Thomas Wolf · 2019
Cited alongside, same era.
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Roy Schwartz, Jesse Dodge, Noah A. Smith, and Oren Etzioni · 2019
Cited alongside, same era.
Cross-lingual ability of multilingual bert: An empirical study
Karthikeyan K, Zihan Wang, Stephen Mayhew, and Dan Roth · 2020
Closest in time.
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Closest in time.
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2020
Closest in time.
GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Closest in time.
MLQA: Evaluating Cross-lingual Extractive Question Answering
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk · 2020
Closest in time.
Mogrifier LSTM
Gábor Melis, Tomáš Kočiský, and Phil Blunsom · 2020
Closest in time.
XtremeDistil : Multi-stage Distillation for Massive Multilingual Models
Subhabrata Mukherjee and Ahmed Hassan Awadallah · 2020
Closest in time.
MAD-X: An Adapter-based Framework for Multi-task Cross-lingual Transfer
Jonas Pfeiffer, Ivan Vuli, Iryna Gurevych, and Sebastian Ruder · 2020
Closest in time.
English intermediate-task training improves zero-shot cross-lingual transfer too
Jason Phang, Phu Mon Htut, Yada Pruksachatkun, Haokun Liu, Clara Vania, Katharina Kann, Iacer Calixto, and Samuel R Bowman · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Closest in time.
MobileBERT : a Compact Task-Agnostic BERT for Resource-Limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Closest in time.
Investigating Transferability in Pretrained Language Models
Alex Tamkin, Trisha Singh, Davide Giovanardi, and Noah Goodman · 2020
Closest in time.
Revisiting Few-sample BERT Fine-tuning
Tianyi Zhang, Felix Wu, Arzoo Katiyar, Kilian Q. Weinberger, and Yoav Artzi · 2020
Closest in time.
HULK: An Energy Efficiency Benchmark Platform for Responsible Natural Language Processing
Xiyou Zhou, Zhiyu Chen, Xiaoyong Jin, and William Yang Wang · 2020
Closest in time.
{VECO}: Variable encoder-decoder pre-training for cross-lingual understanding and generation
Anonymous · 2021
Closest in time.
Using the output embedding to improve language models
Ofir Press and Lior Wolf · 2025
Closest in time.