Fetching the paper…
Reading the bibliography…
Multilingual transformer-based models demonstrate remarkable zero and few-shot transfer across languages by learning and reusing language-agnostic features.
A value for n-person games , page 31–40. Cambridge University Press
Lloyd S. Shapley. 1953 · 1953
Earlier work this paper cites.
Contributions of transformer attention heads in multi- and cross-lingual tasks
Weicheng Ma, Kai Zhang, Renze Lou, Lili Wang, and Soroush Vosoughi. 2021 · 1966
Earlier work this paper cites.
Multipath translation lexicon induction via bridge languages
Gideon S. Mann and David Yarowsky. 2001 · 2001
Earlier work this paper cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2003
Earlier work this paper cites.
Empirical bernstein stopping
Volodymyr Mnih, Csaba Szepesvári, and Jean-Yves Audibert. 2008 · 2008
Earlier work this paper cites.
Reusable tagset conversion using tagset drivers
Daniel Zeman. 2008 · 2008
Earlier work this paper cites.
Polynomial calculation of the shapley value based on sampling
Javier Castro, Daniel Gómez, and Juan Tejada. 2009 · 2009
Earlier work this paper cites.
Empirical bernstein bounds and sample-variance penalization
Andreas Maurer and Massimiliano Pontil. 2009 · 2009
Earlier work this paper cites.
Multi-source transfer of delexicalized dependency parsers
Ryan McDonald, Slav Petrov, and Keith Hall. 2011 · 2011
Earlier work this paper cites.
Treebank translation for cross-lingual parser induction
Jörg Tiedemann, Željko Agić, and Joakim Nivre. 2014 · 2014
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Cheap translation for cross-lingual named entity recognition
Stephen Mayhew, Chen-Tse Tsai, and Dan Roth. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Loss in translation: Learning bilingual word mapping with a retrieval criterion
Armand Joulin, Piotr Bojanowski, Tomas Mikolov, Hervé Jégou, and Edouard Grave. 2018 · 2018
Earlier work this paper cites.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Cited alongside, same era.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 2019
Cited alongside, same era.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Cited alongside, same era.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Zou. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Neuron shapley: Discovering the responsible neurons
Amirata Ghorbani and James Y Zou. 2020 · 2020
Later among the works it cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models
Benjamin Hoover, Hendrik Strobelt, and Sebastian Gehrmann. 2020 · 2020
Later among the works it cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Universal Dependencies v2: An evergrowing multilingual treebank collection
Joakim Nivre, Marie-Catherine de Marneffe, Filip Ginter, Jan Hajič, Christopher D. Manning, Sampo Pyysalo, Sebastian Schuster, Francis Tyers, and Daniel Zeman. 2020 · 2020
Later among the works it cites.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Cited alongside, same era.
SNIP: Single-shot network pruning based on connection sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip Torr. 2019 · 2019
Cited alongside, same era.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig. 2019 · 2019
Cited alongside, same era.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Cited alongside, same era.
A multiscale visualization of attention in the transformer model
Jesse Vig. 2019 · 2019
Cited alongside, same era.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
When BERT Plays the Lottery, All Tickets Are Winning
Sai Prasanna, Anna Rogers, and Anna Rumshisky. 2020 · 2020
Later among the works it cites.
On negative interference in multilingual models: Findings and a meta-learning treatment
Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020 · 2020
Later among the works it cites.
MAD-G: Multilingual adapter generation for efficient cross-lingual transfer
Alan Ansell, Edoardo Maria Ponti, Jonas Pfeiffer, Sebastian Ruder, Goran Glavaš, Ivan Vulić, and Anna Korhonen. 2021 · 2021
Later among the works it cites.
Explicit alignment objectives for multilingual bidirectional encoders
Junjie Hu, Melvin Johnson, Orhan Firat, Aditya Siddhant, and Graham Neubig. 2021 · 2021
Later among the works it cites.
Datasets: A community library for natural language processing
Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, and Thomas Wolf. 2021 · 2021
Later among the works it cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Discrete and soft prompting for multilingual models
Mengjie Zhao and Hinrich Schütze. 2021 · 2021
Later among the works it cites.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić. 2022 · 2022
Closest in time.
Make the best of cross-lingual transfer: Evidence from POS tagging with over 100 languages
Wietse de Vries, Martijn Wieling, and Malvina Nissim. 2022 · 2022
Closest in time.
Structured pruning learns compact and accurate models
Mengzhou Xia, Zexuan Zhong, and Danqi Chen. 2022 · 2022
Closest in time.