Fetching the paper…
Reading the bibliography…
Modular deep learning is the state-of-the-art solution for lifting the curse of multilinguality, preventing the impact of negative interference and enabling cross-lingual performance in Multilingual Pre-trained Language Models.
What the [mask]? making sense of language-specific bert models
Debora Nozza, Federico Bianchi, and Dirk Hovy. 2020 · 2003
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Hypernetworks
David Ha, Andrew M. Dai, and Quoc V. Le. 2017 · 2017
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
Learning language representations for typology prediction
Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017 · 2017
Earlier work this paper cites.
Learning multiple visual domains with residual adapters
Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. 2017 · 2017
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
Simple, scalable adaptation for neural machine translation
Ankur Bapna and Orhan Firat. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for NLP
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019 · 2019
Earlier work this paper cites.
Massively multilingual transfer for NER
Afshin Rahimi, Yuan Li, and Trevor Cohn. 2019 · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Earlier work this paper cites.
Don’t stop pretraining: Adapt language models to domains and tasks
Suchin Gururangan, Ana Marasović, Swabha Swayamdipta, Kyle Lo, Iz Beltagy, Doug Downey, and Noah A. Smith. 2020 · 2020
Earlier work this paper cites.
CamemBERT: a tasty French language model
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, and Benoît Sagot. 2020 · 2020
Earlier work this paper cites.
AdapterHub: A framework for adapting transformers
Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulić, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020a · 2020
Earlier work this paper cites.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Jonas Pfeiffer, Ivan Vulić, Iryna Gurevych, and Sebastian Ruder. 2020b · 2020
Cited alongside, same era.
A study of residual adapters for multi-domain neural machine translation
Minh Quang Pham, Josep Maria Crego, François Yvon, and Jean Senellart. 2020 · 2020
Cited alongside, same era.
Monolingual adapters for zero-shot neural machine translation
Jerin Philip, Alexandre Berard, Matthias Gallé, and Laurent Besacier. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
UDapter: Language adaptation for truly Universal Dependency parsing
Ahmet Üstün, Arianna Bisazza, Gosse Bouma, and Gertjan van Noord. 2020 · 2020
Cited alongside, same era.
On negative interference in multilingual models: Findings and a meta-learning treatment
Unifying cross-lingual transfer across scenarios of resource scarcity
Alan Ansell, Marinela Parović, Ivan Vulić, Anna Korhonen, and Edoardo Ponti. 2023 · 2023
Later among the works it cites.
Language-family adapters for low-resource multilingual neural machine translation
Alexandra Chronopoulou, Dario Stojanovski, and Alexander Fraser. 2023b · 2023
Later among the works it cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Túlio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023 · 2023
Later among the works it cites.
Gated adapters for multi-domain neural machine translation
Mateusz Klimaszewski, Zeno Belligoli, Satendra Kumar, and Emmanouil Stergiadis. 2023 · 2023
Later among the works it cites.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward Raff, and Colin Raffel. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zirui Wang, Zachary C. Lipton, and Yulia Tsvetkov. 2020 · 2020
Cited alongside, same era.
The Reality of Multi-Lingual Machine Translation , volume 21 of Studies in Computational and Theoretical Linguistics
Tom Kocmi, Dominik Macháček, and Ondřej Bojar. 2021 · 2021
Cited alongside, same era.
AdapterFusion: Non-destructive task composition for transfer learning
Jonas Pfeiffer, Aishwarya Kamath, Andreas Rücklé, Kyunghyun Cho, and Iryna Gurevych. 2021 · 2021
Cited alongside, same era.
IndicXNLI: Evaluating multilingual inference for Indian languages
Divyanshu Aggarwal, Vivek Gupta, and Anoop Kunchukuttan. 2022 · 2022
Cited alongside, same era.
Composable sparse fine-tuning for cross-lingual transfer
Alan Ansell, Edoardo Ponti, Anna Korhonen, and Ivan Vulić. 2022 · 2022
Cited alongside, same era.
Multilingual machine translation with hyper-adapters
Christos Baziotis, Mikel Artetxe, James Cross, and Shruti Bhosale. 2022 · 2022
Cited alongside, same era.
Survey of low-resource machine translation
Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindřich Helcl, and Alexandra Birch. 2022 · 2022
Cited alongside, same era.
Modular deep learning
Jonas Pfeiffer, Sebastian Ruder, Ivan Vulić, and Edoardo Ponti. 2023 · 2023
Later among the works it cites.
Zipit! merging models from different tasks without training
George Stoica, Daniel Bolya, Jakob Bjorner, Taylor Hearn, and Judy Hoffman. 2023 · 2023
Later among the works it cites.
AdapterDistillation: Non-destructive task composition with knowledge distillation
Junjie Wang, Yicheng Chen, Wangshu Zhang, Sen Hu, Teng Xu, and Jing Zheng. 2023 · 2023
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model
BigScience Workshop. 2023 · 2023
Later among the works it cites.
TIES-merging: Resolving interference when merging models
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal. 2023 · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operation
Jinghan Zhang, Shiqi Chen, Junteng Liu, and Junxian He. 2023 · 2023
Later among the works it cites.
Tower: An open multilingual large language model for translation-related tasks
Duarte M. Alves, José Pombal, Nuno M. Guerreiro, Pedro H. Martins, João Alves, Amin Farajian, Ben Peters, Ricardo Rei, Patrick Fernandes, Sweta Agrawal, Pierre Colombo, José G. C. de Souza, and André F. T. Martins. 2024 · 2024
Closest in time.
Lorahub: Efficient cross-task generalization via dynamic lora composition
Chengsong Huang, Qian Liu, Bill Yuchen Lin, Tianyu Pang, Chao Du, and Min Lin. 2024 · 2024
Closest in time.
Investigating the potential of task arithmetic for cross-lingual transfer
Marinela Parović, Ivan Vulić, and Anna Korhonen. 2024 · 2024
Closest in time.
AdapterSoup: Weight averaging to improve generalization of pretrained language models
Alexandra Chronopoulou, Matthew Peters, Alexander Fraser, and Jesse Dodge. 2023a · 2063
Closest in time.