Fetching the paper…
Reading the bibliography…
Instruction tuning a large language model with multiple languages can prepare it for multilingual downstream tasks.
The Winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
URIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R. Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Earlier work this paper cites.
Learning language representations for typology prediction
Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017 · 2017
Earlier work this paper cites.
Huggingface’s transformers: State-of-the-art natural language processing
Thomas Wolf et al. 2019 · 2019
Earlier work this paper cites.
Stochastic approach to worldwide language classification: the signals and the noise towards long-range exploration
Vincent Beaufils and Johannes Tomin. 2020 · 2020
Earlier work this paper cites.
XCOPA: A multilingual dataset for causal commonsense reasoning
Edoardo Maria Ponti, Goran Glavaš, Olga Majewska, Qianchu Liu, Ivan Vulić, and Anna Korhonen. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Few-shot learning with multilingual language models
Xi Victoria Lin et al. 2021 · 2021
Earlier work this paper cites.
It’s All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning
Alexey Tikhonov and Max Ryabinin. 2021 · 2021
Cited alongside, same era.
BLOOM: A 176B-parameter open-access multilingual language model
Teven Le Scao et al. 2022 · 2022
Cited alongside, same era.
mGPT: Few-shot learners go multilingual
Oleh Shliazhko, Alena Fenogenova, Maria Tikhonova, Vladislav Mikhailov, Anastasia Kozlova, and Tatiana Shavrina. 2022 · 2022
Cited alongside, same era.
The Falcon series of open language models
Ebtesam Almazrouei et al. 2023 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, et al. 2023 · 2023
Cited alongside, same era.
Bactrian-X: A multilingual replicable instruction-following model with low-rank adaptation
LLaMA: Open and efficient foundation language models
Hugo Touvron et al. 2023 · 2023
Later among the works it cites.
Monolingual or multilingual instruction tuning: Which makes a better alpaca
Pinzhen Chen, Shaoxiong Ji, Nikolay Bogoychev, Andrey Kutuzov, Barry Haddow, and Kenneth Heafield. 2024 · 2024
Closest in time.
Zero-shot cross-lingual transfer in instruction tuning of large language model
Nadezhda Chirkova and Vassilina Nikoulina. 2024 · 2024
Closest in time.
The Llama 3 herd of models
Aaron Grattafiori et al. 2024 · 2024
Closest in time.
Turning English-centric LLMs into polyglots: How much multilinguality is needed?
Tannon Kew, Florian Schottmann, and Rico Sennrich. 2024 · 2024
Closest in time.
Soft prompt tuning for cross-lingual transfer: When less is more
Fred Philippy, Siwen Guo, Shohreh Haddadan, Cedric Lothritz, Jacques Klein, and Tegawendé F. Bissyandé. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Haonan Li, Fajri Koto, Minghao Wu, Alham Fikri Aji, and Timothy Baldwin. 2023 · 2023
Cited alongside, same era.
Crosslingual generalization through multitask finetuning
Niklas Muennighoff et al. 2023 · 2023
Cited alongside, same era.
Closest in time.
Multilingual instruction tuning with just a pinch of multilinguality
Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. 2024 · 2024
Closest in time.