Fetching the paper…
Reading the bibliography…
While large language models (LLMs) have been pre-trained on multilingual corpora, their performance still lags behind in most languages compared to a few resource-rich languages.
Visualizing data using t-sne
Laurens Van der Maaten and Geoffrey Hinton. 2008 · 2008
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean. 2015 · 2015
Earlier work this paper cites.
Sequence-level knowledge distillation
Yoon Kim and Alexander M. Rush. 2016 · 2016
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
Explaining sequence-level knowledge distillation as data-augmentation for neural machine translation
Mitchell A. Gordon and Kevin Duh. 2019 · 2019
Earlier work this paper cites.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 2019
Earlier work this paper cites.
Snapshot distillation: Teacher-student optimization in one generation
Chenglin Yang, Lingxi Xie, Chi Su, and Alan L. Yuille. 2019 · 2019
Earlier work this paper cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma. 2019 · 2019
Earlier work this paper cites.
Distilling knowledge learned in bert for text generation
Yen-Chun Chen, Zhe Gan, Yu Cheng, Jingzhou Liu, and Jingjing Liu. 2020 · 2020
Earlier work this paper cites.
Mkqa: A linguistically diverse benchmark for multilingual open domain question answering
Shayne Longpre, Yi Lu, and Joachim Daiber. 2020 · 2020
Earlier work this paper cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Knowledge distillation for multilingual unsupervised neural machine translation
Haipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama, Eiichiro Sumita, and Tiejun Zhao. 2020 · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Earlier work this paper cites.
Improving massively multilingual neural machine translation and zero-shot translation
Biao Zhang, Philip Williams, Ivan Titov, and Rico Sennrich. 2020 · 2020
Earlier work this paper cites.
Question translation training for better multilingual reasoning
Wenhao Zhu, Shujian Huang, Fei Yuan, Shuaijie She, Jiajun Chen, and Alexandra Birch. 2024 · 2020
Earlier work this paper cites.
Improving pretrained cross-lingual language models via self-labeled word alignment
Zewen Chi, Li Dong, Bo Zheng, Shaohan Huang, Xian-Ling Mao, Heyan Huang, and Furu Wei. 2021 · 2021
Earlier work this paper cites.
Knowledge distillation: A survey
Jianping Gou, Baosheng Yu, Stephen J Maybank, and Dacheng Tao. 2021 · 2021
Earlier work this paper cites.
XL-sum: Large-scale multilingual abstractive summarization for 44 languages
Tahmid Hasan, Abhik Bhattacharjee, Md. Saiful Islam, Kazi Mubasshir, Yuan-Fang Li, Yong-Bin Kang, M. Sohel Rahman, and Rifat Shahriyar. 2021 · 2021
Earlier work this paper cites.
Pre-training multilingual neural machine translation by leveraging alignment information
Zehui Lin, Xiao Pan, Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou, and Lei Li. 2021 · 2021
Earlier work this paper cites.
mt5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Earlier work this paper cites.
WeChat neural machine translation systems for WMT21
Xianfeng Zeng, Yijin Liu, Ernan Li, Qiu Ran, Fandong Meng, Peng Li, Jinan Xu, and Jie Zhou. 2021 · 2021
Earlier work this paper cites.
Directquote: A dataset for direct quotation extraction and attribution in news articles
Yuanchi Zhang and Yang Liu. 2021 · 2021
Cited alongside, same era.
No language left behind: Scaling human-centered machine translation
Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022 · 2022
Cited alongside, same era.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Cited alongside, same era.
Zero-shot cross-lingual transfer of prompt-based tuning with a unified multilingual prompt
Lianzhe Huang, Shuming Ma, Dongdong Zhang, Furu Wei, and Houfeng Wang. 2022 · 2022
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Later among the works it cites.
Turning english-centric llms into polyglots: How much multilinguality is needed?
Tannon Kew, Florian Schottmann, and Rico Sennrich. 2023 · 2023
Later among the works it cites.
Seallms – large language models for southeast asia
Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li, Mahani Aljunied, Qingyu Tan, Liying Cheng, Guanzheng Chen, Yue Deng, Sen Yang, Chaoqun Liu, Hang Zhang, and Lidong Bing. 2023 · 2023
Later among the works it cites.
Lextreme: A multi-lingual and multi-task benchmark for the legal domain
Joel Niklaus, Veton Matoshi, Pooja Rani, Andrea Galassi, Matthias Stürmer, and Ilias Chalkidis. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2022 · 2022
Cited alongside, same era.
Few-shot learning with multilingual language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2022 · 2022
Cited alongside, same era.
Cpt: Cross-modal prefix-tuning for speech-to-text translation
Yukun Ma, Trung Hieu Nguyen, and Bin Ma. 2022 · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Cited alongside, same era.
Minh Pham, Minsu Cho, Ameya Joshi, and Chinmay Hegde. 2022 · 2022
Cited alongside, same era.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022a · 2022
Cited alongside, same era.
CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. 2022b · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
OpenAI. 2023 · 2023
Later among the works it cites.
Several categories of large language models (llms): A short survey
Saurabh Pahune and Manoj Chandrasekharan. 2023 · 2023
Later among the works it cites.
Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages
Libo Qin, Qiguang Chen, Fuxuan Wei, Shijue Huang, and Wanxiang Che. 2023 · 2023
Later among the works it cites.
Does the english matter? elicit cross-lingual abilities of large language models
Leonardo Ranaldi and Giulia Pucci. 2023 · 2023
Later among the works it cites.
Empowering multi-step reasoning across languages via tree-of-thoughts
Leonardo Ranaldi and Fabio Massimo Zanzotto. 2023 · 2023
Later among the works it cites.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023 · 2023
Later among the works it cites.
Hyperpolyglot llms: Cross-lingual interpretability in token embeddings
Andrea Wen-Yi and David Mimno. 2023 · 2023
Later among the works it cites.
Wen Yang, Chong Li, Jiajun Zhang, and Chengqing Zong. 2023 · 2023
Later among the works it cites.
Continual knowledge distillation for neural machine translation
Yuanchi Zhang, Peng Li, Maosong Sun, and Yang Liu. 2023 · 2023
Later among the works it cites.
xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning
Linzheng Chai, Jian Yang, Tao Sun, Hongcheng Guo, Jiaheng Liu, Bing Wang, Xiannian Liang, Jiaqi Bai, Tongliang Li, Qiyao Peng, and Zhoujun Li. 2024 · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Closest in time.
cld3: Google’s Compact Language Detector 3
Jeroen Ooms. 2024 · 2024
Closest in time.
Advfusion: Multilingual adapter-based knowledge transfer for code summarization
Iman Saberi, Fatemeh Fard, and Fuxiang Chen. 2024 · 2024
Closest in time.
Multilingual instruction tuning with just a pinch of multilinguality
Uri Shaham, Jonathan Herzig, Roee Aharoni, Idan Szpektor, Reut Tsarfaty, and Matan Eyal. 2024 · 2024
Closest in time.
Langbridge: Multilingual reasoning without multilingual supervision
Dongkeun Yoon, Joel Jang, Sungdong Kim, Seungone Kim, Sheikh Shafayat, and Minjoon Seo. 2024 · 2024
Closest in time.
Llama beyond english: An empirical study on language capability transfer
Jun Zhao, Zhihao Zhang, Luhui Gao, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024 · 2024
Closest in time.