Fetching the paper…
Reading the bibliography…
The capabilities of Large Language Models (LLMs) in low-resource languages lag far behind those in English, making their universal accessibility a significant challenge.
A new algorithm for data compression
Gage, P · 1994
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2014
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Sennrich, R., Haddow, B., and Birch, A · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y · 2019
Earlier work this paper cites.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Clark, C., Lee, K., Chang, M.-W., Kwiatkowski, T., Collins, M., and Toutanova, K · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 2019
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Artetxe, M., Ruder, S., and Yogatama, D · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
Language ID in the wild: Unexpected challenges on the path to a thousand-language web text corpus
Caswell, I., Breiner, T., van Esch, D., and Bapna, A · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., Grave, E., Ott, M., Zettlemoyer, L., and Stoyanov, V · 2020
Earlier work this paper cites.
MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer
Pfeiffer, J., Vulić, I., Gurevych, I., and Ruder, S · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Earlier work this paper cites.
As good as new. how to successfully recycle English GPT-2 to make models for other languages
de Vries, W. and Nissim, M · 2021
Earlier work this paper cites.
Towards continual learning for multilingual machine translation via vocabulary substitution
Garcia, X., Constant, N., Parikh, A., and Firat, O · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2021
Earlier work this paper cites.
UNKs everywhere: Adapting multilingual language models to new scripts
Pfeiffer, J., Vulić, I., Gurevych, I., and Ruder, S · 2021
Earlier work this paper cites.
Winogrande: an adversarial winograd schema challenge at scale
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 2021
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Xue, L., Constant, N., Roberts, A., Kale, M., Al-Rfou, R., Siddhant, A., Barua, A., and Raffel, C · 2021
Earlier work this paper cites.
Composable sparse fine-tuning for cross-lingual transfer
Ansell, A., Ponti, E., Korhonen, A., and Vulić, I · 2022
Earlier work this paper cites.
Building machine translation systems for the next thousand languages, 2022
Bapna, A., Caswell, I., Kreutzer, J., Firat, O., van Esch, D., Siddhant, A., Niu, M., Baljekar, P., Garcia, X., Macherey, W., Breiner, T., Axelrod, V., Riesa, J., Cao, Y., Chen, M. X., Macherey, K., Krikun, M., Wang, P., Gutkin, A., Shah, A., Huang, Y., Chen, Z., Wu, Y., and Hughes, M · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation, 2022
Costa-jussà, M. R. et al · 2022
Cited alongside, same era.
Fast vocabulary transfer for language model compression
Gee, L., Zugarini, A., Rigutini, L., and Torroni, P · 2022
Cited alongside, same era.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Goyal, N., Gao, C., Chaudhary, V., Chen, P.-J., Wenzek, G., Ju, D., Krishnan, S., Ranzato, M., Guzmán, F., and Fan, A · 2022
Cited alongside, same era.
An embarrassingly simple method to mitigate undesirable properties of pretrained language model tokenizers
Hofmann, V., Schuetze, H., and Pierrehumbert, J · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., yelong shen, Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
SIB-200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects
Adelani, D., Liu, H., Shen, X., Vassilyev, N., Alabi, J., Mao, Y., Gao, H., and Lee, E.-S · 2024
Later among the works it cites.
Gemini: A family of highly capable multimodal models, 2024
Anil, R. et al · 2024
Later among the works it cites.
Aya 23: Open weight releases to further multilingual progress, 2024
Aryabumi, V., Dang, J., Talupuru, D., Dash, S., Cairuz, D., Lin, H., Venkitesh, B., Smith, M., Campos, J. A., Tan, Y. C., Marchisio, K., Bartolo, M., Ruder, S., Locatelli, A., Kreutzer, J., Frosst, N., Gomez, A., Blunsom, P., Fadaee, M., Üstün, A., and Hooker, S · 2024
Later among the works it cites.
The belebele benchmark: a parallel reading comprehension dataset in 122 language variants
Bandarkar, L., Liang, D., Muller, B., Artetxe, M., Shukla, S. N., Husa, D., Goyal, N., Krishnan, A., Zettlemoyer, L., and Khabsa, M · 2024
Later among the works it cites.
LLM augmented LLMs: Expanding capabilities through composition
Bansal, R., Samanta, B., Dalmia, S., Gupta, N., Ganapathy, S., Bapna, A., Jain, P., and Talukdar, P · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning
Liu, H., Tam, D., Mohammed, M., Mohta, J., Huang, T., Bansal, M., and Raffel, C · 2022
Cited alongside, same era.
Lifting the curse of multilinguality by pre-training modular transformers
Pfeiffer, J., Goyal, N., Lin, X., Li, X., Cross, J., Riedel, S., and Artetxe, M · 2022
Cited alongside, same era.
Finetuned language models are zero-shot learners
Wei, J., Bosma, M., Zhao, V., Guu, K., Yu, A. W., Lester, B., Du, N., Dai, A. M., and Le, Q. V · 2022
Cited alongside, same era.
How robust is neural machine translation to language imbalance in multilingual tokenizer training?
Zhang, S., Chaudhary, V., Goyal, N., Cross, J., Wenzek, G., Bansal, M., and Guzman, F · 2022
Cited alongside, same era.
Do all languages cost the same? tokenization in the era of commercial language models
Ahia, O., Kumar, S., Gonen, H., Kasai, J., Mortensen, D., Smith, N., and Tsvetkov, Y · 2023
Cited alongside, same era.
MEGA: Multilingual evaluation of generative AI
Ahuja, K., Diddee, H., Hada, R., Ochieng, M., Ramesh, K., Jain, P., Nambi, A., Ganu, T., Segal, S., Ahmed, M., Bali, K., and Sitaram, S · 2023
Cited alongside, same era.
Palm 2 technical report, 2023
Anil, R. et al · 2023
Cited alongside, same era.
Later among the works it cites.
LoRA learns less and forgets less
Biderman, D., Portes, J., Ortiz, J. J. G., Paul, M., Greengard, P., Jennings, C., King, D., Havens, S., Chiley, V., Frankle, J., Blakeney, C., and Cunningham, J. P · 2024
Later among the works it cites.
Monolingual or multilingual instruction tuning: Which makes a better alpaca
Chen, P., Ji, S., Bogoychev, N., Kutuzov, A., Haddow, B., and Heafield, K · 2024
Later among the works it cites.
Efficient and effective text encoding for chinese llama and alpaca, 2024
Cui, Y., Yang, Z., and Yao, X · 2024
Later among the works it cites.
The llama 3 herd of models, 2024
Dubey, A. et al · 2024
Later among the works it cites.
Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities
Fujii, K., Nakamura, T., Loem, M., Iida, H., Ohi, M., Hattori, K., Shota, H., Mizuki, S., Yokota, R., and Okazaki, N · 2024
Later among the works it cites.
Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., and Zhang, Y · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size, 2024
Riviere, M. et al · 2024
Later among the works it cites.
Continual learning of large language models: A comprehensive survey, 2024
Shi, H., Xu, Z., Wang, H., Qin, W., Wang, W., Wang, Y., Wang, Z., Ebrahimi, S., and Wang, H · 2024
Later among the works it cites.
IndicGenBench: A multilingual benchmark to evaluate generation capabilities of LLMs on Indic languages
Singh, H., Gupta, N., Bharadwaj, S., Tewari, D., and Talukdar, P · 2024
Later among the works it cites.
Language-specific neurons: The key to multilingual capabilities in large language models
Tang, T., Luo, W., Huang, H., Zhang, D., Wang, X., Zhao, X., Wei, F., and Wen, J.-R · 2024
Later among the works it cites.
Scaling laws with vocabulary: Larger models deserve larger vocabularies
Tao, C., Liu, Q., Dou, L., Muennighoff, N., Wan, Z., Luo, P., Lin, M., and Wong, N · 2024
Later among the works it cites.
Aya model: An instruction finetuned open-access multilingual language model
Üstün, A., Aryabumi, V., Yong, Z., Ko, W.-Y., D’souza, D., Onilude, G., Bhandari, N., Singh, S., Ooi, H.-L., Kayid, A., Vargus, F., Blunsom, P., Longpre, S., Muennighoff, N., Fadaee, M., Kreutzer, J., and Hooker, S · 2024
Later among the works it cites.
Do llamas work in English? on the latent language of multilingual transformers
Wendler, C., Veselovsky, V., Monea, G., and West, R · 2024
Later among the works it cites.
Yang, A. et al · 2024
Later among the works it cites.
MAmmoTH2: Scaling instructions from the web
Yue, X., Zheng, T., Zhang, G., and Chen, W · 2024
Later among the works it cites.
Breaking language barriers: Cross-lingual continual pre-training at scale
Zheng, W., Pan, W., Xu, X., Qin, L., Yue, L., and Zhou, M · 2024
Later among the works it cites.