Fetching the paper…
Reading the bibliography…
The development of Large Language Models (LLMs) relies on extensive text corpora, which are often unevenly distributed across languages.
Hellaswag: Can a machine really finish your sentence?
Zellers, R.; Holtzman, A.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 1905
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Conneau, A.; Khandelwal, K.; Goyal, N.; Chaudhary, V.; Wenzek, G.; Guzmán, F.; Grave, E.; Ott, M.; Zettlemoyer, L.; and Stoyanov, V. 2019 · 1911
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020 · 2009
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Clark, P.; Cowhey, I.; Etzioni, O.; Khot, T.; Sabharwal, A.; Schoenick, C.; and Tafjord, O. 2018 · 2018
Earlier work this paper cites.
The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT
Tiedemann, J. 2020 · 2020
Earlier work this paper cites.
Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation
Zhang, B.; Williams, P.; Titov, I.; and Sennrich, R. 2020 · 2020
Earlier work this paper cites.
Systematic Inequalities in Language Technology Performance across the World’s Languages
Blasi, D.; Anastasopoulos, A.; and Neubig, G. 2021 · 2021
Earlier work this paper cites.
Ammus: A survey of transformer-based pretrained models in natural language processing
Kalyan, K. S.; Rajasekharan, A.; and Sangeetha, S. 2021 · 2021
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. 2022 · 2022
Earlier work this paper cites.
Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023 · 2023
Earlier work this paper cites.
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023 · 2023
Cited alongside, same era.
Language models represent space and time
Gurnee, W.; and Tegmark, M. 2023 · 2023
Cited alongside, same era.
Huang, H.; Tang, T.; Zhang, D.; Zhao, W. X.; Song, T.; Xia, Y.; and Wei, F. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
How Large Language Models Encode Context Knowledge? A Layer-Wise Probing Study
Ju, T.; Sun, W.; Du, W.; Yuan, X.; Ren, Z.; and Liu, G. 2024 · 2024
Closest in time.
Transformers for Low-Resource Languages: Is F \ \backslash ’eidir Linn!
Lankford, S.; Afli, H.; and Way, A. 2024 · 2024
Closest in time.
Inference-time intervention: Eliciting truthful answers from a language model
Li, K.; Patel, O.; Viégas, F.; Pfister, H.; and Wattenberg, M. 2024 · 2024
Closest in time.
Is Translation All You Need? A Study on Solving Multilingual Tasks with Large Language Models
Liu, C.; Zhang, W.; Zhao, Y.; Luu, A. T.; and Bing, L. 2024 · 2024
Closest in time.
Multilingual large language model: A survey of resources, taxonomy and frontiers
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Marks, S.; and Tegmark, M. 2023 · 2023
Cited alongside, same era.
GPT-3 Dataset Statistics
OpenAI. 2023 · 2023
Cited alongside, same era.
Cross-lingual prompting: Improving zero-shot chain-of-thought reasoning across languages
Qin, L.; Chen, Q.; Wei, F.; Huang, S.; and Che, W. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Cited alongside, same era.
Don’t trust ChatGPT when your question is not in English: A study of multilingual abilities and types of LLMs
Zhang, X.; Li, S.; Hauer, B.; Shi, N.; and Kondrak, G. 2023 · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency
Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.-K.; et al. 2023 · 2023
Cited alongside, same era.
MEGAVERSE: Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks
Ahuja, S.; Aggarwal, D.; Gumma, V.; Watts, I.; Sathe, A.; Ochieng, M.; Hada, R.; Jain, P.; Axmed, M.; Bali, K.; and Sitaram, S. 2024 · 2024
Cited alongside, same era.
Qin, L.; Chen, Q.; Zhou, Y.; Chen, Z.; Li, Y.; Liao, L.; Li, M.; Che, W.; and Yu, P. S. 2024 · 2024
Closest in time.
Language Imbalance Can Boost Cross-lingual Generalisation
Schäfer, A.; Ravfogel, S.; Hofmann, T.; Pimentel, T.; and Schlag, I. 2024 · 2024
Closest in time.
The language barrier: Dissecting safety challenges of llms in multilingual contexts
Shen, L.; Tan, W.; Chen, S.; Chen, Y.; Zhang, J.; Xu, H.; Zheng, B.; Koehn, P.; and Khashabi, D. 2024 · 2024
Closest in time.
Is Cosine-Similarity of Embeddings Really About Similarity?
Steck, H.; Ekanadham, C.; and Kallus, N. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Team, G.; Mesnard, T.; Hardin, C.; Dadashi, R.; Bhupatiraju, S.; Pathak, S.; Sifre, L.; Rivière, M.; Kale, M. S.; Love, J.; et al. 2024 · 2024
Closest in time.
Do Llamas Work in English? On the Latent Language of Multilingual Transformers
Wendler, C.; Veselovsky, V.; Monea, G.; and West, R. 2024 · 2024
Closest in time.
Weakly supervised scene text generation for low-resource languages
Xie, Y.; Chen, X.; Zhan, H.; Shivakumara, P.; Yin, B.; Liu, C.; and Lu, Y. 2024 · 2024
Closest in time.