Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have multilingual capabilities and can solve tasks across various languages.
WordNet: A lexical database for English
G. A. Miller · 1994
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing
M. Honnibal and I. Montani · 2017
Earlier work this paper cites.
interpreting gpt: the logit lens, Aug 2020
nostalgebraist · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. Shieber · 2020
Earlier work this paper cites.
Extrapolating large language models to non-english by aligning languages, 2023
W. Zhu, Y. Lv, Q. Dong, F. Yuan, J. Xu, S. Huang, L. Kong, J. Chen, and L. Li · 2020
Earlier work this paper cites.
When being unseen from mBERT is just the beginning: Handling new languages with multilingual language models
B. Muller, A. Anastasopoulos, B. Sagot, and D. Seddah · 2021
Earlier work this paper cites.
How good is your tokenizer? on the monolingual performance of multilingual language models
P. Rust, J. Pfeiffer, I. Vulić, S. Ruder, and I. Gurevych · 2021
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
L. Xue, N. Constant, A. Roberts, M. Kale, R. Al-Rfou, A. Siddhant, A. Barua, and C. Raffel · 2021
Earlier work this paper cites.
Few-shot learning with multilingual language models, 2022
X. V. Lin, T. Mihaylov, M. Artetxe, T. Wang, S. Chen, D. Simig, M. Ott, N. Goyal, S. Bhosale, J. Du, R. Pasunuru, S. Shleifer, P. S. Koura, V. Chaudhary, B. O’Horo, J. Wang, L. Zettlemoyer, Z. Kozareva, M. Diab, V. Stoyanov, and X. Li · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
K. Meng, D. Bau, A. Andonian, and Y. Belinkov · 2022
Earlier work this paper cites.
Language models are multilingual chain-of-thought reasoners
F. Shi, M. Suzgun, M. Freitag, X. Wang, S. Srivats, S. Vosoughi, H. W. Chung, Y. Tay, S. Ruder, D. Zhou, et al · 2022
Earlier work this paper cites.
Extracting latent steering vectors from pretrained language models
N. Subramani, N. Suresh, and M. E. Peters · 2022
Earlier work this paper cites.
Mega: Multilingual evaluation of generative ai, 2023
K. Ahuja, H. Diddee, R. Hada, M. Ochieng, K. Ramesh, P. Jain, A. Nambi, T. Ganu, S. Segal, M. Axmed, K. Bali, and S. Sitaram · 2023
Earlier work this paper cites.
Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, Q. V. Do, Y. Xu, and P. Fung · 2023
Earlier work this paper cites.
What would be the most safety-relevant features in language models?
A. J. Chris Olah · 2023
Earlier work this paper cites.
Do multilingual language models think better in english?, 2023
J. Etxaniz, G. Azkune, A. Soroa, O. L. de Lacalle, and M. Artetxe · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
M. Geva, J. Bastings, K. Filippova, and A. Globerson · 2023
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun · 2023
Cited alongside, same era.
H. Huang, T. Tang, D. Zhang, W. X. Zhao, T. Song, Y. Xia, and F. Wei · 2023
Cited alongside, same era.
The hydra effect: Emergent self-repair in language model computations, 2023
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024
T. Lieberum, S. Rajamanoharan, A. Conmy, L. Smith, N. Sonnerat, V. Varma, J. Kramár, A. Dragan, R. Shah, and N. Nanda · 2024
Later among the works it cites.
C. C. Liu, F. Koto, T. Baldwin, and I. Gurevych · 2024
Later among the works it cites.
Understanding and mitigating language confusion in llms, 2024
K. Marchisio, W.-Y. Ko, A. Bérard, T. Dehaze, and S. Ruder · 2024
Later among the works it cites.
Eurollm: Multilingual language models for europe, 2024
P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. C. de Souza, A. Birch, and A. F. T. Martins · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. McGrath, M. Rahtz, J. Kramar, V. Mikulik, and S. Legg · 2023
Cited alongside, same era.
Future lens: Anticipating subsequent tokens from a single hidden state
K. Pal, J. Sun, A. Yuan, B. Wallace, and D. Bau · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
A. M. Turner, L. Thiergart, G. Leech, D. Udell, J. J. Vazquez, U. Mini, and M. MacDiarmid · 2023
Cited alongside, same era.
Llm-jp: A cross-organizational project for the research and development of fully open japanese llms
A. Aizawa, E. Aramaki, B. Chen, F. Cheng, H. Deguchi, R. Enomoto, K. Fujii, K. Fukumoto, T. Fukushima, N. Han, et al · 2024
Cited alongside, same era.
Aya 23: Open weight releases to further multilingual progress, 2024
V. Aryabumi, J. Dang, D. Talupuru, S. Dash, D. Cairuz, H. Lin, B. Venkitesh, M. Smith, K. Marchisio, S. Ruder, A. Locatelli, J. Kreutzer, N. Frosst, P. Blunsom, M. Fadaee, A. Üstün, and S. Hooker · 2024
Cited alongside, same era.
xcot: Cross-lingual instruction tuning for cross-lingual chain-of-thought reasoning, 2024
L. Chai, J. Yang, T. Sun, H. Guo, J. Liu, B. Wang, X. Liang, J. Bai, T. Li, Q. Peng, and Z. Li · 2024
Cited alongside, same era.
Y. Y. Chiu, L. Jiang, B. Y. Lin, C. Y. Park, S. S. Li, S. Ravi, M. Bhatia, M. Antoniak, Y. Tsvetkov, V. Shwartz, and Y. Choi · 2024
Cited alongside, same era.
Multilingual jailbreak challenges in large language models, 2024
Y. Deng, W. Zhang, S. J. Pan, and L. Bing · 2024
Cited alongside, same era.
T. Naous, M. J. Ryan, A. Ritter, and W. Xu · 2024
Later among the works it cites.
Steering llama 2 via contrastive activation addition, 2024
N. Panickssery, N. Gabrieli, J. Schulz, M. Tong, E. Hubinger, and A. M. Turner · 2024
Later among the works it cites.
Language model tokenizers introduce unfairness between languages
A. Petrov, E. La Malfa, P. Torr, and A. Bibi · 2024
Later among the works it cites.
Multi-fact: Assessing factuality of multilingual llms using factscore, 2024
S. Shafayat, E. Kim, J. Oh, and A. Oh · 2024
Later among the works it cites.
Is cosine-similarity of embeddings really about similarity?
H. Steck, C. Ekanadham, and N. Kallus · 2024
Later among the works it cites.
Analyzing the generalization and reliability of steering vectors, 2024
D. Tan, D. Chanin, A. Lynch, D. Kanoulas, B. Paige, A. Garriga-Alonso, and R. Kirk · 2024
Later among the works it cites.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan · 2024
Later among the works it cites.
Do llamas work in english? on the latent language of multilingual transformers, 2024
C. Wendler, V. Veselovsky, G. Monea, and R. West · 2024
Later among the works it cites.
Z. Wu, X. V. Yu, D. Yogatama, J. Lu, and Y. Kim · 2024
Later among the works it cites.
Beyond english-centric llms: What language do multilingual language models think in?
C. Zhong, F. Cheng, Q. Liu, J. Jiang, Z. Wan, C. Chu, Y. Murawaki, and S. Kurohashi · 2024
Later among the works it cites.
J. Brinkmann, C. Wendler, C. Bartelt, and A. Mueller · 2025
Closest in time.
How do multilingual language models remember facts?, 2025
C. Fierro, N. Foroutan, D. Elliott, and A. Søgaard · 2025
Closest in time.