Fetching the paper…
Reading the bibliography…
Recently, large language models (LLMs) have achieved tremendous breakthroughs in the field of NLP, but still lack understanding of their internal neuron activities when processing different languages.
Language control in the bilingual brain
Jenny Crinion, Robert Turner, Alice Grogan, Takashi Hanakawa, Uta Noppeney, Joseph T Devlin, Toshihiko Aso, Shinichi Urayama, Hidenao Fukuyama, Katharine Stockton, et al. 2006 · 2006
Earlier work this paper cites.
Speaking in multiple languages: Neural correlates of language proficiency in multilingual word production
Gerda Videsott, Bärbel Herrnberger, Klaus Hoenig, Edgar Schilly, Jo Grothe, Werner Wiater, Manfred Spitzer, and Markus Kiefer. 2010 · 2010
Earlier work this paper cites.
The brain basis of language processing: from structure to function
Angela D Friederici. 2011 · 2011
Earlier work this paper cites.
Balanced k-means for clustering
Mikko I Malinen and Pasi Fränti. 2014 · 2014
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021 · 2021
Earlier work this paper cites.
Common sense beyond english: Evaluating and improving multilingual language models for commonsense reasoning
Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. 2021a · 2021
Earlier work this paper cites.
Importance-based neuron allocation for multilingual neural machine translation
Wanying Xie, Yang Feng, Shuhao Gu, and Dong Yu. 2021 · 2021
Earlier work this paper cites.
Share or not? learning to schedule language-specific capacity for multilingual translation
Biao Zhang, Ankur Bapna, Rico Sennrich, and Orhan Firat. 2021 · 2021
Cited alongside, same era.
On the multilingual capabilities of very large-scale English language models
Jordi Armengol-Estapé, Ona de Gibert Bonet, and Maite Melero. 2022 · 2022
Cited alongside, same era.
The bigscience roots corpus: A 1.6 tb composite multilingual dataset
Hugo Laurençon, Lucile Saulnier, Thomas Wang, Christopher Akiki, Albert Villanova del Moral, Teven Le Scao, Leandro Von Werra, Chenghao Mou, Eduardo González Ponferrada, Huu Nguyen, et al. 2022 · 2022
Cited alongside, same era.
Language models are multilingual chain-of-thought reasoners
Freda Shi, Mirac Suzgun, Markus Freitag, Xuezhi Wang, Suraj Srivats, Soroush Vosoughi, Hyung Won Chung, Yi Tay, Sebastian Ruder, Denny Zhou, et al. 2022 · 2022
Cited alongside, same era.
MoEfication: Transformer feed-forward layers are mixtures of experts
Zhengyan Zhang, Yankai Lin, Zhiyuan Liu, Peng Li, Maosong Sun, and Jie Zhou. 2022 · 2022
Cited alongside, same era.
Relu strikes back: Exploiting activation sparsity in large language models
Seyed Iman Mirzadeh, Keivan Alizadeh-Vahid, Sachin Mehta, Carlo C del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar. 2023 · 2023
Later among the works it cites.
Llama-moe: Building mixture-of-experts from llama with continual pre-training
LLaMA-MoE Team. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023 · 2023
Later among the works it cites.
A survey of large language models
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, et al. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Unveiling multilinguality in transformer models: Exploring language specificity in feed-forward networks
Sunit Bhattacharya and Ondřej Bojar. 2023 · 2023
Cited alongside, same era.
Interpretability illusions in the generalization of simplified models
Dan Friedman, Andrew Kyle Lampinen, Lucas Dixon, Danqi Chen, and Asma Ghandeharioun. 2023 · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023 · 2023
Cited alongside, same era.
Learning language specific sub-network for multilingual machine translation
Zehui Lin, Liwei Wu, Mingxuan Wang, and Lei Li. 2021b
Cited in the paper.
Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, and Yutaka Matsuo. 2024 · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen. 2024 · 2024
Closest in time.
Unveiling linguistic regions in large language models
Zhihao Zhang, Jun Zhao, Qi Zhang, Tao Gui, and Xuanjing Huang. 2024 · 2024
Closest in time.
How do large language models handle multilingualism?
Yiran Zhao, Wenxuan Zhang, Guizhen Chen, Kenji Kawaguchi, and Lidong Bing. 2024 · 2024
Closest in time.