Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) store and retrieve vast amounts of factual knowledge acquired during pre-training.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Scientific and lay communities: Earning epistemic trust through knowledge sharing
Heidi E. Grasswick. 2010 · 2010
Earlier work this paper cites.
64Knowledge and expertise
Katherine Hawley. 2012 · 2012
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019 · 2019
Earlier work this paper cites.
X-FACTR: Multilingual factual knowledge retrieval from pretrained language models
Zhengbao Jiang, Antonios Anastasopoulos, Jun Araki, Haibo Ding, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Earlier work this paper cites.
Multilingual LAMA: Investigating knowledge in multilingual pretrained language models
Nora Kassner, Philipp Dufter, and Hinrich Schütze. 2021 · 2021
Earlier work this paper cites.
Few-shot learning with multilingual language models
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, et al. 2021 · 2021
Earlier work this paper cites.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Earlier work this paper cites.
Factual consistency of multilingual pretrained language models
Constanza Fierro and Anders Søgaard. 2022 · 2022
Earlier work this paper cites.
Discovering language-neutral sub-networks in multilingual language models
Negar Foroutan, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut, and Karl Aberer. 2022 · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
Trust as an unquestioning attitude
C. Thi Nguyen. 2022 · 2022
Cited alongside, same era.
Direct and indirect effects
Judea Pearl. 2022 · 2022
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson. 2023 · 2023
Cited alongside, same era.
Mamba: Linear-time sequence modeling with selective state spaces
Albert Gu and Tri Dao. 2023 · 2023
Patchscopes: A unifying framework for inspecting hidden representations of language models
Asma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon, and Mor Geva. 2024 · 2024
Closest in time.
A comparison of language modeling and translation as multilingual pretraining objectives
Zihao Li, Shaoxiong Ji, Timothee Mickus, Vincent Segonne, and Jörg Tiedemann. 2024 · 2024
Closest in time.
Eurollm: Multilingual language models for europe
Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M Guerreiro, Ricardo Rei, Duarte M Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, et al. 2024 · 2024
Closest in time.
Locating and editing factual associations in mamba
Arnab Sen Sharma, David Atkinson, and David Bau. 2024 · 2024
Closest in time.
Mass-editing memory with attention in transformers: A cross-lingual exploration of knowledge
Daniel Tamayo, Aitor Gonzalez-Agirre, Javier Hernando, and Marta Villegas. 2024 · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Cross-lingual consistency of factual knowledge in multilingual language models
Jirui Qi, Raquel Fernández, and Arianna Bisazza. 2023 · 2023
Cited alongside, same era.
Summing up the facts: Additive mechanisms behind factual recall in llms
Bilal Chughtai, Alan Cooney, and Neel Nanda. 2024 · 2024
Cited alongside, same era.
How do llamas process multilingual text? a latent exploration through activation patching
Clément Dumas, Veniamin Veselovsky, Giovanni Monea, Robert West, and Chris Wendler. 2024 · 2024
Cited alongside, same era.
On the similarity of circuits across languages: a case study on the subject-verb agreement task
Javier Ferrando and Marta R. Costa-jussà. 2024 · 2024
Cited alongside, same era.
Defining knowledge: Bridging epistemology and large language models
Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Anders Søgaard, and Nicolas Garneau. 2024 · 2024
Cited alongside, same era.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau. 2024 · 2024
Closest in time.
Locating and extracting relational concepts in large language models
Zijian Wang, Britney Whyte, and Chang Xu. 2024 · 2024
Closest in time.
Do llamas work in English? on the latent language of multilingual transformers
Chris Wendler, Veniamin Veselovsky, Giovanni Monea, and Robert West. 2024 · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda. 2024 · 2024
Closest in time.
The same but different: Structural similarities and differences in multilingual language modeling
Ruochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff, and Ellie Pavlick. 2025 · 2025
Closest in time.
GeoMLAMA: Geo-diverse commonsense probing on multilingual pre-trained language models
Da Yin, Hritik Bansal, Masoud Monajatipoor, Liunian Harold Li, and Kai-Wei Chang. 2022 · 2055
Closest in time.