Fetching the paper…
Reading the bibliography…
Large language models (LLMs) store extensive factual knowledge, but the mechanisms behind how they store and express this knowledge remain unclear.
Cell adhesion molecules: signalling functions at the synapse
Matthew B Dalva, Andrew C McClelland, and Matthew S Kayser · 2007
Earlier work this paper cites.
Investigating saturation effects in integrated gradients
Vivek Miglani, Narine Kokhlikyan, Bilal Alsallakh, Miguel Martin, and Orion Reblitz-Richardson · 2010
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Synapse development organized by neuronal activity-regulated immediate-early genes
Seungjoon Kim, Hyeonho Kim, and Ji Won Um · 2018
Earlier work this paper cites.
Memory formation depends on both synapse-specific modifications of synaptic strength and cell-specific increases in excitability
John Lisman, Katherine Cooper, Megha Sehgal, and Alcino J Silva · 2018
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Measuring and improving consistency in pretrained language models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy · 2021
Earlier work this paper cites.
Discretized integrated gradients for explaining language models
Soumya Sanyal and Xiang Ren · 2021
Earlier work this paper cites.
Integrated directional gradients: Feature interaction attribution for neural NLP models
Sandipan Sikdar, Parantapa Bhattacharya, and Kieran Heese · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Earlier work this paper cites.
Organic electrochemical neurons and synapses with ion mediated spiking
Padinhare Cholakkal Harikesh, Chi-Yuan Yang, Deyu Tu, Jennifer Y Gerasimov, Abdul Manan Dar, Adam Armada-Moreira, Matteo Massetti, Renee Kroon, David Bliman, Roger Olsson, et al · 2022
Earlier work this paper cites.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel Lundström, Tianjian Huang, and Meisam Razaviyayn · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex J Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Discovering low-rank subspaces for language-agnostic multilingual representations
Zhihui Xie, Handong Zhao, Tong Yu, and Shuai Li · 2022
Cited alongside, same era.
Distributed representations: Composition & superposition
Anthropic · 2023
Cited alongside, same era.
Towards monosemanticity: Decomposing language models with dictionary learning
Trenton Bricken, Adly Templeton, Joshua Batson, Brian Chen, Adam Jermyn, Tom Conerly, Nick Turner, Cem Anil, Carson Denison, Amanda Askell, Robert Lasenby, Yifan Wu, Shauna Kravec, Nicholas Schiefer, Tim Maxwell, Nicholas Joseph, Zac Hatfield-Dodds, Alex Tamkin, Karina Nguyen, Brayden McLean, Josiah E Burke, Tristan Hume, Shan Carter, Tom Henighan, and Christopher Olah · 2023
Cited alongside, same era.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2023
Cited alongside, same era.
Editing large language models: Problems, methods, and opportunities
Yunzhi Yao, Peng Wang, Bozhong Tian, Siyuan Cheng, Zhoubo Li, Shumin Deng, Huajun Chen, and Ningyu Zhang · 2023
Later among the works it cites.
Unveiling a core linguistic region in large language models, 2023
Jun Zhao, Zhihao Zhang, Yide Ma, Qi Zhang, Tao Gui, Luhui Gao, and Xuanjing Huang · 2023
Later among the works it cites.
The life cycle of knowledge in big language models: A survey
Boxi Cao, Hongyu Lin, Xianpei Han, and Le Sun · 2024
Closest in time.
Chenhui Hu, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao · 2024
Closest in time.
Wenyue Hua, Jiang Guo, Mingwen Dong, Henghui Zhu, Patrick Ng, and Zhiguo Wang · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sequential integrated gradients: a simple but effective method for explaining language models
Joseph Enguehard · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Cited alongside, same era.
Detecting edit failures in large language models: An improved specificity benchmark
Jason Hoelscher-Obermaier, Julia Persson, Esben Kran, Ioannis Konstas, and Fazl Barez · 2023
Cited alongside, same era.
Provenance documentation to enable explainable and trustworthy ai: A literature review
Amruta Kale, Tin Nguyen, Jr. Harris, Frederick C., Chenhao Li, Jiyin Zhang, and Xiaogang Ma · 2023
Cited alongside, same era.
Evaluation on chatgpt for chinese language understanding
Linhan Li, Huaping Zhang, Chunjin Li, Haowen You, and Wenyao Cui · 2023
Cited alongside, same era.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
Cited alongside, same era.
Closest in time.
Takeshi Kojima, Itsuki Okimura, Yusuke Iwasawa, Hitomi Yanaka, and Yutaka Matsuo · 2024
Closest in time.
Unveiling the pitfalls of knowledge editing for large language models
Zhoubo Li, Ningyu Zhang, Yunzhi Yao, Mengru Wang, Xi Chen, and Huajun Chen · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date, 2024
MetaAI · 2024
Closest in time.
What does the knowledge neuron thesis have to do with knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn · 2024
Closest in time.
Understanding neural circuit function through synaptic engineering
Ithai Rabinowitch, Daniel A Colón-Ramos, and Michael Krieg · 2024
Closest in time.
Identifying semantic induction heads to understand in-context learning
Jie Ren, Qipeng Guo, Hang Yan, Dongrui Liu, Xipeng Qiu, and Dahua Lin · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models
Tianyi Tang, Wenyang Luo, Haoyang Huang, Dongdong Zhang, Xiaolei Wang, Xin Zhao, Furu Wei, and Ji-Rong Wen · 2024
Closest in time.
MULFE: A multi-level benchmark for free text model editing
Chenhao Wang, Pengfei Cao, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, and Jun Zhao · 2024
Closest in time.
Unveiling factual recall behaviors of large language models through knowledge neurons
Yifei Wang, Yuheng Chen, Wanting Wen, Yu Sheng, Linjing Li, and Daniel Dajun Zeng · 2024
Closest in time.
Knowledge circuits in pretrained transformers, 2024
Yunzhi Yao, Ningyu Zhang, Zekun Xi, Mengru Wang, Ziwen Xu, Shumin Deng, and Huajun Chen · 2024
Closest in time.
Unveiling linguistic regions in large language models
Zhihao Zhang, Jun Zhao, Qi Zhang, Tao Gui, and Xuanjing Huang · 2024
Closest in time.