Fetching the paper…
Reading the bibliography…
Knowledge editing aims to update outdated or incorrect knowledge in large language models (LLMs).
A simple neural network generating an interactive memory
Anderson, J. A. 1972 · 1972
Earlier work this paper cites.
Correlation matrix memories
Kohonen, T. 1972 · 1972
Earlier work this paper cites.
The feasibility of data whitening to improve performance of weather radar
Koivunen, A.; and Kostinski, A. 1999 · 1999
Earlier work this paper cites.
Face recognition based on whitening transformation of distribution of subspaces
Kawahara, T.; Nishiyama, M.; Kozakaya, T.; and Yamaguchi, O. 2007 · 2007
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M.; Schuster, R.; Berant, J.; and Levy, O. 2020 · 2012
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I.; et al. 2019 · 2019
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D.; Dong, L.; Hao, Y.; Sui, Z.; Chang, B.; and Wei, F. 2021 · 2021
Earlier work this paper cites.
Editing Factual Knowledge in Language Models
De Cao, N.; et al. 2021 · 2021
Earlier work this paper cites.
Towards continual knowledge learning of language models
Jang, J.; Ye, S.; Yang, S.; Shin, J.; Han, J.; Kim, G.; Choi, S. J.; and Seo, M. 2021 · 2021
Earlier work this paper cites.
Datasets: A community library for natural language processing
Lhoest, Q.; Del Moral, A. V.; Jernite, Y.; Thakur, A.; Von Platen, P.; Patil, S.; Chaumond, J.; Drame, M.; Plu, J.; Tunstall, L.; et al. 2021 · 2021
Earlier work this paper cites.
Mitchell, E.; Lin, C.; Bosselut, A.; Finn, C.; and Manning, C. D. 2021 · 2021
Earlier work this paper cites.
GPT-J-6B: a 6 billion parameter autoregressive language model (2021)
Wang, B.; and Komatsuzaki, A. 2022 · 2021
Cited alongside, same era.
Calibrating factual knowledge in pretrained language models
Dong, Q.; Dai, D.; Song, Y.; Xu, J.; Sui, Z.; and Li, L. 2022 · 2022
Cited alongside, same era.
Softmax Linear Units
Elhage, N.; Hume, T.; Olsson, C.; Nanda, N.; Henighan, T.; Johnston, S.; ElShowk, S.; Joseph, N.; DasSarma, N.; Mann, B.; Hernandez, D.; Askell, A.; Ndousse, K.; Jones, A.; Drain, D.; Chen, A.; Bai, Y.; Ganguli, D.; Lovitt, L.; Hatfield-Dodds, Z.; Kernion, J.; Conerly, T.; Kravec, S.; Fort, S.; Kadavath, S.; Jacobson, J.; Tran-Johnson, E.; Kaplan, J.; Clark, J.; Brown, T.; McCandlish, S.; Amodei, D.; and Olah, C. 2022a · 2022
Cited alongside, same era.
Memory-based model editing at scale
Mitchell, E.; Lin, C.; Bosselut, A.; Manning, C. D.; and Finn, C. 2022 · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S.; Schoelkopf, H.; Anthony, Q. G.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M. A.; Purohit, S.; Prashanth, U. S.; Raff, E.; et al. 2023 · 2023
PMET: Precise Model Editing in a Transformer
Li, X.; Li, S.; Song, S.; Yang, J.; Ma, J.; and Yu, J. 2023 · 2023
Later among the works it cites.
Massive editing for large language models via meta learning
Tan, C.; Zhang, G.; and Fu, J. 2023 · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023 · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Yao, Y.; Wang, P.; Tian, B.; Cheng, S.; Li, Z.; Deng, S.; Chen, H.; and Zhang, N. 2023 · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards Monosemanticity: Decomposing Language Models With Dictionary Learning
Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; Lasenby, R.; Wu, Y.; Kravec, S.; Schiefer, N.; Maxwell, T.; Joseph, N.; Hatfield-Dodds, Z.; Tamkin, A.; Nguyen, K.; McLean, B.; Burke, J. E.; Hume, T.; Carter, S.; Henighan, T.; and Olah, C. 2023 · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H.; Ewart, A.; Riggs, L.; Huben, R.; and Sharkey, L. 2023 · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Geva, M.; Bastings, J.; Filippova, K.; and Globerson, A. 2023 · 2023
Cited alongside, same era.
Finding neurons in a haystack: Case studies with sparse probing
Gurnee, W.; Nanda, N.; Pauly, M.; Harvey, K.; Troitskii, D.; and Bertsimas, D. 2023 · 2023
Cited alongside, same era.
Transformer-patcher: One mistake worth one neuron
Huang, Z.; Shen, Y.; Zhang, X.; Zhou, J.; Rong, W.; and Xiong, Z. 2023 · 2023
Cited alongside, same era.
Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; et al. 2022b
Cited in the paper.
Locating and editing factual associations in GPT
Meng, K.; Bau, D.; Andonian, A.; and Belinkov, Y. 2022a
Cited in the paper.
Zhong, Z.; Wu, Z.; Manning, C. D.; Potts, C.; and Chen, D. 2023 · 2023
Later among the works it cites.
Scaling and evaluating sparse autoencoders
Gao, L.; la Tour, T. D.; Tillman, H.; Goh, G.; Troll, R.; Radford, A.; Sutskever, I.; Leike, J.; and Wu, J. 2024 · 2024
Closest in time.
Aging with grace: Lifelong model editing with discrete key-value adaptors
Hartvigsen, T.; Sankaranarayanan, S.; Palangi, H.; Kim, Y.; and Ghassemi, M. 2024 · 2024
Closest in time.
Wilke: Wise-layer knowledge editor for lifelong knowledge editing
Hu, C.; Cao, P.; Chen, Y.; Liu, K.; and Zhao, J. 2024 · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders
Rajamanoharan, S.; Conmy, A.; Smith, L.; Lieberum, T.; Varma, V.; Kramár, J.; Shah, R.; and Nanda, N. 2024 · 2024
Closest in time.
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Templeton, A.; Conerly, T.; Marcus, J.; Lindsey, J.; Bricken, T.; Chen, B.; Pearce, A.; Citro, C.; Ameisen, E.; Jones, A.; Cunningham, H.; Turner, N. L.; McDougall, C.; MacDiarmid, M.; Freeman, C. D.; Sumers, T. R.; Rees, E.; Batson, J.; Jermyn, A.; Carter, S.; Olah, C.; and Henighan, T. 2024 · 2024
Closest in time.