Fetching the paper…
Reading the bibliography…
The locate-then-edit paradigm has shown significant promise for knowledge editing (KE) in Large Language Models (LLMs).
Solving Least Squares Problems
Lawson, C. L. and Hanson, R. J · 1995
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Belinkov, Y. and Glass, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Belinkov, Y · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2022
Earlier work this paper cites.
Analyzing transformers in embedding space
Dar, G., Geva, M., Gupta, A., and Berant, J · 2022
Earlier work this paper cites.
The state of the art in open domain complex question answering: a survey
Etezadi, R. and Shamsfard, M · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Geva, M., Caciularu, A., Wang, K., and Goldberg, Y · 2022
Earlier work this paper cites.
Locating and editing factual associations in GPT
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2022
Earlier work this paper cites.
Eliciting latent predictions from transformers with the tuned lens
Belrose, N., Furman, Z., Smith, L., Halawi, D., Ostrovsky, I., McKinney, L., Biderman, S., and Steinhardt, J · 2023
Earlier work this paper cites.
Analyzing transformers in embedding space
Dar, G., Geva, M., Gupta, A., and Berant, J · 2023
Cited alongside, same era.
In-context learning creates task vectors, 2023
Hendel, R., Geva, M., and Globerson, A · 2023
Cited alongside, same era.
Towards a mechanistic interpretation of multi-step reasoning capabilities of language models
Hou, Y., Li, J., Fei, Y., Stolfo, A., Zhou, W., Zeng, G., Bosselut, A., and Sachan, M · 2023
Cited alongside, same era.
Decoderlens: Layerwise interpretation of encoder-decoder transformers
Langedijk, A., Mohebbi, H., Sarti, G., Zuidema, W., and Jumelet, J · 2023
Cited alongside, same era.
A survey on knowledge editing of neural networks
Mazzia, V., Pedrani, A., Caciolai, A., Rottmann, K., and Bernardi, D · 2023
Cited alongside, same era.
A unified framework for model editing, 2024
Gupta, A., Sajnani, D., and Anumanchipalli, G · 2024
Closest in time.
Model editing with canonical examples
Hewitt, J., Chen, S., Xie, L. L., Adams, E., Liang, P., and Manning, C. D · 2024
Closest in time.
Wilke: Wise-layer knowledge editor for lifelong knowledge editing
Hu, C., Cao, P., Chen, Y., Liu, K., and Zhao, J · 2024
Closest in time.
Jin, Z., Cao, P., Yuan, H., Chen, Y., Xu, J., Li, H., Jiang, X., Liu, K., and Zhao, J · 2024
Closest in time.
Investigating multi-hop factual shortcuts in knowledge editing of large language models, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mass-editing memory in a transformer
Meng, K., Sharma, A. S., Andonian, A. J., Belinkov, Y., and Bau, D · 2023
Cited alongside, same era.
A mechanism for solving relational tasks in transformer language models
Merullo, J., Eickhoff, C., and Pavlick, E · 2023
Cited alongside, same era.
Massive editing for large language models via meta learning
Tan, C., Zhang, G., and Fu, J · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K. R., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., Bikel, D. M., Blecher, L., Ferrer, C. C., Chen, M., Cucurull, G., Esiobu, D., Fernandes, J., Fu, J., Fu, W., Fuller, B., Gao, C., Goswami, V., Goyal, N., Hartshorn, A. S., Hosseini, S., Hou, R., Inan, H., Kardas, M., Kerkez, V., Khabsa, M., Kloumann, I. M., Korenev, A. V., Koura, P. S., Lachaux, M.-A., Lavril, T., Lee, J., Liskovich, D., Lu, Y., Mao, Y., Martinet, X., Mihaylov, T., Mishra, P., Molybog, I., Nie, Y., Poulton, A., Reizenstein, J., Rungta, R., Saladi, K., Schelten, A., Silva, R., Smith, E. M., Subramanian, R., Tan, X., Tang, B., Taylor, R., Williams, A., Kuan, J. X., Xu, P., Yan, Z., Zarov, I., Zhang, Y., Fan, A., Kambadur, M., Narang, S., Rodriguez, A., Stojnic, R., Edunov, S., and Scialom, T · 2023
Cited alongside, same era.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., and Chen, D · 2023
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2024
Cited alongside, same era.
Hopping too late: Exploring the limitations of large language models on multi-hop queries, 2024
Biran, E., Gottesman, D., Yang, S., Geva, M., and Globerson, A · 2024
Cited alongside, same era.
Ju, T., Chen, Y., Yuan, X., Zhang, Z., Du, W., Zheng, Y., and Liu, G · 2024
Closest in time.
The devil is in the neurons: Interpreting and mitigating social biases in language models
Liu, Y., Liu, Y., Chen, X., Chen, P.-Y., Zan, D., Kan, M.-Y., and Ho, T.-Y · 2024
Closest in time.
Language models implement simple word2vec-style vector arithmetic, 2024
Merullo, J., Eickhoff, C., and Pavlick, E · 2024
Closest in time.
Function vectors in large language models, 2024
Todd, E., Li, M. L., Sharma, A. S., Mueller, A., Wallace, B. C., and Bau, D · 2024
Closest in time.
Cognitive overload attack:prompt injection for long context, 2024
Upadhayay, B., Behzadan, V., and Karbasi, A · 2024
Closest in time.
Gaussian process probes (gpp) for uncertainty-aware probing
Wang, Z., Ku, A., Baldridge, J., Griffiths, T., and Kim, B · 2024
Closest in time.
Efficient streaming language models with attention sinks, 2024
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
History matters: Temporal knowledge editing in large language model
Yin, X., Jiang, J., Yang, L., and Wan, X · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models, 2024
Zhang, N., Yao, Y., Tian, B., Wang, P., Deng, S., Wang, M., Xi, Z., Mao, S., Zhang, J., Ni, Y., Cheng, S., Xu, Z., Xu, X., Gu, J.-C., Jiang, Y., Xie, P., Huang, F., Liang, L., Zhang, Z., Zhu, X., Zhou, J., and Chen, H · 2024
Closest in time.