Fetching the paper…
Reading the bibliography…
We analyze the storage and recall of factual associations in autoregressive transformer language models, finding evidence that these associations correspond to localized, directly-editable computations.
A simple neural network generating an interactive memory
Anderson, J. A · 1972
Earlier work this paper cites.
Correlation matrix memories
Kohonen, T · 1972
Earlier work this paper cites.
Direct and indirect effects
Pearl, J · 2001
Earlier work this paper cites.
Causal mediation analysis for interpreting neural NLP: The case of gender bias
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Sakenis, S., Huang, J., Singer, Y., and Shieber, S · 2004
Earlier work this paper cites.
Causality: Models, Reasoning and Inference
Pearl, J · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Ettinger, A., Elgohary, A., and Resnik, P · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Adi, Y., Kermany, E., Belinkov, Y., Lavi, O., and Goldberg, Y · 2017
Earlier work this paper cites.
Probing Classifiers: Promises, Shortcomings, and Advances
Belinkov, Y · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Belinkov, Y., Durrani, N., Dalvi, F., Sajjad, H., and Glass, J · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Levy, O., Seo, M., Choi, E., and Zettlemoyer, L · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Conneau, A., Kruszewski, G., Lample, G., Barrault, L., and Baroni, M · 2018
Earlier work this paper cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Hupkes, D., Veldhoen, S., and Zuidema, W · 2018
Cited alongside, same era.
Generating informative and diverse conversational responses via adversarial information maximization
Zhang, Y., Galley, M., Gao, J., Gan, Z., Li, X., Brockett, C., and Dolan, W. B · 2018
Cited alongside, same era.
Analysis methods in neural language processing: A survey
Belinkov, Y. and Glass, J · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Modifying memories in transformer models
Zhu, C., Rawat, A. S., Zaheer, M., Bhojanapalli, S., Li, D., Yu, F., and Kumar, S · 2020
Later among the works it cites.
Editing factual knowledge in language models
De Cao, N., Aziz, W., and Titov, I · 2021
Later among the works it cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Later among the works it cites.
CausaLM: Causal model explanation through counterfactual language models
Feder, A., Oved, N., Shalit, U., and Reichart, R · 2021
Later among the works it cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Finlayson, M., Mueller, A., Gehrmann, S., Shieber, S., Linzen, T., and Belinkov, Y · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Cited alongside, same era.
Rewriting a deep generative model
Bau, D., Liu, S., Wang, T., Zhu, J.-Y., and Torralba, A · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
How can we know what language models know?
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Lewis, M., Liu, Y., Goyal, N., Ghazvininejad, M., Mohamed, A., Levy, O., Stoyanov, V., and Zettlemoyer, L · 2020
Cited alongside, same era.
How context affects language models’ factual predictions
Petroni, F., Lewis, P., Piktus, A., Rocktäschel, T., Wu, Y., Miller, A. H., and Riedel, S · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2021
Later among the works it cites.
Do language models have beliefs? methods for detecting, updating, and visualizing model beliefs
Hase, P., Diab, M., Celikyilmaz, A., Li, X., Kozareva, Z., Stoyanov, V., Bansal, M., and Iyer, S · 2021
Later among the works it cites.
Fast model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Finn, C., and Manning, C. D · 2021
Later among the works it cites.
GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model
Wang, B. and Komatsuzaki, A · 2021
Later among the works it cites.
Of non-linearity and commutativity in BERT
Zhao, S., Pascual, D., Brunner, G., and Wattenhofer, R · 2021
Later among the works it cites.
Factual probing is [MASK]: Learning vs. learning to recall
Zhong, Z., Friedman, D., and Chen, D · 2021
Later among the works it cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2022
Closest in time.
Mass-editing memory in a transformer
Meng, K., Sen Sharma, A., Andonian, A., Belinkov, Y., and Bau, D · 2022
Closest in time.