Fetching the paper…
Reading the bibliography…
How do transformer-based large language models (LLMs) store and retrieve knowledge? We focus on the most basic form of this task -- factual recall, where the model is tasked with explicitly surfacing stored facts in prompts of form `Fact: The Colosseum is in the country of'.
Language Models as Knowledge Bases?, September 2019
Petroni, F., Rocktäschel, T., Lewis, P., Bakhtin, A., Wu, Y., Miller, A. H., and Riedel, S · 1909
Earlier work this paper cites.
How Can We Know What Language Models Know?, May 2020
Jiang, Z., Xu, F. F., Araki, J., and Neubig, G · 1911
Earlier work this paper cites.
How Much Knowledge Can You Pack Into the Parameters of a Language Model?, October 2020
Roberts, A., Raffel, C., and Shazeer, N · 2002
Earlier work this paper cites.
Language Models are Few-Shot Learners, July 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories, September 2021
Geva, M., Schuster, R., Berant, J., and Levy, O · 2012
Earlier work this paper cites.
Collaborative data science, 2015
Inc., P. T · 2015
Earlier work this paper cites.
Feature Visualization
Olah, C., Mordvintsev, A., and Schubert, L · 2017
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Kopf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., and Chintala, S · 2019
Earlier work this paper cites.
interpreting GPT: the logit lens — LessWrong, January 2020
nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., Goh, G., Petrov, M., and Carter, S · 2020
Earlier work this paper cites.
Curve circuits
Cammarata, N., Goh, G., Carter, S., Voss, C., Schubert, L., and Olah, C · 2021
Earlier work this paper cites.
Measuring and Improving Consistency in Pretrained Language Models, May 2021
Elazar, Y., Kassner, N., Ravfogel, S., Ravichander, A., Hovy, E., Schütze, H., and Goldberg, Y · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits, 2021
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Causal Abstractions of Neural Networks, October 2021
Geiger, A., Lu, H., Icard, T., and Potts, C · 2021
Earlier work this paper cites.
GPT-J-6B: A 6 billion parameter autoregressive language model, May 2021
Wang, B. and Komatsuzaki, A · 2021
Cited alongside, same era.
Factual Probing Is [MASK]: Learning vs. Learning to Recall, December 2021
Zhong, Z., Friedman, D., and Chen, D · 2021
Cited alongside, same era.
Toy models of superposition
Elhage, N., Hume, T., Olsson, C., Schiefer, N., Henighan, T., Kravec, S., Hatfield-Dodds, Z., Lasenby, R., Drain, D., Chen, C., Grosse, R., McCandlish, S., Kaplan, J., Amodei, D., Wattenberg, M., and Olah, C · 2022
Cited alongside, same era.
Inducing Causal Structure for Interpretable Neural Networks, July 2022
Geiger, A., Wu, Z., Lu, H., Rozner, J., Kreiss, E., Icard, T., Goodman, N. D., and Potts, C · 2022
Cited alongside, same era.
Language models can explain neurons in language models, 2023
Bills, S., Cammarata, N., Mossing, D., Tillman, H., Gao, L., Goh, G., Sutskever, I., Leike, J., Wu, J., and Saunders, W · 2023
Later among the works it cites.
A Toy Model of Universality: Reverse Engineering How Networks Learn Group Operations, May 2023
Chughtai, B., Chan, L., and Nanda, N · 2023
Later among the works it cites.
Towards Automated Circuit Discovery for Mechanistic Interpretability, July 2023
Conmy, A., Mavor-Parker, A. N., Lynch, A., Heimersheim, S., and Garriga-Alonso, A · 2023
Later among the works it cites.
Sparse Autoencoders Find Highly Interpretable Features in Language Models, September 2023
Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L · 2023
Later among the works it cites.
Studying Large Language Model Generalization with Influence Functions, August 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Geva, M., Caciularu, A., Wang, K. R., and Goldberg, Y · 2022
Cited alongside, same era.
TransformerLens/further_comments.md at main · neelnanda-io/TransformerLens, 2022
Nanda, N · 2022
Cited alongside, same era.
TransformerLens, January 2023
Nanda, N · 2022
Cited alongside, same era.
In-context learning and induction heads
Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Johnston, S., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2022
Cited alongside, same era.
But is it really in Rome? An investigation of the ROME model editing technique — AI Alignment Forum, 2022
Thibodeau, J · 2022
Cited alongside, same era.
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Emergent Abilities of Large Language Models, October 2022
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Cited alongside, same era.
The Reversal Curse: LLMs trained on ”A is B” fail to learn ”B is A”, September 2023
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2023
Cited alongside, same era.
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., Hubinger, E., Lukošiūtė, K., Nguyen, K., Joseph, N., McCandlish, S., Kaplan, J., and Bowman, S. R · 2023
Later among the works it cites.
Finding Neurons in a Haystack: Case Studies with Sparse Probing, June 2023
Gurnee, W., Nanda, N., Pauly, M., Harvey, K., Troitskii, D., and Bertsimas, D · 2023
Later among the works it cites.
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2023
Later among the works it cites.
Linearity of Relation Decoding in Transformer Language Models, August 2023
Hernandez, E., Sharma, A. S., Haklay, T., Meng, K., Wattenberg, M., Andreas, J., Belinkov, Y., and Bau, D · 2023
Later among the works it cites.
Attention Head Superposition, 2023
Jermyn, A., Olah, C., and Henighan, T · 2023
Later among the works it cites.
The Hydra Effect: Emergent Self-repair in Language Model Computations, July 2023
McGrath, T., Rahtz, M., Kramar, J., Mikulik, V., and Legg, S · 2023
Later among the works it cites.
Progress measures for grokking via mechanistic interpretability, January 2023
Nanda, N., Chan, L., Lieberum, T., Smith, J., and Steinhardt, J · 2023
Later among the works it cites.
An adversarial example for Direct Logit Attribution: memory management in gelu-4l
Rager, C., Lau, Y.-T., Dao, J., and Jett · 2023
Later among the works it cites.
Universal and Transferable Adversarial Attacks on Aligned Language Models, July 2023
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.