Fetching the paper…
Reading the bibliography…
Pretrained language models (LMs) can generalize to implications of facts that they are finetuned on.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D · 2017
Earlier work this paper cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., and Chen, D · 2017
Earlier work this paper cites.
The building blocks of interpretability
Olah, C., Satyanarayan, A., Johnson, I., Carter, S., Schubert, L., Ye, K., and Mordvintsev, A · 2018
Earlier work this paper cites.
Intrinsic dimensionality explains the effectiveness of language model fine-tuning
Aghajanyan, A., Zettlemoyer, L., and Gupta, S · 2020
Earlier work this paper cites.
Thread: circuits
Cammarata, N., Carter, S., Goh, G., Olah, C., Petrov, M., Schubert, L., Voss, C., Egan, B., and Lim, S. K · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Elhage, N., Nanda, N., Olsson, C., Henighan, T., Joseph, N., Mann, B., Askell, A., Bai, Y., Chen, A., Conerly, T., DasSarma, N., Drain, D., Ganguli, D., Hatfield-Dodds, Z., Hernandez, D., Jones, A., Kernion, J., Lovitt, L., Ndousse, K., Amodei, D., Brown, T., Clark, J., Kaplan, J., McCandlish, S., and Olah, C · 2021
Earlier work this paper cites.
Causal scrubbing: A method for rigorously testing interpretability hypotheses
Chan, L., Garriga-Alonso, A., Goldowsky-Dill, N., Greenblatt, R., Nitishinskaya, J., Radhakrishnan, A., Shlegeris, B., and Thomas, N · 2022
Earlier work this paper cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Earlier work this paper cites.
Interpretability in the wild: A circuit for indirect object identification in gpt-2 small
Wang, K., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2022
Cited alongside, same era.
Physics of language models: Part 3.2, knowledge manipulation
Allen-Zhu, Z. and Li, Y · 2023
Cited alongside, same era.
Dissecting recall of factual associations in auto-regressive language models
Geva, M., Bastings, J., Filippova, K., and Globerson, A · 2023
Cited alongside, same era.
Studying large language model generalization with influence functions
Grosse, R., Bae, J., Anil, C., Elhage, N., Tamkin, A., Tajdini, A., Steiner, B., Li, D., Durmus, E., Perez, E., et al · 2023
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs
How do large language models acquire factual knowledge during pretraining?
Chang, H., Park, J., Ye, S., Yang, S., Seo, Y., Chang, D.-S., and Seo, M · 2024
Closest in time.
Evaluating the ripple effects of knowledge editing in language models
Cohen, R., Biran, E., Yoran, O., Globerson, A., and Geva, M · 2024
Closest in time.
Olmo: Accelerating the science of language models
Groeneveld, D., Beltagy, I., Walsh, P., Bhagia, A., Kinney, R., Tafjord, O., Jha, A. H., Ivison, H., Magnusson, I., Wang, Y., et al · 2024
Closest in time.
Simple and scalable strategies to continually pre-train large language models
Ibrahim, A., Thérien, B., Gupta, K., Richter, M. L., Anthony, Q., Lesort, T., Belilovsky, E., and Rish, I · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hase, P., Bansal, M., Kim, B., and Ghandeharioun, A · 2023
Cited alongside, same era.
Large language models struggle to learn long-tail knowledge
Kandpal, N., Deng, H., Roberts, A., Wallace, E., and Raffel, C · 2023
Cited alongside, same era.
Implicit meta-learning may lead language models to trust more reliable sources
Krasheninnikov, D., Krasheninnikov, E., Mlodozeniec, B. K., Maharaj, T., and Krueger, D · 2023
Cited alongside, same era.
Can lms learn new entities from descriptions? challenges in propagating injected knowledge
Onoe, Y., Zhang, M. J., Padmanabhan, S., Durrett, G., and Choi, E · 2023
Cited alongside, same era.
The two-hop curse: Llms trained on a-¿ b, b-¿ c fail to learn a–¿ c
Balesni, M., Korbak, T., and Evans, O · 2024
Cited alongside, same era.
Hopping too late: Exploring the limitations of large language models on multi-hop queries
Biran, E., Gottesman, D., Yang, S., Geva, M., and Globerson, A · 2024
Cited alongside, same era.
Taken out of context: On measuring situational awareness in llms
Berglund, L., Stickland, A. C., Balesni, M., Kaufmann, M., Tong, M., Korbak, T., Kokotajlo, D., and Evans, O
Cited in the paper.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O
Cited in the paper.
Kramár, J., Lieberum, T., Shah, R., and Nanda, N · 2024
Closest in time.
Why does new knowledge create messy ripple effects in llms?
Qin, J., Zhang, Z., Han, C., Li, M., Yu, P., and Ji, H · 2024
Closest in time.
Connecting the dots: Llms can infer and verbalize latent structure from disparate training data
Treutlein, J., Choi, D., Betley, J., Anil, C., Marks, S., Grosse, R. B., and Evans, O · 2024
Closest in time.
Grokked transformers are implicit reasoners: A mechanistic journey to the edge of generalization
Wang, B., Yue, X., Su, Y., and Sun, H · 2024
Closest in time.
Do large language models latently perform multi-hop reasoning?
Yang, S., Gribovskaya, E., Kassner, N., Geva, M., and Riedel, S · 2024
Closest in time.
Co-occurrence is not factual association in language models
Zhang, X., Li, M., and Wu, J · 2024
Closest in time.