Fetching the paper…
Reading the bibliography…
Understanding the internal mechanisms of transformer-based language models remains challenging.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2012
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
Zoom in: An introduction to circuits
Olah, C., Cammarata, N., Schubert, L., et al · 2020
Earlier work this paper cites.
Visualizing the impact of feature attribution baselines
Sturmfels, P., Lundberg, S., and Lee, S.-I · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Vig, J., Gehrmann, S., Belinkov, Y., Qian, S., Nevo, D., Singer, Y., and Shieber, S · 2020
Earlier work this paper cites.
Causal abstractions of neural networks
Geiger, A., Lu, H., Icard, T., Smith, J., and Doe, J · 2021
Earlier work this paper cites.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., et al · 2022
Earlier work this paper cites.
Mechanistic interpretability, variables, and the importance of interpretable bases
Olah, C · 2022
Earlier work this paper cites.
Achiam, J., Adler, S., Agarwal, S., and et al · 2023
Earlier work this paper cites.
Identifying and adapting transformer-components responsible for gender bias in an english language model
Chintam, A., Beloch, R., Zuidema, W., Hanna, M., and van der Wal, O · 2023
Earlier work this paper cites.
Towards automated circuit discovery for mechanistic interpretability
Conmy, A., Mavor-Parker, A., Lynch, A., et al · 2023
Cited alongside, same era.
Sparse autoencoders find highly interpretable features in language models
Cunningham, H., Ewart, A., Riggs, L., et al · 2023
Cited alongside, same era.
Lieberum, T., Rahtz, M., Kramár, J., et al · 2023
Cited alongside, same era.
Mechanistic interpretability quickstart guide
Nanda, N · 2023
Cited alongside, same era.
Attribution patching outperforms automated circuit discovery
Syed, A., Rager, C., and Conmy, A · 2023
Cited alongside, same era.
He, Z., Ge, X., Tang, Q., et al · 2024
Later among the works it cites.
Dissecting fine-tuning unlearning in large language models
Hong, Y., Zou, Y., Hu, L., Zeng, Z., Wang, D., and Yang, H · 2024
Later among the works it cites.
A hopfieldian view-based interpretation for chain-of-thought reasoning
Hu, L., Liu, L., Yang, S., Chen, X., Xiao, H., Li, M., Zhou, P., Ali, M. A., and Wang, D · 2024
Later among the works it cites.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Marks, S., Rager, C., Michaud, E. J., et al · 2024
Later among the works it cites.
Circuit component reuse across tasks in transformer language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanford alpaca: An instruction-following llama model
Taori, R., Gulrajani, I., Zhang, T., Dubois, Y., Li, X., Guestrin, C., Liang, P., and Hashimoto, T. B · 2023
Cited alongside, same era.
Linear representations of sentiment in large language models
Tigges, C., Hollinsworth, O. J., Geiger, A., et al · 2023
Cited alongside, same era.
Interpretability in the wild: A circuit for indirect object identification in gpt-2 small
Wang, K. R., Variengien, A., Conmy, A., Shlegeris, B., and Steinhardt, J · 2023
Cited alongside, same era.
Locate-then-edit for multi-hop factual recall under knowledge editing
Zhang, Z., Li, Y., Kan, Z., Cheng, K., Hu, L., and Wang, D · 2023
Cited alongside, same era.
Mechanistic interpretability for ai safety–a review
Bereska, L. and Gavves, E · 2024
Cited alongside, same era.
Leveraging logical rules in knowledge editing: A cherry on the top
Cheng, K., Ali, M. A., Yang, S., Lin, G., Zhai, Y., Fei, H., Xu, K., Yu, L., Hu, L., and Wang, D · 2024
Cited alongside, same era.
Finding alignments between interpretable causal variables and distributed neural representations
Geiger, A., Wu, Z., Potts, C., Icard, T., and Goodman, N · 2024
Cited alongside, same era.
Merullo, J., Eickhoff, C., and Pavlick, E · 2024
Later among the works it cites.
Transformer circuit faithfulness metrics are not robust
Miller, J., Chughtai, B., and Saunders, W · 2024
Later among the works it cites.
Nainani, J · 2024
Later among the works it cites.
Sparse autoencoders enable scalable and reliable circuit identification in language models
O’Neill, C. and Bui, T · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Prakash, N., Shaham, T. R., Haklay, T., et al · 2024
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Wu, Z., Geiger, A., Icard, T., et al · 2024
Later among the works it cites.
What makes your model a low-empathy or warmth person: Exploring the origins of personality in llms
Yang, S., Zhu, S., Bao, R., Liu, L., Cheng, Y., Hu, L., Li, M., and Wang, D · 2024
Later among the works it cites.