Fetching the paper…
Reading the bibliography…
In this paper, we introduce DiscoGP, a novel framework for extracting self-contained modular units, or sheaves, within neural language models (LMs).
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 1910
Earlier work this paper cites.
DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2020 · 1910
Earlier work this paper cites.
Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville. 2013 · 2013
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Learning Sparse Neural Networks through L_0 Regularization
Christos Louizos, Max Welling, and Diederik P. Kingma. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin. 2019 · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Zoom In: An Introduction to Circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter. 2020 · 2020
Earlier work this paper cites.
Winning the Lottery with Continuous Sparsification
Pedro Savarese, Hugo Silva, and Michael Maire. 2020 · 2020
Earlier work this paper cites.
Investigating Gender Bias in Language Models Using Causal Mediation Analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber. 2020 · 2020
Earlier work this paper cites.
BLiMP: The Benchmark of Linguistic Minimal Pairs for English
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman. 2020 · 2020
Earlier work this paper cites.
Low-Complexity Probing via Finding Subnetworks
Steven Cao, Victor Sanh, and Alexander Rush. 2021 · 2021
Earlier work this paper cites.
Are Neural Nets Modular? Inspecting Functional Modularity Through Differentiable Weight Masks
Róbert Csordás, Sjoerd van Steenkiste, and Jürgen Schmidhuber. 2021 · 2021
Earlier work this paper cites.
Measuring and Improving Consistency in Pretrained Language Models
Yanai Elazar, Nora Kassner, Shauli Ravfogel, Abhilasha Ravichander, Eduard Hovy, Hinrich Schütze, and Yoav Goldberg. 2021 · 2021
Earlier work this paper cites.
A Mathematical Framework for Transformer Circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, and 6 others. 2021 · 2021
Earlier work this paper cites.
Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov. 2021 · 2021
Earlier work this paper cites.
Causal Abstractions of Neural Networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. 2021 · 2021
Earlier work this paper cites.
Transformer Feed-Forward Layers Are Key-Value Memories
Mor Geva, Roei Schuster, Jonathan Berant, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Parameter-Efficient Transfer Learning with Diff Pruning
Demi Guo, Alexander Rush, and Yoon Kim. 2021 · 2021
Cited alongside, same era.
Can Subnetwork Structure be the Key to Out-of-Distribution Generalization?
Dinghuai Zhang, Kartik Ahuja, Yilun Xu, Yisen Wang, and Aaron Courville. 2021 · 2021
Cited alongside, same era.
Knowledge Neurons in Pretrained Transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei. 2022 · 2022
Cited alongside, same era.
Sparse Interventions in Language Models with Differentiable Masking
Nicola De Cao, Leon Schmid, Dieuwke Hupkes, and Ivan Titov. 2022 · 2022
Cited alongside, same era.
Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space
OpenAI. 2023 · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, Wenyin Fu, and 49 others. 2023 · 2023
Later among the works it cites.
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman. 2023 · 2023
Later among the works it cites.
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Fred Zhang and Neel Nanda. 2023 · 2023
Later among the works it cites.
Finding Transformer Circuits With Edge Pruning
Adithya Bhaskar, Alexander Wettig, Dan Friedman, and Danqi Chen. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mor Geva, Avi Caciularu, Kevin Wang, and Yoav Goldberg. 2022 · 2022
Cited alongside, same era.
Locating and Editing Factual Associations in GPT
Kevin Meng, David Bau, Alex J. Andonian, and Yonatan Belinkov. 2022 · 2022
Cited alongside, same era.
A Comprehensive Mechanistic Interpretability Explainer & Glossary
Neel Nanda. 2022 · 2022
Cited alongside, same era.
TransformerLens
Neel Nanda and Joseph Bloom. 2022 · 2022
Cited alongside, same era.
Interpretability in the Wild: A Circuit for Indirect Object Identification in GPT-2 Small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2022 · 2022
Cited alongside, same era.
Discovering Knowledge-Critical Subnetworks in Pretrained Language Models
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss, and Antoine Bosselut. 2023 · 2023
Cited alongside, same era.
Towards Automated Circuit Discovery for Mechanistic Interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso. 2023 · 2023
Cited alongside, same era.
Closest in time.
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. 2024 · 2024
Closest in time.
Linearity of Relation Decoding in Transformer Language Models
Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. 2024 · 2024
Closest in time.
Intrinsic Evaluation of Unlearning Using Parametric Knowledge Traces
Yihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel, and Mor Geva. 2024 · 2024
Closest in time.
Sparse Autoencoders Find Highly Interpretable Features in Language Models
Robert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, and Lee Sharkey. 2024 · 2024
Closest in time.
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
What does the Knowledge Neuron Thesis Have to do with Knowledge?
Jingcheng Niu, Andrew Liu, Zining Zhu, and Gerald Penn. 2024 · 2024
Closest in time.
Hypothesis Testing the Circuit Hypothesis in LLMs
Claudia Shi, Nicolas Beltran-Velez, Achille Nazaret, Carolina Zheng, Adrià Garriga-Alonso, Andrew Jesson, Maggie Makar, and David Blei. 2024 · 2024
Closest in time.
Attribution Patching Outperforms Automated Circuit Discovery
Aaquib Syed, Can Rager, and Arthur Conmy. 2024 · 2024
Closest in time.
Mechanistic Understanding and Mitigation of Language Model Non-Factual Hallucinations
Lei Yu, Meng Cao, Jackie CK Cheung, and Yue Dong. 2024a · 2024
Closest in time.
The Computational Complexity of Circuit Discovery for Inner Interpretability
Federico Adolfi, Martina G. Vilas, and Todd Wareham. 2025 · 2025
Closest in time.
Circuit Compositions: Exploring Modular Structures in Transformer-Based Language Models
Philipp Mondorf, Sondre Wold, and Barbara Plank. 2025 · 2025
Closest in time.
Llama See, Llama Do: A Mechanistic Perspective on Contextual Entrainment and Distraction in LLMs
Jingcheng Niu, Xingdi Yuan, Tong Wang, Hamidreza Saghir, and Amir H. Abdi. 2025 · 2025
Closest in time.
EAP-GP: Mitigating Saturation Effect in Gradient-based Automated Circuit Identification
Lin Zhang, Wenshuo Dong, Zhuoran Zhang, Shu Yang, Lijie Hu, Ninghao Liu, Pan Zhou, and Di Wang. 2025 · 2025
Closest in time.