Fetching the paper…
Reading the bibliography…
The rise of the term "mechanistic interpretability" has accompanied increasing interest in understanding neural models -- particularly language models.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 1901
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R Bowman. 2019 · 1911
Earlier work this paper cites.
From understanding computation to understanding neural circuitry
David Marr and Tomaso Poggio. 1976 · 1976
Earlier work this paper cites.
The problem of causal selection
Germund Hesslow. 1988 · 1988
Earlier work this paper cites.
Encyclopedia of Social Science Research Methods , chapter Causal Mechanisms
Daniel Little. 2004 · 2004
Earlier work this paper cites.
Measuring semantic similarity by latent relational analysis
Peter D. Turney. 2005 · 2005
Earlier work this paper cites.
The structure and function of explanations
Tania Lombrozo. 2006 · 2006
Earlier work this paper cites.
Language encodes geographical information
Max M Louwerse and Rolf A Zwaan. 2009 · 2009
Earlier work this paper cites.
Pretrained language model embryology: The birth of albert
Cheng-Han Chiang, Sung-Feng Huang, and Hung yi Lee. 2020 · 2010
Earlier work this paper cites.
SemEval-2012 task 2: Measuring degrees of relational similarity
David Jurgens, Saif Mohammad, Peter Turney, and Keith Holyoak. 2012 · 2012
Earlier work this paper cites.
Domain and function: A dual-space model of semantic relations and compositions
P. D. Turney. 2012 · 2012
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
Tomas Mikolov, Wen-tau Yih, and Geoffrey Zweig. 2013c · 2013
Earlier work this paper cites.
Selective effects of explanation on learning during early childhood
Cristine H Legare and Tania Lombrozo. 2014 · 2014
Earlier work this paper cites.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg. 2014 · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
How transferable are features in deep neural networks?
Jason Yosinski, Jeff Clune, Yoshua Bengio, and Hod Lipson. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Neural networks, types, and functional programming – colah’s blog
Chris Olah. 2015 · 2015
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. 2015 · 2015
Earlier work this paper cites.
Explanations and causal judgments are differentially sensitive to covariation and mechanism information
Nadya Vasilyeva and Tania Lombrozo. 2015 · 2015
Earlier work this paper cites.
Explaining predictions of non-linear classifiers in NLP
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek. 2016 · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Allyson Ettinger, Ahmed Elgohary, and Philip Resnik. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Investigating the influence of noise and distractors on the interpretation of neural networks
Pieter-Jan Kindermans, Kristof Schütt, Klaus-Robert Müller, and Sven Dähne. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
Assessing the ability of LSTMs to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Explanatory preferences shape learning and inference
Tania Lombrozo. 2016 · 2016
Earlier work this paper cites.
Proceedings of the 1st Workshop on Evaluating Vector-Space Representations for NLP . Association for Computational Linguistics, Berlin, Germany
RepEval, editor. 2016 · 2016
Earlier work this paper cites.
"Why should i trust you?" Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Does string-based neural mt learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight. 2016 · 2016
Earlier work this paper cites.
Attention-based LSTM for aspect-level sentiment classification
Yequan Wang, Minlie Huang, Xiaoyan Zhu, and Li Zhao. 2016 · 2016
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg. 2017 · 2017
Earlier work this paper cites.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Visualizing and understanding neural machine translation
Yanzhuo Ding, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
The promise and peril of human evaluation for model interpretability
Bernease Herman. 2017 · 2017
Earlier work this paper cites.
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema. 2017 · 2017
Earlier work this paper cites.
Representation of linguistic form and function in recurrent neural networks
Ákos Kádár, Grzegorz Chrupała, and Afra Alishahi. 2017 · 2017
Earlier work this paper cites.
The strange geometry of skip-gram with negative sampling
David Mimno and Laure Thompson. 2017 · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
Alec Radford, Rafal Jozefowicz, and Ilya Sutskever. 2017 · 2017
Earlier work this paper cites.
SVCCA: Singular Vector Canonical Correlation Analysis for Deep Learning Dynamics and Interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
On internal language representations in deep learning: An analysis of machine translation and speech recognition
Yonatan Belinkov. 2018 · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability
Zachary C. Lipton. 2018 · 2018
Earlier work this paper cites.
SparseMAP: Differentiable sparse structured inference
Vlad Niculae, Andre Martins, Mathieu Blondel, and Claire Cardie. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
SPINE: SParse Interpretable Neural Embeddings
Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, and Eduard Hovy. 2018 · 2018
Earlier work this paper cites.
On the practical computational power of finite precision RNNs for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav. 2018 · 2018
Earlier work this paper cites.
Language modeling teaches you more than translation does: Lessons learned through auxiliary syntactic task analysis
Kelly Zhang and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel, and Max Welling. 2017 · 2018
Earlier work this paper cites.
Analyzing and interpreting neural networks for nlp: A report on the first blackboxnlp workshop
Afra Alishahi, Grzegorz Chrupała, and Tal Linzen. 2019 · 2019
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass. 2019 · 2019
Earlier work this paper cites.
Abstracting causal models
Sander Beckers and Joseph Y. Halpern. 2019 · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
What Is One Grain of Sand in the Desert? Analyzing Individual Neurons in Deep NLP Models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
Kawin Ethayarajh. 2019 · 2019
Earlier work this paper cites.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Earlier work this paper cites.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon. 2019 · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning. 2019 · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Earlier work this paper cites.
The (Un)reliability of Saliency Methods , pages 267–280
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2019 · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky. 2019 · 2019
Earlier work this paper cites.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni. 2019 · 2019
Earlier work this paper cites.
Discovery of natural language concepts in individual units of cnns
Seil Na, Yo Joong Choe, Dong-Hyun Lee, and Gunhee Kim. 2019 · 2019
Earlier work this paper cites.
Word2sense: Sparse interpretable word embeddings
Abhishek Panigrahi, Harsha Vardhan Simhadri, and Chiranjib Bhattacharyya. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Earlier work this paper cites.
Visualizing and measuring the geometry of bert
Emily Reif, Ann Yuan, Martin Wattenberg, Fernanda B Viegas, Andy Coenen, Adam Pearce, and Been Kim. 2019 · 2019
Cited alongside, same era.
Human-centered artificial intelligence and machine learning
Mark O Riedl. 2019 · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Cited alongside, same era.
Understanding learning dynamics of language models with SVCCA
Naomi Saphra and Adam Lopez. 2019b · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Cited alongside, same era.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Cited alongside, same era.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt. 2023 · 2023
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell. 2023 · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda. 2023 · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid. 2023 · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in gpt-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Cited alongside, same era.
No training required: Exploring random encoders for sentence classification
John Wieting and Douwe Kiela. 2019 · 2019
Cited alongside, same era.
Quantifying attention flow in transformers
Samira Abnar and Willem Zuidema. 2020 · 2020
Cited alongside, same era.
Approximate causal abstractions
Sander Beckers, Frederick Eberhardt, and Joseph Y. Halpern. 2020 · 2020
Cited alongside, same era.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer. 2020 · 2020
Cited alongside, same era.
Later among the works it cites.
How elite schools like stanford became fixated on the AI apocalypse
Washington Post. 2023 · 2023
Later among the works it cites.
A reply to makelov et al.(2023)’s" interpretability illusion" arguments
Zhengxuan Wu, Atticus Geiger, Jing Huang, Aryaman Arora, Thomas Icard, Christopher Potts, and Noah D Goodman. 2024 · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas. 2023 · 2023
Later among the works it cites.
Acl policies for review and citation
ACL Executive Committee. 2024 · 2024
Closest in time.
“Excited to see that it’s that time of the year when we reinvent probing again.”
Jacob Andreas. 2023 · 2024
Closest in time.
“I still don’t totally understand the difference between “mechanistic” and “non-mechanistic” interpretability but it seems to be mainly a distinction of the authors’ social network?”
Jacob Andreas. 2024 · 2024
Closest in time.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda. 2024 · 2024
Closest in time.
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts. 2024 · 2024
Closest in time.
“Distributional semantics? Reminds me of the "florida" example in the @omerlevy_ and @yoavgo paper from 2014. Granted, contemporary LLMs probably do it much better, but the ability is likely not new”
Yoav Artzi. 2023 · 2024
Closest in time.
Mechanistic Interpretability Workshop at the 41st International Conference on Machine Learning (ICML)
Fazl Barez, Mor Geva, Lawrence Chan, Atticus Geiger, Kayo Yin, Neel Nanda, and Max Tegmark. 2024 · 2024
Closest in time.
New England Mechanistic Interpretability (NEMI) Workshop Series
David Bau, Max Tegmark, Koyena Pal, Kenneth Li, Eric Michaud, and Jannik Brinkmann. 2024 · 2024
Closest in time.
“Excited to see important work from @andyzou_jiaming, @DanHendrycks…, on interpreting & controlling language models at representation level, to improve fairness & safety of LMs. Unfortunately it fails to engage with a large body of work on these topics from the past 5 years.”
Yonatan Belinkov. 2023a · 2024
Closest in time.
“We are interested! #blackboxNLP has been the largest #nlproc workshop for several years now. And we have an interpretability track in all main #nlproc confs! Please submit your work to be reviewed in such venues…Even if you disagree with other approaches to interpretability, I think engagement through common conferences would help the community grow.”
Yonatan Belinkov. 2023b · 2024
Closest in time.
“The @AiEleuther interpretability team is releasing a set of top-k sparse autoencoders for every layer of Llama 3 8B: https://huggingface.co/EleutherAI/sae-llama-3-8b-32x
Nora Belrose. 2024 · 2024
Closest in time.
“is mechanic [sic] interpretability a sexier way of saying interpretability?”
Nathan Beniach. 2024 · 2024
Closest in time.
Mechanistic interpretability for AI safety - a review
Leonard Bereska and Efstratios Gavves. 2024 · 2024
Closest in time.
Identifying functionally important features with end-to-end sparse dictionary learning
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey. 2024 · 2024
Closest in time.
Sudden drops in the loss: Syntax acquisition, phase transitions, and simplicity bias in MLMs
Angelica Chen, Ravid Shwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra. 2024 · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James R. Glass, and Pengcheng He. 2024 · 2024
Closest in time.
Recurrent neural networks learn to store and generate sequences using non-linear representations
Róbert Csordás, Christopher Potts, Christopher D. Manning, and Atticus Geiger. 2024 · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Aidan Ewart, Logan Riggs Smith, Robert Huben, and Lee Sharkey. 2024 · 2024
Closest in time.
“watching the mechanistic interpretability community rediscover manifolds with non-trivial topologies in real time is simultaneously amazing/exciting and concerning — highly recommend anyone in (mech-int) ML read @naturecomputes stunning work below to skip some steps!”
Tim Davidson. 2024 · 2024
Closest in time.
Transcoders find interpretable llm feature circuits
Jacob Dunefsky, Philippe Chlenski, and Neel Nanda. 2024 · 2024
Closest in time.
A primer on the inner workings of transformer-based language models
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà. 2024 · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
Javier Ferrando and Elena Voita. 2024 · 2024
Closest in time.
PhD fellowships
Future of Life Institute · 2024
Closest in time.
Scaling and evaluating sparse autoencoders
Leo Gao, Tom Dupré la Tour, Henk Tillman, Gabriel Goh, Rajan Troll, Alec Radford, Ilya Sutskever, Jan Leike, and Jeffrey Wu. 2024 · 2024
Closest in time.
The missing curve detectors of inceptionv1: Applying sparse autoencoders to inceptionv1 early vision
Liv Gorton. 2024 · 2024
Closest in time.
What is mechanistic interpretability? You’re not the only one asking!
Michael Hanna. 2024 · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov. 2024 · 2024
Closest in time.
Should we publish mechanistic interpretability research?
Marius Hobbhahn and Lawrence Chan. 2023 · 2024
Closest in time.
Measuring progress in dictionary learning for language model interpretability with board game models
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Riggs Smith, Claudio Mayrink Verdun, David Bau, and Samuel Marks. 2024 · 2024
Closest in time.
Probing the category of verbal aspect in transformer language models
Anisia Katinskaia and Roman Yangarber. 2024 · 2024
Closest in time.
Interpreting attention layer outputs with sparse autoencoders
Connor Kissane, Robert Krzyzanowski, Joseph Isaac Bloom, Arthur Conmy, and Neel Nanda. 2024 · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Why and when interpretability work is dangerous
Nicholas / Heather Kross. 2023 · 2024
Closest in time.
Interpretability and analysis of models for NLP @ ACL 2020
Carolin Lawrence. 2020 · 2024
Closest in time.
https://github.com/ruizheliuoa/awesome-interpretability-in-large-language-models
Ruizhe Li. 2024 · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
Tom Lieberum, Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Nicolas Sonnerat, Vikrant Varma, János Kramár, Anca Dragan, Rohin Shah, and Neel Nanda. 2024 · 2024
Closest in time.
Interpretability needs a new paradigm
Andreas Madsen, Himabindu Lakkaraju, Siva Reddy, and Sarath Chandar. 2024 · 2024
Closest in time.
Sparse autoencoders match supervised features for model steering on the IOI task
Aleksandar Makelov. 2024 · 2024
Closest in time.
Is this the subspace you are looking for? an interpretability illusion for subspace activation patching
Aleksandar Makelov, Georg Lange, Atticus Geiger, and Neel Nanda. 2024 · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller. 2024 · 2024
Closest in time.
Samuel Marks and Max Tegmark. 2024 · 2024
Closest in time.
The illusion of state in state-space models
William Merrill, Jackson Petty, and Ashish Sabharwal. 2024 · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
Jack Merullo, Carsten Eickhoff, and Ellie Pavlick. 2024 · 2024
Closest in time.
Missed causes and ambiguous effects: Counterfactuals pose challenges for interpreting neural networks
Aaron Mueller. 2024 · 2024
Closest in time.
Aaron Mueller, Jannik Brinkmann, Millicent Li, Samuel Marks, Koyena Pal, Nikhil Prakash, Can Rager, Aruna Sankaranarayanan, Arnab Sen Sharma, Jiuding Sun, et al. 2024 · 2024
Closest in time.
A comprehensive mechanistic interpretability explainer & glossary
Neel Nanda. 2022 · 2024
Closest in time.
Attribution patching: Activation patching at industrial scale
Neel Nanda. 2023a · 2024
Closest in time.
“We’re currently refining my grokking work to better fit academic interests and language, and are submitting it to a peer-reviewed AI venue. This was a bunch of effort, but I’m really hoping this can get more awareness and interest in mech interp from academic communities!”
Neel Nanda. 2023b · 2024
Closest in time.
“A bunch of people who don’t seem to identify as EA at all use the term to describe their work nowadays. I’m happy the field is growing outside of the EA bubble!”
Neel Nanda. 2024a · 2024
Closest in time.
An extremely opinionated annotated list of my favourite mechanistic interpretability papers v2
Neel Nanda. 2024b · 2024
Closest in time.
“I introduced the term to distinguish the work I was doing on circuits from a lot of other work that was going on circa 2018, notably saliency maps…I was motivated by many of my colleagues at Google Brain being deeply skeptical of things like saliency maps. When I started the OpenAI interpretability team, I used it to distinguish our goal: understand how the weights of a neural network map to algorithms…Since then, I think it’s become an umbrella term for a variety of other things.”
Chris Olah. 2024a · 2024
Closest in time.
“The motivating moment for me was that I went on a walk with a senior colleague shortly before I left Google Brain, and they matter of factly told me that “all interpretability is bullshit” because they’d been so turned off by saliency maps…I wanted a way to get people who were skeptical in this way to realize that I was talking about something pretty different than saliency maps and be willing to give it a second look.”
Chris Olah. 2024b · 2024
Closest in time.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch. 2024 · 2024
Closest in time.
Improving sparse decomposition of language model activations with gated sparse autoencoders
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, Janos Kramar, Rohin Shah, and Neel Nanda. 2024a · 2024
Closest in time.
“Many claims, modulo the autoencoder bit (which I strongly suspect is unnecessary for many of the findings), were published in the past. The post does not acknowledge the existence of much of the previous work. This is not atypical.”
Shauli Ravfogel. 2023 · 2024
Closest in time.
“I recently asked pre-PhD researchers what area they were most excited about, and overwhelmingly the answer was “mechanistic interpretability”. Not sure how that happened, but I am interested how it came about.”
Sasha Rush. 2024 · 2024
Closest in time.
Against monodomainism
Naomi Saphra. 2021 · 2024
Closest in time.
Quantifying the plausibility of context reliance in neural machine translation
Gabriele Sarti, Grzegorz Chrupała, Malvina Nissim, and Arianna Bisazza. 2024 · 2024
Closest in time.
“This is what happens when a significant contingent of people studying LLMs don’t meaningfully engage with *ACL literature. This is why we need policies that don’t drive out ECRs by disadvantaging their ability to preprint and publicize against ICLR/CVPR/NeurIPS-primary ECRs.”
Michael Saxon. 2023 · 2024
Closest in time.
Against almost every theory of impact of interpretability
Charbel-Raphaël Segerie. 2023 · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, et al. 2024 · 2024
Closest in time.
Llm circuit analyses are consistent across training and scale
Curt Tigges, Michael Hanna, Qinan Yu, and Stella Biderman. 2024 · 2024
Closest in time.
“anthropic is making a mistake by not better grounding their interpretability research to existing topics like compressed sensing or frames. engaging smart academics who’ve spent years working on these areas in a different context is far more valuable than lesswrong posters”
@typedfemale. 2023 · 2024
Closest in time.
Answer, assemble, ace: Understanding how transformers answer multiple choice questions
Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hannaneh Hajishirzi, and Ashish Sabharwal. 2024 · 2024
Closest in time.
Explanation in the era of large language models
Zining Zhu, Hanjie Chen, Xi Ye, Qing Lyu, Chenhao Tan, Ana Marasovic, and Sarah Wiegreffe. 2024 · 2024
Closest in time.