Fetching the paper…
Reading the bibliography…
Interpretability provides a toolset for understanding how and why neural networks behave in certain ways.
Explaining classifiers with causal concept effect (CaCE)
Yash Goyal, Amir Feder, Uri Shalit, and Been Kim · 1907
Earlier work this paper cites.
RoBERTa: A robustly optimized BERT pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
Counterfactuals
David K. Lewis · 1973
Earlier work this paper cites.
Distributed memory and the representation of general and specific information
James L. McClelland and David E. Rumelhart · 1985
Earlier work this paper cites.
Causation
David Lewis · 1986
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Michael C. Mozer and Paul Smolensky · 1988
Earlier work this paper cites.
Representation and structure in connectionist models
Jeffrey L Elman · 1989
Earlier work this paper cites.
Finding structure in time
Jeffrey L. Elman · 1990
Earlier work this paper cites.
A neural expert system with automated extraction of fuzzy if-then rules and its application to medical diagnosis
Yoichi Hayashi · 1990
Earlier work this paper cites.
A simple procedure for pruning back-propagation trained neural networks
E. D. Karnin · 1990
Earlier work this paper cites.
Distributed representations, simple recurrent networks, and grammatical structure
Jeffrey L. Elman · 1991
Earlier work this paper cites.
Identifiability and exchangeability for direct and indirect effects
James M. Robins and Sander Greenland · 1992
Earlier work this paper cites.
Using sampling and queries to extract rules from trained neural networks
Mark W. Craven and Jude W. Shavlik · 1994
Earlier work this paper cites.
Extracting tree-structured representations of trained networks
Mark Craven and Jude Shavlik · 1995
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Extracting decision trees from trained neural networks
R. Krishnan, G. Sivakumar, and P. Bhattacharya · 1999
Earlier work this paper cites.
Causation as influence
David Lewis · 2000
Earlier work this paper cites.
Causality: Models, Reasoning, and Inference
Judea Pearl · 2000
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Direct and indirect effects
Judea Pearl · 2001
Earlier work this paper cites.
Extracting decision trees from trained neural networks
Olcay Boz · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
Opening the black box - Data driven visualization of neural networks
F.-Y. Tzeng and K.-L. Ma · 2005
Earlier work this paper cites.
Visualizing higher-layer features of a deep network
Dumitru Erhan, Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2009
Earlier work this paper cites.
Recurrent neural network based language model
Tomas Mikolov, Martin Karafiát, Lukáš Burget, Jan Černocký, and Sanjeev Khudanpur · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Distributional vectors encode referential attributes
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó · 2015
Earlier work this paper cites.
What’s in an embedding? Analyzing word embeddings through multilingual evaluation
Arne Köhn · 2015
Earlier work this paper cites.
A latent variable model approach to PMI-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei · 2016
Earlier work this paper cites.
Does string-based neural MT learn source syntax?
Xing Shi, Inkit Padhi, and Kevin Knight · 2016
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Yonatan Belinkov, Nadir Durrani, Fahim Dalvi, Hassan Sajjad, and James Glass · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Understanding neural networks through representation erasure, 2017
Jiwei Li, Will Monroe, and Dan Jurafsky · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M. Lundberg and Su-In Lee · 2017
Earlier work this paper cites.
Towards faithful model explanation in NLP: A survey
Qing Lyu, Marianna Apidianaki, and Chris Callison-Burch · 2017
Earlier work this paper cites.
LSTMVis: A tool for visual analysis of hidden state dynamics in recurrent neural networks
Hendrik Strobelt, Sebastian Gehrmann, Hanspeter Pfister, and Alexander M Rush · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Earlier work this paper cites.
Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information
Mario Giulianelli, Jack Harding, Florian Mohnert, Dieuwke Hupkes, and Willem Zuidema · 2018
Earlier work this paper cites.
Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
Dieuwke Hupkes, Sara Veldhoen, and Willem Zuidema · 2018
Earlier work this paper cites.
On the importance of single directions for generalization
Ari S. Morcos, David G.T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Earlier work this paper cites.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Earlier work this paper cites.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov · 2019
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James R. Glass · 2019
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Yonatan Belinkov and James Glass · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Cited alongside, same era.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Cited alongside, same era.
The emergence of number and syntax units in LSTM language models
Yair Lakretz, German Kruszewski, Theo Desbordes, Dieuwke Hupkes, Stanislas Dehaene, and Marco Baroni · 2019
Cited alongside, same era.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Uncovering intermediate variables in transformers using circuit probing
Michael A. Lepori, Thomas Serre, and Ellie Pavlick · 2023
Later among the works it cites.
Emergent world representations: Exploring a sequence model trained on a synthetic task
Kenneth Li, Aspen K Hopkins, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark · 2023
Later among the works it cites.
The hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Later among the works it cites.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Understanding the role of individual units in a deep neural network
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Agata Lapedriza, Bolei Zhou, and Antonio Torralba · 2020
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Analyzing redundancy in pretrained transformer models
Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov · 2020
Cited alongside, same era.
A survey of the state of explainable AI for natural language processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen · 2020
Cited alongside, same era.
How do decisions emerge across layers in neural models? Interpretation with differentiable masking
Nicola De Cao, Michael Sejr Schlichtkrull, Wilker Aziz, and Ivan Titov · 2020
Cited alongside, same era.
Concept bottleneck models
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Cited alongside, same era.
Later among the works it cites.
The quantization model of neural scaling
Eric J Michaud, Ziming Liu, Uzay Girit, and Max Tegmark · 2023
Later among the works it cites.
Emergent linear representations in world models of self-supervised sequence models
Neel Nanda, Andrew Lee, and Martin Wattenberg · 2023
Later among the works it cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Later among the works it cites.
Toward transparent AI: A survey on interpreting the inner structures of deep neural networks, 2023
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2023
Later among the works it cites.
On the effect of dropping layers of pre-trained transformer models
Hassan Sajjad, Fahim Dalvi, Nadir Durrani, and Preslav Nakov · 2023
Later among the works it cites.
Taking features out of superposition with sparse autoencoders, 2023
Lee Sharkey, Dan Braun, and Beren Millidge · 2023
Later among the works it cites.
Corrupting neuron explanations of deep visual features
Divyansh Srivastava, Tuomas Oikarinen, and Tsui-Wei Weng · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Later among the works it cites.
Codebook features: Sparse and discrete interpretability for neural networks
Alex Tamkin, Mohammad Taufeeque, and Noah D. Goodman · 2023
Later among the works it cites.
Learning silhouettes with group sparse autoencoders
Emmanouil Theodosis and Demba Ba · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
Curt Tigges, Oskar John Hollinsworth, Atticus Geiger, and Neel Nanda · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization
Alexander Matt Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen · 2023
Later among the works it cites.
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
Aryaman Arora, Dan Jurafsky, and Christopher Potts · 2024
Closest in time.
Mechanistic interpretability for AI safety – A review
Leonard Bereska and Efstratios Gavves · 2024
Closest in time.
Identifying functionally important features with end-to-end sparse dictionary learning
Dan Braun, Jordan Taylor, Nicholas Goldowsky-Dill, and Lee Sharkey · 2024
Closest in time.
A mechanistic analysis of a transformer trained on a symbolic multi-step reasoning task
Jannik Brinkmann, Abhay Sheshadri, Victor Levoso, Paul Swoboda, and Christian Bartelt · 2024
Closest in time.
Evaluating the Ripple Effects of Knowledge Editing in Language Models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2024
Closest in time.
Sparse autoencoders find highly interpretable features in language models
Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart, Robert Huben, and Lee Sharkey · 2024
Closest in time.
Transcoders find interpretable LLM feature circuits
Jacob Dunefsky, Philippe Chlenski, and Neel Nanda · 2024
Closest in time.
Not all language model features are linear, 2024
Joshua Engels, Isaac Liao, Eric J. Michaud, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
A primer on the inner workings of transformer-based language models, 2024
Javier Ferrando, Gabriele Sarti, Arianna Bisazza, and Marta R. Costa-jussà · 2024
Closest in time.
Nnsight and ndif: Democratizing access to foundation model internals, 2024
Jaden Fiotto-Kaufman, Alexander R Loftus, Eric Todd, Jannik Brinkmann, Caden Juang, Koyena Pal, Can Rager, Aaron Mueller, Samuel Marks, Arnab Sen Sharma, Francesca Lucchetti, Michael Ripa, Adam Belfki, Nikhil Prakash, Sumeet Multani, Carla Brodley, Arjun Guha, Jonathan Bell, Byron Wallace, and David Bau · 2024
Closest in time.
Unified concept editing in diffusion models
Rohit Gandikota, Hadas Orgad, Yonatan Belinkov, Joanna Materzyńska, and David Bau · 2024
Closest in time.
Jorge García-Carrasco, Alejandro Maté, and Juan Trujillo · 2024
Closest in time.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah Goodman · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms
Michael Hanna, Sandro Pezzelle, and Yonatan Belinkov · 2024
Closest in time.
RAVEL: Evaluating interpretability methods on disentangling language model representations
Jing Huang, Zhengxuan Wu, Christopher Potts, Mor Geva, and Atticus Geiger · 2024
Closest in time.
Polysemantic attention head in a 4-layer transformer, 2023
Jett Janiak, cmathw, and Stefan Heimersheim · 2024
Closest in time.
Measuring progress in dictionary learning for language model interpretability with board game models
Adam Karvonen, Benjamin Wright, Can Rager, Rico Angell, Jannik Brinkmann, Logan Riggs Smith, Claudio Mayrink Verdun, David Bau, and Samuel Marks · 2024
Closest in time.
AtP*: An efficient and scalable method for localizing llm behaviour to components
János Kramár, Tom Lieberum, Rohin Shah, and Neel Nanda · 2024
Closest in time.
The remarkable robustness of LLMs: Stages of inference?, 2024
Vedang Lad, Wes Gurnee, and Max Tegmark · 2024
Closest in time.
Circuit breaking: Removing model behaviors with targeted ablation
Maximilian Li, Xander Davies, and Max Nadeau · 2024
Closest in time.
Towards principled evaluations of sparse autoencoders for interpretability and control
Aleksandar Makelov, George Lange, and Neel Nanda · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
Samuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov, David Bau, and Aaron Mueller · 2024
Closest in time.
Transformer circuit faithfulness metrics are not robust
Joseph Miller, Bilal Chughtai, and William Saunders · 2024
Closest in time.
Missed causes and ambiguous effects: Counterfactuals pose challenges for interpreting neural networks
Aaron Mueller · 2024
Closest in time.
Interpreting context look-ups in transformers: Investigating attention-MLP interactions
Clement Neo, Shay B Cohen, and Fazl Barez · 2024
Closest in time.
Steering Llama 2 via contrastive activation addition
Nina Panickssery, Nick Gabrieli, Julian Schulz, Meg Tong, Evan Hubinger, and Alexander Matt Turner · 2024
Closest in time.
Does transformer interpretability transfer to RNNs?
Gonçalo Paulo, Thomas Marshall, and Nora Belrose · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Phenomenal yet puzzling: Testing inductive reasoning capabilities of language models with hypothesis refinement
Linlu Qiu, Liwei Jiang, Ximing Lu, Melanie Sclar, Valentina Pyatkin, Chandra Bhagavatula, Bailin Wang, Yoon Kim, Yejin Choi, Nouha Dziri, and Xiang Ren · 2024
Closest in time.
Improving dictionary learning with gated sparse autoencoders, 2024
Senthooran Rajamanoharan, Arthur Conmy, Lewis Smith, Tom Lieberum, Vikrant Varma, János Kramár, Rohin Shah, and Neel Nanda · 2024
Closest in time.
Causality for trustworthy artificial intelligence: Status, challenges and perspectives
Atul Rawal, Adrienne Raglin, Danda B. Rawat, Brian M. Sadler, and James McCoy · 2024
Closest in time.
First tragedy, then parse: History repeats itself in the new era of large language models
Naomi Saphra, Eve Fleisig, Kyunghyun Cho, and Adam Lopez · 2024
Closest in time.
A multimodal automated interpretability agent
Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang, Achyuta Rajaram, Evan Hernandez, Jacob Andreas, and Antonio Torralba · 2024
Closest in time.
Locating and editing factual associations in Mamba
Arnab Sen Sharma, David Atkinson, and David Bau · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
Function vectors in large language models
Eric Todd, Millicent Li, Arnab Sen Sharma, Aaron Mueller, Byron C Wallace, and David Bau · 2024
Closest in time.
pyvene: A library for understanding and improving PyTorch models via interventions
Zhengxuan Wu, Atticus Geiger, Aryaman Arora, Jing Huang, Zheng Wang, Noah Goodman, Christopher Manning, and Christopher Potts · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
Fred Zhang and Neel Nanda · 2024
Closest in time.