Fetching the paper…
Reading the bibliography…
The rapid progress of research aimed at interpreting the inner workings of advanced language models has highlighted a need for contextualizing the insights gained from years of work in this area.
Jumprelu: A retrofit defense strategy for adversarial attacks
N. B. Erichson, Z. Yao, and M. W. Mahoney · 1904
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?
P. M. Htut, J. Phang, S. Bordia, and S. R. Bowman · 1911
Earlier work this paper cites.
A value for n-person games
L. S. Shapley · 1953
Earlier work this paper cites.
Neural and conceptual interpretation of PDP models , pp. 390–431
P. Smolensky · 1986
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
R. Tibshirani · 1996
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: A strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Scaling laws for neural language models, 2020
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Direct and indirect effects
J. Pearl · 2001
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin · 2003
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
N. Kokhlikyan, V. Miglani, M. Martin, E. Wang, B. Alsallakh, J. Reynolds, A. Melnikov, N. Kliushkina, C. Araya, S. Yan, and O. Reblitz-Richardson · 2009
Earlier work this paper cites.
Causality
J. Pearl · 2009
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Earlier work this paper cites.
A. Makhzani and B. Frey · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
K. Simonyan, A. Vedaldi, and A. Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
S. Bach, A. Binder, G. Montavon, F. Klauschen, K.-R. Müller, and W. Samek · 2015
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
M. Denil, A. Demiraj, and N. de Freitas · 2015
Earlier work this paper cites.
Distributional vectors encode referential attributes
A. Gupta, G. Boleda, M. Baroni, and S. Padó · 2015
Earlier work this paper cites.
What’s in an embedding? analyzing word embeddings through multilingual evaluation
A. Köhn · 2015
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
G. Alain and Y. Bengio · 2016
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
J. Li, X. Chen, E. Hovy, and D. Jurafsky · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
A. Nguyen, A. Dosovitskiy, J. Yosinski, T. Brox, and J. Clune · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Residual networks behave like ensembles of relatively shallow networks
A. Veit, M. Wilber, and S. Belongie · 2016
Earlier work this paper cites.
The shattered gradients problem: If resnets are the answer, then what is the question?
D. Balduzzi, M. Frean, L. Leary, J. P. Lewis, K. W.-D. Ma, and B. McWilliams · 2017
Earlier work this paper cites.
What do neural machine translation models learn about morphology?
Y. Belinkov, N. Durrani, F. Dalvi, H. Sajjad, and J. Glass · 2017
Earlier work this paper cites.
Language modeling with gated convolutional networks
Y. N. Dauphin, A. Fan, M. Auli, and D. Grangier · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
F. Doshi-Velez and B. Kim · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
P. W. Koh and P. Liang · 2017
Earlier work this paper cites.
Understanding neural networks through representation erasure, 2017
J. Li, W. Monroe, and D. Jurafsky · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
S. M. Lundberg and S.-I. Lee · 2017
Earlier work this paper cites.
Towards Faithful Model Explanation in NLP: A Survey
Q. Lyu, M. Apidianaki, and C. Callison-Burch · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
A. Radford, R. Jozefowicz, and I. Sutskever · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
A. Shrikumar, P. Greenside, and A. Kundaje · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise, 2017
D. Smilkov, N. Thorat, B. Kim, F. Viégas, and M. Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
M. Sundararajan, A. Taly, and Q. Yan · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim · 2018
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
S. Arora, Y. Li, Y. Liang, T. Ma, and A. Risteski · 2018
Earlier work this paper cites.
Visualisation and ‘diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure
D. Hupkes, S. Veldhoen, and W. Zuidema · 2018
Earlier work this paper cites.
Residual connections encourage iterative inference, 2018
S. Jastrzębski, D. Arpit, N. Ballas, V. Verma, T. Che, and Y. Bengio · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
B. Kim, M. Wattenberg, J. Gilmer, C. Cai, J. Wexler, F. Viegas, and R. sayres · 2018
Earlier work this paper cites.
Influence-directed explanations for deep convolutional networks
K. Leino, S. Sen, A. Datta, M. Fredrikson, and L. Li · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Z. C. Lipton · 2018
Earlier work this paper cites.
Dissecting contextual word embeddings: Architecture and representation
M. E. Peters, M. Neumann, L. Zettlemoyer, and W.-t. Yih · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
A. Raganato and J. Tiedemann · 2018
Earlier work this paper cites.
Computationally efficient measures of internal neuron importance
A. Shrikumar, J. Su, and A. Kundaje · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
A. Bau, Y. Belinkov, H. Sajjad, N. Durrani, F. Dalvi, and J. Glass · 2019
Earlier work this paper cites.
Analysis methods in neural language processing: A survey
Y. Belinkov and J. Glass · 2019
Earlier work this paper cites.
Understanding the origins of bias in word embeddings
M.-E. Brunet, C. Alkalay-Houlihan, A. Anderson, and R. Zemel · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
K. Clark, U. Khandelwal, O. Levy, and C. D. Manning · 2019
Earlier work this paper cites.
Adaptively sparse transformers
G. M. Correia, V. Niculae, and A. F. T. Martins · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
F. Dalvi, N. Durrani, H. Sajjad, Y. Belinkov, A. Bau, and J. Glass · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
How important is a neuron
K. Dhamdhere, M. Sundararajan, and Q. Yan · 2019
Earlier work this paper cites.
How contextual are contextualized word representations? Comparing the geometry of BERT, ELMo, and GPT-2 embeddings
K. Ethayarajh · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
J. Hewitt and P. Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
J. Hewitt and C. D. Manning · 2019
Earlier work this paper cites.
Attention is not Explanation
S. Jain and B. C. Wallace · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
O. Kovaleva, A. Romanov, A. Rogers, and A. Rumshisky · 2019
Earlier work this paper cites.
Open sesame: Getting inside BERT’s linguistic knowledge
Y. Lin, Y. C. Tan, and R. Frank · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
N. F. Liu, M. Gardner, Y. Belinkov, M. E. Peters, and N. A. Smith · 2019
Earlier work this paper cites.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
T. McCoy, E. Pavlick, and T. Linzen · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
P. Michel, O. Levy, and G. Neubig · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever · 2019
Earlier work this paper cites.
A multiscale visualization of attention in the transformer model
J. Vig · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
J. Vig and Y. Belinkov · 2019
Earlier work this paper cites.
The bottom-up evolution of representations in the transformer: A study with machine translation and language modeling objectives
E. Voita, R. Sennrich, and I. Titov · 2019
Earlier work this paper cites.
Root mean square layer normalization
B. Zhang and R. Sennrich · 2019
Earlier work this paper cites.
Quantifying attention flow in transformers
S. Abnar and W. Zuidema · 2020
Earlier work this paper cites.
Debugging tests for model explanations
J. Adebayo, M. Muelly, I. Liccardi, and B. Kim · 2020
Earlier work this paper cites.
A diagnostic study of explainability techniques for text classification
P. Atanasova, J. G. Simonsen, C. Lioma, and I. Augenstein · 2020
Earlier work this paper cites.
The elephant in the interpretability room: Why use attention as explanation when we have saliency methods?
J. Bastings and K. Filippova · 2020
Earlier work this paper cites.
Understanding the role of individual units in a deep neural network
D. Bau, J.-Y. Zhu, H. Strobelt, A. Lapedriza, B. Zhou, and A. Torralba · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
On identifiability in transformers
G. Brunner, Y. Liu, D. Pascual, O. Richter, M. Ciaramita, and R. Wattenhofer · 2020
Earlier work this paper cites.
Thread: Circuits
N. Cammarata, S. Carter, G. Goh, C. Olah, M. Petrov, L. Schubert, C. Voss, B. Egan, and S. K. Lim · 2020
Earlier work this paper cites.
How do decisions emerge across layers in neural models? interpretation with differentiable masking
N. De Cao, M. S. Schlichtkrull, W. Aziz, and I. Titov · 2020
Earlier work this paper cites.
ERASER: A benchmark to evaluate rationalized NLP models
J. DeYoung, S. Jain, N. F. Rajani, E. Lehman, C. Xiong, R. Socher, and B. C. Wallace · 2020
Earlier work this paper cites.
Unsupervised quality estimation for neural machine translation
M. Fomicheva, S. Sun, L. Yankovskaya, F. Blain, F. Guzmán, M. Fishel, N. Aletras, V. Chaudhary, and L. Specia · 2020
Earlier work this paper cites.
Neural natural language inference models partially embed theories of lexical entailment and negation
A. Geiger, K. Richardson, and C. Potts · 2020
Earlier work this paper cites.
Explaining black box predictions and unveiling data artifacts through influence functions
X. Han, B. C. Wallace, and Y. Tsvetkov · 2020
Earlier work this paper cites.
exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformer Models
B. Hoover, H. Strobelt, and S. Gehrmann · 2020
Earlier work this paper cites.
Attention is not only a weight: Analyzing transformers with vector norms
G. Kobayashi, T. Kuribayashi, S. Yokoi, and K. Inui · 2020
Earlier work this paper cites.
Questioning the ai: Informing design practices for explainable ai user experiences
Q. V. Liao, D. Gruen, and S. Miller · 2020
Earlier work this paper cites.
A brief prehistory of double descent
M. Loog, T. Viering, A. Mey, J. H. Krijthe, and D. M. J. Tax · 2020
Earlier work this paper cites.
Interpreting GPT: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
An overview of early vision in inceptionv1
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
C. Olah, N. Cammarata, L. Schubert, G. Goh, M. Petrov, and S. Carter · 2020
Earlier work this paper cites.
Information-theoretic probing for linguistic structure
T. Pimentel, J. Valvoda, R. H. Maudslay, R. Zmigrod, A. Williams, and R. Cotterell · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
S. Ravfogel, Y. Elazar, H. Gonen, M. Twiton, and Y. Goldberg · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
When explanations lie: Why many modified BP attributions fail
L. Sixt, M. Granz, and T. Landgraf · 2020
Earlier work this paper cites.
Finding experts in transformer models, 2020
X. Suau, L. Zappella, and N. Apostoloff · 2020
Earlier work this paper cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models
I. Tenney, J. Wexler, J. Bastings, T. Bolukbasi, A. Coenen, S. Gehrmann, E. Jiang, M. Pushkarna, C. Radebaugh, E. Reif, and A. Yuan · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
J. Vig, S. Gehrmann, Y. Belinkov, S. Qian, D. Nevo, Y. Singer, and S. Shieber · 2020
Earlier work this paper cites.
Information-theoretic probing with minimum description length
E. Voita and I. Titov · 2020
Earlier work this paper cites.
Transformers: State-of-the-art natural language processing
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, J. Davison, S. Shleifer, P. von Platen, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. Rush · 2020
Earlier work this paper cites.
On layer normalization in the transformer architecture
R. Xiong, Y. Yang, D. He, K. Zheng, S. Zheng, C. Xing, H. Zhang, Y. Lan, L. Wang, and T.-Y. Liu · 2020
Earlier work this paper cites.
Ecco: An open source library for the explainability of transformer language models
J. Alammar · 2021
Earlier work this paper cites.
An interpretability illusion for bert, 2021
T. Bolukbasi, A. Pearce, A. Yuan, A. Coenen, E. Reif, F. Viégas, and M. Wattenberg · 2021
Earlier work this paper cites.
Transformer interpretability beyond attention visualization
H. Chefer, S. Gur, and L. Wolf · 2021
Earlier work this paper cites.
Explaining by removing: A unified framework for model explanation
I. Covert, S. Lundberg, and S.-I. Lee · 2021
Earlier work this paper cites.
Are neural nets modular? inspecting functional modularity through differentiable weight masks
R. Csordás, S. van Steenkiste, and J. Schmidhuber · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
N. De Cao, W. Aziz, and I. Titov · 2021
Earlier work this paper cites.
Who needs to know what, when?: Broadening the explainable ai (xai) design space by looking at explanations across the ai lifecycle
S. Dhanorkar, C. T. Wolf, K. Qian, A. Xu, L. Popa, and Y. Li · 2021
Earlier work this paper cites.
Evaluating saliency methods for neural language models
S. Ding and P. Koehn · 2021
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Y. Elazar, S. Ravfogel, A. Jacovi, and Y. Goldberg · 2021
Earlier work this paper cites.
A mathematical framework for transformer circuits
N. Elhage, N. Nanda, C. Olsson, T. Henighan, N. Joseph, B. Mann, A. Askell, Y. Bai, A. Chen, T. Conerly, N. DasSarma, D. Drain, D. Ganguli, Z. Hatfield-Dodds, D. Hernandez, A. Jones, J. Kernion, L. Lovitt, K. Ndousse, D. Amodei, T. Brown, J. Clark, J. Kaplan, S. McCandlish, and C. Olah · 2021
Earlier work this paper cites.
Attention flows are shapley value explanations
K. Ethayarajh and D. Jurafsky · 2021
Earlier work this paper cites.
Attention weights in transformer NMT fail aligning words between sequences but largely explain model predictions
J. Ferrando and M. R. Costa-jussà · 2021
Earlier work this paper cites.
Causal abstractions of neural networks
A. Geiger, H. Lu, T. Icard, and C. Potts · 2021
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
M. Geva, R. Schuster, J. Berant, and O. Levy · 2021
Earlier work this paper cites.
Surface form competition: Why the highest probability answer isn’t always right
A. Holtzman, P. West, V. Shwartz, Y. Choi, and L. Zettlemoyer · 2021
Earlier work this paper cites.
Incorporating Residual and Normalization Layers into Analysis of Masked Language Models
G. Kobayashi, T. Kuribayashi, S. Yokoi, and K. Inui · 2021
Earlier work this paper cites.
BERT busters: Outlier dimensions that disrupt transformers
O. Kovaleva, S. Kulshreshtha, A. Rogers, and A. Rumshisky · 2021
Earlier work this paper cites.
InterpreT: An interactive visualization tool for interpreting transformers
V. Lal, A. Ma, E. Aflalo, P. Howard, A. Simoes, D. Korat, O. Pereg, G. Singer, and M. Wasserblat · 2021
Earlier work this paper cites.
Positional artefacts propagate through masked language model embeddings
Z. Luo, A. Kulmizev, and X. Mao · 2021
Earlier work this paper cites.
Effects of parameter norm growth during transformer training: Inductive bias from gradient descent
W. Merrill, V. Ramanujan, Y. Goldberg, R. Schwartz, and N. A. Smith · 2021
Earlier work this paper cites.
Investigating the limitations of transformers with simple arithmetic tasks, 2021
R. Nogueira, Z. Jiang, and J. Lin · 2021
Earlier work this paper cites.
Transformers Interpret, February 2021
C. Pierse · 2021
Earlier work this paper cites.
A Primer in BERTology: What We Know About How BERT Works
A. Rogers, O. Kovaleva, and A. Rumshisky · 2021
Earlier work this paper cites.
Discretized integrated gradients for explaining language models
S. Sanyal and X. Ren · 2021
Earlier work this paper cites.
All bark and no bite: Rogue dimensions in transformer language models obscure representational quality
W. Timkey and M. van Schijndel · 2021
Earlier work this paper cites.
Analyzing the source and target contributions to predictions in neural machine translation
E. Voita, R. Sennrich, and I. Titov · 2021
Earlier work this paper cites.
Thinking like transformers
G. Weiss, Y. Goldberg, and E. Yahav · 2021
Earlier work this paper cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
J. Adebayo, M. Muelly, H. Abelson, and B. Kim · 2022
Cited alongside, same era.
Towards tracing knowledge in language models back to the training data
E. Akyurek, T. Bolukbasi, F. Liu, B. Xiong, I. Tenney, J. Andreas, and K. Guu · 2022
Cited alongside, same era.
XAI for transformers: Better explanations through conservative propagation
A. Ali, T. Schnake, O. Eberle, G. Montavon, K.-R. Müller, and L. Wolf · 2022
Cited alongside, same era.
Exploring length generalization in large language models
C. Anil, Y. Wu, A. J. Andreassen, A. Lewkowycz, V. Misra, V. V. Ramasesh, A. Slone, G. Gur-Ari, E. Dyer, and B. Neyshabur · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Y. Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan, N. Joseph, S. Kadavath, J. Kernion, T. Conerly, S. El-Showk, N. Elhage, Z. Hatfield-Dodds, D. Hernandez, T. Hume, S. Johnston, S. Kravec, L. Lovitt, N. Nanda, C. Olsson, D. Amodei, T. Brown, J. Clark, S. McCandlish, C. Olah, B. Mann, and J. Kaplan · 2022
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
T. Räuker, A. Ho, S. Casper, and D. Hadfield-Menell · 2023
Later among the works it cites.
Attention lens: A tool for mechanistically interpreting the attention head information retrieval mechanism, 2023
M. Sakarvadia, A. Khan, A. Ajith, D. Grzenda, N. Hudson, A. Bauer, K. Chard, and I. Foster · 2023
Later among the works it cites.
Inseq: An interpretability toolkit for sequence generation models
G. Sarti, N. Feldhus, L. Sickert, O. van der Wal, M. Nissim, and A. Bisazza · 2023
Later among the works it cites.
A mechanistic interpretation of arithmetic reasoning in language models using causal mediation analysis
A. Stolfo, Y. Belinkov, and M. Sachan · 2023
Later among the works it cites.
Attribution patching outperforms automated circuit discovery
A. Syed, C. Rager, and A. Conmy · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“will you find these shortcuts?” a protocol for evaluating the faithfulness of input salience methods for text classification
J. Bastings, S. Ebert, P. Zablotskaia, A. Sandholm, and K. Filippova · 2022
Cited alongside, same era.
Probing classifiers: Promises, shortcomings, and advances
Y. Belinkov · 2022
Cited alongside, same era.
Is attention explanation? an introduction to the debate
A. Bibal, R. Cardon, D. Alfter, R. Wilkens, X. Wang, T. François, and P. Watrin · 2022
Cited alongside, same era.
Causal scrubbing, a method for rigorously testing interpretability hypotheses
L. Chan, A. Garriga-Alonso, N. Goldwosky-Dill, R. Greenblatt, J. Nitishinskaya, A. Radhakrishnan, B. Shlegeris, and N. Thomas · 2022
Cited alongside, same era.
CircuitVis, December 2022
A. Cooney · 2022
Cited alongside, same era.
Knowledge neurons in pretrained transformers
D. Dai, L. Dong, Y. Hao, Z. Sui, B. Chang, and F. Wei · 2022
Cited alongside, same era.
Sparse interventions in language models with differentiable masking
N. De Cao, L. Schmid, D. Hupkes, and I. Titov · 2022
Cited alongside, same era.
Later among the works it cites.
B2T connection: Serving stability and performance in deep transformers
S. Takase, S. Kiyono, S. Kobayashi, and J. Suzuki · 2023
Later among the works it cites.
Scan and snap: Understanding training dynamics and token composition in 1-layer transformer, 2023
Y. Tian, Y. Wang, B. Chen, and S. Du · 2023
Later among the works it cites.
Linear representations of sentiment in large language models
C. Tigges, O. J. Hollinsworth, A. Geiger, and N. Nanda · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, D. Bikel, L. Blecher, C. C. Ferrer, M. Chen, G. Cucurull, D. Esiobu, J. Fernandes, J. Fu, W. Fu, B. Fuller, C. Gao, V. Goswami, N. Goyal, A. Hartshorn, S. Hosseini, R. Hou, H. Inan, M. Kardas, V. Kerkez, M. Khabsa, I. Kloumann, A. Korenev, P. S. Koura, M.-A. Lachaux, T. Lavril, J. Lee, D. Liskovich, Y. Lu, Y. Mao, X. Martinet, T. Mihaylov, P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizenstein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. Smith, R. Subramanian, X. E. Tan, B. Tang, R. Taylor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Rodriguez, R. Stojnic, S. Edunov, and T. Scialom · 2023
Later among the works it cites.
Activation addition: Steering language models without optimization, 2023
A. M. Turner, L. Thiergart, D. Udell, G. Leech, U. Mini, and M. MacDiarmid · 2023
Later among the works it cites.
M. Turpin, J. Michael, E. Perez, and S. Bowman · 2023
Later among the works it cites.
Some common confusion about induction heads
A. Variengien · 2023
Later among the works it cites.
Look before you leap: A universal emergent decomposition of retrieval tasks in language models, 2023
A. Variengien and E. Winsor · 2023
Later among the works it cites.
Explaining grokking through circuit efficiency
V. Varma, R. Shah, Z. Kenton, J. Kram’ar, and R. Kumar · 2023
Later among the works it cites.
N. Varshney, W. Yao, H. Zhang, J. Chen, and D. Yu · 2023
Later among the works it cites.
Explanations can reduce overreliance on ai systems during decision-making
H. Vasconcelos, M. Jörke, M. Grunde-McLaughlin, T. Gerstenberg, M. S. Bernstein, and R. Krishna · 2023
Later among the works it cites.
Neurons in large language models: Dead, n-gram, positional
E. Voita, J. Ferrando, and C. Nalmpantis · 2023
Later among the works it cites.
Transformers learn in-context by gradient descent
J. Von Oswald, E. Niklasson, E. Randazzo, J. Sacramento, A. Mordvintsev, A. Zhmoginov, and M. Vladymyrov · 2023
Later among the works it cites.
Transformers are uninterpretable with myopic methods: a case study with bounded dyck grammars
K. Wen, Y. Li, B. Liu, and A. Risteski · 2023
Later among the works it cites.
Interpretability at scale: Identifying causal mechanisms in alpaca
Z. Wu, A. Geiger, T. Icard, C. Potts, and N. Goodman · 2023
Later among the works it cites.
A reply to makelov et al. (2023)’s "interpretability illusion" arguments, 2024d
Z. Wu, A. Geiger, J. Huang, A. Arora, T. Icard, C. Potts, and N. D. Goodman · 2023
Later among the works it cites.
Efficient streaming language models with attention sinks
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis · 2023
Later among the works it cites.
Local interpretation of transformer based on linear decomposition
S. Yang, S. Huang, W. Zou, J. Zhang, X. Dai, and J. Chen · 2023
Later among the works it cites.
Editing large language models: Problems, methods, and opportunities
Y. Yao, P. Wang, B. Tian, S. Cheng, Z. Li, S. Deng, H. Chen, and N. Zhang · 2023
Later among the works it cites.
Characterizing mechanisms for factual recall in language models
Q. Yu, J. Merullo, and E. Pavlick · 2023
Later among the works it cites.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Z. Zhong, Z. Liu, M. Tegmark, and J. Andreas · 2023
Later among the works it cites.
Representation engineering: A top-down approach to ai transparency
A. Zou, L. Phan, S. Chen, J. Campbell, P. Guo, R. Ren, A. Pan, X. Yin, M. Mazeika, A.-K. Dombrowski, S. Goel, N. Li, M. J. Byun, Z. Wang, A. Mallen, S. Basart, S. Koyejo, D. Song, M. Fredrikson, J. Z. Kolter, and D. Hendrycks · 2023
Later among the works it cites.
AttnLRP: Attention-aware layer-wise relevance propagation for transformers, 2024
R. Achtibat, S. M. V. Hatefi, M. Dreyer, A. Jain, T. Wiegand, S. Lapuschkin, and W. Samek · 2024
Closest in time.
Faithfulness vs. plausibility: On the (un)reliability of explanations from large language models
C. Agarwal, S. H. Tanneru, and H. Lakkaraju · 2024
Closest in time.
In-context language learning: Architectures and algorithms, 2024
E. Akyürek, B. Wang, Y. Kim, and J. Andreas · 2024
Closest in time.
The hidden attention of mamba models, 2024
A. Ali, I. Zimerman, and L. Wolf · 2024
Closest in time.
Syntaxshap: Syntax-aware explainability method for text generation
K. Amara, R. Sevastjanova, and M. El-Assady · 2024
Closest in time.
Introducing the next generation of claude
Anthropic · 2024
Closest in time.
Refusal in language models is mediated by a single direction
A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Rimsky, W. Gurnee, and N. Nanda · 2024
Closest in time.
Causalgym: Benchmarking causal interpretability methods on linguistic tasks, 2024
A. Arora, D. Jurafsky, and C. Potts · 2024
Closest in time.
What changed? converting representational interventions to natural language
M. Avitan, R. Cotterell, Y. Goldberg, and S. Ravfogel · 2024
Closest in time.
Sparse autoencoders
N. Belrose · 2024
Closest in time.
Mechanistic interpretability for ai safety – a review
L. Bereska and E. Gavves · 2024
Closest in time.
Impossibility theorems for feature attribution
B. Bilodeau, N. Jaques, P. W. Koh, and B. Kim · 2024
Closest in time.
Open source sparse autoencoders for all residual stream layers of GPT2 small
J. Bloom · 2024
Closest in time.
Understanding SAE features with the logit lens
J. Bloom and J. Lin · 2024
Closest in time.
Batchtopk: A simple improvement for topk-saes
B. Bussmann, P. Leask, and N. Nanda · 2024
Closest in time.
Spectral filters, dark signals, and attention sinks, 2024
N. Cancedda · 2024
Closest in time.
Black-box access is insufficient for rigorous ai audits
S. Casper, C. Ezell, C. Siegmann, N. Kolt, T. L. Curtis, B. Bucknall, A. A. Haupt, K. Wei, J. Scheurer, M. Hobbhahn, L. Sharkey, S. Krishna, M. von Hagen, S. Alberti, A. Chan, Q. Sun, M. Gerovitch, D. Bau, M. Tegmark, D. Krueger, and D. Hadfield-Menell · 2024
Closest in time.
Breaking down the defenses: A comparative survey of attacks on large language models, 2024
A. G. Chowdhury, M. M. Islam, V. Kumar, F. H. Shezan, V. Kumar, V. Jain, and A. Chadha · 2024
Closest in time.
Dola: Decoding by contrasting layers improves factuality in large language models
Y.-S. Chuang, Y. Xie, H. Luo, Y. Kim, J. R. Glass, and P. He · 2024
Closest in time.
Summing up the facts: Additive mechanisms behind factual recall in llms, 2024
B. Chughtai, A. Cooney, and N. Nanda · 2024
Closest in time.
Circuits updates - april 2024. update on how we train saes
T. Conerly, A. Templeton, T. Bricken, J. Marcus, and T. Henighan · 2024
Closest in time.
Vision transformers need registers
T. Darcet, M. Oquab, J. Mairal, and P. Bojanowski · 2024
Closest in time.
On the similarity of circuits across languages: a case study on the subject-verb agreement task
J. Ferrando and M. R. Costa-jussà · 2024
Closest in time.
Information flow routes: Automatically interpreting language models at scale
J. Ferrando and E. Voita · 2024
Closest in time.
nnsight: The package for interpreting and manipulating the internals of deep learned models. , 2024
J. Fiotto-Kaufman · 2024
Closest in time.
Scaling and evaluating sparse autoencoders, 2024
L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
G. Gemma Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Rivière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. Tsai, B. Shahriari, C. L. Lan, C. A. Choquette-Choo, C. Crepy, D. Cer, D. Ippolito, D. Reid, E. Buchatskaya, E. Ni, E. Noland, G. Yan, G. Tucker, G.-C. Muraru, G. Rozhdestvenskiy, H. Michalewski, I. Tenney, I. Grishchenko, J. Austin, J. Keeling, J. Labanowski, J.-B. Lespiau, J. Stanway, J. Brennan, J. Chen, J. Ferret, J. Chiu, J. Mao-Jones, K. Lee, K. Yu, K. Millican, L. L. Sjoesund, L. Lee, L. Dixon, M. Reid, M. Mikuła, M. Wirth, M. Sharman, N. Chinaev, N. Thain, O. Bachem, O. Chang, O. Wahltinez, P. Bailey, P. Michel, P. Yotov, R. Chaabouni, R. Comanescu, R. Jana, R. Anil, R. McIlroy, R. Liu, R. Mullins, S. L. Smith, S. Borgeaud, S. Girgin, S. Douglas, S. Pandya, S. Shakeri, S. De, T. Klimenko, T. Hennigan, V. Feinberg, W. Stokowiec, Y. hui Chen, Z. Ahmed, Z. Gong, T. Warkentin, L. Peran, M. Giang, C. Farabet, O. Vinyals, J. Dean, K. Kavukcuoglu, D. Hassabis, Z. Ghahramani, D. Eck, J. Barral, F. Pereira, E. Collins, A. Joulin, N. Fiedel, E. Senter, A. Andreev, and K. Kenealy · 2024
Closest in time.
Patchscopes: A unifying framework for inspecting hidden representations of language models
A. Ghandeharioun, A. Caciularu, A. Pearce, L. Dixon, and M. Geva · 2024
Closest in time.
Successor heads: Recurring, interpretable attention heads in the wild
R. Gould, E. Ong, G. Ogden, and A. Conmy · 2024
Closest in time.
Olmo: Accelerating the science of language models, 2024
D. Groeneveld, I. Beltagy, P. Walsh, A. Bhagia, R. Kinney, O. Tafjord, A. H. Jha, H. Ivison, I. Magnusson, Y. Wang, S. Arora, D. Atkinson, R. Authur, K. R. Chandu, A. Cohan, J. Dumas, Y. Elazar, Y. Gu, J. Hessel, T. Khot, W. Merrill, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. E. Peters, V. Pyatkin, A. Ravichander, D. Schwenk, S. Shah, W. Smith, E. Strubell, N. Subramani, M. Wortsman, P. Dasigi, N. Lambert, K. Richardson, L. Zettlemoyer, J. Dodge, K. Lo, L. Soldaini, N. A. Smith, and H. Hajishirzi · 2024
Closest in time.
Model editing can hurt general abilities of large language models
J.-C. Gu, H. Xu, J.-Y. Ma, P. Lu, Z.-H. Ling, K. wei Chang, and N. Peng · 2024
Closest in time.
Sae reconstruction errors are (empirically) pathological
W. Gurnee · 2024
Closest in time.
Language models represent space and time
W. Gurnee and M. Tegmark · 2024
Closest in time.
Universal neurons in gpt2 language models, 2024
W. Gurnee, T. Horsley, Z. C. Guo, T. R. Kheirkhah, Q. Sun, W. Hathaway, N. Nanda, and D. Bertsimas · 2024
Closest in time.
Parameter-efficient fine-tuning for large models: A comprehensive survey, 2024
Z. Han, C. Gao, J. Liu, J. Zhang, and S. Q. Zhang · 2024
Closest in time.
Have faith in faithfulness: Going beyond circuit overlap when finding model mechanisms, 2024
M. Hanna, S. Pezzelle, and Y. Belinkov · 2024
Closest in time.
Z. He, X. Ge, Q. Tang, T. Sun, Q. Cheng, and X. Qiu · 2024
Closest in time.
How to use and interpret activation patching
S. Heimersheim and N. Nanda · 2024
Closest in time.
Linearity of relation decoding in transformer language models
E. Hernandez, A. S. Sharma, T. Haklay, K. Meng, M. Wattenberg, J. Andreas, Y. Belinkov, and D. Bau · 2024
Closest in time.
Enhanced hallucination detection in neural machine translation through simple detector aggregation, 2024
A. Himmi, G. Staerman, M. Picot, P. Colombo, and N. M. Guerreiro · 2024
Closest in time.
Outlier-efficient hopfield layers for large transformer-based models
J. Y.-C. Hu, P.-H. Chang, R. Luo, H.-Y. Chen, W. Li, W.-P. Wang, and H. Liu · 2024
Closest in time.
Trillion parameter ai serving infrastructure for scientific discovery: A survey and vision
N. Hudson, J. G. Pauloski, M. Baughman, A. Kamatar, M. Sakarvadia, L. Ward, R. Chard, A. Bauer, M. Levental, W. Wang, W. Engler, O. P. Skelly, B. Blaiszik, R. Stevens, K. Chard, and I. Foster · 2024
Closest in time.
What happens when you fine-tuning your model? mechanistic analysis of procedurally generated tasks
S. Jain, R. Kirk, E. S. Lubana, R. P. Dick, H. Tanaka, T. Rocktäschel, E. Grefenstette, and D. Krueger · 2024
Closest in time.
Sparse autoencoders
T. D. L. T. Jeffrey Wu, Leo Gao · 2024
Closest in time.
Circuits updates - jnauary 2024. ghost grads: An improvement on resampling
A. Jermyn and A. Templeton · 2024
Closest in time.
On the origins of linear representations in large language models, 2024
Y. Jiang, G. Rajendran, P. Ravikumar, B. Aragam, and V. Veitch · 2024
Closest in time.
A. Karvonen, B. Wright, C. Rager, R. Angell, J. Brinkmann, L. Smith, C. M. Verdun, D. Bau, and S. Marks · 2024
Closest in time.
Backward lens: Projecting language model gradients into the vocabulary space, 2024
S. Katz, Y. Belinkov, M. Geva, and L. Wolf · 2024
Closest in time.
Analyzing feed-forward blocks in transformers through the lens of attention map
G. Kobayashi, T. Kuribayashi, S. Yokoi, and K. Inui · 2024
Closest in time.
Atp*: An efficient and scalable method for localizing llm behaviour to components, 2024
J. Kramár, T. Lieberum, R. Shah, and N. Nanda · 2024
Closest in time.
The disagreement problem in explainable machine learning: A practitioner’s perspective
S. Krishna, T. Han, A. Gu, S. Wu, S. Jabbari, and H. Lakkaraju · 2024
Closest in time.
We inspected every head in GPT-2 small using saes so you don’t have to
R. Krzyzanowski, C. Kissane, A. Conmy, and N. Nanda · 2024
Closest in time.
Datainf: Efficiently estimating data influence in loRA-tuned LLMs and diffusion models
Y. Kwon, E. Wu, K. Wu, and J. Zou · 2024
Closest in time.
Unveiling the pitfalls of knowledge editing for large language models
Z. Li, N. Zhang, Y. Yao, M. Wang, X. Chen, and H. Chen · 2024
Closest in time.
Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2
T. Lieberum, S. Rajamanoharan, A. Conmy, L. Smith, N. Sonnerat, V. Varma, J. Kramár, A. Dragan, R. Shah, and N. Nanda · 2024
Closest in time.
Announcing neuronpedia: Platform for accelerating research into sparse autoencoders
J. Lin and J. Bloom · 2024
Closest in time.
On training data influence of gpt models
Q. Liu, Y. Chai, S. Wang, Y. Sun, K. Wang, and H. Wu · 2024
Closest in time.
Explainable artificial intelligence (xai) 2.0: A manifesto of open challenges and interdisciplinary research directions
L. Longo, M. Brcic, F. Cabitza, J. Choi, R. Confalonieri, J. D. Ser, R. Guidotti, Y. Hayashi, F. Herrera, A. Holzinger, R. Jiang, H. Khosravi, F. Lecue, G. Malgieri, A. Páez, W. Samek, J. Schneider, T. Speith, and S. Stumpf · 2024
Closest in time.
Attention meets post-hoc interpretability: A mathematical perspective
G. Lopardo, F. Precioso, and D. Garreau · 2024
Closest in time.
Interpreting key mechanisms of factual recall in transformer-based language models
A. Lv, K. Zhang, Y. Chen, Y. Wang, L. Liu, J.-R. Wen, J. Xie, and R. Yan · 2024
Closest in time.
Simple probes can catch sleeper agents
M. MacDiarmid, T. Maxwell, N. Schiefer, J. Mu, J. Kaplan, D. Duvenaud, S. Bowman, A. Tamkin, E. Perez, M. Sharma, C. Denison, and E. Hubinger · 2024
Closest in time.
Are self-explanations from large language models faithful?
A. Madsen, S. Chandar, and S. Reddy · 2024
Closest in time.
Sparse feature circuits: Discovering and editing interpretable causal graphs in language models
S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller · 2024
Closest in time.
Sae-vis: Announcement post
C. McDougall and J. Bloom · 2024
Closest in time.
Circuit component reuse across tasks in transformer language models
J. Merullo, C. Eickhoff, and E. Pavlick · 2024
Closest in time.
Large language models: A survey, 2024
S. Minaee, T. Mikolov, N. Nikzad, M. Chenaghlu, R. Socher, X. Amatriain, and J. Gao · 2024
Closest in time.
A glitch in the matrix? locating and detecting language model grounding with fakepedia, 2024
G. Monea, M. Peyrard, M. Josifoski, V. Chaudhary, J. Eisner, E. Kıcıman, H. Palangi, B. Patra, and R. West · 2024
Closest in time.
Transformer debugger
D. Mossing, S. Bills, H. Tillman, T. Dupré la Tour, N. Cammarata, L. Gao, J. Achiam, C. Yeh, J. Leike, J. Wu, and W. Saunders · 2024
Closest in time.
Interpreting context look-ups in transformers: Investigating attention-mlp interactions, 2024
C. Neo, S. B. Cohen, and F. Barez · 2024
Closest in time.
Competition of mechanisms: Tracing how language models handle facts and counterfactuals
F. Ortu, Z. Jin, D. Doimo, M. Sachan, A. Cazzaniga, and B. Schölkopf · 2024
Closest in time.
Does transformer interpretability transfer to rnns?, 2024
G. Paulo, T. Marshall, and N. Belrose · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
N. Prakash, T. R. Shaham, T. Haklay, Y. Belinkov, and D. Bau · 2024
Closest in time.
Progress update 1 from the gdm mech interp team. improving ghost grads
S. Rajamanoharan · 2024
Closest in time.
Geometry and dynamics of layernorm
P. M. Riechers · 2024
Closest in time.
Explorations of self-repair in language models, 2024
C. Rushing and N. Nanda · 2024
Closest in time.
Quantifying the plausibility of context reliance in neural machine translation
G. Sarti, G. Chrupała, M. Nissim, and A. Bisazza · 2024
Closest in time.
Decomposing and editing predictions by modeling model computation
H. Shah, A. Ilyas, and A. Madry · 2024
Closest in time.
A multimodal automated interpretability agent
T. R. Shaham, S. Schwettmann, F. Wang, A. Rajaram, E. Hernandez, J. Andreas, and A. Torralba · 2024
Closest in time.
The probabilities also matter: A more faithful metric for faithfulness of free-text explanations in large language models, 2024
N. Y. Siegel, O.-M. Camburu, N. Heess, and M. Perez-Ortiz · 2024
Closest in time.
Localizing paragraph memorization in language models, 2024
N. Stoehr, M. Gordon, C. Zhang, and O. Lewis · 2024
Closest in time.
Confidence regulation neurons in language models
A. Stolfo, B. Wu, W. Gurnee, Y. Belinkov, X. Song, M. Sachan, and N. Nanda · 2024
Closest in time.
Massive activations in large language models, 2024
M. Sun, X. Chen, J. Z. Kolter, and Z. Liu · 2024
Closest in time.
Language-specific neurons: The key to multilingual capabilities in large language models, 2024
T. Tang, W. Luo, H. Huang, D. Zhang, X. Wang, X. Zhao, F. Wei, and J.-R. Wen · 2024
Closest in time.
Transformers as support vector machines, 2024
D. A. Tarzanagh, Y. Li, C. Thrampoulidis, and S. Oymak · 2024
Closest in time.
Gemma 2: Improving open language models at a practical size
G. Team, M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, L. Hussenot, T. Mesnard, B. Shahriari, A. Ramé, J. Ferret, P. Liu, P. Tafti, A. Friesen, M. Casbon, S. Ramos, R. Kumar, C. L. Lan, S. Jerome, A. Tsitsulin, N. Vieillard, P. Stanczyk, S. Girgin, N. Momchev, M. Hoffman, S. Thakoor, J.-B. Grill, B. Neyshabur, O. Bachem, A. Walton, A. Severyn, A. Parrish, A. Ahmad, A. Hutchison, A. Abdagic, A. Carl, A. Shen, A. Brock, A. Coenen, A. Laforge, A. Paterson, B. Bastian, B. Piot, B. Wu, B. Royal, C. Chen, C. Kumar, C. Perry, C. Welty, C. A. Choquette-Choo, D. Sinopalnikov, D. Weinberger, D. Vijaykumar, D. Rogozińska, D. Herbison, E. Bandy, E. Wang, E. Noland, E. Moreira, E. Senter, E. Eltyshev, F. Visin, G. Rasskin, G. Wei, G. Cameron, G. Martins, H. Hashemi, H. Klimczak-Plucińska, H. Batra, H. Dhand, I. Nardini, J. Mein, J. Zhou, J. Svensson, J. Stanway, J. Chan, J. P. Zhou, J. Carrasqueira, J. Iljazi, J. Becker, J. Fernandez, J. van Amersfoort, J. Gordon, J. Lipschultz, J. Newlan, J. yeong Ji, K. Mohamed, K. Badola, K. Black, K. Millican, K. McDonell, K. Nguyen, K. Sodhia, K. Greene, L. L. Sjoesund, L. Usui, L. Sifre, L. Heuermann, L. Lago, L. McNealus, L. B. Soares, L. Kilpatrick, L. Dixon, L. Martins, M. Reid, M. Singh, M. Iverson, M. Görner, M. Velloso, M. Wirth, M. Davidow, M. Miller, M. Rahtz, M. Watson, M. Risdal, M. Kazemi, M. Moynihan, M. Zhang, M. Kahng, M. Park, M. Rahman, M. Khatwani, N. Dao, N. Bardoliwalla, N. Devanathan, N. Dumai, N. Chauhan, O. Wahltinez, P. Botarda, P. Barnes, P. Barham, P. Michel, P. Jin, P. Georgiev, P. Culliton, P. Kuppala, R. Comanescu, R. Merhej, R. Jana, R. A. Rokni, R. Agarwal, R. Mullins, S. Saadat, S. M. Carthy, S. Perrin, S. M. R. Arnold, S. Krause, S. Dai, S. Garg, S. Sheth, S. Ronstrom, S. Chan, T. Jordan, T. Yu, T. Eccles, T. Hennigan, T. Kocisky, T. Doshi, V. Jain, V. Yadav, V. Meshram, V. Dharmadhikari, W. Barkley, W. Wei, W. Ye, W. Han, W. Kwon, X. Xu, Z. Shen, Z. Gong, Z. Wei, V. Cotruta, P. Kirk, A. Rao, M. Giang, L. Peran, T. Warkentin, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, D. Sculley, J. Banks, A. Dragan, S. Petrov, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, S. Borgeaud, N. Fiedel, A. Joulin, K. Kenealy, R. Dadashi, and A. Andreev · 2024
Closest in time.
Circuits updates - february 2024. update on dictionary learning improvements
A. Templeton, T. Conerly, J. Marcus, T. Henighan, A. Golubeva, and T. Bricken · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones, H. Cunningham, N. L. Turner, C. McDougall, M. MacDiarmid, C. D. Freeman, T. R. Sumers, E. Rees, J. Batson, A. Jermyn, S. Carter, C. Olah, and T. Henighan · 2024
Closest in time.
Interactive prompt debugging with sequence salience
I. Tenney, R. Mullins, B. Du, S. Pandya, M. Kahng, and L. Dixon · 2024
Closest in time.
JoMA: Demystifying multilayer transformers via joint dynamics of MLP and attention
Y. Tian, Y. Wang, Z. Zhang, B. Chen, and S. S. Du · 2024
Closest in time.
LLMs represent contextual tasks as compact function vectors
E. Todd, M. Li, A. S. Sharma, A. Mueller, B. C. Wallace, and D. Bau · 2024
Closest in time.
Lm transparency tool: Interactive tool for analyzing transformer language models
I. Tufanov, K. Hambardzumyan, J. Ferrando, and E. Voita · 2024
Closest in time.
Llmcheckup: Conversational examination of large language models via interpretability tools, 2024
Q. Wang, T. Anikina, N. Feldhus, J. van Genabith, L. Hennig, and S. Möller · 2024
Closest in time.
Gradient-based language model red teaming, 2024
N. Wichers, C. Denison, and A. Beirami · 2024
Closest in time.
Addressing feature suppression in saes
B. Wright and L. Sharkey · 2024
Closest in time.
Z. Yu and S. Ananiadou · 2024
Closest in time.
Attention satisfies: A constraint-satisfaction lens on factual errors of language models
M. Yuksekgonul, V. Chandrasekaran, E. Jones, S. Gunasekar, R. Naik, H. Palangi, E. Kamar, and B. Nushi · 2024
Closest in time.
Towards best practices of activation patching in language models: Metrics and methods
F. Zhang and N. Nanda · 2024
Closest in time.
Reagent: A model-agnostic feature attribution method for generative language models, 2024
Z. Zhao and B. Shan · 2024
Closest in time.
On prompt-driven safeguarding for large language models
C. Zheng, F. Yin, H. Zhou, F. Meng, J. Zhou, K.-W. Chang, M. Huang, and N. Peng · 2024
Closest in time.
What algorithms can transformers learn? a study in length generalization
H. Zhou, A. Bradley, E. Littwin, N. Razin, O. Saremi, J. M. Susskind, S. Bengio, and P. Nakkiran · 2024
Closest in time.
The explainability of transformers: Current status and directions
P. Fantozzi and M. Naldi · 2073
Closest in time.