Fetching the paper…
Reading the bibliography…
Interpretability methods aim to understand the algorithm implemented by a trained model (e.g., a Transofmer) by examining various aspects of the model, such as the weight matrices or the attention patterns.
On context-free languages and push-down automata
M.P. Schützenberger · 1963
Earlier work this paper cites.
On the computational power of neural nets
Hava T. Siegelmann and Eduardo D. Sontag · 1992
Earlier work this paper cites.
Lstm recurrent networks learn simple context-free and context-sensitive languages
F. Gers and J. Schmidhuber · 2001
Earlier work this paper cites.
Constraints on multiple center-embedding of clauses
Fred Karlsson · 2007
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky · 2016
Earlier work this paper cites.
The mythos of model interpretability, 2017
Zachary C. Lipton · 2017
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Unsupervised grammar induction with depth-bounded pcfg
Lifeng Jin, Finale Doshi-Velez, Timothy Miller, William Schuler, and Lane Schwartz · 2018
Earlier work this paper cites.
An analysis of encoder representations in transformer-based machine translation
Alessandro Raganato and Jörg Tiedemann · 2018
Earlier work this paper cites.
On the practical computational power of finite precision rnns for language recognition
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2018
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning · 2019
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
Do attention heads in bert track syntactic dependencies?, 2019
Phu Mon Htut, Jason Phang, Shikha Bordia, and Samuel R. Bowman · 2019
Earlier work this paper cites.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace · 2019
Earlier work this paper cites.
Revealing the dark secrets of BERT
Olga Kovaleva, Alexey Romanov, Anna Rogers, and Anna Rumshisky · 2019
Earlier work this paper cites.
Open sesame: Getting inside BERT’s linguistic knowledge
Yongjie Lin, Yi Chern Tan, and Robert Frank · 2019
Earlier work this paper cites.
Linguistic knowledge and transferability of contextual representations
Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith · 2019
Earlier work this paper cites.
Sequential neural networks as automata
William Merrill · 2019
Earlier work this paper cites.
Are sixteen heads really better than one?
Paul Michel, Omer Levy, and Graham Neubig · 2019
Earlier work this paper cites.
Transformers without tears: Improving the normalization of self-attention
Toan Q. Nguyen and Julian Salazar · 2019
Earlier work this paper cites.
Is attention interpretable?
Sofia Serrano and Noah A. Smith · 2019
Earlier work this paper cites.
LSTM networks can perform dynamic counting
Mirac Suzgun, Yonatan Belinkov, Stuart Shieber, and Sebastian Gehrmann · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov · 2019
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter · 2019
Cited alongside, same era.
On the Ability and Limitations of Transformers to Recognize Formal Languages
Satwik Bhattamishra, Kabir Ahuja, and Navin Goyal · 2020
Cited alongside, same era.
On the computational power of transformers and its implications in sequence modeling
Satwik Bhattamishra, Arkil Patel, and Navin Goyal · 2020
Cited alongside, same era.
On identifiability in transformers
Gino Brunner, Yang Liu, Damian Pascual, Oliver Richter, Massimiliano Ciaramita, and Roger Wattenhofer · 2020
Cited alongside, same era.
Thread: Circuits
Nick Cammarata, Shan Carter, Gabriel Goh, Chris Olah, Michael Petrov, Ludwig Schubert, Chelsea Voss, Ben Egan, and Swee Kiat Lim · 2020
Cited alongside, same era.
How can self-attention networks recognize Dyck-n languages?
Javid Ebrahimi, Dhruv Gelda, and Wei Zhang · 2020
Cited alongside, same era.
Attention is turing-complete
Jorge Perez, Pablo Barcelo, and Javier Marinkovic · 2021
Later among the works it cites.
Effective attention sheds light on interpretability
Kaiser Sun and Ana Marasović · 2021
Later among the works it cites.
Colin Wei, Yining Chen, and Tengyu Ma · 2021
Later among the works it cites.
Thinking like transformers
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2021
Later among the works it cites.
Self-attention networks can process bounded hierarchical languages
Shunyu Yao, Binghui Peng, Christos Papadimitriou, and Karthik Narasimhan · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Why attention is not explanation: Surgical intervention and causal reasoning about neural models
Christopher Grimsley, Elijah Mayfield, and Julia R.S. Bursten · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Rnns can generate bounded hierarchical languages with optimal memory
John Hewitt, Michael Hahn, Surya Ganguli, Percy Liang, and Christopher D Manning · 2020
Cited alongside, same era.
Attention is not only a weight: Analyzing transformers with vector norms
Goro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, and Kentaro Inui · 2020
Cited alongside, same era.
Understanding the difficulty of training transformers
Liyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen, and Jiawei Han · 2020
Cited alongside, same era.
Later among the works it cites.
Impossibility theorems for feature attribution
Blair Bilodeau, Natasha Jaques, Pang Wei Koh, and Been Kim · 2022
Later among the works it cites.
Interpretable machine learning: Moving from mythos to diagnostics
Valerie Chen, Jeffrey Li, Joon Sik Kim, Gregory Plumb, and Ameet Talwalkar · 2022
Later among the works it cites.
Analyzing transformers in embedding space, 2022
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2022
Later among the works it cites.
Inductive biases and variable creation in self-attention mechanisms
Benjamin L Edelman, Surbhi Goel, Sham Kakade, and Cyril Zhang · 2022
Later among the works it cites.
Vision transformers provably learn spatial structure
Samy Jelassi, Michael Eli Sander, and Yuanzhi Li · 2022
Later among the works it cites.
Are representations built from the ground up? an empirical examination of local composition in language models
Emmy Liu and Graham Neubig · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Later among the works it cites.
Unveiling transformers with lego: a synthetic reasoning task, 2022
Yi Zhang, Arturs Backurs, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, and Tal Wagner · 2022
Later among the works it cites.
A toy model of universality: Reverse engineering how networks learn group operations
B. Chughtai, Lawrence Chan, and Neel Nanda · 2023
Closest in time.
Attention scheme inspired softmax regression, 2023
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
Tinystories: How small can language models be and still speak coherent english?, 2023
Ronen Eldan and Yuanzhi Li · 2023
Closest in time.
An over-parameterized exponential regression, 2023
Yeqi Gao, Sridhar Mahadevan, and Zhao Song · 2023
Closest in time.
Is attention interpretation? a quantitative assessment on sets
Jonathan Haab, Nicolas Deutschmann, and María Rodríguez Martínez · 2023
Closest in time.
How do transformers learn topic structure: Towards a mechanistic understanding
Yuchen Li, Yuanzhi Li, and Andrej Risteski · 2023
Closest in time.
Transformers learn shortcuts to automata
Bingbin Liu, Jordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Interpretability in the wild: a circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Closest in time.
Do transformers parse while predicting the masked word?, 2023
Haoyu Zhao, Abhishek Panigrahi, Rong Ge, and Sanjeev Arora · 2023
Closest in time.
The clock and the pizza: Two stories in mechanistic explanation of neural networks
Ziqian Zhong, Ziming Liu, Max Tegmark, and Jacob Andreas · 2023
Closest in time.