Fetching the paper…
Reading the bibliography…
Recent research in mechanistic interpretability has attempted to reverse-engineer Transformer models by carefully inspecting network weights and activations.
Inductive logic programming: Theory and methods
Stephen Muggleton and Luc De Raedt · 1994
Earlier work this paper cites.
Building a question answering test collection
Ellen M. Voorhees and Dawn M. Tice · 2000
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik Tjong Kim Sang and Fien De Meulder · 2003
Earlier work this paper cites.
A sentimental education: Sentiment analysis using subjectivity summarization based on minimum cuts
Bo Pang and Lillian Lee · 2004
Earlier work this paper cites.
Rule extraction from recurrent neural networks: A taxonomy and review
Henrik Jacobsson · 2005
Earlier work this paper cites.
Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Neural programmer-interpreters
Scott Reed and Nando De Freitas · 2015
Earlier work this paper cites.
Falling rule lists
Fulton Wang and Cynthia Rudin · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Making neural programming architectures generalize via recursion
Jonathon Cai, Richard Shin, and Dawn Song · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
seqeval: A Python framework for sequence labeling evaluation, 2018
Hiroki Nakayama · 2018
Earlier work this paper cites.
An empirical evaluation of rule extraction from recurrent neural networks
Qinglong Wang, Kaixuan Zhang, Alexander G Ororbia II, Xinyu Xing, Xue Liu, and C Lee Giles · 2018
Earlier work this paper cites.
Extracting automata from recurrent neural networks using queries and counterexamples
Gail Weiss, Yoav Goldberg, and Eran Yahav · 2018
Earlier work this paper cites.
This looks like that: Deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Cited alongside, same era.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Cited alongside, same era.
Attention is not explanation
Sarthak Jain and Byron C Wallace · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Sparse attention with linear units
Biao Zhang, Ivan Titov, and Rico Sennrich · 2021
Later among the works it cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Later among the works it cites.
Inductive logic programming at 30: A new introduction
Andrew Cropper and Sebastijan Dumančić · 2022
Later among the works it cites.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, et al · 2022
Later among the works it cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Later among the works it cites.
Formal language recognition by hard attention transformers: Perspectives from circuit complexity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ali Payani and Faramarz Fekri · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Cited alongside, same era.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Discovering symbolic models from deep learning with inductive biases
Miles Cranmer, Alvaro Sanchez Gonzalez, Peter Battaglia, Rui Xu, Kyle Cranmer, David Spergel, and Shirley Ho · 2020
Cited alongside, same era.
Theoretical limitations of self-attention in neural sequence models
Michael Hahn · 2020
Cited alongside, same era.
Neurosymbolic Transformers for multi-agent communication
Jeevana Priya Inala, Yichen Yang, James Paulos, Yewen Pu, Osbert Bastani, Vijay Kumar, Martin Rinard, and Armando Solar-Lezama · 2020
Cited alongside, same era.
Yiding Hao, Dana Angluin, and Robert Frank · 2022
Later among the works it cites.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Later among the works it cites.
Transformers implement first-order logic with majority quantifiers
William Merrill and Ashish Sabharwal · 2022
Later among the works it cites.
Saturated Transformers are constant-depth threshold circuits
William Merrill, Ashish Sabharwal, and Noah A Smith · 2022
Later among the works it cites.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, et al · 2022
Later among the works it cites.
Deep differentiable logic gate networks
Felix Petersen, Christian Borgelt, Hilde Kuehne, and Oliver Deussen · 2022
Later among the works it cites.
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah Smith, and Mike Lewis · 2022
Later among the works it cites.
Finding alignments between interpretable causal variables and distributed neural representations
Atticus Geiger, Zhengxuan Wu, Christopher Potts, Thomas Icard, and Noah D Goodman · 2023
Closest in time.
Looped Transformers as programmable computers
Angeliki Giannou, Shashank Rajput, Jy yong Sohn, Kangwook Lee, Jason D. Lee, and Dimitris Papailiopoulos · 2023
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik · 2023
Closest in time.
The Hydra effect: Emergent self-repair in language model computations
Thomas McGrath, Matthew Rahtz, Janos Kramar, Vladimir Mikulik, and Shane Legg · 2023
Closest in time.
Progress measures for grokking via mechanistic interpretability
Neel Nanda, Lawrence Chan, Tom Lieberum, Jess Smith, and Jacob Steinhardt · 2023
Closest in time.
Interpreting GPT: The logit lens
Nostalgebraist · 2023
Closest in time.
Interpretability in the wild: A circuit for indirect object identification in GPT-2 small
Kevin Ro Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt · 2023
Closest in time.