Fetching the paper…
Reading the bibliography…
Artificial neural networks have long been understood as "black boxes": though we know their computation graphs and learned parameters, the knowledge encoded by these weights and functions they perform are not inherently interpretable.
“Large language models in medicine”
Arun Thirunavukarasu et al · 1940
Earlier work this paper cites.
“The magical number seven, plus or minus two: Some limits on our capacity for processing information.”
George Miller · 1956
Earlier work this paper cites.
“Verbal behavior”
Burrhus Skinner · 1957
Earlier work this paper cites.
“ELIZA—a computer program for the study of natural language communication between man and machine”
Joseph Weizenbaum · 1966
Earlier work this paper cites.
“A review of BF Skinner’s Verbal Behavior”
Noam Chomsky · 1980
Earlier work this paper cites.
“Vision: A Computational Investigation Into the Human Representation and Processing of Visual Information”
D. Marr · 1982
Earlier work this paper cites.
“The impact and promise of the cognitive revolution.”
Roger Sperry · 1993
Earlier work this paper cites.
“Learning Overcomplete Representations”
Michael. Lewicki and Terrence. Sejnowski · 2000
Earlier work this paper cites.
“Statistical modeling: The two cultures (with comments and a rejoinder by the author)”
Leo Breiman · 2001
Earlier work this paper cites.
“Origins of the cognitive (r) evolution”
George Mandler · 2002
Earlier work this paper cites.
“Syntactic structures”
Noam Chomsky · 2002
Earlier work this paper cites.
“Genealogy of the “grandmother cell””
Charles Gross · 2002
Earlier work this paper cites.
“The cognitive revolution: a historical perspective”
George Miller · 2003
Earlier work this paper cites.
“An elementary proof of a theorem of Johnson and Lindenstrauss”
Sanjoy Dasgupta and Anupam Gupta · 2003
Earlier work this paper cites.
“Efficient sparse coding algorithms”
Honglak Lee, Alexis Battle, Rajat Raina and Andrew Ng · 2006
Earlier work this paper cites.
“Authenticity in the age of digital companions”
Sherry Turkle · 2007
Earlier work this paper cites.
“Visualizing higher-layer features of a deep network”
Dumitru Erhan, Yoshua Bengio, Aaron Courville and Pascal Vincent · 2009
Earlier work this paper cites.
“Deep inside convolutional networks: Visualising image classification models and saliency maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2013
Earlier work this paper cites.
“Gnostic cells in the 21st century”
Rodrigo Quiroga · 2013
Earlier work this paper cites.
“Deep inside convolutional networks: Visualising image classification models and saliency maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2013
Earlier work this paper cites.
“Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps”
Karen Simonyan, Andrea Vedaldi and Andrew Zisserman · 2014
Earlier work this paper cites.
“Visualizing and understanding convolutional networks”
Matthew Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
“Object detectors emerge in deep scene cnns”
Bolei Zhou et al · 2014
Earlier work this paper cites.
“On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation”
Sebastian Bach et al · 2015
Earlier work this paper cites.
“Human-level control through deep reinforcement learning”
Volodymyr Mnih et al · 2015
Earlier work this paper cites.
“Visualizing and understanding recurrent networks”
Andrej Karpathy, Justin Johnson and Li Fei-Fei · 2015
Earlier work this paper cites.
“Sparse Overcomplete Word Vector Representations”
Manaal Faruqui et al · 2015
Earlier work this paper cites.
“”Why should i trust you?” Explaining the predictions of any classifier”
Marco Ribeiro, Sameer Singh and Carlos Guestrin · 2016
Earlier work this paper cites.
“Learning deep features for discriminative localization”
Bolei Zhou et al · 2016
Earlier work this paper cites.
“Reinforcement learning with Marr”
Yael Niv and Angela Langdon · 2016
Earlier work this paper cites.
“Synthesizing the preferred inputs for neurons in neural networks via deep generator networks”
Anh Nguyen et al · 2016
Earlier work this paper cites.
“Understanding intermediate layers using linear classifier probes”
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
“Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings”
Tolga Bolukbasi et al · 2016
Earlier work this paper cites.
“A study of thinking”
Jerome Bruner · 2017
Earlier work this paper cites.
“What does explainable AI really mean? A new conceptualization of perspectives”
Derek Doran, Sarah Schulz and Tarek Besold · 2017
Earlier work this paper cites.
“A unified approach to interpreting model predictions”
Scott. Lundberg and Su Lee · 2017
Earlier work this paper cites.
“Axiomatic attribution for deep networks”
Mukund Sundararajan, Ankur Taly and Qiqi Yan · 2017
Earlier work this paper cites.
“Grad-cam: Visual explanations from deep networks via gradient-based localization”
Ramprasaath Selvaraju et al · 2017
Earlier work this paper cites.
“Counterfactual explanations without opening the black box: Automated decisions and the GDPR”
Sandra Wachter, Brent Mittelstadt and Chris Russell · 2017
Earlier work this paper cites.
“Feature visualization”
Chris Olah, Alexander Mordvintsev and Ludwig Schubert · 2017
Earlier work this paper cites.
“Network dissection: Quantifying interpretability of deep visual representations”
David Bau et al · 2017
Earlier work this paper cites.
“Towards a rigorous science of interpretable machine learning”
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
“The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery.”
Zachary Lipton · 2018
Earlier work this paper cites.
“Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks”
Ruth Fong and Andrea Vedaldi · 2018
Earlier work this paper cites.
“Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)”
Been Kim et al · 2018
Cited alongside, same era.
“Spine: Sparse interpretable neural embeddings”
Anant Subramanian et al · 2018
Cited alongside, same era.
“Peeking inside the black-box: a survey on explainable artificial intelligence (XAI)”
Amina Adadi and Mohammed Berrada · 2018
Cited alongside, same era.
“Understanding deep networks via extremal perturbations and smooth masks”
Ruth Fong, Mandela Patrick and Andrea Vedaldi · 2019
Cited alongside, same era.
“Visualizing and Understanding Generative Adversarial Networks”
David Bau et al · 2019
Cited alongside, same era.
“What is one grain of sand in the desert? analyzing individual neurons in deep nlp models”
Fahim Dalvi et al · 2019
“Psychological foundations of explainability and interpretability in artificial intelligence”
David Broniatowski · 2021
Later among the works it cites.
“Probing Classifiers: Promises, Shortcomings, and Advances”
Yonatan Belinkov · 2022
Later among the works it cites.
“Do explanations explain? model knows best”
Ashkan Khakzar, Pedram Khorsandi, Rozhin Nobahari and Nassir Navab · 2022
Later among the works it cites.
“Mechanistic Interpretability, Variables, and the Importance of Interpretable Bases”
Christopher Olah · 2022
Later among the works it cites.
“Explainability and artificial intelligence in medicine”
Sandeep Reddy · 2022
Later among the works it cites.
“Opening the black box: the promise and limitations of explainable machine learning in cardiology”
Jeremy Petch, Shuang Di and Walter Nelson · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned”
Elena Voita et al · 2019
Cited alongside, same era.
“Analyzing the Structure of Attention in a Transformer Language Model”
Jesse Vig and Yonatan Belinkov · 2019
Cited alongside, same era.
“What Does BERT Look at? An Analysis of BERT’s Attention”
Kevin Clark, Urvashi Khandelwal, Omer Levy and Christopher. Manning · 2019
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2019
Cited alongside, same era.
“BERT Rediscovers the Classical NLP Pipeline”
Ian Tenney, Dipanjan Das and Ellie Pavlick · 2019
Cited alongside, same era.
“What does BERT learn about the structure of language?”
Ganesh Jawahar, Benoit Sagot and Djame Seddah · 2019
Cited alongside, same era.
Later among the works it cites.
“Knowledge Neurons in Pretrained Transformers”
Damai Dai et al · 2022
Later among the works it cites.
Nelson Elhage et al · 2022
Later among the works it cites.
“Locating and editing factual associations in GPT”
Kevin Meng, David Bau, Alex Andonian and Yonatan Belinkov · 2022
Later among the works it cites.
“Does BERT Rediscover a Classical NLP Pipeline?”
Jingcheng Niu, Wenjie Lu and Gerald Penn · 2022
Later among the works it cites.
“The Architectural Bottleneck Principle”
Tiago Pimentel, Josef Valvoda, Niklas Stoehr and Ryan Cotterell · 2022
Later among the works it cites.
“Concept activation regions: A generalized framework for concept-based explanations”
Jonathan Crabbé and Mihaela van Schaar · 2022
Later among the works it cites.
“LiT: Zero-Shot Transfer With Locked-Image Text Tuning”
Xiaohua Zhai et al · 2022
Later among the works it cites.
“Probing for the Usage of Grammatical Number”
Karim Lasri et al · 2022
Later among the works it cites.
“Linear adversarial concept erasure”
Shauli Ravfogel, Michael Twiton, Yoav Goldberg and Ryan Cotterell · 2022
Later among the works it cites.
“Adversarial Concept Erasure in Kernel Space”
Shauli Ravfogel, Francisco Vargas, Yoav Goldberg and Ryan Cotterell · 2022
Later among the works it cites.
“Fundamental limits and tradeoffs in invariant representation learning”
Han Zhao et al · 2022
Later among the works it cites.
“Polysemanticity and capacity in neural networks”
Adam Scherlis et al · 2022
Later among the works it cites.
“In-context learning and induction heads”
Catherine Olsson et al · 2022
Later among the works it cites.
“Inducing Causal Structure for Interpretable Neural Networks”
Atticus Geiger et al · 2022
Later among the works it cites.
“Dissociating language and thought in large language models: a cognitive perspective”
Kyle Mahowald et al · 2023
Later among the works it cites.
“Towards Automated Circuit Discovery for Mechanistic Interpretability”
Arthur Conmy et al · 2023
Later among the works it cites.
“Competence-Based Analysis of Language Models”
Adam Davies, Jize Jiang and ChengXiang Zhai · 2023
Later among the works it cites.
“Causal Abstraction for Faithful Model Interpretation”
Atticus Geiger, Chris Potts and Thomas Icard · 2023
Later among the works it cites.
“Representation engineering: A top-down approach to ai transparency”
Andy Zou et al · 2023
Later among the works it cites.
“Towards Monosemanticity: Decomposing Language Models With Dictionary Learning” https://transformer-circuits.pub/2023/monosemantic-features/index.html
Trenton Bricken et al · 2023
Later among the works it cites.
“Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 Small”
Kevin Wang et al · 2023
Later among the works it cites.
“Multimodal neurons in pretrained text-only transformers”
Sarah Schwettmann et al · 2023
Later among the works it cites.
“Mass-Editing Memory in a Transformer”
Kevin Meng et al · 2023
Later among the works it cites.
“The linear representation hypothesis and the geometry of large language models”
Kiho Park, Yo Choe and Victor Veitch · 2023
Later among the works it cites.
Samuel Marks and Max Tegmark · 2023
Later among the works it cites.
“Linear representations of sentiment in large language models”
Curt Tigges, Oskar Hollinsworth, Atticus Geiger and Neel Nanda · 2023
Later among the works it cites.
“Emergent Linear Representations in World Models of Self-Supervised Sequence Models”
Neel Nanda, Andrew Lee and Martin Wattenberg · 2023
Later among the works it cites.
“Gold Doesn’t Always Glitter: Spectral Removal of Linear and Nonlinear Guarded Attribute Information”
Shun Shao, Yftah Ziser and Shay. Cohen · 2023
Later among the works it cites.
“Sparse autoencoders find highly interpretable features in language models”
Hoagy Cunningham et al · 2023
Later among the works it cites.
“Language models can explain neurons in language models”, https://openaipublic.blob.core.windows.net/neuron-explainer/paper/index.html , 2023
Steven Bills et al · 2023
Later among the works it cites.
“The clock and the pizza: Two stories in mechanistic explanation of neural networks”
Ziqian Zhong, Ziming Liu, Max Tegmark and Jacob Andreas · 2023
Later among the works it cites.
“Progress measures for grokking via mechanistic interpretability”
Neel Nanda et al · 2023
Later among the works it cites.
“Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet”
Adly Templeton et al · 2024
Closest in time.
“Missed Causes and Ambiguous Effects: Counterfactuals Pose Challenges for Interpreting Neural Networks”
Aaron Mueller · 2024
Closest in time.
“Hypothesis Testing the Circuit Hypothesis in LLMs”
Claudia Shi et al · 2024
Closest in time.
“Finding alignments between interpretable causal variables and distributed neural representations”
Atticus Geiger et al · 2024
Closest in time.
“Interpretability at scale: Identifying causal mechanisms in alpaca”
Zhengxuan Wu et al · 2024
Closest in time.