Distributed representations
Geoffrey E Hinton · 1984
Earlier work this paper cites.
A value for n -person games
Lloyd S. Shapley · 1988
Earlier work this paper cites.
Sparse coding with an overcomplete basis set: a strategy employed by v1?
B. A. Olshausen and D. J. Field · 1997
Earlier work this paper cites.
Conceptual spaces: The geometry of thought
Peter Gardenfors · 2004
Earlier work this paper cites.
Pattern recognition and machine learning
Christopher M. Bishop · 2006
Earlier work this paper cites.
Causality
Judea Pearl · 2009
Earlier work this paper cites.
Algebraic Geometry and Statistical Learning Theory
Sumio Watanabe · 2009
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Original
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Original
Matthew D. Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Polysemy: Current perspectives and approaches
Ingrid Lossius Falkum and Agustin Vicente · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
Causal inference using invariant prediction: identification and confidence intervals
Original
Jonas Peters, Peter Bühlmann, and Nicolai Meinshausen · 2015
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Original
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
"why should i trust you?": Explaining the predictions of any classifier
Original
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Grad-cam: Why did you say that? visual explanations from deep networks via gradient-based localization
Original
Ramprasaath R. Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Batra · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Original
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Original
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Research debt
Chris Olah and Shan Carter · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Elements of Causal Inference: Foundations and Learning Algorithms
Jonas Peters, Dominik Janzing, and Bernhard Schlkopf · 2017
Earlier work this paper cites.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients
Original
Andrew Slavin Ross and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Original
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Original
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Original
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Original
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Linear algebraic structure of word senses, with applications to polysemy
Original
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski · 2018
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks
Original
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba · 2018
Earlier work this paper cites.
Visualizing the feature importance for black box models
Original
Giuseppe Casalicchio, Christoph Molnar, and Bernd Bischl · 2018
Earlier work this paper cites.
Recurrent world models facilitate policy evolution
David R. Ha and J. Schmidhuber · 2018
Earlier work this paper cites.
Towards a definition of disentangled representations
Original
Irina Higgins, David Amos, David Pfau, Sebastien Racaniere, Loic Matthey, Danilo Rezende, and Alexander Lerchner · 2018
Earlier work this paper cites.
Domain adaptation by using causal inference to predict invariant conditional distributions
Original
Sara Magliacane, Thijs van Ommen, Tom Claassen, Stephan Bongers, Philip Versteeg, and Joris M. Mooij · 2018
Earlier work this paper cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Earlier work this paper cites.
The Book of Why: The New Science of Cause and Effect
Judea Pearl and Dana Mackenzie · 2018
Earlier work this paper cites.
Invariant models for causal transfer learning
Mateo Rojas-Carulla, Bernhard Scholkopf, Richard Turner, and Jonas Peters · 2018
Earlier work this paper cites.
Mathematical Theory of Bayesian Statistics
Sumio Watanabe · 2018
Earlier work this paper cites.
An introduction to systems biology: design principles of biological circuits
Uri Alon · 2019
Earlier work this paper cites.
A meta-transfer objective for learning to disentangle causal mechanisms
Original
Yoshua Bengio, Tristan Deleu, Nasim Rahaman, Rosemary Ke, Sébastien Lachapelle, Olexa Bilaniuk, Anirudh Goyal, and Christopher Pal · 2019
Earlier work this paper cites.
Counterfactuals uncover the modular structure of deep generative models
Original
Michel Besserve, Arash Mehrjou, Rémy Sun, and Bernhard Schölkopf · 2019
Earlier work this paper cites.
Activation atlas
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Earlier work this paper cites.
What is one grain of sand in the desert? analyzing individual neurons in deep nlp models
Fahim Dalvi, Nadir Durrani, Hassan Sajjad, Yonatan Belinkov, Anthony Bau, and James Glass · 2019
Earlier work this paper cites.
Adversarial robustness as a prior for learned representations
Original
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Original
Jonathan Frankle and Michael Carbin · 2019
Earlier work this paper cites.
On the uniqueness and stability of dictionaries for sparse representation of noisy signals
Charles J. Garfinkle and Christopher J. Hillar · 2019
Earlier work this paper cites.
Accelerating convolutional neural networks via activation map compression
Original
Georgios Georgiadis · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D. Manning · 2019
Earlier work this paper cites.
Risks from learned optimization in advanced machine learning systems
Original
Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, and Scott Garrabrant · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
Original
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
Similarity of neural network representations revisited
Original
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Earlier work this paper cites.
Challenging common assumptions in the unsupervised learning of disentangled representations
Original
Francesco Locatello, Stefan Bauer, Mario Lucic, Gunnar Rätsch, Sylvain Gelly, Bernhard Schölkopf, and Olivier Bachem · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Earlier work this paper cites.
Sgd on neural networks learns functions of increasing complexity
Original
Preetum Nakkiran, Gal Kaplun, Dimitris Kalimeris, Tristan Yang, Benjamin L. Edelman, Fred Zhang, and Boaz Barak · 2019
Earlier work this paper cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Original
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 2019
Earlier work this paper cites.
Semi-supervised learning, causality, and the conditional cluster assumption
Original
Julius von Kügelgen, M. Loog, A. Mey, and B. Scholkopf · 2019
Earlier work this paper cites.
Interpreting cnns via decision trees
Original
Quanshi Zhang, Yu Yang, Haotian Ma, and Ying Nian Wu · 2019
Earlier work this paper cites.
Curve detectors
Nick Cammarata, Gabriel Goh, Shan Carter, Ludwig Schubert, Michael Petrov, and Chris Olah · 2020
Earlier work this paper cites.
Analyzing individual neurons in pre-trained language models
Original
Nadir Durrani, Hassan Sajjad, Fahim Dalvi, and Yonatan Belinkov · 2020
Earlier work this paper cites.
Neuron shapley: Discovering the responsible neurons
Original
Amirata Ghorbani and James Zou · 2020
Earlier work this paper cites.
Recurrent independent mechanisms
Original
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2020
Earlier work this paper cites.
Let’s agree to agree: Neural networks share classification order on real datasets
Original
Guy Hacohen, Leshem Choshen, and Daphna Weinshall · 2020
Earlier work this paper cites.
Understanding rl vision
Jacob Hilton, Nick Cammarata, Shan Carter, Gabriel Goh, and Chris Olah · 2020
Earlier work this paper cites.
An overview of 11 proposals for building safe advanced ai
Original
Evan Hubinger · 2020
Earlier work this paper cites.
Towards faithfully interpretable nlp systems: How should we define and evaluate faithfulness?
Original
Alon Jacovi and Yoav Goldberg · 2020
Earlier work this paper cites.
Towards falsifiable interpretability research
Original
Matthew L. Leavitt and Ari Morcos · 2020
Earlier work this paper cites.
Compositional explanations of neurons
Original
Jesse Mu and Jacob Andreas · 2020
Earlier work this paper cites.
interpreting gpt: the logit lens
nostalgebraist · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Shauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton, and Yoav Goldberg · 2020
Earlier work this paper cites.
Logical neural networks
Original
Ryan Riegel, Alexander Gray, Francois Luus, Naweed Khan, Ndivhuwo Makondo, Ismail Yunus Akhalwaya, Haifeng Qian, Ronald Fagin, Francisco Barahona, Udit Sharma, Shajith Ikbal, Hima Karanam, Sumit Neelam, Ankita Likhyani, and Santosh Srivastava · 2020
Earlier work this paper cites.
Investigating gender bias in language models using causal mediation analysis
Jesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian, Daniel Nevo, Yaron Singer, and Stuart Shieber · 2020
Earlier work this paper cites.
Information-theoretic probing with minimum description length
Original
Elena Voita and Ivan Titov · 2020
Earlier work this paper cites.
Blimp: The benchmark of linguistic minimal pairs for english
Alex Warstadt, Alicia Parrish, Haokun Liu, Anhad Mohananey, Wei Peng, Sheng-Fu Wang, and Samuel R. Bowman · 2020
Earlier work this paper cites.
Revisiting model stitching to compare neural representations
Original
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Original
Yonatan Belinkov · 2021
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell · 2021
Earlier work this paper cites.
An interpretability illusion for bert
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Vi’egas, and M. Wattenberg · 2021
Earlier work this paper cites.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2021
Earlier work this paper cites.
Low-complexity probing via finding subnetworks
Original
Steven Cao, Victor Sanh, and Alexander M. Rush · 2021
Earlier work this paper cites.
Robust feature-level adversaries are interpretability tools
Original
Stephen Casper, Max Nadeau, Dylan Hadfield-Menell, and Gabriel Kreiman · 2021
Earlier work this paper cites.
Eliciting latent knowledge , January 2021
Paul Christiano, Ajeya Cotra, and Mark Xu · 2021
Earlier work this paper cites.
Explaining by removing: a unified framework for model explanation
Original
Ian C. Covert, Scott Lundberg, and Su-In Lee · 2021
Earlier work this paper cites.
Fighting adversarial images with interpretable gradients
Keke Du, Shan Chang, Huixiang Wen, and Hao Zhang · 2021
Earlier work this paper cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Original
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Earlier work this paper cites.
Causalm: Causal model explanation through counterfactual language models
Original
Amir Feder, Nadav Oved, Uri Shalit, and Roi Reichart · 2021
Earlier work this paper cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Earlier work this paper cites.
Elite backprop: Training sparse interpretable neurons
Theodoros Kasioumis, Joe Townsend, and Hiroya Inakoshi · 2021
Earlier work this paper cites.
Systematic evaluation of causal discovery in visual model based reinforcement learning
Original
Nan Rosemary Ke, Aniket Didolkar, Sarthak Mittal, Anirudh Goyal, Guillaume Lajoie, Stefan Bauer, Danilo Rezende, Yoshua Bengio, Michael Mozer, and Christopher Pal · 2021
Earlier work this paper cites.
Implicit representations of meaning in neural language models
Belinda Z. Li, Maxwell Nye, and Jacob Andreas · 2021
Earlier work this paper cites.
Probing across time: What does roberta know and when?
Original
Leo Z. Liu, Yizhong Wang, Jungo Kasai, Hannaneh Hajishirzi, and Noah A. Smith · 2021
Earlier work this paper cites.
Towards causal representation learning
Original
Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, and Yoshua Bengio · 2021
Earlier work this paper cites.
Learning to synthesize programs as interpretable and generalizable policies
Original
Dweep Trivedi, Jesse Zhang, Shao-Hua Sun, and Joseph J. Lim · 2021
Earlier work this paper cites.
Branch specialization
Chelsea Voss, Gabriel Goh, Nick Cammarata, Michael Petrov, Ludwig Schubert, and Chris Olah · 2021
Earlier work this paper cites.
Sparse attention with linear units
Original
Biao Zhang, Ivan Titov, and Rico Sennrich · 2021
Earlier work this paper cites.
How well do feature visualizations support causal understanding of cnn activations?
Original
Roland S. Zimmermann, Judy Borowski, Robert Geirhos, Matthias Bethge, Thomas S. A. Wallis, and Wieland Brendel · 2021
Earlier work this paper cites.
Vl-interpret: An interactive visualization tool for interpreting vision-language transformers
Estelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu, Chenfei Wu, Nan Duan, and Vasudev Lal · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Original
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Original
Boaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Earlier work this paper cites.
Interpreting neural networks through the polytope lens
Original
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, and Connor Leahy · 2022
Earlier work this paper cites.
Broken neural scaling laws
Original
Ethan Caballero, Kshitij Gupta, Irina Rish, and David Krueger · 2022
Earlier work this paper cites.
Diagnostics for deep neural networks with automated copy/paste attacks
Original
Stephen Casper, Kaivalya Hariharan, and Dylan Hadfield-Menell · 2022
Earlier work this paper cites.
Causal scrubbing: a method for rigorously testing interpretability hypotheses [redwood research]
Lawrence Chan, Adrià Garriga-alonso, Nicholas Goldowsky-Dill, ryan_greenblatt, jenny, Ansh Radhakrishnan, Buck, and Nate Thomas · 2022
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, Baobao Chang, and Furu Wei · 2022
Earlier work this paper cites.
Analyzing transformers in embedding space
Original
Guy Dar, Mor Geva, Ankit Gupta, and Jonathan Berant · 2022
Earlier work this paper cites.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Original
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Earlier work this paper cites.
X-risk analysis for ai research
Original
Dan Hendrycks and Mantas Mazeika · 2022
Earlier work this paper cites.