Fetching the paper…
Reading the bibliography…
The last decade of machine learning has seen drastic increases in scale and capabilities.
Scaling Laws for Neural Language Models, 2020
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Representation learning: A review and new perspectives
Yoshua Bengio, Aaron Courville, and Pascal Vincent · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2014
Earlier work this paper cites.
Superintelligence: Paths, Dangers, Strategies
N. Bostrom · 2014
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
Been Kim, Cynthia Rudin, and Julie A Shah · 2014
Earlier work this paper cites.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba · 2014
Earlier work this paper cites.
Distributional vectors encode referential attributes
Abhijeet Gupta, Gemma Boleda, Marco Baroni, and Sebastian Padó · 2015
Earlier work this paper cites.
What’s in an embedding? analyzing word embeddings through multilingual evaluation
Arne Köhn · 2015
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John Hopcroft · 2015
Earlier work this paper cites.
Understanding deep image representations by inverting them
Aravindh Mahendran and Andrea Vedaldi · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Data-free parameter pruning for deep neural networks
Suraj Srinivas and R Venkatesh Babu · 2015
Earlier work this paper cites.
Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Yossi Adi, Einat Kermany, Yonatan Belinkov, Ofer Lavi, and Yoav Goldberg · 2016
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes
Guillaume Alain and Yoshua Bengio · 2016
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Earlier work this paper cites.
Infogan: Interpretable representation learning by information maximizing generative adversarial nets
Xi Chen, Yan Duan, Rein Houthooft, John Schulman, Ilya Sutskever, and Pieter Abbeel · 2016
Earlier work this paper cites.
Probing for semantic evidence of composition by means of simple classification tasks
Allyson Ettinger, Ahmed Elgohary, and Philip Resnik · 2016
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell · 2016
Earlier work this paper cites.
beta-vae: Learning basic visual concepts with a constrained variational framework
Irina Higgins, Loic Matthey, Arka Pal, Christopher Burgess, Xavier Glorot, Matthew Botvinick, Shakir Mohamed, and Alexander Lerchner · 2016
Earlier work this paper cites.
Network trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hengyuan Hu, Rui Peng, Yu-Wing Tai, and Chi-Keung Tang · 2016
Earlier work this paper cites.
Future progress in artificial intelligence: A survey of expert opinion
Vincent C Müller and Nick Bostrom · 2016
Earlier work this paper cites.
Synthesizing the preferred inputs for neurons in neural networks via deep generator networks
Anh Nguyen, Alexey Dosovitskiy, Jason Yosinski, Thomas Brox, and Jeff Clune · 2016
Earlier work this paper cites.
Plug & play generative networks: Conditional iterative generation of images in latent space
Anh Nguyen, Jason Yosinski, Yoshua Bengio, Alexey Dosovitskiy, and Jeff Clune · 2016
Earlier work this paper cites.
Multifaceted feature visualization: Uncovering the different types of features learned by each neuron in deep neural networks, 2016
Anh Nguyen, Jason Yosinski, and Jeff Clune · 2016
Earlier work this paper cites.
Andrei A Rusu, Neil C Rabinowitz, Guillaume Desjardins, Hubert Soyer, James Kirkpatrick, Koray Kavukcuoglu, Razvan Pascanu, and Raia Hadsell · 2016
Earlier work this paper cites.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2017
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations, 2017
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Towards interpretable deep neural networks by leveraging adversarial examples
Yinpeng Dong, Hang Su, Jun Zhu, and Fan Bao · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Variational inference of disentangled latent concepts from unlabeled observations
Abhishek Kumar, Prasanna Sattigeri, and Avinash Balakrishnan · 2017
Earlier work this paper cites.
Interactive visualization and manipulation of attention-based neural machine translation
Jaesong Lee, Joong-Hwi Shin, and Jun-Seok Kim · 2017
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
The neural lasso: Local linear sparsity for interpretable explanations
Andrew Ross, Isaac Lage, and Finale Doshi-Velez · 2017
Earlier work this paper cites.
Dynamic routing between capsules
Sara Sabour, Nicholas Frosst, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Life 3.0: Being human in the age of artificial intelligence
Max Tegmark · 2017
Earlier work this paper cites.
Lifelong learning with dynamically expandable networks
Jaehong Yoon, Eunho Yang, Jeongtae Lee, and Sung Ju Hwang · 2017
Earlier work this paper cites.
Continual learning through synaptic intelligence
Friedemann Zenke, Ben Poole, and Surya Ganguli · 2017
Earlier work this paper cites.
Peeking inside the black-box: a survey on explainable artificial intelligence (xai)
Amina Adadi and Mohammed Berrada · 2018
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Generating post-hoc rationales of deep visual classification decisions
Zeynep Akata, Lisa Anne Hendricks, Stephan Alaniz, and Trevor Darrell · 2018
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez Melis and Tommi Jaakkola · 2018
Earlier work this paper cites.
Identifying and controlling important neurons in neural machine translation
Anthony Bau, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani, Fahim Dalvi, and James Glass · 2018
Earlier work this paper cites.
Gan dissection: Visualizing and understanding generative adversarial networks, 2018
David Bau, Jun-Yan Zhu, Hendrik Strobelt, Bolei Zhou, Joshua B. Tenenbaum, William T. Freeman, and Antonio Torralba · 2018
Earlier work this paper cites.
Understanding disentangling in beta-vae
Christopher P Burgess, Irina Higgins, Arka Pal, Loic Matthey, Nick Watters, Guillaume Desjardins, and Alexander Lerchner · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Lateral inhibition-inspired convolutional neural network for visual attention and saliency detection
Chunshui Cao, Yongzhen Huang, Zilei Wang, Liang Wang, Ninglong Xu, and Tieniu Tan · 2018
Earlier work this paper cites.
Isolating sources of disentanglement in variational autoencoders
Ricky TQ Chen, Xuechen Li, Roger B Grosse, and David K Duvenaud · 2018
Earlier work this paper cites.
What you can cram into a single vector: Probing sentence embeddings for linguistic properties
Alexis Conneau, German Kruszewski, Guillaume Lample, Loïc Barrault, and Marco Baroni · 2018
Earlier work this paper cites.
Net2vec: Quantifying and explaining how concepts are encoded by filters in deep neural networks
Ruth Fong and Andrea Vedaldi · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness, 2018
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal · 2018
Earlier work this paper cites.
Grounding visual explanations
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al · 2018
Earlier work this paper cites.
Disentangling by factorising
Hyunjik Kim and Andriy Mnih · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata · 2018
Earlier work this paper cites.
Modular networks: Learning to decompose neural computation
Louis Kirsch, Julius Kunze, and David Barber · 2018
Earlier work this paper cites.
Deep learning for case-based reasoning through prototypes: A neural network that explains its predictions
Oscar Li, Hao Liu, Chaofan Chen, and Cynthia Rudin · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Earlier work this paper cites.
Visual interrogation of attention-based models for natural language inference and machine comprehension
Shusen Liu, Tao Li, Zhimin Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer · 2018
Earlier work this paper cites.
Packnet: Adding multiple tasks to a single network by iterative pruning
Arun Mallya and Svetlana Lazebnik · 2018
Earlier work this paper cites.
Insights on representational similarity in neural networks with canonical correlation
Ari Morcos, Maithra Raghu, and Samy Bengio · 2018
Earlier work this paper cites.
On the importance of single directions for generalization, 2018
Ari S. Morcos, David G. T. Barrett, Neil C. Rabinowitz, and Matthew Botvinick · 2018
Earlier work this paper cites.
The building blocks of interpretability
Chris Olah, Arvind Satyanarayan, Ian Johnson, Shan Carter, Ludwig Schubert, Katherine Ye, and Alexander Mordvintsev · 2018
Earlier work this paper cites.
Evaluation of sentence embeddings in downstream and linguistic probing tasks
Christian S Perone, Roberto Silveira, and Thomas S Paula · 2018
Earlier work this paper cites.
Improving the adversarial robustness and interpretability of deep neural networks by regularizing their input gradients
Andrew Ross and Finale Doshi-Velez · 2018
Earlier work this paper cites.
Overcoming catastrophic forgetting with hard attention to the task
Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou · 2018
Earlier work this paper cites.
Analysing neural network topologies: a game theoretic approach
Julian Stier, Gabriele Gianini, Michael Granitzer, and Konstantin Ziegler · 2018
Earlier work this paper cites.
S eq 2s eq-v is: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M Rush · 2018
Earlier work this paper cites.
Spine: Sparse interpretable neural embeddings
Anant Subramanian, Danish Pruthi, Harsh Jhamtani, Taylor Berg-Kirkpatrick, and Eduard Hovy · 2018
Earlier work this paper cites.
Why the failure? how adversarial examples can provide insights for interpretable machine learning
Richard Tomsett, Amy Widdicombe, Tianwei Xing, Supriyo Chakraborty, Simon Julier, Prudhvi Gurram, Raghuveer Rao, and Mani Srivastava · 2018
Earlier work this paper cites.
Robustness may be at odds with accuracy
Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry · 2018
Earlier work this paper cites.
Programmatically interpretable reinforcement learning
Abhinav Verma, Vijayaraghavan Murali, Rishabh Singh, Pushmeet Kohli, and Swarat Chaudhuri · 2018
Earlier work this paper cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Liwei Wang, Lunjia Hu, Jiayuan Gu, Zhiqiang Hu, Yue Wu, Kun He, and John Hopcroft · 2018
Earlier work this paper cites.
Interpret neural networks by identifying critical data routing paths
Yulong Wang, Hang Su, Bo Zhang, and Xiaolin Hu · 2018
Earlier work this paper cites.
Modular representation of layered neural networks
Chihiro Watanabe, Kaoru Hiramatsu, and Kunio Kashino · 2018
Earlier work this paper cites.
Beyond sparsity: Tree regularization of deep models for interpretability
Mike Wu, Michael Hughes, Sonali Parbhoo, Maurizio Zazzi, Volker Roth, and Finale Doshi-Velez · 2018
Earlier work this paper cites.
Visual interpretability for deep learning: a survey
Quan-shi Zhang and Song-Chun Zhu · 2018
Earlier work this paper cites.
Quanshi Zhang, Xin Wang, Ruiming Cao, Ying Nian Wu, Feng Shi, and Song-Chun Zhu · 2018
Earlier work this paper cites.
Interpretable convolutional neural networks
Quanshi Zhang, Ying Nian Wu, and Song-Chun Zhu · 2018
Earlier work this paper cites.
Interpretable basis decomposition for visual explanation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Earlier work this paper cites.
Revisiting the importance of individual units in cnns via ablation
Bolei Zhou, Yiyou Sun, David Bau, and Antonio Torralba · 2018
Earlier work this paper cites.
Uncertainty-based continual learning with adaptive regularization
Hongjoon Ahn, Sungmin Cha, Donggyu Lee, and Taesup Moon · 2019
Earlier work this paper cites.
Task-free continual learning
Rahaf Aljundi, Klaas Kelchtermans, and Tinne Tuytelaars · 2019
Earlier work this paper cites.
A review of modularization techniques in artificial neural networks
Mohammed Amer and Tomás Maul · 2019
Earlier work this paper cites.
Gradient-based attribution methods
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2019
Earlier work this paper cites.
Make up your mind! adversarial generation of inconsistent natural language explanations
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, and Phil Blunsom · 2019
Earlier work this paper cites.
Exploring neural networks with activation atlases
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Earlier work this paper cites.
Frivolous units: Wider networks are not really that wide
Stephen Casper, Xavier Boix, Vanessa D’Amario, Ling Guo, Martin Schrimpf, Kasper Vinken, and Gabriel Kreiman · 2019
Earlier work this paper cites.
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Earlier work this paper cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
Maria De-Arteaga, Alexey Romanov, Hanna Wallach, Jennifer Chayes, Christian Borgs, Alexandra Chouldechova, Sahin Geyik, Krishnaram Kenthapadi, and Adam Tauman Kalai · 2019
Earlier work this paper cites.
An effective hit-or-miss layer favoring feature interpretation as learned prototypes deformations
Adrien Deliège, Anthony Cioppa, and Marc Van Droogenbroeck · 2019
Earlier work this paper cites.
Eraser: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C Wallace · 2019
Earlier work this paper cites.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel · 2019
Earlier work this paper cites.
Adversarial robustness as a prior for learned representations
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
Learning perceptually-aligned representations via adversarial robustness
Logan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
On the connection between adversarial robustness and saliency map interpretability
Christian Etmann, Sebastian Lunz, Peter Maass, and Carola-Bibiane Schönlieb · 2019
Earlier work this paper cites.
Deep neural model inspection and comparison via functional neuron pathways
James Fiacco, Samridhi Choudhary, and Carolyn Penstein Rosé · 2019
Cited alongside, same era.
Scaleable input gradient regularization for adversarial robustness
Chris Finlay and Adam M Oberman · 2019
Cited alongside, same era.
Dissecting pruned neural networks
Jonathan Frankle and David Bau · 2019
Cited alongside, same era.
Recurrent independent mechanisms
Anirudh Goyal, Alex Lamb, Jordan Hoffmann, Shagun Sodhani, Sergey Levine, Yoshua Bengio, and Bernhard Schölkopf · 2019
Cited alongside, same era.
Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems
Neural execution engines: Learning to execute subroutines
Yujun Yan, Kevin Swersky, Danai Koutra, Parthasarathy Ranganathan, and Milad Hashemi · 2020
Later among the works it cites.
Extraction of an explanatory graph to interpret a cnn
Quanshi Zhang, Xin Wang, Ruiming Cao, Ying Nian Wu, Feng Shi, and Song-Chun Zhu · 2020
Later among the works it cites.
Masking as an efficient alternative to finetuning for pretrained language models
Mengjie Zhao, Tao Lin, Fei Mi, Martin Jaggi, and Hinrich Schütze · 2020
Later among the works it cites.
Lirex: Augmenting language inference with relevant explanation
Xinyan Zhao and VG Vydiswaran · 2020
Later among the works it cites.
One network fits all? modular versus monolithic task formulations in neural networks
Atish Agarwala, Abhimanyu Das, Brendan Juba, Rina Panigrahy, Vatsal Sharan, Xin Wang, and Qiuyi Zhang · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
Sarthak Jain and Byron C Wallace · 2019
Cited alongside, same era.
Self-assembling modular networks for interpretable multi-hop reasoning
Yichen Jiang and Mohit Bansal · 2019
Cited alongside, same era.
Are perceptually-aligned gradients a general property of robust classifiers?
Simran Kaur, Jeremy Cohen, and Zachary C Lipton · 2019
Cited alongside, same era.
Bridging adversarial robustness and gradient interpretability
Beomsu Kim, Junghoon Seo, and Taegyun Jeon · 2019
Cited alongside, same era.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton · 2019
Cited alongside, same era.
Unsupervised learning by competing hidden units
Dmitry Krotov and John J Hopfield · 2019
Cited alongside, same era.
Later among the works it cites.
On the pitfalls of analyzing individual neurons in language models, 2021
Omer Antverg and Yonatan Belinkov · 2021
Later among the works it cites.
Revisiting model stitching to compare neural representations
Yamini Bansal, Preetum Nakkiran, and Boaz Barak · 2021
Later among the works it cites.
Representation topology divergence: A method for comparing neural network representations
Serguei Barannikov, Ilya Trofimov, Nikita Balabin, and Evgeny Burnaev · 2021
Later among the works it cites.
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova · 2021
Later among the works it cites.
Extreme sparsity gives rise to functional specialization
Gabriel Béna and Dan FM Goodman · 2021
Later among the works it cites.
An interpretability illusion for bert, 2021
Tolga Bolukbasi, Adam Pearce, Ann Yuan, Andy Coenen, Emily Reif, Fernanda Viégas, and Martin Wattenberg · 2021
Later among the works it cites.
Toward a unified framework for debugging gray-box models
Andrea Bontempelli, Fausto Giunchiglia, Andrea Passerini, and Stefano Teso · 2021
Later among the works it cites.
Robust feature level adversaries are interpretability tools
Stephen Casper, Max Nadeau, Dylan Hadfield-Menell, and Gabriel Kreiman · 2021
Later among the works it cites.
Transformer interpretability beyond attention visualization
Hila Chefer, Shir Gur, and Lior Wolf · 2021
Later among the works it cites.
Similarity and matching of neural network representations
Adrián Csiszárik, Péter Kőrösi-Szabó, Ákos Matszangosz, Gergely Papp, and Dániel Varga · 2021
Later among the works it cites.
Knowledge neurons in pretrained transformers
Damai Dai, Li Dong, Yaru Hao, Zhifang Sui, and Furu Wei · 2021
Later among the works it cites.
Grounding representation similarity with statistical testing
Frances Ding, Jean-Stanislas Denain, and Jacob Steinhardt · 2021
Later among the works it cites.
Fighting adversarial images with interpretable gradients
Keke Du, Shan Chang, Huixiang Wen, and Hao Zhang · 2021
Later among the works it cites.
Topkconv: Increased adversarial robustness through deeper interpretability
Henry Eigen and Amir Sadovnik · 2021
Later among the works it cites.
Amnesic probing: Behavioral explanation with amnesic counterfactuals
Yanai Elazar, Shauli Ravfogel, Alon Jacovi, and Yoav Goldberg · 2021
Later among the works it cites.
A mathematical framework for transformer circuits, 2021
N Elhage, N Nanda, C Olsson, T Henighan, N Joseph, B Mann, A Askell, Y Bai, A Chen, T Conerly, et al · 2021
Later among the works it cites.
Clusterability in neural networks
Daniel Filan, Stephen Casper, Shlomi Hod, Cody Wild, Andrew Critch, and Stuart Russell · 2021
Later among the works it cites.
Causal analysis of syntactic agreement mechanisms in neural language models
Matthew Finlayson, Aaron Mueller, Sebastian Gehrmann, Stuart Shieber, Tal Linzen, and Yonatan Belinkov · 2021
Later among the works it cites.
Design and evaluation of a multi-domain trojan detection method on deep neural networks
Yansong Gao, Yeonjae Kim, Bao Gia Doan, Zhi Zhang, Gongxuan Zhang, Surya Nepal, Damith Ranasinghe, and Hyoungshick Kim · 2021
Later among the works it cites.
Causal abstractions of neural networks
Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts · 2021
Later among the works it cites.
Multimodal neurons in artificial neural networks
Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah · 2021
Later among the works it cites.
Self-attention attribution: Interpreting information interactions inside transformer
Yaru Hao, Li Dong, Furu Wei, and Ke Xu · 2021
Later among the works it cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2021
Later among the works it cites.
Quantifying local specialization in deep neural networks
Shlomi Hod, Stephen Casper, Daniel Filan, Cody Wild, Andrew Critch, and Stuart Russell · 2021
Later among the works it cites.
Adrian Hoffmann, Claudio Fanconi, Rahul Rade, and Jonas Kohler · 2021
Later among the works it cites.
Probing neural networks with t-sne, class-specific projections and a guided tour
Christopher R Hoyt and Art B Owen · 2021
Later among the works it cites.
Automating auditing: An ambitious concrete technical research proposal, Aug 2021
Evan Hubinger · 2021
Later among the works it cites.
3db: A framework for debugging computer vision models, 2021
Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, Ashish Kapoor, and Aleksander Madry · 2021
Later among the works it cites.
Implicit representations of meaning in neural language models
Belinda Z Li, Maxwell Nye, and Jacob Andreas · 2021
Later among the works it cites.
Probing multimodal embeddings for linguistic properties: the visual-semantic case
Adam Dahlgren Lindström, Suna Bensch, Johanna Björklund, and Frank Drewes · 2021
Later among the works it cites.
Semantic bottlenecks: Quantifying and improving inspectability of deep representations
Max Losch, Mario Fritz, and Bernt Schiele · 2021
Later among the works it cites.
Promises and pitfalls of black-box concept learning models
Anita Mahinpei, Justin Clark, Isaac Lage, Finale Doshi-Velez, and Weiwei Pan · 2021
Later among the works it cites.
Is sparse attention more interpretable?
Clara Meister, Stefan Lazov, Isabelle Augenstein, and Ryan Cotterell · 2021
Later among the works it cites.
Probing tasks under pressure
Alessio Miaschi, Chiara Alzetta, Dominique Brunato, Felice Dell’Orletta, and Giulia Venturi · 2021
Later among the works it cites.
Identifiable variational autoencoders via sparse decoding
Gemma E Moran, Dhanya Sridhar, Yixin Wang, and David M Blei · 2021
Later among the works it cites.
Robust explainability: A tutorial on gradient-based attribution methods for deep neural networks
Ian E Nielsen, Dimah Dera, Ghulam Rasool, Nidhal Bouaynaya, and Ravi P Ramachandran · 2021
Later among the works it cites.
An empirical study on the relation between network interpretability and adversarial robustness
Adam Noack, Isaac Ahern, Dejing Dou, and Boyang Li · 2021
Later among the works it cites.
Optimism in the face of adversity: Understanding and improving deep learning through adversarial robustness
Guillermo Ortiz-Jiménez, Apostolos Modas, Seyed-Mohsen Moosavi-Dezfooli, and Pascal Frossard · 2021
Later among the works it cites.
Weight banding
Michael Petrov, Chelsea Voss, Ludwig Schubert, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Later among the works it cites.
Do vision transformers see like convolutional neural networks?
Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang, and Alexey Dosovitskiy · 2021
Later among the works it cites.
Towards axiomatic, hierarchical, and symbolic explanation for deep models
Jie Ren, Mingjie Li, Qihan Ren, Huiqi Deng, and Quanshi Zhang · 2021
Later among the works it cites.
Attention-based interpretability with concept transformers
Mattia Rigotti, Christoph Miksovic, Ioana Giurgiu, Thomas Gschwind, and Paolo Scotton · 2021
Later among the works it cites.
Interpretable image classification with differentiable prototypes assignment
Dawid Rymarczyk, Łukasz Struski, Michał Górszczak, Koryna Lewandowska, Jacek Tabor, and Bartosz Zieliński · 2021
Later among the works it cites.
Neuron-level interpretation of deep nlp models: A survey
Hassan Sajjad, Nadir Durrani, and Fahim Dalvi · 2021
Later among the works it cites.
Explaining deep neural networks and beyond: A review of methods and applications
Wojciech Samek, Grégoire Montavon, Sebastian Lapuschkin, Christopher J Anders, and Klaus-Robert Müller · 2021
Later among the works it cites.
Editing a classifier by rewriting its prediction rules
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Later among the works it cites.
Anindya Sarkar, Anirban Sarkar, Sowrya Gali, and Vineeth N Balasubramanian · 2021
Later among the works it cites.
Explaining neural networks by decoding layer activations
Johannes Schneider and Michalis Vlachos · 2021
Later among the works it cites.
High-low frequency detectors
Ludwig Schubert, Chelsea Voss, Nick Cammarata, Gabriel Goh, and Chris Olah · 2021
Later among the works it cites.
\emphParameter Counts in Machine Learning, June 2021
Jaime Sevilla, Pablo Villalobos, and Juan Felipe Cerón · 2021
Later among the works it cites.
Marco Valentino, Ian Pratt-Hartmann, and André Freitas · 2021
Later among the works it cites.
Visualizing weights
Chelsea Voss, Nick Cammarata, Gabriel Goh, Michael Petrov, Ludwig Schubert, Ben Egan, Swee Kiat Lim, and Chris Olah · 2021
Later among the works it cites.
Branch specialization
Chelsea Voss, Gabriel Goh, Nick Cammarata, Michael Petrov, Ludwig Schubert, and Chris Olah · 2021
Later among the works it cites.
Leveraging sparse linear layers for debuggable deep networks
Eric Wong, Shibani Santurkar, and Aleksander Madry · 2021
Later among the works it cites.
Deep neural network compression through interpretability-based filter pruning
Kaixuan Yao, Feilong Cao, Yee Leung, and Jiye Liang · 2021
Later among the works it cites.
Pruning by explaining: A novel criterion for deep neural network pruning
Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Alexander Binder, Simon Wiedemann, Klaus-Robert Müller, and Wojciech Samek · 2021
Later among the works it cites.
Hypothesis-driven online video stream learning with augmented memory
Mengmi Zhang, Rohil Badkundri, Morgan B Talbot, Rushikesh Zawar, and Gabriel Kreiman · 2021
Later among the works it cites.
Topological detection of trojaned neural networks
Songzhu Zheng, Yikai Zhang, Hubert Wagner, Mayank Goswami, and Chao Chen · 2021
Later among the works it cites.
Meaningfully debugging model mistakes using conceptual counterfactual explanations
Abubakar Abid, Mert Yuksekgonul, and James Zou · 2022
Closest in time.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Closest in time.
The values encoded in machine learning research
Abeba Birhane, Pratyusha Kalluri, Dallas Card, William Agnew, Ravit Dotan, and Michelle Bao · 2022
Closest in time.
Interpreting neural networks through the polytope lens
Sid Black, Lee Sharkey, Leo Grinsztajn, Eric Winsor, Dan Braun, Jacob Merizian, Kip Parker, Carlos Ramón Guevara, Beren Millidge, Gabriel Alfour, et al · 2022
Closest in time.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Closest in time.
Diagnostics for deep neural networks with automated copy/paste attacks
Stephen Casper, Kaivalya Hariharan, and Dylan Hadfield-Menell · 2022
Closest in time.
Graphical clusterability and local specialization in deep neural networks
Stephen Casper, Shlomi Hod, Daniel Filan, Cody Wild, Andrew Critch, and Stuart Russell · 2022
Closest in time.
Dall-eval: Probing the reasoning skills and social biases of text-to-image generative transformers
Jaemin Cho, Abhay Zala, and Mohit Bansal · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2022
Closest in time.
Eliciting latent knowledge: How to tell if your eyes deceive you
Paul Christiano, Ajeya Cotra, and Mark Xu · 2022
Closest in time.
Deconfounded representation similarity for comparison of neural networks
Tianyu Cui, Yogesh Kumar, Pekka Marttinen, and Samuel Kaski · 2022
Closest in time.
Auditing visualizations: Transparency methods struggle to detect anomalous behavior, 2022
Jean-Stanislas Denain and Jacob Steinhardt · 2022
Closest in time.
Softmax linear units
Nelson Elhage, Tristan Hume, Catherine Olsson, Neel Nanda, Tom Henighan, Scott Johnston, Sheer ElShowk, Nicholas Joseph, Nova DasSarma, Ben Mann, Danny Hernandez, Amanda Askell, Kamal Ndousse, Andy Jones, Dawn Drain, Anna Chen, Yuntao Bai, Deep Ganguli, Liane Lovitt, Zac Hatfield-Dodds, Jackson Kernion, Tom Conerly, Shauna Kravec, Stanislav Fort, Saurav Kadavath, Josh Jacobson, Eli Tran-Johnson, Jared Kaplan, Jack Clark, Tom Brown, Sam McCandlish, Dario Amodei, and Christopher Olah · 2022
Closest in time.
Nelson Elhage, Tristan Hume, Catherine Olsson, Nicholas Schiefer, Tom Henighan, Shauna Kravec, Zac Hatfield-Dodds, Robert Lasenby, Dawn Drain, Carol Chen, et al · 2022
Closest in time.
Protoformer: Embedding prototypes for transformers
Ashkan Farhangi, Ning Sui, Nan Hua, Haiyan Bai, Arthur Huang, and Zhishan Guo · 2022
Closest in time.
Attribution-based explanations that provide recourse cannot be robust
Hidde Fokkema, Rianne de Heide, and Tim van Erven · 2022
Closest in time.
Inducing causal structure for interpretable neural networks
Atticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner, Elisa Kreiss, Thomas Icard, Noah Goodman, and Christopher Potts · 2022
Closest in time.
Mor Geva, Avi Caciularu, Guy Dar, Paul Roit, Shoval Sadde, Micah Shlain, Bar Tamir, and Yoav Goldberg · 2022
Closest in time.
Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space
Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg · 2022
Closest in time.
On the symmetries of deep learning models and their internal representations
Charles Godfrey, Davis Brown, Tegan Emerson, and Henry Kvinge · 2022
Closest in time.
Pruning for interpretable, feature-preserving circuits in cnns
Chris Hamblin, Talia Konkle, and George Alvarez · 2022
Closest in time.
Exploring linear feature disentanglement for neural networks
Tiantian He, Zhibin Li, Yongshun Gong, Yazhou Yao, Xiushan Nie, and Yilong Yin · 2022
Closest in time.
Towards benchmarking explainable artificial intelligence methods, 2022
Lars Holmberg · 2022
Closest in time.
Distilling model failures as directions in latent space, 2022
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry · 2022
Closest in time.
On the robustness of explanations of deep neural network models: A survey
Amlan Jyoti, Karthik Balaji Ganesh, Manoj Gayala, Nandita Lakshmi Tunuguntla, Sandesh Kamath, and Vineeth N Balasubramanian · 2022
Closest in time.
Language models (mostly) know what they know, 2022
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez, Nicholas Schiefer, Zac Hatfield Dodds, Nova DasSarma, Eli Tran-Johnson, Scott Johnston, Sheer El-Showk, Andy Jones, Nelson Elhage, Tristan Hume, Anna Chen, Yuntao Bai, Sam Bowman, Stanislav Fort, Deep Ganguli, Danny Hernandez, Josh Jacobson, Jackson Kernion, Shauna Kravec, Liane Lovitt, Kamal Ndousse, Catherine Olsson, Sam Ringer, Dario Amodei, Tom Brown, Jack Clark, Nicholas Joseph, Ben Mann, Sam McCandlish, Chris Olah, and Jared Kaplan · 2022
Closest in time.
Clustering units in neural networks: upstream vs downstream information
Richard D Lange, David S Rolnick, and Konrad P Kording · 2022
Closest in time.
Probing via prompting, 2022
Jiaoda Li, Ryan Cotterell, and Mrinmaya Sachan · 2022
Closest in time.
A rigorous study of integrated gradients method and extensions to internal neuron attributions
Daniel Lundstrom, Tianjian Huang, and Meisam Razaviyayn · 2022
Closest in time.
Locating and editing factual associations in gpt
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Closest in time.
Is a modular architecture enough?, 2022
Sarthak Mittal, Yoshua Bengio, and Guillaume Lajoie · 2022
Closest in time.
Interpretable Machine Learning
Christoph Molnar · 2022
Closest in time.
Clip-dissect: Automatic description of neuron representations in deep vision networks, 2022
Tuomas Oikarinen and Tsui-Wei Weng · 2022
Closest in time.
In-context learning and induction heads
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Closest in time.
Capsule networks–a survey
Mensah Kwabena Patrick, Adebayo Felix Adekoya, Ayidzoe Abra Mighty, and Baagyire Y Edward · 2022
Closest in time.
Linear adversarial concept erasure
Shauli Ravfogel, Michael Twiton, Yoav Goldberg, and Ryan D Cotterell · 2022
Closest in time.
Interpretable machine learning: Fundamental principles and 10 grand challenges
Cynthia Rudin, Chaofan Chen, Zhi Chen, Haiyang Huang, Lesia Semenova, and Chudi Zhong · 2022
Closest in time.
Compute trends across three eras of machine learning, 2022
Jaime Sevilla, Lennart Heim, Anson Ho, Tamay Besiroglu, Marius Hobbhahn, and Pablo Villalobos · 2022
Closest in time.
A closer look at rehearsal-free continual learning
James Seale Smith, Junjiao Tian, Yen-Chang Hsu, and Zsolt Kira · 2022
Closest in time.
A survey of neural trojan attacks and defenses in deep learning
Jie Wang, Ghulam Mubashar Hassan, and Naveed Akhtar · 2022
Closest in time.
Discovering bugs in vision models using off-the-shelf image generation and captioning, 2022
Olivia Wiles, Isabela Albuquerque, and Sven Gowal · 2022
Closest in time.
Post-hoc concept bottleneck models
Mert Yuksekgonul, Maggie Wang, and James Zou · 2022
Closest in time.
Exsum: From local explanations to model understanding
Yilun Zhou, Marco Tulio Ribeiro, and Julie Shah · 2022
Closest in time.
Adversarial training for high-stakes reliability
Daniel M Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, et al · 2022
Closest in time.