Fetching the paper…
Reading the bibliography…
Model visualizations provide information that outputs alone might miss.
Measuring robustness to natural distribution shifts in image classification
Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, and Ludwig Schmidt · 2007
Earlier work this paper cites.
Applied logistic regression , volume 398
David W Hosmer Jr., Stanley Lemeshow, and Rodney X Sturdivant · 2013
Earlier work this paper cites.
Striving for simplicity: The all convolutional net, 2015
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller · 2015
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks, 2017
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Research: Caricatures, 2018
Chris Olah · 2018
Earlier work this paper cites.
The unreasonable effectiveness of deep features as a perceptual metric, 2018
Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang · 2018
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing, 2019
Jeremy M Cohen, Elan Rosenfeld, and J. Zico Kolter · 2019
Earlier work this paper cites.
ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness, 2019
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features, 2019
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks, 2019
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu · 2019
Earlier work this paper cites.
PyTorch: An imperative style, high-performance deep learning library, 2019
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Grad-CAM: Visual explanations from deep networks via gradient-based localization
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra · 2019
Cited alongside, same era.
Debugging tests for model explanations, 2020
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim · 2020
Cited alongside, same era.
Explainable machine learning in deployment, 2020
Umang Bhatt, Alice Xiang, Shubham Sharma, Adrian Weller, Ankur Taly, Yunhan Jia, Joydeep Ghosh, Ruchir Puri, José M. F. Moura, and Peter Eckersley · 2020
Cited alongside, same era.
Curve circuits
Nick Cammarata, Gabriel Goh, Shan Carter, Chelsea Voss, Ludwig Schubert, and Chris Olah · 2020
Cited alongside, same era.
Pytorch library for CAM methods
Jacob Gildenblat and contributors · 2021
Later among the works it cites.
Lucent, Lucid library adapted for pytorch, 2021
Lim Swee Kiat · 2021
Later among the works it cites.
Backdoor learning: A survey, 2021
Yiming Li, Baoyuan Wu, Yong Jiang, Zhifeng Li, and Shu-Tao Xia · 2021
Later among the works it cites.
Editing a classifier by rewriting its prediction rules, 2021
Shibani Santurkar, Dimitris Tsipras, Mahalaxmi Elango, David Bau, Antonio Torralba, and Aleksander Madry · 2021
Later among the works it cites.
Salient ImageNet: How to discover spurious features in deep learning?, 2021
Sahil Singla and Soheil Feizi · 2021
Later among the works it cites.
Leveraging sparse linear layers for debuggable deep networks, 2021
Eric Wong, Shibani Santurkar, and Aleksander Mądry · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Underspecification presents challenges for credibility in modern machine learning, 2020
Alexander D’Amour, Katherine Heller, Dan Moldovan, Ben Adlam, Babak Alipanahi, Alex Beutel, Christina Chen, Jonathan Deaton, Jacob Eisenstein, Matthew D. Hoffman, Farhad Hormozdiari, Neil Houlsby, Shaobo Hou, Ghassen Jerfel, Alan Karthikesalingam, Mario Lucic, Yian Ma, Cory McLean, Diana Mincu, Akinori Mitani, Andrea Montanari, Zachary Nado, Vivek Natarajan, Christopher Nielson, Thomas F. Osborne, Rajiv Raman, Kim Ramasamy, Rory Sayres, Jessica Schrouff, Martin Seneviratne, Shannon Sequeira, Harini Suresh, Victor Veitch, Max Vladymyrov, Xuezhi Wang, Kellie Webster, Steve Yadlowsky, Taedong Yun, Xiaohua Zhai, and D. Sculley · 2020
Cited alongside, same era.
Concept bottleneck models, 2020
Pang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann, Emma Pierson, Been Kim, and Percy Liang · 2020
Cited alongside, same era.
Captum: A unified and generic model interpretability library for PyTorch, 2020
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz-Richardson · 2020
Cited alongside, same era.
What do you see? evaluation of explainable artificial intelligence (XAI) interpretability through neural backdoors, 2020
Yi-Shan Lin, Wen-Chuan Lee, and Z. Berkay Celik · 2020
Cited alongside, same era.
Do adversarially robust ImageNet models transfer better?, 2020
Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor, and Aleksander Madry · 2020
Cited alongside, same era.
Blind backdoors in deep learning models
Eugene Bagdasaryan and Vitaly Shmatikov · 2021
Cited alongside, same era.
"will you find these shortcuts?" a protocol for evaluating the faithfulness of input salience methods for text classification, 2021
Jasmijn Bastings, Sebastian Ebert, Polina Zablotskaia, Anders Sandholm, and Katja Filippova · 2021
Cited alongside, same era.
A study of face obfuscation in ImageNet, 2021
Kaiyu Yang, Jacqueline Yau, Li Fei-Fei, Jia Deng, and Olga Russakovsky · 2021
Later among the works it cites.
Do feature attribution methods correctly attribute features?, 2021
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, and Julie Shah · 2021
Later among the works it cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim · 2022
Closest in time.
Natural language descriptions of deep visual features, 2022
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Closest in time.
Interpretable Machine Learning
Christoph Molnar · 2022
Closest in time.
Benchmarking interpretability tools for deep neural networks, 2023
Stephen Casper, Yuxiao Li, Jiawei Li, Tong Bu, Kevin Zhang, and Dylan Hadfield-Menell · 2023
Closest in time.