Fetching the paper…
Reading the bibliography…
Interpretable AI tools are often motivated by the goal of understanding model behavior in out-of-distribution (OOD) contexts.
Adversarial machine learning
Ling Huang, Anthony D Joseph, Blaine Nelson, Benjamin IP Rubinstein, and J Doug Tygar · 2011
Earlier work this paper cites.
For better or worse, benchmarks shape a field
David Patterson · 2012
Earlier work this paper cites.
OpenSurfaces: A richly annotated catalog of surface appearance
Sean Bell, Paul Upchurch, Noah Snavely, and Kavita Bala · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus · 2014
Earlier work this paper cites.
Tiny imagenet visual recognition challenge
Ya Le and Xuan Yang · 2015
Earlier work this paper cites.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Earlier work this paper cites.
Image style transfer using convolutional neural networks
Leon A Gatys, Alexander S Ecker, and Matthias Bethge · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Network dissection: Quantifying interpretability of deep visual representations, 2017
David Bau, Bolei Zhou, Aditya Khosla, Aude Oliva, and Antonio Torralba · 2017
Earlier work this paper cites.
Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer · 2017
Earlier work this paper cites.
Targeted backdoor attacks on deep learning systems using data poisoning
Xinyun Chen, Chang Liu, Bo Li, Kimberly Lu, and Dawn Song · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning, 2017
Finale Doshi-Velez and Been Kim · 2017
Earlier work this paper cites.
Feature visualization
Chris Olah, Alexander Mordvintsev, and Ludwig Schubert · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Earlier work this paper cites.
Large scale gan training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2018
Earlier work this paper cites.
Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
Jinfeng Li, Shouling Ji, Tianyu Du, Bo Li, and Ting Wang · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton · 2018
Earlier work this paper cites.
Differentiable image parameterizations
Alexander Mordvintsev, Nicola Pezzotti, Ludwig Schubert, and Chris Olah · 2018
Earlier work this paper cites.
Collision between vehicle controlled by developmental automated driving system and pedestrian, 2018
National Transportation Safety Board NTSB · 2018
Earlier work this paper cites.
Exploring neural networks with activation atlases
Shan Carter, Zan Armstrong, Ludwig Schubert, Ian Johnson, and Chris Olah · 2019
Earlier work this paper cites.
This looks like that: deep learning for interpretable image recognition
Chaofan Chen, Oscar Li, Daniel Tao, Alina Barnett, Cynthia Rudin, and Jonathan K Su · 2019
Cited alongside, same era.
Badnets: Evaluating backdooring attacks on deep neural networks
Tianyu Gu, Kang Liu, Brendan Dolan-Gavitt, and Siddharth Garg · 2019
Cited alongside, same era.
Tabor: A highly accurate approach to inspecting and restoring trojan backdoors in ai systems
Wenbo Guo, Lun Wang, Xinyu Xing, Min Du, and Dawn Song · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim · 2019
Cited alongside, same era.
Adversarial examples are not bugs, they are features
Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry · 2019
Cited alongside, same era.
3db: A framework for debugging computer vision models
Guillaume Leclerc, Hadi Salman, Andrew Ilyas, Sai Vemprala, Logan Engstrom, Vibhav Vineet, Kai Xiao, Pengchuan Zhang, Shibani Santurkar, Greg Yang, et al · 2021
Later among the works it cites.
The effectiveness of feature attribution methods and its correlation with automatic evaluation scores
Giang Nguyen, Daeyoung Kim, and Anh Nguyen · 2021
Later among the works it cites.
Ai and the everything in the whole wide world benchmark
Inioluwa Deborah Raji, Emily M Bender, Amandalynne Paullada, Emily Denton, and Alex Hanna · 2021
Later among the works it cites.
Auditing visualizations: Transparency methods struggle to detect anomalous behavior, 2022
Jean-Stanislas Denain and Jacob Steinhardt · 2022
Later among the works it cites.
Domino: Discovering systematic errors with cross-modal embeddings
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller · 2019
Cited alongside, same era.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che · 2019
Cited alongside, same era.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin · 2019
Cited alongside, same era.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Bolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li, Bimal Viswanath, Haitao Zheng, and Ben Y Zhao · 2019
Cited alongside, same era.
Debugging tests for model explanations
Julius Adebayo, Michael Muelly, Ilaria Liccardi, and Been Kim · 2020
Cited alongside, same era.
Exemplary natural images explain cnn activations better than state-of-the-art feature visualization
Judy Borowski, Roland S Zimmermann, Judith Schepers, Robert Geirhos, Thomas SA Wallis, Matthias Bethge, and Wieland Brendel · 2020
Cited alongside, same era.
Robustbench: a standardized adversarial robustness benchmark
Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein · 2020
Cited alongside, same era.
Sabri Eyuboglu, Maya Varma, Khaled Saab, Jean-Benoit Delbrouck, Christopher Lee-Messer, Jared Dunnmon, James Zou, and Christopher Ré · 2022
Later among the works it cites.
A bird’s eye view of the ml field [pragmatic ai safety no. 2], 2022
Dan Hendrycks and Thomas Woodside · 2022
Later among the works it cites.
Natural language descriptions of deep visual features
Evan Hernandez, Sarah Schwettmann, David Bau, Teona Bagashvili, Antonio Torralba, and Jacob Andreas · 2022
Later among the works it cites.
Towards benchmarking explainable artificial intelligence methods, 2022
Lars Holmberg · 2022
Later among the works it cites.
Distilling model failures as directions in latent space, 2022
Saachi Jain, Hannah Lawrence, Ankur Moitra, and Aleksander Madry · 2022
Later among the works it cites.
Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi · 2022
Later among the works it cites.
Robust explainability: A tutorial on gradient-based attribution methods for deep neural networks
Ian E Nielsen, Dimah Dera, Ghulam Rasool, Ravi P Ramachandran, and Nidhal Carla Bouaynaya · 2022
Later among the works it cites.
Red-teaming the stable diffusion safety filter
Javier Rando, Daniel Paleka, David Lindner, Lennard Heim, and Florian Tramèr · 2022
Later among the works it cites.
Toward transparent ai: A survey on interpreting the inner structures of deep neural networks
Tilman Räuker, Anson Ho, Stephen Casper, and Dylan Hadfield-Menell · 2022
Later among the works it cites.
Chatgpt: Optimizing language models for dialogue, 2022
J Schulman, B Zoph, C Kim, J Hilton, J Menick, J Weng, JFC Uribe, L Fedus, L Metz, M Pokorny, et al · 2022
Later among the works it cites.
Emily Wenger, Roma Bhattacharjee, Arjun Nitin Bhagoji, Josephine Passananti, Emilio Andere, Haitao Zheng, and Ben Y Zhao · 2022
Later among the works it cites.
Backdoorbench: A comprehensive benchmark of backdoor learning
Baoyuan Wu, Hongrui Chen, Mingda Zhang, Zihao Zhu, Shaokui Wei, Danni Yuan, Chao Shen, and Hongyuan Zha · 2022
Later among the works it cites.
Adversarial training for high-stakes reliability
Daniel M Ziegler, Seraphina Nix, Lawrence Chan, Tim Bauman, Peter Schmidt-Nielsen, Tao Lin, Adam Scherlis, Noa Nabeshima, Ben Weinstein-Raun, Daniel de Haas, et al · 2022
Later among the works it cites.
Evaluating post-hoc interpretability with intrinsic interpretability, 2023
José Pereira Amorim, Pedro Henriques Abreu, João Santos, and Henning Müller · 2023
Closest in time.
Funnybirds: A synthetic vision dataset for a part-based analysis of explainable ai methods, 2023
Robin Hesse, Simone Schaub-Meyer, and Stefan Roth · 2023
Closest in time.
Rethinking backdoor attacks
Alaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov, Kristian Georgiev, Hadi Salman, Andrew Ilyas, and Aleksander Madry · 2023
Closest in time.
Tracr: Compiled transformers as a laboratory for interpretability
David Lindner, János Kramár, Matthew Rahtz, Thomas McGrath, and Vladimir Mikulik · 2023
Closest in time.
Benchmarking robustness to adversarial image obfuscations, 2023
Florian Stimberg, Ayan Chakrabarti, Chun-Ta Lu, Hussein Hazimeh, Otilia Stretcu, Wei Qiao, Yintao Liu, Merve Kaya, Cyrus Rashtchian, Ariel Fuxman, Mehmet Tek, and Sven Gowal · 2023
Closest in time.