Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
Interpreting and explaining deep models in computer vision
Wojciech Samek, Grégoire Montavon, and Klaus-Robert Müller · 2018
Later among the works it cites.
Towards robust interpretability with self-explaining neural networks
David Alvares-Melis and Tommi S. Jaakkola · 2018
Later among the works it cites.
Interpretable deep learning under fire
Original
Xinyang Zhang, Ningfei Wang, Shouling Ji, Hua Shen, and Ting Wang · 2018
Later among the works it cites.
RISE: Randomized input sampling for explanation of black-box models
Vitali Petsiuk, Abir Das, and Kate Saenko · 2018
Later among the works it cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Anish Athalye, Nicholas Carlini, and David Wagner · 2018
Later among the works it cites.
Sanity chekcs for saliency maps
Julius Adebayo, J. Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim · 2018
Later among the works it cites.
Turning your weakness into a strength: Watermarking deep neural networks by backdooring
Yossi Adi, Carsten Baum, Moustapha Cisse, Benny Pinkas, and Joseph Keshet · 2018
Later among the works it cites.
Interpretable machine learning: definitions, methods, and applications
Original
W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu · 2019
Closest in time.
The (un) reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T Schütt, Sven Dähne, Dumitru Erhan, and Been Kim · 2019
Closest in time.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou · 2019
Closest in time.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2019
Closest in time.
The odds are odd: A statistical test for detecting adversarial examples
Kevin Roth, Yannic Kilcher, and Thomas Hofmann · 2019
Closest in time.
Adversarial training for free!
Original
Ali Shafahi, Mahyar Najibi, Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein · 2019
Closest in time.