Fetching the paper…
Reading the bibliography…
As machine learning models are increasingly deployed in high-stakes domains such as legal and financial decision-making, there has been growing interest in post-hoc methods for generating counterfactual explanations.
A theory of the learnable
Leslie G Valiant · 1984
Earlier work this paper cites.
Predicting recidivism in north carolina, 1978 and 1980
Peter Schmidt and Ann D. Witte · 1988
Earlier work this paper cites.
An introduction to computational learning theory
Michael J Kearns, Umesh Virkumar Vazirani, and Umesh Vazirani · 1994
Earlier work this paper cites.
Intelligible models for classification and regression
Yin Lou, Rich Caruana, and Johannes Gehrke · 2012
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus · 2014
Earlier work this paper cites.
Falling rule lists
Fulton Wang and Cynthia Rudin · 2015
Earlier work this paper cites.
How we analyzed the compas recidivism algorithm
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner · 2016
Earlier work this paper cites.
Measuring neural net robustness with constraints
Osbert Bastani, Yani Ioannou, Leonidas Lampropoulos, Dimitrios Vytiniotis, Aditya Nori, and Antonio Criminisi · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Earlier work this paper cites.
Deepfool: A simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
Interpreting blackbox models via model extraction
Osbert Bastani, Carolyn Kim, and Hamsa Bastani · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Learning cost-effective and interpretable treatment regimes
Himabindu Lakkaraju and Cynthia Rudin · 2017
Cited alongside, same era.
Counterfactual explanations without opening the black box: Automated decisions and the gdpr
S. Wachter, Brent D. Mittelstadt, and Chris Russell · 2017
Cited alongside, same era.
Interpretable classification models for recidivism prediction
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Alibi: Algorithms for monitoring and explaining machine learning models, 2019
Janis Klaise, Arnaud Van Looveren, Giovanni Vacanti, and Alexandru Coca · 2019
Later among the works it cites.
Faithful and customizable explanations of black box models
Himabindu Lakkaraju, Ece Kamar, Rich Caruana, and Jure Leskovec · 2019
Later among the works it cites.
Interpretable counterfactual explanations guided by prototypes, 2019
Arnaud Van Looveren and Janis Klaise · 2019
Later among the works it cites.
Actionable recourse in linear classification
Berk Ustun, Alexander Spangher, and Yang Liu · 2019
Later among the works it cites.
Huggingface’s transformers: State-of-the-art natural language processing, 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jiaming Zeng, Berk Ustun, and Cynthia Rudin · 2017
Cited alongside, same era.
Generating counterfactual explanations with natural language
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata · 2018
Cited alongside, same era.
Understanding adversarial training: Increasing local stability of supervised models through robust optimization
Uri Shaham, Yutaro Yamada, and Sahand Negahban · 2018
Cited alongside, same era.
Interpreting neural network judgments via minimal, stable, and symbolic corrections
Xin Zhang, Armando Solar-Lezama, and Rishabh Singh · 2018
Cited alongside, same era.
Generating natural adversarial examples
Zhengli Zhao, Dheeru Dua, and Sameer Singh · 2018
Cited alongside, same era.
Explaining and harnessing adversarial examples
Ian Goodfellow, Jonathon Shlens, and Christian Szegedy
Cited in the paper.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy
Cited in the paper.
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush · 2019
Later among the works it cites.
Explaining machine learning classifiers through diverse counterfactual explanations
Ramaravind K Mothilal, Amit Sharma, and Chenhao Tan · 2020
Closest in time.
Pac confidence sets for deep neural networks via calibrated prediction
Sangdon Park, Osbert Bastani, Nikolai Matni, and Insup Lee · 2020
Closest in time.
Face: Feasible and actionable counterfactual explanations
Rafael Poyiadzi, Kacper Sokol, Raul Santos-Rodriguez, Tijl De Bie, and Peter Flach · 2020
Closest in time.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju · 2020
Closest in time.