Fetching the paper…
Reading the bibliography…
Feature attribution methods, proposed recently, help users interpret the predictions of complex models.
Three naive bayes approaches for discrimination-free classification
Toon Calders and Sicco Verwer. 2010 · 2010
Earlier work this paper cites.
Intelligible models for classification and regression
Yin Lou, Rich Caruana, and Johannes Gehrke. 2012 · 2012
Earlier work this paper cites.
The bayesian case model: A generative approach for case-based reasoning and prototype classification
Been Kim, Cynthia Rudin, and Julie Shah. 2014 · 2014
Earlier work this paper cites.
Convolutional neural networks for sentence classification
Yoon Kim. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, Wojciech Samek, and Oscar Deniz Suarez. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Machine bias: There’s software used across the country to predict future criminals. and it’s biased against blacks
Julia Angwin, Jeff Larson, Surya Mattu, and Lauren Kirchner. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Zou, Venkatesh Saligrama, and Adam Kalai. 2016 · 2016
Earlier work this paper cites.
European union regulations on algorithmic decision-making and a “right to explanation”
Bryce Goodman and Seth Flaxman. 2016 · 2016
Earlier work this paper cites.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nathan Srebro. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
”why should i trust you?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Hateful symbols or hateful people? predictive features for hate speech detection on twitter
Zeerak Waseem and Dirk Hovy. 2016 · 2016
Earlier work this paper cites.
Automated hate speech detection and the problem of offensive language
Thomas Davidson, Dana Warmsley, Michael W. Macy, and Ingmar Weber. 2017 · 2017
Cited alongside, same era.
Distilling a neural network into a soft decision tree
Nicholas Frosst and Geoffrey E. Hinton. 2017 · 2017
Cited alongside, same era.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Cited alongside, same era.
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2017 · 2017
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Pai-Shun Ting, Karthikeyan Shanmugam, and Payel Das. 2018 · 2018
Later among the works it cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Later among the works it cites.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Later among the works it cites.
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed Huai hsin Chi, and Alex Beutel. 2018 · 2018
Later among the works it cites.
Explaining explanations: An overview of interpretability of machine learning
Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Cited alongside, same era.
Right for the right reasons: Training differentiable models by constraining their explanations
Andrew Slavin Ross, Michael C. Hughes, and Finale Doshi-Velez. 2017 · 2017
Cited alongside, same era.
A survey on hate speech detection using natural language processing
Anna Schmidt and Michael Wiegand. 2017 · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Cited alongside, same era.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda B. Viégas, and Martin Wattenberg. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Cited alongside, same era.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda B. Viégas, and Rory Sayres. 2018 · 2018
Later among the works it cites.
The mythos of model interpretability
Zachary C. Lipton. 2018 · 2018
Later among the works it cites.
Consistent individualized feature attribution for tree ensembles
Scott M. Lundberg, Gabriel G. Erion, and Su-In Lee. 2018 · 2018
Later among the works it cites.
Learning adversarially fair and transferable representations
David Madras, Elliot Creager, Toniann Pitassi, and Richard S. Zemel. 2018 · 2018
Later among the works it cites.
Did the model understand the question?
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018 · 2018
Later among the works it cites.
Beyond word importance: Contextual decomposition to extract interactions from lstms
W. James Murdoch, Peter J. Liu, and Bin Yu. 2018 · 2018
Later among the works it cites.
Reducing gender bias in abusive language detection
Ji Ho Park, Jamin Shin, and Pascale Fung. 2018 · 2018
Later among the works it cites.
Using a deep learning algorithm and integrated gradients explanation to assist grading for diabetic retinopathy
Rory Sayres, Ankur Taly, Ehsan Rahimy, Katy Blumer, David Coz, Naama Hammel, Jonathan Krause, Arunachalam Narayanaswamy, Zahra Rastegar, Derek Wu, Shawn Xu, Scott Barb, Anthony Joseph, Michael Shumski, Jesse Smith, Arjun B. Sood, Greg S. Corrado, Lily Peng, and Dale R. Webster. 2018 · 2018
Later among the works it cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Closest in time.