Fetching the paper…
Reading the bibliography…
Motivations for methods in explainable artificial intelligence (XAI) often include detecting, quantifying and mitigating bias, and contributing to making machine learning models fairer.
Intrinsic bias metrics do not correlate with application bias
Seraphina Goldfarb-Tarrant, Rebecca Marchant, Ricardo Muñoz Sánchez, Mugdha Pandya, and Adam Lopez. 2021 · 1940
Earlier work this paper cites.
What constitutes fairness in work settings? a four-component model of procedural justice
Steven L Blader and Tom R Tyler. 2003 · 2003
Earlier work this paper cites.
Opportunities and challenges in explainable artificial intelligence (xai): A survey
Arun Das and Paul Rad. 2020 · 2006
Earlier work this paper cites.
Causality
Judea Pearl. 2009 · 2009
Earlier work this paper cites.
Explainability for fair machine learning
Tom Begley, Tobias Schwedes, Christopher Frye, and Ilya Feige. 2020 · 2010
Earlier work this paper cites.
Fairness in machine learning: A survey
Simon Caton and Christian Haas. 2020 · 2010
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyung Hyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst. 2016 · 2016
Earlier work this paper cites.
Demographic dialectal variation in social media: A case study of African-American English
Su Lin Blodgett, Lisa Green, and Brendan O’Connor. 2016 · 2016
Earlier work this paper cites.
Man is to computer programmer as woman is to homemaker? debiasing word embeddings
Tolga Bolukbasi, Kai-Wei Chang, James Y Zou, Venkatesh Saligrama, and Adam T Kalai. 2016 · 2016
Earlier work this paper cites.
Retain: An interpretable predictive model for healthcare using reverse time attention mechanism
Edward Choi, Mohammad Taha Bahadori, Jimeng Sun, Joshua Kulas, Andy Schuetz, and Walter Stewart. 2016 · 2016
Earlier work this paper cites.
“Why should i trust you?” Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Racial disparity in natural language processing: A case study of social media African-American English
Su Lin Blodgett and Brendan O’Connor. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Counterfactual fairness
Matt J Kusner, Joshua Loftus, Chris Russell, and Ricardo Silva. 2017 · 2017
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Smoothgrad: removing noise by adding noise
Daniel Smilkov, Nikhil Thorat, Been Kim, Fernanda Viégas, and Martin Wattenberg. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Peeking inside the black-box: A survey on explainable artificial intelligence (XAI)
Amina Adadi and Mohammed Berrada. 2018 · 2018
Earlier work this paper cites.
Beyond distributive fairness in algorithmic decision making: Feature selection for procedurally fair learning
Nina Grgić-Hlača, Muhammad Bilal Zafar, Krishna P Gummadi, and Adrian Weller. 2018 · 2018
Earlier work this paper cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi. 2018 · 2018
Earlier work this paper cites.
Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (tcav)
Been Kim, Martin Wattenberg, Justin Gilmer, Carrie Cai, James Wexler, Fernanda Viegas, et al. 2018 · 2018
Earlier work this paper cites.
Examining gender and race bias in two hundred sentiment analysis systems
Svetlana Kiritchenko and Saif Mohammad. 2018 · 2018
Earlier work this paper cites.
Causal reasoning for algorithmic fairness
Joshua R Loftus, Chris Russell, Matt J Kusner, and Ricardo Silva. 2018 · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rachel Rudinger, Jason Naradowsky, Brian Leonard, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Distill-and-compare: Auditing black-box models using transparent model distillation
Sarah Tan, Rich Caruana, Giles Hooker, and Yin Lou. 2018 · 2018
Earlier work this paper cites.
Representer point selection for explaining deep neural networks
Chih-Kuan Yeh, Joon Kim, Ian En-Hsu Yen, and Pradeep K Ravikumar. 2018 · 2018
Earlier work this paper cites.
Fairness in decision-making—the causal explanation formula
Junzhe Zhang and Elias Bareinboim. 2018 · 2018
Cited alongside, same era.
Fairwashing: the risk of rationalization
Ulrich Aïvodji, Hiromi Arai, Olivier Fortineau, Sébastien Gambs, Satoshi Hara, and Alain Tapp. 2019 · 2019
Cited alongside, same era.
Differential privacy has disparate impact on model accuracy
Eugene Bagdasaryan, Omid Poursaeed, and Vitaly Shmatikov. 2019 · 2019
Cited alongside, same era.
Bias and fairness in natural language processing
Kai-Wei Chang, Vinod Prabhakaran, and Vicente Ordonez. 2019 · 2019
Cited alongside, same era.
What does BERT look at? An analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning. 2019 · 2019
Cited alongside, same era.
Certified adversarial robustness via randomized smoothing
Jeremy Cohen, Elan Rosenfeld, and Zico Kolter. 2019 · 2019
Cited alongside, same era.
Interpreting predictions of NLP models
Eric Wallace, Matt Gardner, and Sameer Singh. 2020 · 2020
Later among the works it cites.
Impact of politically biased data on hate speech classification
Maximilian Wich, Jan Bauer, and Georg Groh. 2020 · 2020
Later among the works it cites.
Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making
Yunfeng Zhang, Q Vera Liao, and Rachel KE Bellamy. 2020b · 2020
Later among the works it cites.
Fine-grained classification of political bias in German news: A data set and initial experiments
Dmitrii Aksenov, Peter Bourgonje, Karolina Zaczynska, Malte Ostendorff, Julian Moreno Schneider, and Georg Rehm. 2021 · 2021
Later among the works it cites.
Can explainable AI explain unfairness? A framework for evaluating explainable AI
Kiana Alikhademi, Brianna Richardson, Emma Drobina, and Juan E Gilbert. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Counterfactual fairness in text classification through robustness
Sahaj Garg, Vincent Perot, Nicole Limtiaco, Ankur Taly, Ed H Chi, and Alex Beutel. 2019 · 2019
Cited alongside, same era.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang. 2019 · 2019
Cited alongside, same era.
Attention is not explanation
Sarthak Jain and Byron C Wallace. 2019 · 2019
Cited alongside, same era.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2019 · 2019
Cited alongside, same era.
Perturbation sensitivity analysis to detect unintended model biases
Vinodkumar Prabhakaran, Ben Hutchinson, and Margaret Mitchell. 2019 · 2019
Cited alongside, same era.
The risk of racial bias in hate speech detection
Maarten Sap, Dallas Card, Saadia Gabriel, Yejin Choi, and Noah A Smith. 2019 · 2019
Cited alongside, same era.
Can we improve model robustness through secondary attribute counterfactuals?
Ananth Balashankar, Xuezhi Wang, Ben Packer, Nithum Thain, Ed Chi, and Alex Beutel. 2021 · 2021
Later among the works it cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, and Colin Raffel. 2021 · 2021
Later among the works it cites.
Quantifying social biases in NLP: A generalization and empirical comparison of extrinsic fairness metrics
Paula Czarnowska, Yogarshi Vyas, and Kashif Shah. 2021 · 2021
Later among the works it cites.
Moving beyond “algorithmic bias is a data problem”
Sara Hooker. 2021 · 2021
Later among the works it cites.
Explanation-based human debugging of NLP models: A survey
Piyawat Lertvittayakumjorn and Francesca Toni. 2021 · 2021
Later among the works it cites.
Metamorphic testing and certified mitigation of fairness violations in NLP models
Pingchuan Ma, Shuai Wang, and Jin Liu. 2021 · 2021
Later among the works it cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021 · 2021
Later among the works it cites.
A survey on bias and fairness in machine learning
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. 2021 · 2021
Later among the works it cites.
From what to how: an initial review of publicly available AI ethics tools, methods and research to translate principles into practices
Jessica Morley, Luciano Floridi, Libby Kinsey, and Anat Elhalal. 2021 · 2021
Later among the works it cites.
Do the ends justify the means? variation in the distributive and procedural fairness of machine learning algorithms
Lily Morse, Mike Horia M Teodorescu, Yazeed Awwad, and Gerald C Kane. 2021 · 2021
Later among the works it cites.
Understanding and interpreting the impact of user context in hate speech detection
Edoardo Mosca, Maximilian Wich, and Georg Groh. 2021 · 2021
Later among the works it cites.
On fairness and interpretability
Deepak P, Sanil V, and Joemon M. Jose. 2021 · 2021
Later among the works it cites.
Deep causal graphs for causal inference, black-box explainability and fairness
Alvaro Parafita and Jordi Vitria. 2021 · 2021
Later among the works it cites.
Does robustness improve fairness? approaching fairness with word substitution robustness methods for text classification
Yada Pruksachatkun, Satyapriya Krishna, Jwala Dhamala, Rahul Gupta, and Kai-Wei Chang. 2021 · 2021
Later among the works it cites.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew E Peters. 2021 · 2021
Later among the works it cites.
Hatecheck: Functional tests for hate speech detection models
Paul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem, Helen Margetts, and Janet Pierrehumbert. 2021 · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Exploring the efficacy of automatically generated counterfactuals for sentiment analysis
Linyi Yang, Jiazheng Li, Pádraig Cunningham, Yue Zhang, Barry Smyth, and Ruihai Dong. 2021 · 2021
Later among the works it cites.
Hildif: Interactive debugging of NLI models using influence functions
Hugo Zylberajch, Piyawat Lertvittayakumjorn, and Francesca Toni. 2021 · 2021
Later among the works it cites.
Necessity and sufficiency for explaining text classifiers: A case study in hate speech detection
Esma Balkır, Isar Nejadgholi, Kathleen C Fraser, and Svetlana Kiritchenko. 2022 · 2022
Closest in time.
A clarification of the nuances in the fairness metrics landscape
Alessandro Castelnovo, Riccardo Crupi, Greta Greco, Daniele Regoli, Ilaria Giuseppina Penco, and Andrea Claudio Cosentini. 2022 · 2022
Closest in time.
Marrying fairness and explainability in supervised learning
Przemyslaw Grabowicz, Nicholas Perello, and Aarshee Mishra. 2022 · 2022
Closest in time.
Ethics sheet for automatic emotion recognition and sentiment analysis
Saif M Mohammad. 2022 · 2022
Closest in time.
Improving generalizability in implicitly abusive language detection with concept activation vectors
Isar Nejadgholi, Kathleen C Fraser, and Svetlana Kiritchenko. 2022 · 2022
Closest in time.
Interpretable data-based explanations for fairness debugging
Romila Pradhan, Jiongli Zhu, Boris Glavic, and Babak Salimi. 2022 · 2022
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.