Fetching the paper…
Reading the bibliography…
Feature attribution a.k.a.
Did the model understand the question?
Pramod Kaushik Mudrakarta, Ankur Taly, Mukund Sundararajan, and Kedar Dhamdhere. 2018 · 1906
Earlier work this paper cites.
Benchmarking Attribution Methods with Relative Feature Importance
Mengjiao Yang and Been Kim. 2019 · 1907
Earlier work this paper cites.
Bidirectional recurrent neural networks
Mike Schuster, Kuldip K. Paliwal, and A. General. 1997 · 1997
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, A. Binder, G. Montavon, F. Klauschen, Klaus-Robert Müller, and W. Samek. 2015 · 2015
Earlier work this paper cites.
Extraction of salient sentences from labelled documents
Misha Denil, Alban Demiraj, and Nando de Freitas. 2015 · 2015
Earlier work this paper cites.
Visualizing and understanding neural models in NLP
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2017 · 2017
Earlier work this paper cites.
Deeper attention to abusive user content moderation
John Pavlopoulos, Prodromos Malakasiotis, and Ion Androutsopoulos. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017a · 2017
Earlier work this paper cites.
Ex machina: Personal attacks seen at scale
Ellery Wulczyn, Nithum Thain, and Lucas Dixon. 2017 · 2017
Earlier work this paper cites.
Noel Codella, Veronica Rotemberg, Philipp Tschandl, M. Emre Celebi, Stephen Dusza, David Gutman, Brian Helba, Aadi Kalloo, Konstantinos Liopyris, Michael Marchetti, Harald Kittler, and Allan Halpern. 2019 · 2018
Earlier work this paper cites.
Measuring and mitigating unintended bias in text classification
Lucas Dixon, John Li, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Evaluating feature importance estimates
Sara Hooker, Dumitru Erhan, Pieter jan Kindermans, and Been Kim. 2018 · 2018
Earlier work this paper cites.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Nina Poerner, Hinrich Schütze, and Benjamin Roth. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
Explaining therapy predictions with layer-wise relevance propagation in neural networks
Yinchong Yang, Volker Tresp, Marius Wunderle, and Peter A. Fasching. 2018 · 2018
Cited alongside, same era.
Evaluating recurrent neural network explanations
Leila Arras, Ahmed Osman, Klaus-Robert Müller, and Wojciech Samek. 2019 · 2019
Cited alongside, same era.
Don’t take the premise for granted: Mitigating artifacts in natural language inference
Yonatan Belinkov, Adam Poliak, Stuart Shieber, Benjamin Van Durme, and Alexander Rush. 2019 · 2019
Cited alongside, same era.
Can I trust the explainer? verifying post-hoc explanatory methods
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Later among the works it cites.
Shortcut learning in deep neural networks
R. Geirhos, J.-H. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann. 2020 · 2020
Later among the works it cites.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C. Wallace, and Yulia Tsvetkov. 2020 · 2020
Later among the works it cites.
Evaluating attribution methods using white-box LSTMs
Yiding Hao. 2020 · 2020
Later among the works it cites.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What can ai do for me? evaluating machine learning interpretations in cooperative play
Shi Feng and Jordan Boyd-Graber. 2019 · 2019
Cited alongside, same era.
Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets
Mor Geva, Yoav Goldberg, and Jonathan Berant. 2019 · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Right for the wrong reasons: Diagnosing syntactic heuristics in natural language inference
Tom McCoy, Ellie Pavlick, and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Towards Explainable Artificial Intelligence , pages 5–22. Springer International Publishing, Cham
Wojciech Samek and Klaus-Robert Müller. 2019 · 2019
Cited alongside, same era.
Quantifying interpretability and trust in machine learning systems
Philipp Schmidt and Felix Biessmann. 2019 · 2019
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard H. Hovy, and Zachary Chase Lipton. 2020 · 2020
Later among the works it cites.
Interpretation of NLP models through input marginalization
Siwon Kim, Jihun Yi, Eunji Kim, and Sungroh Yoon. 2020 · 2020
Later among the works it cites.
Exposing Shallow Heuristics of Relation Extraction Models with Challenge Data
Shachar Rosenman, Alon Jacovi, and Yoav Goldberg. 2020 · 2020
Later among the works it cites.
Data staining: A method for comparing faithfulness of explainers
Jacob Sippy, Gagan Bansal, and Daniel S. Weld. 2020 · 2020
Later among the works it cites.
The language interpretability tool: Extensible, interactive visualizations and analysis for NLP models
Ian Tenney, James Wexler, Jasmijn Bastings, Tolga Bolukbasi, Andy Coenen, Sebastian Gehrmann, Ellen Jiang, Mahima Pushkarna, Carey Radebaugh, Emily Reif, and Ann Yuan. 2020 · 2020
Later among the works it cites.
Evaluating saliency methods for neural language models
Shuoyang Ding and Philipp Koehn. 2021 · 2021
Closest in time.
Towards interpreting and mitigating shortcut learning behavior of NLU models
Mengnan Du, Varun Manjunatha, Rajiv Jain, Ruchi Deshpande, Franck Dernoncourt, Jiuxiang Gu, Tong Sun, and Xia Hu. 2021 · 2021
Closest in time.
Towards benchmarking the utility of explanations for model debugging
Maximilian Idahl, Lijun Lyu, Ujwal Gadiraju, and Avishek Anand. 2021 · 2021
Closest in time.
Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2021 · 2021
Closest in time.
Combining feature and instance attribution to detect artifacts
Pouya Pezeshkpour, Sarthak Jain, Sameer Singh, and Byron C. Wallace. 2021 · 2021
Closest in time.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew Peters. 2021 · 2021
Closest in time.
Integrated directional gradients: Feature interaction attribution for neural NLP models
Sandipan Sikdar, Parantapa Bhattacharya, and Kieran Heese. 2021 · 2021
Closest in time.
Do feature attribution methods correctly attribute features?
Yilun Zhou, Serena Booth, Marco Tílio Ribeiro, and Julie Shah. 2021 · 2021
Closest in time.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022 · 2022
Closest in time.
On the sensitivity and stability of model interpretations in NLP
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022 · 2022
Closest in time.