Fetching the paper…
Reading the bibliography…
Explanations of neural models aim to reveal a model's decision-making process for its predictions.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
e-SNLI: Natural Language Inference with Natural Language Explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Can I Trust the Explainer? Verifying Post-hoc Explanatory Methods
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2019 · 2019
Earlier work this paper cites.
Explanations can be manipulated and geometry is to blame
Ann-Kathrin Dombrowski, Maximillian Alber, Christopher Anders, Marcel Ackermann, Klaus-Robert Müller, and Pan Kessel. 2019 · 2019
Earlier work this paper cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Rethinking cooperative rationalization: Introspective extraction and complement control
Mo Yu, Shiyu Chang, Yang Zhang, and Tommi Jaakkola. 2019 · 2019
Earlier work this paper cites.
Fairwashing explanations with off-manifold detergent
Christopher Anders, Plamen Pasliev, Ann-Kathrin Dombrowski, Klaus-Robert Müller, and Pan Kessel. 2020 · 2020
Earlier work this paper cites.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020a · 2020
Earlier work this paper cites.
Make up your mind! adversarial generation of inconsistent natural language explanations
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, and Phil Blunsom. 2020 · 2020
Earlier work this paper cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
Robustness in machine learning explanations: does it matter?
Leif Hancox-Li. 2020 · 2020
Cited alongside, same era.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
spaCy: Industrial-strength Natural Language Processing in Python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Cited alongside, same era.
WT5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Cited alongside, same era.
Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing
Sarah Wiegreffe and Ana Marasovic. 2021 · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel Weld. 2021 · 2021
Later among the works it cites.
Post hoc explanations may be ineffective for detecting unknown spurious correlation
Julius Adebayo, Michael Muelly, Harold Abelson, and Been Kim. 2022 · 2022
Later among the works it cites.
Counterfactual explanations and how to find them: literature review and benchmarking
Riccardo Guidotti. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Cited alongside, same era.
SemEval-2020 task 4: Commonsense validation and explanation
Cunxiang Wang, Shuailong Liang, Yili Jin, Yilong Wang, Xiaodan Zhu, and Yue Zhang. 2020 · 2020
Cited alongside, same era.
The struggles of feature-based explanations: Shapley values vs. minimal sufficient subsets
Oana-Maria Camburu, Eleonora Giunchiglia, Jakob Foerster, Thomas Lukasiewicz, and Phil Blunsom. 2021 · 2021
Cited alongside, same era.
Evaluations and methods for explanation through robustness analysis
Cheng-Yu Hsieh, Chih-Kuan Yeh, Xuanqing Liu, Pradeep Kumar Ravikumar, Seungyeon Kim, Sanjiv Kumar, and Cho-Jui Hsieh. 2021 · 2021
Cited alongside, same era.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Cited alongside, same era.
e-ViL: A dataset and benchmark for natural language explanations in vision-language tasks
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021 · 2021
Cited alongside, same era.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew Peters. 2021 · 2021
Cited alongside, same era.
Shailza Jolly, Pepa Atanasova, and Isabelle Augenstein. 2022 · 2022
Later among the works it cites.
Explaining chest x-ray pathologies in natural language
Maxime Kayser, Cornelius Emde, Oana-Maria Camburu, Guy Parsons, Bartlomiej Papiez, and Thomas Lukasiewicz. 2022 · 2022
Later among the works it cites.
Evaluating the Faithfulness of Importance Measures in NLP by Recursively Masking Allegedly Important Tokens and Retraining
Andreas Madsen, Nicholas Meade, Vaibhav Adlakha, and Siva Reddy. 2022 · 2022
Later among the works it cites.
Knowledge-grounded self-rationalization via extractive and natural language explanations
Bodhisattwa Prasad Majumder, Oana Camburu, Thomas Lukasiewicz, and Julian Mcauley. 2022 · 2022
Later among the works it cites.
Few-shot self-rationalization with natural language prompts
Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2022 · 2022
Later among the works it cites.
Investigating the benefits of free-form rationales
Jiao Sun, Swabha Swayamdipta, Jonathan May, and Xuezhe Ma. 2022 · 2022
Later among the works it cites.
On the Sensitivity and Stability of Model Interpretations in NLP
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022 · 2022
Later among the works it cites.
Few-Shot Out-of-Domain Transfer of Natural Language Explanations
Yordan Yordanov, Vid Kocijan, Thomas Lukasiewicz, and Oana-Maria Camburu. 2022 · 2022
Later among the works it cites.