Fetching the paper…
Reading the bibliography…
In interpretable NLP, we require faithful rationales that reflect the model's decision-making process for an explained instance.
An information bottleneck approach for controlling conciseness in rationale extraction
Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 1952
Earlier work this paper cites.
Newsweeder: Learning to filter netnews
Ken Lang. 1995 · 1995
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Explaining question answering models through text generation
Veronica Latcinnik and Jonathan Berant. 2020 · 2004
Earlier work this paper cites.
WT5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Ranking a stream of news
Gianna M Del Corso, Antonio Gulli, and Francesco Romani. 2005 · 2005
Earlier work this paper cites.
Ontonotes: A unified relational semantic representation
Sameer S Pradhan, Eduard Hovy, Mitch Marcus, Martha Palmer, Lance Ramshaw, and Ralph Weischedel. 2007 · 2007
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner. 2008 · 2008
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert Müller. 2010 · 2010
Earlier work this paper cites.
Learning attitudes and attributes from multi-aspect reviews
Julian McAuley, Jure Leskovec, and Dan Jurafsky. 2012 · 2012
Earlier work this paper cites.
Hidden factors and hidden topics: Understanding rating dimensions with review text
Julian McAuley and Jure Leskovec. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D. Manning, Andrew Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Been Kim, Rajiv Khanna, and Oluwasanmi Koyejo. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Assessing the ability of lstms to learn syntax-sensitive dependencies
Tal Linzen, Emmanuel Dupoux, and Yoav Goldberg. 2016 · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
e-SNLI: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. 2018 · 2018
Earlier work this paper cites.
Rationalization: A neural machine translation approach to generating natural language explanations
Upol Ehsan, Brent Harrison, Larry Chan, and Mark O Riedl. 2018 · 2018
Earlier work this paper cites.
Training classifiers with natural language explanations
Braden Hancock, Paroma Varma, Stephanie Wang, Martin Bringmann, Percy Liang, and Christopher Ré. 2018 · 2018
Earlier work this paper cites.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata. 2018 · 2018
Cited alongside, same era.
Towards robust interpretability with self-explaining neural networks
David Alvarez Melis and Tommi Jaakkola. 2018 · 2018
Cited alongside, same era.
Bridging CNNs, RNNs, and weighted finite-state machines
Roy Schwartz, Sam Thomson, and Noah A. Smith. 2018 · 2018
Cited alongside, same era.
FEVER: a large-scale dataset for fact extraction and VERification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal. 2018 · 2018
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
BoolQ: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. 2019 · 2019
Faithful multimodal explanation for visual question answering
Jialin Wu and Raymond Mooney. 2019 · 2019
Later among the works it cites.
Analyzing the interpretability robustness of self-explaining models
Haizhong Zheng, Earlence Fernandes, and Atul Prakash. 2019 · 2019
Later among the works it cites.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Closest in time.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Closest in time.
e-SNLI-VE-2.0: Corrected visual-textual entailment with natural language explanations
Virginie Do, Oana-Maria Camburu, Zeynep Akata, and Thomas Lukasiewicz. 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Saliency-driven word alignment interpretation for neural machine translation
Shuoyang Ding, Hainan Xu, and Philipp Koehn. 2019 · 2019
Cited alongside, same era.
Interpretation of neural networks is fragile
Amirata Ghorbani, Abubakar Abid, and James Zou. 2019 · 2019
Cited alongside, same era.
Towards understanding neural machine translation with word importance
Shilin He, Zhaopeng Tu, Xing Wang, Longyue Wang, Michael Lyu, and Shuming Shi. 2019 · 2019
Cited alongside, same era.
Fooling neural network interpretations via adversarial model manipulation
Juyeon Heo, Sunghwan Joo, and Taesup Moon. 2019 · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Closest in time.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020 · 2020
Closest in time.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Closest in time.
Learning to faithfully rationalize by construction
Sarthak Jain, Sarah Wiegreffe, Yuval Pinter, and Byron C. Wallace. 2020 · 2020
Closest in time.
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha Talukdar. 2020 · 2020
Closest in time.
Evaluating explanation methods for neural machine translation
Jierui Li, Lemao Liu, Huayang Li, Guanlin Li, Guoping Huang, and Shuming Shi. 2020 · 2020
Closest in time.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Closest in time.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Closest in time.
F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering
Hendrik Schuff, Heike Adel, and Ngoc Thang Vu. 2020 · 2020
Closest in time.
Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Closest in time.
Exploring interpretability in event extraction: Multitask learning of a neural event classifier and an explanation decoder
Zheng Tang, Gus Hahn-Powell, and Mihai Surdeanu. 2020 · 2020
Closest in time.
Staying true to your word: (how) can attention become explanation?
Martin Tutek and Jan Snajder. 2020 · 2020
Closest in time.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Closest in time.
Sequential explanations with mental model-based policies
Arnold Yeung, Shalmali Joshi, Joseph Jay Williams, and Frank Rudzicz. 2020 · 2020
Closest in time.
Interpretable deep learning under fire
Xinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji, Xiapu Luo, and Ting Wang. 2020 · 2020
Closest in time.
A study of automatic metrics for the evaluation of natural language explanations
Miruna-Adriana Clinciu, Arash Eshghi, and Helen Hastie. 2021 · 2021
Closest in time.
Aligning faithful interpretations with their social attribution
Alon Jacovi and Yoav Goldberg. 2021 · 2021
Closest in time.
e-vil: A dataset and benchmark for natural language explanations in vision-language tasks
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021 · 2021
Closest in time.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. 2021 · 2021
Closest in time.
Selfexplain: A self-explaining architecture for neural text classifiers
Dheeraj Rajagopal, Vidhisha Balachandran, Eduard Hovy, and Yulia Tsvetkov. 2021 · 2021
Closest in time.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasović. 2021 · 2021
Closest in time.