Fetching the paper…
Reading the bibliography…
A growing line of work has investigated the development of neural NLP models that can produce rationales--subsets of input that can explain their model predictions.
WordNet: An Electronic Lexical Database
Christiane Fellbaum. 1998 · 1998
Earlier work this paper cites.
WT5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Latent aspect rating analysis on review text data: A rating regression approach
Hongning Wang, Yue Lu, and Chengxiang Zhai. 2010 · 2010
Earlier work this paper cites.
Learning attitudes and attributes from multi-aspect reviews
Julian McAuley, Jure Leskovec, and Dan Jurafsky. 2012 · 2012
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
"Why should i trust you?" explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Deep variational information bottleneck
Alexander Alemi, Ian Fischer, Joshua Dillon, and Kevin Murphy. 2017 · 2017
Earlier work this paper cites.
A causal framework for explaining the predictions of black-box sequence-to-sequence models
David Alvarez-Melis and Tommi Jaakkola. 2017 · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole. 2017 · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Earlier work this paper cites.
On the robustness of interpretability methods
David Alvarez-Melis and Tommi S Jaakkola. 2018 · 2018
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Cited alongside, same era.
Adversarial example generation with syntactically controlled paraphrase networks
Mohit Iyyer, John Wieting, Kevin Gimpel, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Attention is not explanation
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary C. Lipton. 2020 · 2020
Later among the works it cites.
Look at the first sentence: Position bias in question answering
Miyoung Ko, Jinhyuk Lee, Hyunjae Kim, Gangwoo Kim, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
Lp-sparsemap: Differentiable relaxed optimization for sparse structured prediction
Vlad Niculae and F. T. André Martins. 2020 · 2020
Later among the works it cites.
An information bottleneck approach for controlling conciseness in rationale extraction
Bhargavi Paranjape, Mandar Joshi, John Thickstun, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020 · 2020
Later among the works it cites.
Fooling lime and shap: Adversarial attacks on post hoc explanation methods
Dylan Slack, Sophie Hilgard, Emily Jia, Sameer Singh, and Himabindu Lakkaraju. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sarthak Jain and Byron C Wallace. 2019 · 2019
Cited alongside, same era.
Is attention interpretable?
Sofia Serrano and Noah A Smith. 2019 · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019 · 2019
Cited alongside, same era.
Rethinking cooperative rationalization: Introspective extraction and complement control
Mo Yu, Shiyu Chang, Yang Zhang, and Tommi Jaakkola. 2019 · 2019
Cited alongside, same era.
Make up your mind! adversarial generation of inconsistent natural language explanations
Oana-Maria Camburu, Brendan Shillingford, Pasquale Minervini, Thomas Lukasiewicz, and Phil Blunsom. 2020 · 2020
Cited alongside, same era.
Invariant rationalization
Shiyu Chang, Yang Zhang, Mo Yu, and Tommi S. Jaakkola. 2020 · 2020
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized nlp models
Jay DeYoung, Sarthak Jain, Nazneen F. Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Cited alongside, same era.
Rationalizing text matching: Learning sparse alignments via optimal transport
Kyle Swanson, Lili Yu, and Tao Lei. 2020 · 2020
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Rémi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander M. Rush. 2020 · 2020
Later among the works it cites.
Do feature attribution methods correctly attribute features?
Yilun Zhou, Serena Booth, Marco Tulio Ribeiro, and Julie Shah. 2020 · 2020
Later among the works it cites.
SPECTRA: Sparse structured text rationalization
Nuno Miguel Guerreiro and André F. T. Martins. 2021 · 2021
Later among the works it cites.
An empirical study on the relation between network interpretability and adversarial robustness
Adam Noack, Isaac Ahern, Dejing Dou, and Boyang Li. 2021 · 2021
Later among the works it cites.
Learning from the best: Rationalizing prediction by adversarial information calibration
Lei Sha, Oana-Maria Camburu, and Thomas Lukasiewicz. 2021 · 2021
Later among the works it cites.
Rationales for sequential predictions
Keyon Vafa, Yuntian Deng, David Blei, and Alexander M Rush. 2021 · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith. 2021 · 2021
Later among the works it cites.