Fetching the paper…
Reading the bibliography…
Free-form rationales aim to aid model interpretability by supplying the background knowledge that can help understand model decisions.
Zhiquan Ye, Qian Chen, Wen Wang, and Zhenhua Ling. 2019 · 1908
Earlier work this paper cites.
The equivalence of weighted kappa and the intraclass correlation coefficient as measures of reliability
Joseph L. Fleiss and Jacob Cohen. 1973 · 1973
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Wt5?! training text-to-text models to explain their predictions
Sharan Narang, Colin Raffel, Katherine Lee, Adam Roberts, Noah Fiedel, and Karishma Malkan. 2020 · 2004
Earlier work this paper cites.
Interactive and interpretable machine learning models for human machine collaboration
Been Kim. 2015 · 2015
Earlier work this paper cites.
Generating visual explanations
Lisa Anne Hendricks, Zeynep Akata, Marcus Rohrbach, Jeff Donahue, Bernt Schiele, and Trevor Darrell. 2016 · 2016
Earlier work this paper cites.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Earlier work this paper cites.
Towards robust interpretability with self-explaining neural networks
David Alvarez-Melis and T. Jaakkola. 2018 · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom. 2018 · 2018
Earlier work this paper cites.
Textual explanations for self-driving vehicles
Jinkyu Kim, Anna Rohrbach, Trevor Darrell, John Canny, and Zeynep Akata. 2018 · 2018
Earlier work this paper cites.
The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery
Zachary C Lipton. 2018 · 2018
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Fatema Rajani, Bryan McCann, Caiming Xiong, and Richard Socher. 2019 · 2019
Earlier work this paper cites.
Do human rationales improve machine explanations?
Julia Strout, Ye Zhang, and Raymond Mooney. 2019 · 2019
Earlier work this paper cites.
QuaRTz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark. 2019 · 2019
Earlier work this paper cites.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Earlier work this paper cites.
Shortcut learning in deep neural networks
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A. Wichmann. 2020 · 2020
Earlier work this paper cites.
Why do you think that? exploring faithful sentence-level rationales without supervision
Max Glockner, Ivan Habernal, and Iryna Gurevych. 2020 · 2020
Cited alongside, same era.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020 · 2020
Cited alongside, same era.
Summit: Scaling deep learning interpretability by visualizing activation and attribution summarizations
Fred Hohman, Haekyu Park, Caleb Robinson, and Duen Horng Chau. 2020 · 2020
Cited alongside, same era.
Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness?
Alon Jacovi and Yoav Goldberg. 2020 · 2020
Cited alongside, same era.
Contrastive explanations for model interpretability
Alon Jacovi, Swabha Swayamdipta, Shauli Ravfogel, Yanai Elazar, Yejin Choi, and Yoav Goldberg. 2021 · 2021
Later among the works it cites.
e-vil: A dataset and benchmark for natural language explanations in vision-language tasks
Maxime Kayser, Oana-Maria Camburu, Leonard Salewski, Cornelius Emde, Virginie Do, Zeynep Akata, and Thomas Lukasiewicz. 2021 · 2021
Later among the works it cites.
FiD-ex: Improving sequence-to-sequence models for extractive rationale generation
Kushal Lakhotia, Bhargavi Paranjape, Asish Ghoshal, Scott Yih, Yashar Mehdad, and Srini Iyer. 2021 · 2021
Later among the works it cites.
Rationale-inspired natural language explanations with commonsense
Bodhisattwa Prasad Majumder, Oana-Maria Camburu, Thomas Lukasiewicz, and Julian McAuley. 2021 · 2021
Later among the works it cites.
Prompting contrastive explanations for commonsense reasoning tasks
Bhargavi Paranjape, Julian Michael, Marjan Ghazvininejad, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
NILE : Natural language inference with faithful natural language explanations
Sawan Kumar and Partha Talukdar. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Cited alongside, same era.
Social bias frames: Reasoning about social and power implications of language
Maarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Cited alongside, same era.
Explanations for CommonsenseQA: New Dataset and Models
Shourya Aggarwal, Divyanshu Mandowara, Vishwajeet Agrawal, Dinesh Khandelwal, Parag Singla, and Dinesh Garg. 2021 · 2021
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021 · 2021
Cited alongside, same era.
Learning to rationalize for nonmonotonic reasoning with distant supervision
Faeze Brahman, Vered Shwartz, Rachel Rudinger, and Yejin Choi. 2021 · 2021
Cited alongside, same era.
Later among the works it cites.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna M. Wallach. 2021 · 2021
Later among the works it cites.
SELFEXPLAIN: A self-explaining architecture for neural text classifiers
Dheeraj Rajagopal, Vidhisha Balachandran, Eduard H Hovy, and Yulia Tsvetkov. 2021 · 2021
Later among the works it cites.
AESOP: Paraphrase generation with adaptive syntactic control
Jiao Sun, Xuezhe Ma, and Nanyun Peng. 2021 · 2021
Later among the works it cites.
When does text prediction benefit from additional context? an exploration of contextual signals for chat and email messages
Stojan Trajanovski, Chad Atalla, Kunho Kim, Vipul Agarwal, Milad Shokouhi, and Chris Quirk. 2021 · 2021
Later among the works it cites.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasovic. 2021 · 2021
Later among the works it cites.
Measuring association between labels and free-text rationales
Sarah Wiegreffe, Ana Marasović, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Lirex: Augmenting language inference with relevant explanations
Xinyan Zhao and V.G.Vinod Vydiswaran. 2021 · 2021
Later among the works it cites.
When can models learn from explanations? a formal framework for understanding the roles of explanation data
Peter Hase and Mohit Bansal. 2022 · 2022
Closest in time.
Few-shot self-rationalization with natural language prompts
Ana Marasović, Iz Beltagy, Doug Downey, and Matthew E. Peters. 2022 · 2022
Closest in time.
On the diversity and limits of human explanations
Chenhao Tan. 2022 · 2022
Closest in time.
Reframing human-AI collaboration for generating free-text explanations
Sarah Wiegreffe, Jack Hessel, Swabha Swayamdipta, Mark Riedl, and Yejin Choi. 2022 · 2022
Closest in time.