BERT: pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Explanations based on the missing: Towards contrastive explanations with pertinent negatives
Amit Dhurandhar, Pin-Yu Chen, Ronny Luss, Chun-Chen Tu, Pai-Shun Ting, Karthikeyan Shanmugam, and Payel Das. 2018 · 2018
Later among the works it cites.
Pathologies of neural models make interpretation difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan L. Boyd-Graber. 2018 · 2018
Later among the works it cites.
The mythos of model interpretability
Zachary C. Lipton. 2018 · 2018
Later among the works it cites.
Interpretable neural predictions with differentiable binary variables
Jasmijn Bastings, Wilker Aziz, and Ivan Titov. 2019 · 2019
Later among the works it cites.
A game theoretic approach to class-wise selective rationalization
Shiyu Chang, Yang Zhang, Mo Yu, and Tommi S. Jaakkola. 2019 · 2019
Later among the works it cites.
Eraser: A benchmark to evaluate rationalized nlp models
Original
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
The (un)reliability of saliency methods
Pieter-Jan Kindermans, Sara Hooker, Julius Adebayo, Maximilian Alber, Kristof T. Schütt, Sven Dähne, Dumitru Erhan, and Been Kim. 2019 · 2019
Later among the works it cites.
Explanation in artificial intelligence: Insights from the social sciences
Tim Miller. 2019 · 2019
Later among the works it cites.
Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead
Cynthia Rudin. 2019 · 2019
Later among the works it cites.
Is attention interpretable?
Sofia Serrano and Noah A. Smith. 2019 · 2019
Later among the works it cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Leveraging rationales to improve human task performance
Devleena Das and Sonia Chernova. 2020 · 2020
Closest in time.
Why attention is not explanation: Surgical intervention and causal reasoning about neural models
Christopher Grimsley, Elijah Mayfield, and Julia R.S. Bursten. 2020 · 2020
Closest in time.
Explainable reinforcement learning through a causal lens
Prashan Madumal, Tim Miller, Liz Sonenberg, and Frank Vetere. 2020 · 2020
Closest in time.
Contrastive explanation: A structural-model approach
Original
Tim Miller. 2020 · 2020
Closest in time.