Fetching the paper…
Reading the bibliography…
While many methods purport to explain predictions by highlighting salient features, what aims these explanations serve and how they ought to be evaluated often go unstated.
Fine-grained sentiment analysis with faithful attention
Ruiqi Zhong, Steven Shao, and Kathleen McKeown. 2019 · 1908
Earlier work this paper cites.
Deriving machine attention from human rationales
Yujia Bao, Shiyu Chang, Mo Yu, and Regina Barzilay. 2018 · 1913
Earlier work this paper cites.
Conditional random fields: Probabilistic models for segmenting and labeling sequence data
John Lafferty, Andrew McCallum, and Fernando CN Pereira. 2001 · 2001
Earlier work this paper cites.
Using “annotator rationales” to improve machine learning for text categorization
Omar Zaidan, Jason Eisner, and Christine Piatko. 2007 · 2007
Earlier work this paper cites.
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner. 2008 · 2008
Earlier work this paper cites.
Reading tea leaves: How humans interpret topic models
Jonathan Chang, Sean Gerrish, Chong Wang, Jordan Boyd-graber, and David Blei. 2009 · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton. 2016 · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Cited alongside, same era.
The enigma of reason
Hugo Mercier and Dan Sperber. 2017 · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Do explanations make VQA models more predictable to a human?
Arjun Chandrasekaran, Viraj Prabhu, Deshraj Yadav, Prithvijit Chattopadhyay, and Devi Parikh. 2018 · 2018
Meta-learning with implicit gradients
Aravind Rajeswaran, Chelsea Finn, Sham M. Kakade, and Sergey Levine. 2019 · 2019
Later among the works it cites.
Invariant rationalization
Shiyu Chang, Yang Zhang, Mo Yu, and Tommi S. Jaakkola. 2020 · 2020
Closest in time.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Closest in time.
Evaluating explainable AI: Which algorithmic explanations help users predict model behavior?
Peter Hase and Mohit Bansal. 2020 · 2020
Closest in time.
Leakage-adjusted simulatability: Can models generate non-trivial explanations of their behavior in natural language?
Peter Hase, Shiyue Zhang, Harry Xie, and Mohit Bansal. 2020 · 2020
Closest in time.
SpanBERT: Improving pre-training by representing and predicting spans
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Evaluating neural network explanation methods using hybrid documents and morphosyntactic agreement
Nina Poerner, Hinrich Schütze, and Benjamin Roth. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
How important is a neuron
Kedar Dhamdhere, Mukund Sundararajan, and Qiqi Yan. 2019 · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Cited alongside, same era.
Natural questions: a benchmark for question answering research
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al. 2019 · 2019
Cited alongside, same era.
Mandar Joshi, Danqi Chen, Yinhan Liu, Daniel S Weld, Luke Zettlemoyer, and Omer Levy. 2020 · 2020
Closest in time.
Weakly-and semi-supervised evidence extraction
Danish Pruthi, Bhuwan Dhingra, Graham Neubig, and Zachary C Lipton. 2020 · 2020
Closest in time.
The explanation game: Towards prediction explainability through sparse communication
Marcos V. Treviso and André F. T. Martins. 2020 · 2020
Closest in time.
Aligning Faithful Interpretations with their Social Attribution
Alon Jacovi and Yoav Goldberg. 2021 · 2021
Closest in time.
QED: A Framework and Dataset for Explanations in Question Answering
Matthew Lamm, Jennimaria Palomaki, Chris Alberti, Daniel Andor, Eunsol Choi, Livio Baldini Soares, and Michael Collins. 2021 · 2021
Closest in time.
Do context-aware translation models pay the right attention?
Kayo Yin, Patrick Fernandes, Danish Pruthi, Aditi Chaudhary, André F. T. Martins, and Graham Neubig. 2021 · 2021
Closest in time.