Fetching the paper…
Reading the bibliography…
As Transformers are increasingly relied upon to solve complex NLP problems, there is an increased need for their decisions to be humanly interpretable.
Visualizing attention in transformer-based language representation models
Jesse Vig. 2019 · 1904
Earlier work this paper cites.
Captum: A unified and generic model interpretability library for pytorch
Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsallakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, et al. 2020 · 2009
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014a · 2014
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014b · 2014
Earlier work this paper cites.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Earlier work this paper cites.
"Why Should I Trust You?": Explaining the Predictions of Any Classifier
Marco Túlio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee. 2017 · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
Earlier work this paper cites.
Sanity checks for saliency maps
Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. 2018 · 2018
Earlier work this paper cites.
Pathologies of neural models make interpretations difficult
Shi Feng, Eric Wallace, Alvin Grissom II, Mohit Iyyer, Pedro Rodriguez, and Jordan Boyd-Graber. 2018 · 2018
Earlier work this paper cites.
AllenNLP: A deep semantic natural language processing platform
Matt Gardner, Joel Grus, Mark Neumann, Oyvind Tafjord, Pradeep Dasigi, Nelson F. Liu, Matthew Peters, Michael Schmitz, and Luke Zettlemoyer. 2018 · 2018
Earlier work this paper cites.
Visual interrogation of attention-based models for natural language inference and machine comprehension
Shusen Liu, Tao Li, Zhimin Li, Vivek Srikumar, Valerio Pascucci, and Peer-Timo Bremer. 2018 · 2018
Earlier work this paper cites.
Perturbation-based explanations of prediction models
Marko Robnik-Šikonja and Marko Bohanec. 2018 · 2018
Earlier work this paper cites.
S eq 2s eq-v is: A visual debugging tool for sequence-to-sequence models
Hendrik Strobelt, Sebastian Gehrmann, Michael Behrisch, Adam Perer, Hanspeter Pfister, and Alexander M Rush. 2018 · 2018
Earlier work this paper cites.
Evaluating recurrent neural network explanations
Leila Arras, Ahmed Osman, Klaus-Robert Müller, and Wojciech Samek. 2019 · 2019
Earlier work this paper cites.
Towards a deep and unified understanding of deep neural models in nlp
Chaoyu Guan, Xiting Wang, Quanshi Zhang, Runjin Chen, Di He, and Xing Xie. 2019 · 2019
Cited alongside, same era.
A benchmark for interpretability methods in deep neural networks
Sara Hooker, Dumitru Erhan, Pieter-Jan Kindermans, and Been Kim. 2019 · 2019
Cited alongside, same era.
Attention is not Explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Cited alongside, same era.
Towards hierarchical importance attribution: Explaining compositional semantics for neural sequence models
Xisen Jin, Zhongyu Wei, Junyi Du, Xiangyang Xue, and Xiang Ren. 2019 · 2019
Cited alongside, same era.
Human-grounded evaluations of explanation methods for text classification
Piyawat Lertvittayakumjorn and Francesca Toni. 2019 · 2019
Cited alongside, same era.
Explaining black box models by means of local rules
Modeling annotators: A generative approach to learning from annotator rationales
Omar Zaidan and Jason Eisner. 2008 · 2020
Later among the works it cites.
XLM-T: A multilingual language model toolkit for twitter
Francesco Barbieri, Luis Espinosa Anke, and José Camacho-Collados. 2021 · 2021
Later among the works it cites.
Thermostat: A large collection of NLP model explanations and analysis tools
Nils Feldhus, Robert Schwarzenberg, and Sebastian Möller. 2021 · 2021
Later among the works it cites.
Robustness gym: Unifying the NLP evaluation landscape
Karan Goel, Nazneen Fatema Rajani, Jesse Vig, Zachary Taschdjian, Mohit Bansal, and Christopher Ré. 2021 · 2021
Later among the works it cites.
Hatexplain: A benchmark dataset for explainable hate speech detection
Binny Mathew, Punyajoy Saha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Eliana Pastor and Elena Baralis. 2019 · 2019
Cited alongside, same era.
Universal adversarial triggers for attacking and analyzing NLP
Eric Wallace, Shi Feng, Nikhil Kandpal, Matt Gardner, and Sameer Singh. 2019a · 2019
Cited alongside, same era.
AllenNLP interpret: A framework for explaining predictions of NLP models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019b · 2019
Cited alongside, same era.
A diagnostic study of explainability techniques for text classification
Pepa Atanasova, Jakob Grue Simonsen, Christina Lioma, and Isabelle Augenstein. 2020 · 2020
Cited alongside, same era.
Evaluating and characterizing human rationales
Samuel Carton, Anirudh Rathore, and Chenhao Tan. 2020 · 2020
Cited alongside, same era.
A survey of the state of explainable AI for natural language processing
Marina Danilevsky, Kun Qian, Ranit Aharonov, Yannis Katsis, Ban Kawas, and Prithviraj Sen. 2020 · 2020
Cited alongside, same era.
ERASER: A benchmark to evaluate rationalized NLP models
Jay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric Lehman, Caiming Xiong, Richard Socher, and Byron C. Wallace. 2020 · 2020
Cited alongside, same era.
Looking for trouble: Analyzing classifier behavior via pattern divergence
Eliana Pastor, Luca de Alfaro, and Elena Baralis. 2021a · 2021
Later among the works it cites.
Explaining NLP models via minimal contrastive editing (MiCE)
Alexis Ross, Ana Marasović, and Matthew Peters. 2021 · 2021
Later among the works it cites.
Discretized integrated gradients for explaining language models
Soumya Sanyal and Xiang Ren. 2021 · 2021
Later among the works it cites.
TextFlint: Unified multilingual robustness evaluation toolkit for natural language processing
Xiao Wang, Qin Liu, Tao Gui, Qi Zhang, Yicheng Zou, Xin Zhou, Jiacheng Ye, Yongxin Zhang, Rui Zheng, Zexiong Pang, Qinzhuo Wu, Zhengyan Li, Chong Zhang, Ruotian Ma, Zichu Fei, Ruijian Cai, Jun Zhao, Xingwu Hu, Zhiheng Yan, Yiding Tan, Yuan Hu, Qiyuan Bian, Zhihua Liu, Shan Qin, Bolin Zhu, Xiaoyu Xing, Jinlan Fu, Yue Zhang, Minlong Peng, Xiaoqing Zheng, Yaqian Zhou, Zhongyu Wei, Xipeng Qiu, and Xuanjing Huang. 2021 · 2021
Later among the works it cites.
Teach me to explain: A review of datasets for explainable natural language processing
Sarah Wiegreffe and Ana Marasovic. 2021 · 2021
Later among the works it cites.
OpenXAI: Towards a transparent evaluation of model explanations
Chirag Agarwal, Satyapriya Krishna, Eshika Saxena, Martin Pawelczyk, Nari Johnson, Isha Puri, Marinka Zitnik, and Himabindu Lakkaraju. 2022 · 2022
Closest in time.
Benchmarking post-hoc interpretability approaches for transformer-based misogyny detection
Giuseppe Attanasio, Debora Nozza, Eliana Pastor, and Dirk Hovy. 2022 · 2022
Closest in time.
Post-hoc interpretability for neural nlp: A survey
Andreas Madsen, Siva Reddy, and Sarath Chandar. 2022 · 2022
Closest in time.
Contrastive explanations of text classifiers as a service
Lorenzo Malandri, Fabio Mercorio, Mario Mezzanzanica, Navid Nobani, and Andrea Seveso. 2022 · 2022
Closest in time.
On the sensitivity and stability of model interpretations in NLP
Fan Yin, Zhouxing Shi, Cho-Jui Hsieh, and Kai-Wei Chang. 2022 · 2022
Closest in time.
Inseq: An interpretability toolkit for sequence generation models
Gabriele Sarti, Nils Feldhus, Ludwig Sickert, and Oskar van der Wal. 2023 · 2023
Closest in time.