Fetching the paper…
Reading the bibliography…
Evaluation beyond aggregate performance metrics, e.g.
Allennlp interpret: A framework for explaining predictions of nlp models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 1909
Earlier work this paper cites.
De meulder, 2003. tjong kim sang, ef, & de meulder, f.(2003). introduction to the conll-2003 shared task: language-independent named entity recognition
Tjong Kim Sang. 2003 · 2003
Earlier work this paper cites.
Too much, too little, or just right? ways explanations impact end users’ mental models
Todd Kulesza, Simone Stumpf, Margaret Burnett, Sherry Yang, Irwin Kwan, and Weng-Keen Wong. 2013 · 2013
Earlier work this paper cites.
An error analysis tool for natural language processing and applied machine learning
Apoorv Agarwal, Ankit Agarwal, and Deepak Mittal. 2014 · 2014
Earlier work this paper cites.
Visual exploration of machine learning results using data cube analysis
Minsuk Kahng, Dezhi Fang, and Duen Horng Polo Chau. 2016 · 2016
Earlier work this paper cites.
The mythos of model interpretability
Zachary C Lipton. 2016 · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Cited alongside, same era.
Bidirectional attention flow for machine comprehension
Minjoon Seo, Aniruddha Kembhavi, Ali Farhadi, and Hannaneh Hajishirzi. 2016 · 2016
Cited alongside, same era.
Results of the wnut2017 shared task on novel and emerging entity recognition
Leon Derczynski, Eric Nichols, Marieke van Erp, and Nut Limsopatham. 2017 · 2017
Cited alongside, same era.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Cited alongside, same era.
ActiVis: Visual exploration of industry-scale deep neural network models
Minsuk Kahng, Pierre Y Andrews, Aditya Kalro, and Duen Horng Polo Chau. 2017 · 2017
Cited alongside, same era.
Explaining explanation, part 4: a deep dive on deep nets
Robert Hoffman, Tim Miller, Shane T Mueller, Gary Klein, and William J Clancey. 2018 · 2018
Later among the works it cites.
Crowdsourcing a large corpus of clickbait on twitter
Martin Potthast, Tim Gollub, Kristof Komlossy, Sebastian Schuster, Matti Wiegmann, Erika Patricia Garces Fernandez, Matthias Hagen, and Benno Stein. 2018 · 2018
Later among the works it cites.
Manipulating and measuring model interpretability
Forough Poursabzi-Sangdeh, Daniel G Goldstein, Jake M Hofman, Jennifer Wortman Vaughan, and Hanna Wallach. 2018 · 2018
Later among the works it cites.
Manifold: A model-agnostic framework for interpretation and diagnosis of machine learning models
Jiawei Zhang, Yang Wang, Piero Molino, Lezhi Li, and David S Ebert. 2018 · 2018
Later among the works it cites.
Gamut: A design probe to understand how data scientists understand machine learning models
Fred Hohman, Andrew Head, Rich Caruana, Robert DeLine, and Steven M Drucker. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Semi-supervised sequence tagging with bidirectional language models
Matthew E Peters, Waleed Ammar, Chandra Bhagavatula, and Russell Power. 2017 · 2017
Cited alongside, same era.
Named entity recognition on code-switched data: Overview of the calcs 2018 shared task
Gustavo Aguilar, Fahad AlGhamdi, Victor Soto, Mona Diab, Julia Hirschberg, and Thamar Solorio. 2018 · 2018
Cited alongside, same era.
Qadiver: Interactive framework for diagnosing qa models
Gyeongbok Lee, Sungdong Kim, and Seung-won Hwang. 2019 · 2019
Later among the works it cites.
Errudite: Scalable, reproducible, and testable error analysis
Tongshuang Wu, Marco Tulio Ribeiro, Jeffrey Heer, and Daniel S Weld. 2019 · 2019
Later among the works it cites.