Fetching the paper…
Reading the bibliography…
We build on abduction-based explanations for ma-chine learning and develop a method for computing local explanations for neural network models in natural language processing (NLP).
A theory of diagnosis from first principles
Raymond Reiter · 1987
Earlier work this paper cites.
Twitter sentiment classification using distant supervision
Alec Go, Richa Bhayani, and Lei Huang · 2009
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Earlier work this paper cites.
Minimum satisfying assignments for SMT
Isil Dillig, Thomas Dillig, Kenneth L McMillan, and Alex Aiken · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Don’t count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors
Marco Baroni, Georgiana Dinu, and Germán Kruszewski · 2014
Earlier work this paper cites.
Retrofitting word vectors to semantic lexicons
Manaal Faruqui, Jesse Dodge, Sujay K Jauhar, Chris Dyer, Eduard Hovy, and Noah A Smith · 2014
Earlier work this paper cites.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
Sebastian Bach, Alexander Binder, Grégoire Montavon, Frederick Klauschen, Klaus-Robert Müller, and Wojciech Samek · 2015
Earlier work this paper cites.
Interpretable classifiers using rules and bayesian analysis: Building a better stroke prediction model
Benjamin Letham, Cynthia Rudin, Tyler H McCormick, David Madigan, et al · 2015
Earlier work this paper cites.
Visualizing and understanding neural models in nlp
Jiwei Li, Xinlei Chen, Eduard Hovy, and Dan Jurafsky · 2015
Earlier work this paper cites.
Falling rule lists
Fulton Wang and Cynthia Rudin · 2015
Earlier work this paper cites.
Explaining predictions of non-linear classifiers in nlp
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2016
Earlier work this paper cites.
On finding minimum satisfying assignments
Alexey Ignatiev, Alessandro Previti, and Joao Marques-Silva · 2016
Earlier work this paper cites.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M Dai, and Ian Goodfellow · 2016
Earlier work this paper cites.
” why should i trust you?” explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
Explaining recurrent neural network predictions in sentiment analysis
Leila Arras, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2017
Cited alongside, same era.
Interpretability of deep learning models: a survey of results
Supriyo Chakraborty, Richard Tomsett, Ramya Raghavendra, Daniel Harborne, Moustafa Alzantot, Federico Cerutti, Mani Srivastava, Alun Preece, Simon Julier, Raghuveer M Rao, et al · 2017
Cited alongside, same era.
Visualizing and understanding neural machine translation
Yanzhuo Ding, Yang Liu, Huanbo Luan, and Maosong Sun · 2017
Cited alongside, same era.
Towards lower bounds on number of dimensions for word embeddings
Kevin Patel and Pushpak Bhattacharyya · 2017
Cited alongside, same era.
Model-based diagnosis with multiple observations
Alexey Ignatiev, António Morgado, Georg Weissenbacher, Joao Marques-Silva, and ISDCT SB RAS · 2019
Later among the works it cites.
Abduction-based explanations for machine learning models
Alexey Ignatiev, Nina Narodytska, and Joao Marques-Silva · 2019
Later among the works it cites.
On relating explanations and adversarial examples
Alexey Ignatiev, Nina Narodytska, and Joao Marques-Silva · 2019
Later among the works it cites.
On validating, repairing and refining heuristic ml explanations
Alexey Ignatiev, Nina Narodytska, and Joao Marques-Silva · 2019
Later among the works it cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wojciech Samek, Thomas Wiegand, and Klaus-Robert Müller · 2017
Cited alongside, same era.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang · 2018
Cited alongside, same era.
Keras: The python deep learning library
François Chollet et al · 2018
Cited alongside, same era.
Improving the interpretability of deep neural networks with knowledge distillation
Xuan Liu, Xiaoguang Wang, and Stan Matwin · 2018
Cited alongside, same era.
Anchors: High-precision model-agnostic explanations
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2018
Cited alongside, same era.
Reachability analysis of deep neural networks with provable guarantees
Wenjie Ruan, Xiaowei Huang, and Marta Kwiatkowska · 2018
Cited alongside, same era.
Hierarchical interpretations for neural network predictions
Chandan Singh, W James Murdoch, and Bin Yu · 2018
Cited alongside, same era.
The marabou framework for verification and analysis of deep neural networks
Guy Katz, Derek A Huang, Duligur Ibeling, Kyle Julian, Christopher Lazarus, Rachel Lim, Parth Shah, Shantanu Thakoor, Haoze Wu, Aleksandar Zeljić, et al · 2019
Later among the works it cites.
Popqorn: Quantifying robustness of recurrent neural networks
Ching-Yun Ko, Zhaoyang Lyu, Lily Weng, Luca Daniel, Ngai Wong, and Dahua Lin · 2019
Later among the works it cites.
Assessing heuristic machine learning explanations with model counting
Nina Narodytska, Aditya Shrotri, Kuldeep S Meel, Alexey Ignatiev, and Joao Marques-Silva · 2019
Later among the works it cites.
Interpreting cnns via decision trees
Quanshi Zhang, Yu Yang, Haotian Ma, and Ying Nian Wu · 2019
Later among the works it cites.
Sparse-rs: a versatile framework for query-efficient sparse black-box adversarial attacks
Francesco Croce, Maksym Andriushchenko, Naman D Singh, Nicolas Flammarion, and Matthias Hein · 2020
Later among the works it cites.
On the reasons behind decisions
Adnan Darwiche and Auguste Hirth · 2020
Later among the works it cites.
Three modern roles for logic in ai
Adnan Darwiche · 2020
Later among the works it cites.
Assessing robustness of text classification through maximal safe radius computation
Emanuele La Malfa, Min Wu, Luca Laurenti, Benjie Wang, Anthony Hartshorn, and Marta Kwiatkowska · 2020
Later among the works it cites.
On tractable representations of binary neural networks
Weijia Shi, Andy Shih, Adnan Darwiche, and Arthur Choi · 2020
Later among the works it cites.