Fetching the paper…
Reading the bibliography…
Any prediction from a model is made by a combination of learning history and test stimuli.
Interpretable adversarial training for text
Samuel Barham and Soheil Feizi. 2019 · 1905
Earlier work this paper cites.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 1905
Earlier work this paper cites.
Sofia Serrano and Noah A Smith. 2019 · 1906
Earlier work this paper cites.
Analyzing the structure of attention in a transformer language model
Jesse Vig and Yonatan Belinkov. 2019 · 1906
Earlier work this paper cites.
Certified robustness to adversarial word substitutions
Robin Jia, Aditi Raghunathan, Kerem Göksel, and Percy Liang. 2019 · 1909
Earlier work this paper cites.
Attention interpretability across nlp tasks
Shikhar Vashishth, Shyam Upadhyay, Gaurav Singh Tomar, and Manaal Faruqui. 2019 · 1909
Earlier work this paper cites.
Allennlp interpret: A framework for explaining predictions of nlp models
Eric Wallace, Jens Tuyls, Junlin Wang, Sanjay Subramanian, Matt Gardner, and Sameer Singh. 2019 · 1909
Earlier work this paper cites.
Residuals and influence in regression
R Dennis Cook and Sanford Weisberg. 1982 · 1982
Earlier work this paper cites.
Assessment of local influence
R Dennis Cook. 1986 · 1986
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A. Pearlmutter. 1994 · 1994
Earlier work this paper cites.
Explaining black box predictions and unveiling data artifacts through influence functions
Xiaochuang Han, Byron C Wallace, and Yulia Tsvetkov. 2020 · 2005
Earlier work this paper cites.
Defense against adversarial attacks in nlp via dirichlet neighborhood ensemble
Yi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-wei Chang, and Xuanjing Huang. 2020 · 2006
Earlier work this paper cites.
Fast gradient projection method for text adversary generation and adversarial training
Xiaosen Wang, Yichen Yang, Yihe Deng, and Kun He. 2020 · 2008
Earlier work this paper cites.
Deep sparse rectifier neural networks
Xavier Glorot, Antoine Bordes, and Yoshua Bengio. 2011 · 2011
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts. 2011 · 2011
Earlier work this paper cites.
Poisoning attacks against support vector machines
Battista Biggio, Blaine Nelson, and Pavel Laskov. 2012 · 2012
Earlier work this paper cites.
Training and analysing deep recurrent neural networks
Michiel Hermans and Benjamin Schrauwen. 2013 · 2013
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013 · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. 2014 · 2014
Earlier work this paper cites.
Striving for simplicity: The all convolutional net
Jost Tobias Springenberg, Alexey Dosovitskiy, Thomas Brox, and Martin Riedmiller. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
Matthew D Zeiler and Rob Fergus. 2014 · 2014
Earlier work this paper cites.
Visualizing and understanding recurrent networks
Andrej Karpathy, Justin Johnson, and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
When are tree structures necessary for deep learning of representations?
Jiwei Li, Thang Luong, Dan Jurafsky, and Eduard Hovy. 2015b · 2015
Cited alongside, same era.
Using machine teaching to identify optimal training-set attacks on machine learners
Shike Mei and Xiaojin Zhu. 2015 · 2015
Cited alongside, same era.
Improved semantic representations from tree-structured long short-term memory networks
Kai Sheng Tai, Richard Socher, and Christopher D Manning. 2015 · 2015
Cited alongside, same era.
Second-order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan. 2016 · 2016
Cited alongside, same era.
Algorithmic transparency via quantitative input influence: Theory and experiments with learning systems
Anupam Datta, Shayak Sen, and Yair Zick. 2016 · 2016
Cited alongside, same era.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2017
Later among the works it cites.
Towards poisoning of deep learning algorithms with back-gradient optimization
Luis Muñoz-González, Battista Biggio, Ambra Demontis, Andrea Paudice, Vasin Wongrassamee, Emil C Lupu, and Fabio Roli. 2017 · 2017
Later among the works it cites.
Practical black-box attacks against machine learning
Nicolas Papernot, Patrick McDaniel, Ian Goodfellow, Somesh Jha, Z Berkay Celik, and Ananthram Swami. 2017 · 2017
Later among the works it cites.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje. 2017 · 2017
Later among the works it cites.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Klaus Greff, Rupesh K Srivastava, Jan Koutník, Bas R Steunebrink, and Jürgen Schmidhuber. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Adversarial examples in the physical world
Alexey Kurakin, Ian Goodfellow, and Samy Bengio. 2016 · 2016
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky. 2016 · 2016
Cited alongside, same era.
Adversarial training methods for semi-supervised text classification
Takeru Miyato, Andrew M. Dai, and Ian Goodfellow. 2016 · 2016
Cited alongside, same era.
Deepfool: a simple and accurate method to fool deep neural networks
Seyed-Mohsen Moosavi-Dezfooli, Alhussein Fawzi, and Pascal Frossard. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Generating natural adversarial examples
Zhengli Zhao, Dheeru Dua, and Sameer Singh. 2017 · 2017
Later among the works it cites.
Freelb: Enhanced adversarial training for natural language understanding
Chen Zhu, Yu Cheng, Zhe Gan, Siqi Sun, Tom Goldstein, and Jingjing Liu. 2020 · 2017
Later among the works it cites.
Auditing black-box models for indirect influence
Philip Adler, Casey Falk, Sorelle A Friedler, Tionney Nix, Gabriel Rybeck, Carlos Scheidegger, Brandon Smith, and Suresh Venkatasubramanian. 2018 · 2018
Later among the works it cites.
Generating natural language adversarial examples
Moustafa Alzantot, Yash Sharma, Ahmed Elgohary, Bo-Jhang Ho, Mani Srivastava, and Kai-Wei Chang. 2018 · 2018
Later among the works it cites.
Synthesizing robust adversarial examples
Anish Athalye, Logan Engstrom, Andrew Ilyas, and Kevin Kwok. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Later among the works it cites.
Interpreting recurrent and attention-based neural models: a case study on natural language inference
Reza Ghaeini, Xiaoli Z Fern, and Prasad Tadepalli. 2018 · 2018
Later among the works it cites.
Interpretable adversarial perturbation in input embedding space for text
Motoki Sato, Jun Suzuki, Hiroyuki Shindo, and Yuji Matsumoto. 2018 · 2018
Later among the works it cites.
Poison frogs! targeted clean-label poisoning attacks on neural networks
Ali Shafahi, W. Ronny Huang, Mahyar Najibi, Octavian Suciu, Christoph Studer, Tudor Dumitras, and Tom Goldstein. 2018 · 2018
Later among the works it cites.
Unsupervised learning of neural networks to explain neural networks
Quanshi Zhang, Yu Yang, Yuchen Liu, Ying Nian Wu, and Song-Chun Zhu. 2018 · 2018
Later among the works it cites.
What does bert look at? an analysis of bert’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D. Manning. 2019 · 2019
Later among the works it cites.
Attention is not explanation
Sarthak Jain and Byron C. Wallace. 2019 · 2019
Later among the works it cites.
Generating natural language adversarial examples through probability weighted word saliency
Shuhuai Ren, Yihe Deng, Kun He, and Wanxiang Che. 2019 · 2019
Later among the works it cites.
Cxplain: Causal explanations for model interpretation under uncertainty
Patrick Schwab and Walter Karlen. 2019 · 2019
Later among the works it cites.
Full-gradient representation for neural network visualization
Suraj Srinivas and François Fleuret. 2019 · 2019
Later among the works it cites.
Attention is not not explanation
Sarah Wiegreffe and Yuval Pinter. 2019 · 2019
Later among the works it cites.
Generating adversarial examples for holding robustness of source code processing models
Huangzhao Zhang, Zhuo Li, Ge Li, Lei Ma, Yang Liu, and Zhi Jin. 2020 · 2020
Closest in time.