Fetching the paper…
Reading the bibliography…
The impressive performance of neural networks on natural language processing tasks attributes to their ability to model complicated word and phrase compositions.
Bert rediscovers the classical nlp pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick · 1905
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov · 1907
Earlier work this paper cites.
A value for n-person games
Lloyd S Shapley · 1953
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Discovering additive structure in black box functions
Giles Hooker · 2004
Earlier work this paper cites.
Axiomatic characterizations of probabilistic and cardinal-probabilistic interaction indices
Katsushige Fujimoto, Ivan Kojadinovic, and Jean-Luc Marichal · 2006
Earlier work this paper cites.
Detecting statistical interactions with additive groves of trees
Daria Sorokina, Rich Caruana, Mirek Riedewald, and Daniel Fink · 2008
Earlier work this paper cites.
The feature importance ranking measure
Alexander Zien, Nicole Krämer, Sören Sonnenburg, and Gunnar Rätsch · 2009
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Ng, and Christopher Potts · 2013
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Visualizing deep neural network decisions: Prediction difference analysis
Luisa M Zintgraf, Taco S Cohen, Tameem Adel, and Max Welling · 2014
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Character-level convolutional networks for text classification
Xiang Zhang, Junbo Zhao, and Yann LeCun · 2015
Earlier work this paper cites.
Visualizing the effects of predictor variables in black box supervised learning models
Daniel W Apley · 2016
Cited alongside, same era.
Layer-wise relevance propagation for neural networks with local renormalization layers
Alexander Binder, Grégoire Montavon, Sebastian Lapuschkin, Klaus-Robert Müller, and Wojciech Samek · 2016
Cited alongside, same era.
Interpretation of prediction models using the input gradient
Yotam Hechtlinger · 2016
Cited alongside, same era.
Rationalizing neural predictions
Tao Lei, Regina Barzilay, and Tommi S. Jaakkola · 2016
Cited alongside, same era.
Understanding neural networks through representation erasure
Jiwei Li, Will Monroe, and Dan Jurafsky · 2016
Cited alongside, same era.
Position-aware attention and supervised data improve slot filling
Yuhao Zhang, Victor Zhong, Danqi Chen, Gabor Angeli, and Christopher D. Manning · 2017
Later among the works it cites.
Learning to explain: An information-theoretic perspective on model interpretation
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Later among the works it cites.
A survey of methods for explaining black box models
Riccardo Guidotti, Anna Monreale, Salvatore Ruggieri, Franco Turini, Fosca Giannotti, and Dino Pedreschi · 2018
Later among the works it cites.
A survey of evaluation methods and measures for interpretable machine learning
Sina Mohseni, Niloofar Zarei, and Eric D Ragan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
”why should I trust you?”: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
Towards better understanding of gradient-based attribution methods for deep neural networks
Marco Ancona, Enea Ceolini, Cengiz Öztireli, and Markus Gross · 2017
Cited alongside, same era.
Representation of linguistic form and function in recurrent neural networks
Akos Kádár, Grzegorz Chrupała, and Afra Alishahi · 2017
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Cited alongside, same era.
Learning important features through propagating activation differences
Avanti Shrikumar, Peyton Greenside, and Anshul Kundaje · 2017
Cited alongside, same era.
Axiomatic attribution for deep networks
Mukund Sundararajan, Ankur Taly, and Qiqi Yan · 2017
Cited alongside, same era.
Detecting statistical interactions from neural network weights
Michael Tsang, Dehua Cheng, and Yan Liu · 2017
Cited alongside, same era.
Beyond word importance: Contextual decomposition to extract interactions from lstms
W James Murdoch, Peter J Liu, and Bin Yu · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer · 2018
Later among the works it cites.
Evaluating neural network explanation methods using hybrid documents and morphological agreement
Nina Poerner, Hinrich Schütze, and Benjamin Roth · 2018
Later among the works it cites.
Can i trust you more? model-agnostic hierarchical explanations
Michael Tsang, Youbang Sun, Dongxu Ren, and Yan Liu · 2018
Later among the works it cites.
Explaining image classifiers by counterfactual generation
Chun-Hao Chang, Elliot Creager, Anna Goldenberg, and David Duvenaud · 2019
Closest in time.
L-shapley and c-shapley: Efficient model interpretation for structured data
Jianbo Chen, Le Song, Martin J. Wainwright, and Michael I. Jordan · 2019
Closest in time.
Hierarchical interpretations for neural network predictions
Chandan Singh, W James Murdoch, and Bin Yu · 2019
Closest in time.
Bert has a mouth, and it must speak: Bert as a markov random field language model
Alex Wang, Kyunghyun Cho, and CIFAR Azrieli Global Scholar · 2019
Closest in time.