Fetching the paper…
Reading the bibliography…
We propose a novel Shapley value approach to help address neural networks' interpretability and "vanishing gradient" problems.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Stochastic estimation of the maximum of a regression function
Jack Kiefer, Jacob Wolfowitz, et al · 1952
Earlier work this paper cites.
A value for n-person games
L Shapley · 1953
Earlier work this paper cites.
Learning long-term dependencies with gradient descent is difficult
Yoshua Bengio, Patrice Simard, Paolo Frasconi, et al · 1994
Earlier work this paper cites.
The vanishing gradient problem during learning recurrent neural nets and problem solutions
Sepp Hochreiter · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Digital selection and analogue amplification coexist in a cortex-inspired silicon circuit
R. Hahnloser, R. Sarpeshkar, M.A. Mahowald, R.J. Douglas, and H.S. Seung · 2000
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies
S. Hochreiter, Y. Bengjio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
How to explain individual classification decisions
David Baehrens, Timon Schroeter, Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and Klaus-Robert MÞller · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
N. Vinod and G. E. Hinton · 2010
Earlier work this paper cites.
Making machine learning models interpretable
Alfredo Vellido, José David Martín-Guerrero, and Paulo JG Lisboa · 2012
Earlier work this paper cites.
Understanding the exploding gradient problem
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
Rectifier nonlinearities improve neural network acoustic models
A. L. Maas and A. Y. Ng A. Y. Hannun · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
S. Nitish, G. Hinton, A. Kritzevsky, I Sutskever, and R. Salakutdinov · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
An empirical exploration of recurrent network architectures
Rafal Jozefowicz, Wojciech Zaremba, and Ilya Sutskever · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2015
Cited alongside, same era.
A unified approach to interpreting model predictions
Scott M Lundberg and Su-In Lee · 2017
Later among the works it cites.
Learning how to explain neural networks: Patternnet and patternattribution
Pieter-Jan Kindermans, Kristof T Schütt, Maximilian Alber, Klaus-Robert Müller, Dumitru Erhan, Been Kim, and Sven Dähne · 2017
Later among the works it cites.
Explaining recurrent neural network predictions in sentiment analysis
Leila Arras, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2017
Later among the works it cites.
Wojciech Samek, Thomas Wiegand, and Klaus-Robert Müller · 2017
Later among the works it cites.
Hexpo: A vanishing-proof activation function
Shumin Kong and Masahiro Takatsuka · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J Sun · 2015
Cited alongside, same era.
Batch normalization: accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation
S. Bach, A. Binder, G. Montavon, F Klauschen, K-R Müller, and Wojciech Samek · 2015
Cited alongside, same era.
Explaining predictions of non-linear classifiers in nlp
Leila Arras, Franziska Horn, Grégoire Montavon, Klaus-Robert Müller, and Wojciech Samek · 2016
Cited alongside, same era.
Why should i trust you?: Explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Cited alongside, same era.
The mythos of model interpretability
Zachary C Lipton · 2016
Cited alongside, same era.
Layer-wise relevance propogation for deep neural network architectures
A. Binder, S. Bach, G. Montavon, K-R Müller, and Wojciech Samek · 2016
Cited alongside, same era.
Later among the works it cites.
Self-normalizing neural networks
G. Klambauer, T. Unterthiner, A Mayr, and S. Hochreiter · 2017
Later among the works it cites.
Interpretable convolutional neural networks
Quanshi Zhang, Ying Nian Wu, and Song-Chun Zhu · 2018
Later among the works it cites.
Overcoming the vanishing gradient problem in plain recurrent networks
Yuhuang Hu, Adrian Huber, Jithendar Anumula, and Shih-Chii Liu · 2018
Later among the works it cites.
Which neural net architectures give rise to exploding and vanishing gradients?
Boris Hanin · 2018
Later among the works it cites.
Flux: Elegant machine learning with julia
Mike Innes · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Reduced form capital optimization
Y. Li, D. Offengenden, and J. Burgy · 2019
Closest in time.
Explaining deep neural networks with a polynomial time algorithm for shapley value approximation
M. Ancona, C. Oztireli, and M. Gross · 2019
Closest in time.
Keras team · 2019
Closest in time.