Fetching the paper…
Reading the bibliography…
We investigate the generalizability of deep learning based on the sensitivity to input perturbation.
A universal prior for integers and estimation by minimum description length
J. Rissanen · 1983
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1991
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
M. Ranzato, F. J. Huang, Y.-L. Boureau, and Y. LeCun · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning
A. Coates, A. Y. Ng, and H. Lee · 2011
Earlier work this paper cites.
Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups
G. Hinton, L. Deng, D. Yu, G. Dahl, A.-r. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. Sainath, and B. Kingsbury · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Efficient backprop
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller · 2012
Earlier work this paper cites.
Advances in optimizing recurrent networks
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu · 2013
Cited alongside, same era.
Maxout networks
I. J. Goodfellow, D. Warde-Farley, M. Mirza, and A. C. Courville · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Y. N. Dauphin, R. Pascanu, C. Gulcehre, K. Cho, S. Ganguli, and Y. Bengio · 2014
Cited alongside, same era.
Towards deep neural network architectures robust to adversarial examples
S. Gu and L. Rigazio · 2014
Cited alongside, same era.
Network in network
M. Lin, Q. Chen, and S. Yan · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Revisiting distributed synchronous SGD
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Later among the works it cites.
Distributed deep learning using synchronous stochastic gradient descent
D. Das, S. Avancha, D. Mudigere, K. Vaidynathan, S. Sridharan, D. Kalamkar, B. Kaul, and P. Dubey · 2016
Later among the works it cites.
Identity mappings in deep residual networks
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Later among the works it cites.
Densely connected convolutional networks
G. Huang, Z. Liu, K. Q. Weinberger, and L. van der Maaten · 2016
Later among the works it cites.
Exploring the limits of language modeling
R. Jozefowicz, O. Vinyals, M. Schuster, N. Shazeer, and Y. Wu · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. J. Goodfellow, and R. Fergus · 2014
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2015
Cited alongside, same era.
Unitary evolution recurrent neural networks
M. Arjovsky, A. Shah, and Y. Bengio · 2016
Cited alongside, same era.
Deep learning without poor local minima
K. Kawaguchi · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Later among the works it cites.
Full-capacity unitary recurrent neural networks
S. Wisdom, T. Powers, J. Hershey, J. Le Roux, and L. Atlas · 2016
Later among the works it cites.
On large-batch training for deep learning - generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Closest in time.
Uunderstanding deep learning requires rethinking generalization
C. Zhang, S. Bengio, M. Hardt, B. Recht, and O. Vinyals · 2017
Closest in time.