Fetching the paper…
Reading the bibliography…
We demonstrate that there is significant redundancy in the parameterization of several deep learning models.
Learning representations by back-propagating errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Speaker-independent phone recognition using hidden markov models
K.-F. Lee and H.-W. Hon · 1989
Earlier work this paper cites.
Dimensionality reduction and prior knowledge in e-set recognition
K. Lang and G. Hinton · 1990
Earlier work this paper cites.
Optimal brain damage
Y. LeCun, J. S. Denker, S. Solla, R. E. Howard, and L. D. Jackel · 1990
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
A neural support vector network architecture with adaptive kernels
P. Vincent and Y. Bengio · 2000
Earlier work this paper cites.
Kernel Methods for Pattern Analysis
J. Shawe-Taylor and N. Cristianini · 2004
Earlier work this paper cites.
Emergence of complex-like cells in a temporal product network with local receptive fields
K. Gregor and Y. LeCun · 2010
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
V. Nair and G. E. Hinton · 2010
Earlier work this paper cites.
Factored 3-way restricted Boltzmann machines for modeling natural images
M. Ranzato, A. Krizhevsky, and G. E. Hinton · 2010
Cited alongside, same era.
Double sparsity: learning sparse dictionaries for sparse signal approximation
R. Rubinstein, M. Zibulevsky, and M. Elad · 2010
Cited alongside, same era.
High-performance neural networks for visual object classification
D. Cireşan, U. Meier, and J. Masci · 2011
Cited alongside, same era.
Selecting receptive fields in deep networks
A. Coates and A. Y. Ng · 2011
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
A. Coates, A. Y. Ng, and H. Lee · 2011
Cited alongside, same era.
ICA with reconstruction cost for efficient overcomplete feature learning
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, and A. Ng · 2012
Later among the works it cites.
Scalable stacking and learning for building deep architectures
L. Deng, D. Yu, and J. Platt · 2012
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
G. E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Later among the works it cites.
Building high-level features using large scale unsupervised learning
Q. V. Le, M. Ranzato, R. Monga, M. Devin, K. Chen, G. Corrado, J. Dean, and A. Ng · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. V. Le, A. Karpenko, J. Ngiam, and A. Y. Ng · 2011
Cited alongside, same era.
On autoencoders and score matching for energy based models
K. Swersky, M. Ranzato, D. Buchman, B. Marlin, and N. Freitas · 2011
Cited alongside, same era.
Multi-column deep neural networks for image classification
D. Cireşan, U. Meier, and J. Schmidhuber · 2012
Cited alongside, same era.
Emergence of object-selective features in unsupervised feature learning
A. Coates, A. Karpathy, and A. Ng · 2012
Cited alongside, same era.
Y. Bengio · 2013
Closest in time.
Maxout networks
I. J. Goodfellow, D. Warde-Farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Closest in time.
Knowledge matters: Importance of prior information for optimization
C. Gülçehre and Y. Bengio · 2013
Closest in time.
Learning separable filters
R. Rigamonti, A. Sironi, V. Lepetit, and P. Fua · 2013
Closest in time.