Fetching the paper…
Reading the bibliography…
Nowadays, the number of layers and of neurons in each layer of a deep network are typically set manually.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
M. Mozer and P. Smolensky · 1988
Earlier work this paper cites.
Dynamic node creation in backpropagation networks
T. Ash · 1989
Earlier work this paper cites.
Generalizing smoothness constraints from discrete samples
C. Ji, R. R. Snapp, and D. Psaltis · 1990
Earlier work this paper cites.
Optimal brain damage
Y. LeCun, J. S. Denker, and S. A. Solla · 1990
Earlier work this paper cites.
Generalization by weight-elimination with application to forecasting
A. S. Weigend, D. Rumelhart, and B. A. Huberman · 1991
Earlier work this paper cites.
Enhanced training algorithms, and integrated training/architecture selection for multilayer perceptron networks
M. G. Bello · 1992
Earlier work this paper cites.
A simple weight decay can improve generalization
A. Krogh and J. A. Hertz · 1992
Earlier work this paper cites.
Optimal brain surgeon and general network pruning
B. Hassibi, D. G. Stork, and G. J. Wolff · 1993
Earlier work this paper cites.
Pruning algorithms-a survey
R. Reed · 1993
Earlier work this paper cites.
For valid generalization the size of the weights is more important than the size of the network
P. L. Bartlett · 1996
Earlier work this paper cites.
Model selection and estimation in regression with grouped variables
M. Yuan and Y. Lin · 2007
Earlier work this paper cites.
Torch7: A matlab-like environment for machine learning
R. Collobert, K. Kavukcuoglu, and C. Farabet · 2011
Earlier work this paper cites.
Predicting parameters in deep learning
M. Denil, B. Shakibi, L. Dinh, M.A. Ranzato, and N. de Freitas · 2013
Cited alongside, same era.
Maxout networks
I. J. Goodfellow, D. Warde-farley, M. Mirza, A. Courville, and Y. Bengio · 2013
Cited alongside, same era.
A sparse-group lasso
N. Simon, J. Friedman, T. Hastie, and R. Tibshirani · 2013
Cited alongside, same era.
Memory Bounded Deep Convolutional Networks
M. D. Collins and P. Kohli · 2014
Cited alongside, same era.
Exploiting linear structure within convolutional networks for efficient evaluation
E. L Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus · 2014
Cited alongside, same era.
Compressing deep convolutional networks using vector quantization
An exploration of parameter redundancy in deep networks with circulant projections
Yu Cheng, Felix X. Yu, Rogério Schmidt Feris, Sanjiv Kumar, Alok N. Choudhary, and Shih-Fu Chang · 2015
Later among the works it cites.
Deep Residual Learning for Image Recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Later among the works it cites.
Sparse convolutional neural networks
B. Liu, M. Wang, H. Foroosh, M. Tappen, and M. Penksy · 2015
Later among the works it cites.
Auto-sizing neural networks: With applications to n-gram language models
K. Murray and D. Chiang · 2015
Later among the works it cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Gong, L. Liu, M. Yang, and L. D. Bourdev · 2014
Cited alongside, same era.
Distilling the knowledge in a neural network
G. E. Hinton, O. Vinyals, and J. Dean · 2014
Cited alongside, same era.
On the number of linear regions of deep neural networks
G. F Montufar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Cited alongside, same era.
Proximal algorithms
N. Parikh and S. Boyd · 2014
Cited alongside, same era.
Learning block group sparse representation combined with convolutional neural networks for rgb-d object recognition
J. Wang X. Huang X. Zhang S. Tu, Y. Xue · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Cited alongside, same era.
Deep features for text spotting
M. Jaderberg, A. Vedaldi, and A. Zisserman
Cited in the paper.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein, A. C. Berg, and L. Fei-Fei · 2015
Later among the works it cites.
Training very deep networks
R. K Srivastava, K. Greff, and J. Schmidhuber · 2015
Later among the works it cites.
Machine Learning, A Bayesian and Optimization Perspective , volume 8
S. Theodoridis · 2015
Later among the works it cites.
Places: An image database for deep scene understanding
B. Zhou, A. Khosla, A. Lapedriza, A. Torralba, and A. Oliva · 2015
Later among the works it cites.
Decomposeme: Simplifying convnets for end-to-end learning
J.M. Alvarez and L. Petersson · 2016
Closest in time.
Less is more: Towards compact CNNs
H. Zhou, J. M. Alvarez, and F. Porikli · 2016
Closest in time.