Fetching the paper…
Reading the bibliography…
While depth tends to improve network performances, it also makes gradient-based training more difficult since deeper networks tend to be more non-linear.
Gradient-based learning applied to document recognition
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P · 1998
Earlier work this paper cites.
Model compression
Bucila, C., Caruana, R., and Niculescu-Mizil, A · 2006
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y.-W · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Bengio, Y., Lamblin, P., Popovici, D., and Larochelle, H · 2007
Earlier work this paper cites.
An empirical evaluation of deep architectures on problems with many factors of variation
Larochelle, H., Erhan, D., Courville, A., Bergstra, J., and Bengio, Y · 2007
Earlier work this paper cites.
Deep learning via semi-supervised embedding
Weston, J., Ratle, F., and Collobert, R · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Bengio, Y · 2009
Earlier work this paper cites.
The difficulty of training deep architectures and the effect of unsupervised pre-training
Erhan, D., Manzagol, P.A., Bengio, Y., Bengio, S., and Vincent, P · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G · 2009
Earlier work this paper cites.
Approximate nearest neighbor search by residual vector quantization
Chen, Yongjian, Guan, Tao, and Wang, Cheng · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
Jégou, Hervé, Douze, Matthijs, and Schmid, Cordelia · 2011
Cited alongside, same era.
Annotated facial landmarks in the wild: A large-scale, real-world database for facial landmark localization
Koestinger, M., Wohlhart, P., Roth, P.M., and Bischof, H · 2011
Cited alongside, same era.
Reading digits in natural images with unsupervised feature learning
Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A · 2011
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, F., Lamblin, P., Pascanu, R., Bergstra, J., Goodfellow, I., Bergeron, A., Bouchard, N., Warde-Farley, D., and Bengio, Y · 2012
Cited alongside, same era.
A two-stage pretraining algorithm for deep Boltzmann machines
Cho, Kyunghyun, Raiko, Tapani, Ilin, Alexander, and Karhunen, Juha · 2012
Cited alongside, same era.
Lecture 6.5—RmsProp: Divide the gradient by a running average of its recent magnitude
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Dauphin, Y., Pascanu, R., Gulcehre, C., Cho, K., Ganguli, S., and Bengio, Y · 2014
Closest in time.
Exploiting linear structure within convolutional networks for efficient evaluation
Denton, Emily L, Zaremba, Wojciech, Bruna, Joan, LeCun, Yann, and Fergus, Rob · 2014
Closest in time.
Compressing deep convolutional networks using vector quantization
Gong, Yunchao, Liu, Liu, Yang, Min, and Bourdev, Lubomir · 2014
Closest in time.
Distilling knowledge in a neural network
Hinton, G. Vinyals, O. and Dean, J · 2014
Closest in time.
Speeding up convolutional neural networks with low rank expansions
Jaderberg, M., Vedaldi, A., and Zisserman, A · 2014
Closest in time.
On the number of linear regions of deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Representation learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P · 2013
Cited alongside, same era.
Knowledge matters: Importance of prior information for optimization
Gulcehre, C. and Bengio, Y · 2013
Cited alongside, same era.
Do deep nets really need to be deep?
Ba, J. and Caruana, R · 2014
Cited alongside, same era.
Chen-Yu, L., Saining, X., Patrick, G., Zhengyou, Z., and Zhuowen, T · 2014
Cited alongside, same era.
Pylearn2: a machine learning research library
Goodfellow, I. J., Warde-Farley, D., Lamblin, P., Dumoulin, V., Mirza, M., Pascanu, R., Bergstra, J., Bastien, F., and Bengio, Y
Cited in the paper.
Maxout networks
Goodfellow, I.J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y
Cited in the paper.
Montufar, G.F., Pascanu, R., Cho, K., and Bengio, Y · 2014
Closest in time.
ImageNet Large Scale Visual Recognition Challenge, 2014
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., Berg, A. C., and Fei-Fei, L · 2014
Closest in time.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Closest in time.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., D.A., Erhan, D., Vanhoucke, V., and Rabinovich, A · 2014
Closest in time.