Fetching the paper…
Reading the bibliography…
This article exposes the failure of some big neural networks to leverage added capacity to reduce underfitting.
A fast learning algorithm for deep belief nets
Hinton, G. E., Osindero, S., and Teh, Y. (2006) · 2006
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
Le Roux, N., Manzagol, P.-A., and Bengio, Y. (2008) · 2008
Earlier work this paper cites.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, J., Breuleux, O., Bastien, F., Lamblin, P., Pascanu, R., Desjardins, G., Turian, J., Warde-Farley, D., and Bengio, Y. (2010) · 2010
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
Erhan, D., Bengio, Y., Courville, A., Manzagol, P., Vincent, P., and Bengio, S. ((11) 2010) · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X. and Bengio, Y. (2010) · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
Martens, J. (2010) · 2010
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Lee, H., and Ng, A. Y. (2011) · 2011
Cited alongside, same era.
Unsupervised models of images by spike-and-slab RBMs
Courville, A., Bergstra, J., and Bengio, Y. (2011) · 2011
Cited alongside, same era.
An overview of the hdf5 technology suite and its applications
Folk, M., Heber, G., Koziol, Q., Pourmal, E., and Robinson, D. (2011) · 2011
Cited alongside, same era.
Empirical evaluation and combination of advanced language modeling techniques
Mikolov, T., Deoras, A., Kombrink, S., Burget, L., and Cernocky, J. (2011) · 2011
Later among the works it cites.
Higher order contractive auto-encoder
Rifai, S., Mesnil, G., Vincent, P., Muller, X., Bengio, Y., Dauphin, Y., and Glorot, X. (2011) · 2011
Later among the works it cites.
Conversational speech transcription using context-dependent deep neural networks
Seide, F., Li, G., and Yu, D. (2011) · 2011
Later among the works it cites.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012) · 2012
Later among the works it cites.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012) · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…