Fetching the paper…
Reading the bibliography…
Many state-of-the-art results obtained with deep networks are achieved with the largest models that could be trained, and if more computation power was available, we might be able to exploit much larger datasets in order to improve generalization ability.
Interpolated estimation of Markov source parameters from sparse data
Jelinek, F. and Mercer, R. L. (1980) · 1980
Earlier work this paper cites.
Classification and Regression Trees
Breiman, L., Friedman, J. H., Olshen, R. A., and Stone, C. J. (1984) · 1984
Earlier work this paper cites.
Estimation of probabilities from sparse data for the language model component of a speech recognizer
Katz, S. M. (1987) · 1987
Earlier work this paper cites.
Complexity lower bounds for approximation algebraic computation trees
Cucker, F. and Grigoriev, D. (1999) · 1999
Earlier work this paper cites.
Learning deep architectures for AI
Bengio, Y. (2009) · 2009
Earlier work this paper cites.
Large-scale deep unsupervised learning using graphics processors
Raina, R., Madhavan, A., and Ng, A. Y. (2009) · 2009
Earlier work this paper cites.
Decision trees do not generalize to new variations
Bengio, Y., Delalleau, O., and Simard, C. (2010) · 2010
Earlier work this paper cites.
Learning to represent spatial transformations with factored higher-order Boltzmann machines
Memisevic, R. and Hinton, G. E. (2010) · 2010
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
Coates, A., Lee, H., and Ng, A. Y. (2011) · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
Sutskever, I., Martens, J., and Hinton, G. E. (2011) · 2011
Cited alongside, same era.
Multi column deep neural network for traffic sign classification
Ciresan, D., Meier, U., Masci, J., and Schmidhuber, J. (2012) · 2012
Cited alongside, same era.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Le, Q., Mao, M., Ranzato, M., Senior, A., Tucker, P., Yang, K., and Ng, A. Y. (2012) · 2012
Cited alongside, same era.
ImageNet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. (2012) · 2012
Practical recommendations for gradient-based training of deep architectures
Bengio, Y. (2013) · 2013
Later among the works it cites.
Maxout networks
Goodfellow, I. J., Warde-Farley, D., Mirza, M., Courville, A., and Bengio, Y. (2013) · 2013
Later among the works it cites.
Playing atari with deep reinforcement learning
Mnih, V., Kavukcuoglo, K., Silver, D., Graves, A., Antonoglou, I., and Wierstra, D. (2013) · 2013
Later among the works it cites.
Deep autoregressive networks
Gregor, K., Danihelka, I., Mnih, A., Blundell, C., and Wierstra, D. (2014) · 2014
Closest in time.
Neural variational inference and learning in belief networks
Mnih, A. and Gregor, K. (2014) · 2014
Closest in time.
On the number of inference regions of deep feed forward networks with piece-wise linear activations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Le, Q., Ranzato, M., Monga, R., Devin, M., Corrado, G., Chen, K., Dean, J., and Ng, A. (2012) · 2012
Cited alongside, same era.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A. (2013a)
Cited in the paper.
Unsupervised feature learning and deep learning: A review and new perspectives
Bengio, Y., Courville, A., and Vincent, P. (2013b)
Cited in the paper.
Deep neural networks for acoustic modeling in speech recognition
Hinton, G., Deng, L., Dahl, G. E., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B. (2012a)
Cited in the paper.
Improving neural networks by preventing co-adaptation of feature detectors
Hinton, G. E., Srivastava, N., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2012b)
Cited in the paper.
Pascanu, R., Montufar, G., and Bengio, Y. (2014) · 2014
Closest in time.
Techniques for learning binary stochastic feedforward neural networks
Raiko, T., Berglund, M., Alain, G., and Dinh, L. (2014) · 2014
Closest in time.