Stochastic gradient learning in neural networks
Léon Bottou · 1991
Earlier work this paper cites.
Hitting the memory wall: implications of the obvious
W.A. Wulf and S.A. McKee · 1995
Earlier work this paper cites.
Gradient flow in recurrent nets: the difficulty of learning long-term dependencies, 2001
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber · 2001
Earlier work this paper cites.
Visual categorization shapes feature selectivity in the primate temporal cortex
Natasha Sigala and N.K. Logothetis · 2002
Earlier work this paper cites.
Equivalence of backpropagation and contrastive hebbian learning in a layered network
Xiaohui Xie and S.H. Seung · 2003
Earlier work this paper cites.
Energy management for commercial servers
Charles Lefurgy, Karthick Rajamani, Freeman Rawson, Wes Felter, Michael Kistler, and T.W. Keller · 2003
Earlier work this paper cites.
A fast learning algorithm for deep belief nets
G.E. Hinton, S. Osindero, and Y.W. Teh · 2006
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G.E. Hinton and R.R. Salakhutdinov · 2006
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Yoshua Bengio, Pascal Lamblin, Dan Popovici, and Hugo Larochelle · 2007
Earlier work this paper cites.
Unsupervised learning of invariant feature hierarchies with applications to object recognition
Marc’Aurelio Ranzato, F.J. Huang, Y-Lan Boureau, and Yann LeCun · 2007
Earlier work this paper cites.
Feedforward neural network implementation in FPGA using layer multiplexing for effective resource utilization
S. Himavathi, D. Anitha, and A. Muthuramalingam · 2007
Earlier work this paper cites.
Object category structure in response patterns of neuronal population in monkey inferior temporal cortex
Roozbeh Kiani, Hossein Esteky, Koorosh Mirpour, and Keiji Tanaka · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
Pascal Vincent, Hugo Larochelle, Yoshua Bengio, and Pierre-Antoine Manzagol · 2008
Earlier work this paper cites.
Semi-supervised learning of compact document representations with deep networks
Marc’Aurelio Ranzato and Martin Szummer · 2008
Earlier work this paper cites.
Classification using discriminative restricted boltzmann machines
Hugo Larochelle and Yoshua Bengio · 2008
Earlier work this paper cites.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
Honglak Lee, Roger Grosse, Rajesh Ranganath, and Andrew Y. Ng · 2009
Earlier work this paper cites.
What is the best multi-stage architecture for object recognition?
Kevin Jarrett, Koray Kavukcuoglu, Yann LeCun, et al · 2009
Earlier work this paper cites.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. Courville, P.A. Manzagol, P. Vincent, and S. Bengio · 2010
Earlier work this paper cites.
Fast inference in sparse coding algorithms with applications to object recognition
Original
Koray Kavukcuoglu, Marc’Aurelio Ranzato, and Yann LeCun · 2010
Earlier work this paper cites.
Rectified linear units improve restricted Boltzmann machines
Vinod Nair and G.E. Hinton · 2010
Earlier work this paper cites.