Fetching the paper…
Reading the bibliography…
Despite the widespread practical success of deep learning methods, our theoretical understanding of the dynamics of learning in deep neural networks remains quite sparse.
Neural networks and principal component analysis: Learning from examples without local minima
P. Baldi and K. Hornik · 1989
Earlier work this paper cites.
Untersuchungen zu dynamischen neuronalen Netzen
S. Hochreiter · 1991
Earlier work this paper cites.
Learning Long-Term Dependencies with Gradient Descent is Difficult
Y. Bengio, P. Simard, and P. Frasconi · 1994
Earlier work this paper cites.
Efficient BackProp
Y. LeCun, L. Bottou, G.B. Orr, and K.R. Müller · 1998
Earlier work this paper cites.
Effect of Batch Learning In Multilayer Neural Networks
K. Fukumizu · 1998
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G.E. Hinton and R.R. Salakhutdinov · 2006
Earlier work this paper cites.
Scaling learning algorithms towards AI
Y. Bengio and Y. LeCun · 2007
Earlier work this paper cites.
Greedy Layer-Wise Training of Deep Networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Earlier work this paper cites.
A Unified Architecture for Natural Language Processing: Deep Neural Networks with Multitask Learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
The Difficulty of Training Deep Architectures and the Effect of Unsupervised Pre-Training
D. Erhan, P.A. Manzagol, Y. Bengio, S. Bengio, and P. Vincent · 2009
Cited alongside, same era.
Learning Deep Architectures for AI
Y. Bengio · 2009
Cited alongside, same era.
Understanding the difficulty of training deep feedforward neural networks
X. Glorot and Y. Bengio · 2010
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. Courville, P.A. Manzagol, and P. Vincent · 2010
Cited alongside, same era.
Deep learning via Hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Important gains from supervised fine-tuning of deep architectures on large labeled sets
P. Lamblin and Y. Bengio · 2010
Cited alongside, same era.
Building high-level features using large scale unsupervised learning
Q.V. Le, M.A. Ranzato, R. Monga, M. Devin, K. Chen, G.S. Corrado, J. Dean, and A.Y. Ng · 2012
Later among the works it cites.
Multi-column Deep Neural Networks for Image Classification
D. Ciresan, U. Meier, and J. Schmidhuber · 2012
Later among the works it cites.
Acoustic Modeling Using Deep Belief Networks
A. Mohamed, G.E. Dahl, and G. Hinton · 2012
Later among the works it cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2012
Later among the works it cites.
Parsing with Compositional Vector Grammars
R. Socher, J. Bauer, C.D. Manning, and A.Y. Ng · 2013
Closest in time.
Big Neural Networks Waste Capacity
Y.N. Dauphin and Y. Bengio · 2013
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Improved Preconditioner for Hessian Free Optimization
O. Chapelle and D. Erhan · 2011
Cited alongside, same era.
ImageNet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G.E. Hinton · 2012
Cited alongside, same era.
Learning hierarchical category structure in deep neural networks
A.M. Saxe, J.L. McClelland, and S. Ganguli · 2013
Closest in time.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G.E. Hinton · 2013
Closest in time.