Fetching the paper…
Reading the bibliography…
Hessian-free (HF) optimization has been successfully used for training deep autoencoders and recurrent networks.
Fast exact multiplication by the hessian
B.A. Pearlmutter · 1994
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Müller · 1998
Earlier work this paper cites.
Fast curvature matrix-vector products for second-order gradient descent
N.N. Schraudolph · 2002
Earlier work this paper cites.
Greedy layer-wise training of deep networks
Y. Bengio, P. Lamblin, D. Popovici, and H. Larochelle · 2007
Earlier work this paper cites.
Topmoumoute online natural gradient algorithm
N. Le Roux, P.A. Manzagol, and Y. Bengio · 2007
Earlier work this paper cites.
Extracting and composing robust features with denoising autoencoders
P. Vincent, H. Larochelle, Y. Bengio, and P.A. Manzagol · 2008
Earlier work this paper cites.
A deep non-linear feature mapping for large-margin knn classification
R. Min, D.A. Stanley, Z. Yuan, A. Bonner, and Z. Zhang · 2009
Earlier work this paper cites.
Probabilistic dyadic data analysis with local and global consistency
Deng Cai, Xuanhui Wang, and Xiaofei He · 2009
Earlier work this paper cites.
A deep non-linear feature mapping for large-margin knn classification
R. Min, D.A. Stanley, Z. Yuan, A. Bonner, and Z. Zhang · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2010
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Cited alongside, same era.
The tradeoffs of large-scale learning
L. Bottou and O. Bousquet · 2011
Cited alongside, same era.
Learning recurrent neural networks with hessian-free optimization
J. Martens and I. Sutskever · 2011
Cited alongside, same era.
Improved preconditioner for hessian free optimization
O. Chapelle and D. Erhan · 2011
Cited alongside, same era.
Generating text with recurrent neural networks
I. Sutskever, J. Martens, and G. Hinton · 2011
Cited alongside, same era.
Krylov subspace descent for deep learning
O. Vinyals and D. Povey · 2011
Cited alongside, same era.
T. Schaul, S. Zhang, and Y. LeCun · 2012
Later among the works it cites.
Sample size selection in optimization methods for machine learning
R.H. Byrd, G.M. Chin, J. Nocedal, and Y. Wu · 2012
Later among the works it cites.
Practical recommendations for gradient-based training of deep architectures
Y. Bengio · 2012
Later among the works it cites.
Training deep and recurrent networks with hessian-free optimization
J. Martens and I. Sutskever · 2012
Later among the works it cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. Hinton · 2012
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On optimization methods for deep learning
Q.V. Le, J. Ngiam, A. Coates, A. Lahiri, B. Prochnow, and A.Y. Ng · 2011
Cited alongside, same era.
Deep sparse rectifier neural networks
X. Glorot, A. Bordes, and Y. Bengio · 2011
Cited alongside, same era.
Improving neural networks by preventing co-adaptation of feature detectors
G.E. Hinton, N. Srivastava, A. Krizhevsky, I. Sutskever, and R.R. Salakhutdinov · 2012
Cited alongside, same era.
Large scale distributed deep networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. Le, M. Mao, A. Senior, P. Tucker, K. Yang, et al · 2012
Cited alongside, same era.
Sgd-qn: Careful quasi-newton stochastic gradient descent
A. Bordes, L. Bottou, and P. Gallinari
Cited in the paper.
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu · 2012
Later among the works it cites.
Semantic compositionality through recursive matrix-vector spaces
R. Socher, B. Huval, C.D. Manning, and A.Y. Ng · 2012
Later among the works it cites.
Razvan Pascanu and Yoshua Bengio · 2013
Closest in time.
Big neural networks waste capacity
Yann N Dauphin and Yoshua Bengio · 2013
Closest in time.
Training Recurrent Neural Networks
I. Sutskever · 2013
Closest in time.