Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) still is the workhorse for many practical problems.
D. Godard, “Self-recovering equalization carrier tracking in two-dimensional data communications systems,” IEEE Trans. Commun. , vol. 28, no. 11, pp. 1867–1875, Nov. 1980
1980
Earlier work this paper cites.
B. Widrow and S. D. Stearns, Adaptive Signal Processing . Englewood Cliffs, New Jersey: Prentice-Hall, Inc., 1985
1985
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, pp. 533–536, Oct. 1986
1986
Earlier work this paper cites.
P. J. Werbos, “Backpropagation through time: what it does and how to do it,” IEEE Proc. , vol. 78, no. 10, pp. 1550–1560, Oct. 1990
1990
Earlier work this paper cites.
W. L. Buntine and A. S. Weigend, “Computing second derivatives in feed-forward networks: a review,” IEEE Trans. Neural Netw. , vol. 5, no. 3, pp. 480–488, May 1991
1991
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE Trans. Neural Netw. , vol. 5, no. 2, pp. 157–166, Mar. 1994
1994
Earlier work this paper cites.
Y. Li and Z. Ding, “Convergence analysis of finite length blind adaptive equalizers,” IEEE Trans. Signal Process. , vol. 43, no. 9, pp. 2120–2129, Sept. 1995
1995
Earlier work this paper cites.
J.-F. Cardoso and B. Laheld, “Equivariant adaptive source separation,” IEEE Trans. Signal Process. , vol. 44, no. 12, pp. 3017–3030, Dec. 1996
1996
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Computation , vol. 9, no.8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient based learning applied to document recognition,” Proc. IEEE , vol. 86, no. 11, pp. 2278–2324, Nov. 1998
1998
Cited alongside, same era.
S. Amari, “Natural gradient works efficiently in learning,” Neural Computation , vol. 10, no. 2, pp. 251–276, Feb. 1998
1998
Cited alongside, same era.
H. Demuth and M. Beale, Neural Network Toolbox for Use with MATLAB , Natick, MA: The MathWorks, Inc., 2002
2002
Cited alongside, same era.
N. N. Schraudolph, J. Yu, and S. Günter, “A stochastic quasi-Newton method for online convex optimization,” J. Mach. Learn. Res. , vol. 2, pp. 436–443, Jan. 2007
2007
Cited alongside, same era.
B. Antoine, B. Leon, and G. Patrick, “SGD-QN: careful quasi-Newton stochastic gradient descent,” J. Mach. Learn. Res. , vol. 10, pp. 1737–1754, Jul. 2009
2009
R. H. Byrd, S. L. Hansen, J. Nocedal, and Y. Singer, “A stochastic quasi-Newton method for large-scale optimization,” SIAM J. Optimiz. , vol. 26, no. 2, pp. 1008–1031, Jan. 2014
2014
Later among the works it cites.
Y. N. Dauphin, H. Vries, and Y. Bengio, “Equilibrated adaptive learning rates for non-convex optimization,” in Advances in Neural Information Processing Systems , 2015, pp. 1504–1512
2015
Closest in time.
D. E. Carlson, E. Collins, Y. P. Hsieh, L. Carin, and V. Cevher, “Preconditioned spectral descent for deep learning,” in Proc. 28th Int. Conf. Neural Information Processing Systems , Montreal, 2015, pp. 2971–2970
2015
Closest in time.
D. Povey, X. Zhang, and S. Khudanpur, “Parallel training of DNNs with natural gradient and parameter averaging,” in Proc. Int. Conf. Learning Representations , 2015
2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
J. Martens and I. Sutskever, “Training deep and recurrent neural networks with Hessian-free optimization,” In Neural Networks: Tricks of the Trade , 2nd ed., vol. 7700, G. Montavon, G. B. Orr, and K.-R. Müller, Ed. Berlin Heidelberg: Springer, 2012, pp. 479–535
2012
Cited alongside, same era.
I. Sutskever, J. Martens, G. Dahl, and G. E. Hinton, “On the importance of momentum and initialization in deep learning,” In 30th Int. Conf. Machine Learning , Atlanta, 2013, pp. 1139–1147
2013
Cited alongside, same era.
T. Schaul, S. Zhang, and Y. LeCun, “No more pesky learning rates,” arXiv:1206.1106, 2013
2013
Cited alongside, same era.
G. Hinton, Neural Networks for Machine Learning . Retrieved from http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture_slides_lec6.pdf
Cited in the paper.
Y. LeCun, C. Cortes, and C. J. C. Burges, THE MNIST DATABASE . Retrieved from http://yann.lecun.com/exdb/mnist/
Cited in the paper.
J. Martens and R. B. Grosse, “Optimizing neural networks with Kronecker-factored approximate curvature,” in Proc. 32nd Int. Conf. Machine Learning , 2015, pp. 2408–2417
2015
Closest in time.
C. Li, C. Chen, D. Carlson, and L. Carin, “Preconditioned stochastic gradient Langevin dynamics for deep neural networks,” in AAAI Conf. Artificial Intelligence , 2016
2016
Closest in time.
2016
Closest in time.