Fetching the paper…
Reading the bibliography…
Stochastic gradient algorithms are the main focus of large-scale optimization problems and led to important successes in the recent advancement of the deep learning algorithms.
H. Robbins and S. Monro, “A stochastic approximation method,” The annals of mathematical statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
S. Becker and Y. Le Cun, “Improving the convergence of back-propagation learning with second order methods,” in Proceedings of the 1988 connectionist models summer school . San Matteo, CA: Morgan Kaufmann, 1988, pp. 29–37
1988
Earlier work this paper cites.
Y. LeCun, P. Y. Simard, and B. Pearlmutter, “Automatic learning rate maximization by on-line estimation of the hessian’s eigenvectors,” Advances in neural information processing systems , vol. 5, pp. 156–163, 1993
1993
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” IEEE transactions on neural networks , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, no. 8, pp. 1735–1780, Nov. 1997. [Online]. Available: http://dx.doi.org/10.1162/neco.1997.9.8.1735
1997
Earlier work this paper cites.
S. Hochreiter, Y. Bengio, P. Frasconi, and J. Schmidhuber, “Gradient flow in recurrent nets: the difficulty of learning long-term dependencies,” 2001
2001
Earlier work this paper cites.
Y. Levin and A. Ben-Israel, “Directional newton methods in n variables,” Mathematics of Computation , vol. 71, no. 237, pp. 251–262, 2002
2002
Earlier work this paper cites.
N. N. Schraudolph, “Fast curvature matrix-vector products for second-order gradient descent,” Neural computation , vol. 14, no. 7, pp. 1723–1738, 2002
2002
Earlier work this paper cites.
H.-B. An and Z.-Z. Bai, “Directional secant method for nonlinear equations,” Journal of computational and applied mathematics , vol. 175, no. 2, pp. 291–304, 2005
2005
Earlier work this paper cites.
M. Liwicki and H. Bunke, “Iam-ondb - an on-line english sentence database acquired from handwritten text on a whiteboard.” in ICDAR . IEEE Computer Society, 2005, pp. 956–961
2005
Cited alongside, same era.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” The Journal of Machine Learning Research , vol. 12, pp. 2121–2159, 2011
2011
Cited alongside, same era.
2012
Cited alongside, same era.
Y. A. LeCun, L. Bottou, G. B. Orr, and K.-R. Müller, “Efficient backprop,” in Neural networks: Tricks of the trade . Springer, 2012, pp. 9–48
2012
Cited alongside, same era.
M. D. Zeiler, “Adadelta: An adaptive learning rate method,” arXiv preprint arXiv:1212.5701 , 2012
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in Neural Information Processing Systems , 2013, pp. 315–323
2013
Later among the works it cites.
2013
Later among the works it cites.
2013
Later among the works it cites.
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2012
Cited alongside, same era.
T. Mikolov, I. Sutskever, A. Deoras, H. Le, S. Kombrink, and J. Cernocky, “Subword language modeling with neural networks,” preprint , 2012
2012
Cited alongside, same era.
F. Bastien, P. Lamblin, R. Pascanu, J. Bergstra, I. Goodfellow, A. Bergeron, N. Bouchard, D. Warde-Farley, and Y. Bengio, “Theano: new features and speed improvements,” Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop , 2012
2012
Cited alongside, same era.
2013
Cited alongside, same era.
C. Wang, X. Chen, A. Smola, and E. Xing, “Variance reduction for stochastic gradient optimization,” in Advances in Neural Information Processing Systems , 2013, pp. 181–189
2013
Cited alongside, same era.
2014
Later among the works it cites.
2014
Later among the works it cites.
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” in International Conference on Learning Representations , 2015
2015
Later among the works it cites.
B. van Merriënboer, D. Bahdanau, V. Dumoulin, D. Serdyuk, D. Warde-Farley, J. Chorowski, and Y. Bengio, “Blocks and Fuel: Frameworks for deep learning,” ArXiv e-prints , jun 2015
2015
Later among the works it cites.