A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
An algorithm for quadratic programming
Frank, M. and Wolfe, P · 1956
Earlier work this paper cites.
Minimization of unsmooth functionals
Polyak, B. T · 1969
Earlier work this paper cites.
Minimization methods for non-differentiable functions
Shor, N. Z · 1985
Earlier work this paper cites.
A generalized subgradient method with relaxation step
Brännlund, U · 1995
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Earlier work this paper cites.
ADADELTA: an adaptive learning rate method
Zeiler, M · 2012
Earlier work this paper cites.
Block-coordinate Frank-Wolfe optimization for structural SVMs
Lacoste-Julien, S. and Jaggi, M · 2013
Earlier work this paper cites.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y · 2013
Earlier work this paper cites.
Gradient methods for convex minimization: better rates under weaker conditions
Zhang, H. and Yin, W · 2013
Earlier work this paper cites.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Systems , 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Man, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Vi, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Bowman, S. R., Angeli, G., Potts, C., and Manning, C. D · 2015
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
Martens, J. and Grosse, R · 2015
Earlier work this paper cites.
Scale-free algorithms for online linear optimization
Orabona, F. and Pál, D · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., Badia, A. P., Hermann, K. M., Zwols, Y., Ostrovski, G., Cain, A., King, H., Summerfield, C., Blunsom, P., Kavukcuoglu, K., and Hassabis, D · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.