Fetching the paper…
Reading the bibliography…
This paper provides a review and commentary on the past, present, and future of numerical optimization algorithms in the context of machine learning applications.
What is Mathematics?
R. Courant and H. Robbins · 1941
Earlier work this paper cites.
A Stochastic Approximation Method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On a stochastic approximation method
K. L. Chung · 1954
Earlier work this paper cites.
A New Approach to Linear Filtering and Prediction Problems
R. E. Kalman · 1960
Earlier work this paper cites.
Mathematical Statistics
J. E. Freund · 1962
Earlier work this paper cites.
On Convergence Proofs on Perceptrons
A. B. J. Novikoff · 1962
Earlier work this paper cites.
Probability Inequalities for Sums of Bounded Random Variables
W. Hoeffding · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
On stochastic approximations
E. G. Gladyshev · 1965
Earlier work this paper cites.
A theory of adaptive pattern classifiers
S.-I. Amari · 1967
Earlier work this paper cites.
On the uniform convergence of relative frequencies of events to their probabilities
V. N. Vapnik and A. Ya. Chervonenkis · 1968
Earlier work this paper cites.
The Convergence of a Class of Double-Rank Minimization Algorithms
C. G. Broyden · 1970
Earlier work this paper cites.
A New Approach to Variable Metric Algorithms
R. Fletcher · 1970
Earlier work this paper cites.
A Family of Variable Metric Updates Derived by Variational Means
D. Goldfarb · 1970
Earlier work this paper cites.
Conditioning of Quasi-Newton Methods for Function Minimization
D. F. Shanno · 1970
Earlier work this paper cites.
A convergence theorem for non negative almost supermartingales and some applications
H. Robbins and D. Siegmund · 1971
Earlier work this paper cites.
On Search Directions for Minimization Algorithms
M. J. D. Powell · 1973
Earlier work this paper cites.
A Characterization of Superlinear Convergence and Its Application to Quasi-Newton Methods
J. E. Dennis and J. J. Moré · 1974
Earlier work this paper cites.
Sur l’approximation, par elements finis déordre un, et la resolution, par penalisation-dualité, d’une classe de problems de Dirichlet non lineares
R. Glowinski and A. Marrocco · 1975
Earlier work this paper cites.
A Dual Algorithm for the Solution of Nonlinear Variational Problems via Finite Element Approximations
D. Gabay and B. Mercier · 1976
Earlier work this paper cites.
Maximum Likelihood from Incomplete Data via the EM Algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
Comparison of the convergence rates for single-step and multi-step optimization algorithms in the presence of noise
B. T. Polyak · 1977
Earlier work this paper cites.
On Cezaro’s convergence of the steepest descent method for approximating saddle point of convex-concave functions
A. S. Nemirovski and D. B. Yudin · 1978
Earlier work this paper cites.
Theory of Pattern Recognition
V. N. Vapnik and A. Y. Chervonenkis · 1979
Earlier work this paper cites.
Updating Quasi-Newton Matrices With Limited Storage
J. Nocedal · 1980
Earlier work this paper cites.
Inexact Newton Methods
R. S. Dembo, Eisenstat S. C., and T. Steihaug · 1982
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) {\mathcal{O}}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
The Conjugate Gradient Method and Trust Regions in Large Scale Optimization
T. Steihaug · 1983
Earlier work this paper cites.
Estimation of Dependences Based on Empirical Data
V. N. Vapnik · 1983
Earlier work this paper cites.
On the convergence properties of the em algorithm
C. F. J. Wu · 1983
Earlier work this paper cites.
Learning internal representations by error propagation
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Learning Representations by Back-Propagating Errors
D. E. Rumelhart, G. E. Hinton, and R. J. Williams · 1986
Earlier work this paper cites.
Efficient estimations from a slowly convergent Robbins-Monro process
D. Ruppert · 1988
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second-order methods
S. Becker and Y. LeCun · 1989
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis · 1989
Earlier work this paper cites.
Experiments with Time Delay Networks and Dynamic Time Warping for Speaker Independent Isolated Digit Recognition
L. Bottou, F. Fogelman Soulié, P. Blanchet, and J. S. Lienard · 1989
Earlier work this paper cites.
Handwritten Digit Recognition with a Back-Propagation Network
Y. Le Cun, B. E. Boser, J. S. Denker, D. Henderson, R. E. Howard, W. E. Hubbard, and L. D. Jackel · 1989
Earlier work this paper cites.
On the limited memory BFGS method for large scale optimization
D. C. Liu and J. Nocedal · 1989
Earlier work this paper cites.
Stochastic Gradient Learning in Neural Networks
L. Bottou · 1991
Earlier work this paper cites.
Universal Portfolios
T. M. Cover · 1991
Earlier work this paper cites.
New Method of Stochastic Approximation Type
B. T. Polyak · 1991
Earlier work this paper cites.
On the Douglas-Rachford splitting method and the proximal point algorithm for maximal monotone operators
J. Eckstein and D. P. Bertsekas · 1992
Earlier work this paper cites.
Fisher’s method of scoring
M. R. Osborne · 1992
Earlier work this paper cites.
Acceleration of Stochastic Approximation by Averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Optimal stochastic search and adaptive momentum
Todd K. Leen and Genevieve B. Orr · 1993
Earlier work this paper cites.
Fast Exact Multiplication by the Hessian
B. A. Pearlmutter · 1994
Earlier work this paper cites.
Nonlinear Programming
D. P. Bertsekas · 1995
Earlier work this paper cites.
Support-Vector Networks
C. Cortes and V. N. Vapnik · 1995
Earlier work this paper cites.
De-noising by soft-thresholding
D.L. Donoho · 1995
Earlier work this paper cites.
Incremental least squares methods and the extended Kalman filter
D. P. Bertsekas · 1996
Earlier work this paper cites.
What is Mathematics?
R. Courant and H. Robbins · 1996
Earlier work this paper cites.
Numerical Methods for Unconstrained Optimization and Nonlinear Equations
J. E. Dennis, Jr. and R. B. Schnabel · 1996
Earlier work this paper cites.
Regression shrinkage and selection via the lasso
Robert Tibshirani · 1996
Earlier work this paper cites.
Methods of Information Geometry
S.-I. Amari and H. Nagaoka · 1997
Earlier work this paper cites.
Algorithm 78: L-BFGS-B: Fortran subroutines for large-scale bound constrained optimization
C. Zhu, R. H. Byrd, P. Lu, and J. Nocedal · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
S.-I. Amari · 1998
Cited alongside, same era.
Online Algorithms and Stochastic Approximations
L. Bottou · 1998
Cited alongside, same era.
Blind signal separation: statistical principles
Jean-François Cardoso · 1998
Cited alongside, same era.
Inductive Learning Algorithms and Representations for Text Categorization
S. T. Dumais, J. C. Platt, D. Hecherman, and M. Sahami · 1998
Cited alongside, same era.
A Simulation-Based Approach to Two-Stage Stochastic Programming with Recourse
T. Homem-de Mello and A. Shapiro · 1998
Cited alongside, same era.
Text Categorization with Suport Vector Machines: Learning with Many Relevant Features
T. Joachims · 1998
Cited alongside, same era.
Pegasos: primal estimated sub-gradient solver for SVM
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter · 2011
Later among the works it cites.
Optimization for Machine Learning
S. Sra, S. Nowozin, and S.J. Wright · 2011
Later among the works it cites.
Towards optimal one pass large scale learning with averaged stochastic gradient descent
W. Xu · 2011
Later among the works it cites.
Information-Theoretic Lower Bounds on the Oracle Complexity of Stochastic Convex Optimization
A. Agarwal, P. L. Bartlett, P. Ravikumar, and M. J. Wainwright · 2012
Later among the works it cites.
Optimization with sparsity-inducing penalties
F. Bach, R. Jenatton, J. Mairal, and G. Obozinski · 2012
Later among the works it cites.
Sample Size Selection in Optimization Methods for Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gradient Based Learning Applied to Document Recognition
Y. Le Cun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Cited alongside, same era.
Efficient Backprop
Y. Le Cun, L. Bottou, G. B. Orr, and K.-R. Müller · 1998
Cited alongside, same era.
The Importance of Convexity in Learning with Squared Loss
W. S. Lee, P. L. Bartlett, and R. C. Williamson · 1998
Cited alongside, same era.
A Statistical Study of On-line Learning
N. Murata · 1998
Cited alongside, same era.
Statistical Learning Theory
V. N. Vapnik · 1998
Cited alongside, same era.
Uniform Central Limit Theorems
R. M. Dudley · 1999
Cited alongside, same era.
R. H. Byrd, G. M. Chin, J. Nocedal, and Y. Wu · 2012
Later among the works it cites.
A family of second-order methods for convex \ \backslash ell _1-regularized optimization
Richard H Byrd, Gillian M Chin, Jorge Nocedal, and Figen Oztoprak · 2012
Later among the works it cites.
Large Scale Distributed Deep Networks
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, Q. V. Le, M. Z. Mao, M. Ranzato, A. W. Senior, P. A. Tucker, K. Yang, and A. Y. Ng · 2012
Later among the works it cites.
Hybrid Deterministic-Stochastic Methods for Data Fitting
M. P. Friedlander and M. Schmidt · 2012
Later among the works it cites.
Matrix Computations
G. H. Golub and C. F. Van Loan · 2012
Later among the works it cites.
ImageNet Classification with Deep Convolutional Neural Networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Later among the works it cites.
A Stochastic Gradient Method with an Exponential Convergence Rate for Finite Training Sets
N. Le Roux, M. Schmidt, and F. R. Bach · 2012
Later among the works it cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Yu Nesterov · 2012
Later among the works it cites.
Lecture 6.5. RMSPROP: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Later among the works it cites.
ADADELTA: an adaptive learning rate method
Matthew D. Zeiler · 2012
Later among the works it cites.
An inexact successive quadratic approximation method for convex L-1 regularized optimization
Richard H Byrd, Jorge Nocedal, and Figen Oztoprak · 2013
Later among the works it cites.
New types of deep neural network learning for speech recognition and related applications: An overview
L. Deng, G. E. Hinton, and B. Kingsbury · 2013
Later among the works it cites.
Stochastic variational inference
Matthew D. Hoffman, David M. Blei, Chong Wang, and John William Paisley · 2013
Later among the works it cites.
Big & quic: Sparse inverse covariance estimation for a million variables
Cho-Jui Hsieh, Mátyás A Sustik, Inderjit S Dhillon, Pradeep K Ravikumar, and Russell Poldrack · 2013
Later among the works it cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Later among the works it cites.
Normalized online learning
Stéphane Ross, Paul Mineiro, and John Langford · 2013
Later among the works it cites.
Practical inexact proximal quasi-newton method with global complexity analysis
Katya Scheinberg and Xiaocheng Tang · 2013
Later among the works it cites.
Minimizing finite sums with the stochastic average gradient
M. Schmidt, N. Le Roux, and F. Bach · 2013
Later among the works it cites.
Stochastic dual coordinate ascent methods for regularized loss
S. Shalev-Shwartz and T. Zhang · 2013
Later among the works it cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Later among the works it cites.
On stochastic proximal gradient algorithms
Yves F Atchade, Gersende Fort, and Eric Moulines · 2014
Later among the works it cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Çaglar Gülçehre, KyungHyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Later among the works it cites.
SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Later among the works it cites.
Large-scale object classification using label relation graphs
Jia Deng, Nan Ding, Yangqing Jia, Andrea Frome, Kevin Murphy, Samy Bengio, Yuan Li, Hartmut Neven, and Hartwig Adam · 2014
Later among the works it cites.
Qualitatively characterizing neural network optimization problems
Ian J. Goodfellow and Oriol Vinyals · 2014
Later among the works it cites.
Automatic Differentiation
A. Griewank · 2014
Later among the works it cites.
On adaptive sampling rules for stochastic recursions
F. S. Hashemi, S. Ghosh, and R. Pasupathy · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Later among the works it cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Later among the works it cites.
RES: Regularized Stochastic BFGS algorithm
A. Mokhtari and A. Ribeiro · 2014
Later among the works it cites.
CNN features off-the-shelf: an astounding baseline for recognition
A. Razavian, H. Azizpour, J. Sullivan, and S. Carlsson · 2014
Later among the works it cites.
Convex Optimization Algorithms
D. P. Bertsekas · 2015
Later among the works it cites.
The loss surfaces of multilayer networks
Anna Choromanska, Mikael Henaff, Michael Mathieu, Gérard Ben Arous, and Yann LeCun · 2015
Later among the works it cites.
Rmsprop and equilibrated adaptive learning rates for non-convex optimization
Yann N. Dauphin, Harm de Vries, Junyoung Chung, and Yoshua Bengio · 2015
Later among the works it cites.
On the convergence rate of incremental aggregated gradient algorithms
Mert Gurbuzbalaban, Asuman Ozdaglar, and Pablo Parrilo · 2015
Later among the works it cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
An Asynchronous Parallel Stochastic Coordinate Descent Algorithm
J. Liu, S. J. Wright, C. Ré, V. Bittorf, and S. Sridhar · 2015
Later among the works it cites.
Coordinate Descent Converges Faster with the Gauss-Southwell Rule than Random Selection
J. Nutini, M. Schmidt, I. H. Laradji, M. Friedlander, and H. Koepke · 2015
Later among the works it cites.
On Sampling Rates in Stochastic Recursions
R. Pasupathy, P. W. Glynn, S. Ghosh, and F. S. Hashemi · 2015
Later among the works it cites.
Newton sketch: A linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2015
Later among the works it cites.
ImageNet Large Scale Visual Recognition Challenge (ILSVRC)
Stanford Vision Lab · 2015
Later among the works it cites.
Second order stochastic optimization in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2016
Closest in time.
Optimal black-box reductions between optimization objectives
Zeyuan Allen Zhu and Elad Hazan · 2016
Closest in time.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Closest in time.
Manuscript in preparation, 2016
Jean Lafond, Nicolas Vasilache, and Léon Bottou · 2016
Closest in time.
Practical riemannian neural networks
Gaétan Marceau-Caron and Yann Ollivier · 2016
Closest in time.
Advances in neural information processing systems (29 volumes), 1987–2016
NIPS Foundation · 2016
Closest in time.
Sub-sampled Newton methods II: Local convergence rates
Farbod Roosta-Khorasani and Michael W Mahoney · 2016
Closest in time.