Advances in Low-Memory Subgradient Optimization
Original
P. E. Dvurechensky, A. V. Gasnikov, E. A. Nurminski, and F. S. Stonyakin · 1902
Earlier work this paper cites.
Gradient methods for problems with inexact model of the objective
Original
F. S. Stonyakin, D. Dvinskikh, P. Dvurechensky, A. Kroshnin, O. Kuznetsova, A. Agafonov, A. Gasnikov, A. Tyurin, C. A. Uribe, D. Pasechnyuk, and S. Artamonov · 1902
Earlier work this paper cites.
On a combination of alternating minimization and Nesterov’s momentum
Original
S. Guminov, P. Dvurechensky, N. Tupitsa, and A. Gasnikov · 1906
Earlier work this paper cites.
On the line-search gradient methods for stochastic optimization
Original
D. Dvinskikh, A. Ogaltsov, A. Gasnikov, P. Dvurechensky, and V. Spokoiny · 1911
Earlier work this paper cites.
A topological property of real analytic subsets
S. Lojasiewicz · 1963
Earlier work this paper cites.
Gradient methods for the minimisation of functionals
B. Polyak · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
A simplex method for function minimization
J. A. Nelder and R. Mead · 1965
Earlier work this paper cites.
Generalized gradient descent with application to block programming
N. Z. Shor · 1967
Earlier work this paper cites.
Adaptive step size random search
M. Schumer and K. Steiglitz · 1968
Earlier work this paper cites.
Numerical methods for finding global extrema (case of a non-uniform mesh)
Y. G. Evtushenko · 1971
Earlier work this paper cites.
Generalized descent for global optimization
A. O. Griewank · 1981
Earlier work this paper cites.
Orth-method for smooth convex optimization
A. Nemirovski · 1982
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Some np-complete problems in quadratic and nonlinear programming
K. G. Murty and S. N. Kabadi · 1987
Earlier work this paper cites.
Introduction to Optimization
B. Polyak · 1987
Earlier work this paper cites.
Training a 3-node neural network is np-complete
A. Blum and R. L. Rivest · 1989
Earlier work this paper cites.
Black-box complexity of local minimization
S. A. Vavasis · 1993
Earlier work this paper cites.
Improved approximation algorithms for maximum cut and satisfiability problems using semidefinite programming
M. X. Goemans and D. P. Williamson · 1995
Earlier work this paper cites.
Incremental gradient algorithms with stepsizes bounded away from zero
M. V. Solodov · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
P. Tseng · 1998
Earlier work this paper cites.
Trust Region Methods
A. Conn, N. Gould, and P. Toint · 2000
Earlier work this paper cites.
Lectures on Modern Convex Optimization
A. Ben-Tal and A. Nemirovski · 2001
Earlier work this paper cites.
Convex Optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: a basic course
Y. Nesterov · 2004
Earlier work this paper cites.
Decoding by linear programming
E. J. Candes and T. Tao · 2005
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. Polyak · 2006
Earlier work this paper cites.
Numerical optimization
J. Nocedal and S. Wright · 2006
Earlier work this paper cites.
Zeroth-order methods for noisy Hölder-gradient functions
Original
I. Shibaev, P. Dvurechensky, and A. Gasnikov · 2006
Earlier work this paper cites.
Optimal inapproximability results for max-cut and other 2-variable csps?
S. Khot, G. Kindler, E. Mossel, and R. O’Donnell · 2007
Earlier work this paper cites.
Stochastic global optimization
A. Zhigljavsky and A. Zilinskas · 2007
Earlier work this paper cites.
A simple proof of the restricted isometry property for random matrices
R. Baraniuk, M. Davenport, R. DeVore, and M. Wakin · 2008
Earlier work this paper cites.
Enhancing sparsity by reweighted ℓ 1 \ell_{1} minimization
E. J. Candes, M. B. Wakin, and S. P. Boyd · 2008
Earlier work this paper cites.
Encyclopedia of optimization
C. A. Floudas and P. M. Pardalos · 2008
Earlier work this paper cites.
Iterative hard thresholding for compressed sensing
T. Blumensath and M. E. Davies · 2009
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
L. Bottou · 2009
Earlier work this paper cites.
Exact matrix completion via convex optimization
E. J. Candès and B. Recht · 2009
Earlier work this paper cites.
Introduction to Derivative-Free Optimization
A. Conn, K. Scheinberg, and L. Vicente · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky, G. Hinton, et al · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Earlier work this paper cites.
The power of convex relaxation: Near-optimal matrix completion
E. J. Candès and T. Tao · 2010
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Introduction to online optimization
S. Bubeck · 2011
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. part i: motivation, convergence and numerical results
C. Cartis, N. I. Gould, and P. L. Toint · 2011
Earlier work this paper cites.
Adaptive cubic regularisation methods for unconstrained optimization. part ii: worst-case function- and derivative-evaluation complexity
C. Cartis, N. I. M. Gould, and P. L. Toint · 2011
Earlier work this paper cites.
Proximal splitting methods in signal processing
P. L. Combettes and J.-C. Pesquet · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Random gradient-free minimization of convex functions
Y. Nesterov and V. Spokoiny · 2011
Earlier work this paper cites.
Optimization with sparsity-inducing penalties
F. Bach, R. Jenatton, J. Mairal, G. Obozinski, et al · 2012
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Statistical language models based on neural networks
T. Mikolov · 2012
Earlier work this paper cites.
How to make the gradients small
Y. Nesterov · 2012
Earlier work this paper cites.
Parametric estimation. finite sample theory
V. Spokoiny et al · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Estimation, optimization, and parallelism when data is sparse
J. Duchi, M. I. Jordan, and B. McMahan · 2013
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Original
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Low-rank matrix completion using alternating minimization
P. Jain, P. Netrapalli, and S. Sanghavi · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Original
M. Schmidt and N. L. Roux · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien · 2014
Earlier work this paper cites.
Finito: A faster, permutable incremental gradient method for big data problems
A. Defazio, J. Domke, et al · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Original
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
On the computational efficiency of training neural networks
R. Livni, S. Shalev-Shwartz, and O. Shamir · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
E. J. Candes, X. Li, and M. Soltanolkotabi · 2015
Earlier work this paper cites.
Stochastic block mirror descent methods for nonsmooth and stochastic optimization
C. D. Dang and G. Lan · 2015
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Earlier work this paper cites.
Intersecting faces: Non-negative matrix factorization with new guarantees
R. Ge and J. Zou · 2015
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2015
Earlier work this paper cites.
Variance reduced stochastic gradient descent with neighbors
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams · 2015
Earlier work this paper cites.
Incremental majorization-minimization optimization with application to large-scale machine learning
J. Mairal · 2015
Earlier work this paper cites.
Phase retrieval with application to optical imaging: a contemporary overview
Y. Shechtman, Y. C. Eldar, O. Cohen, H. N. Chapman, J. Miao, and M. Segev · 2015
Earlier work this paper cites.
Efficient approaches for escaping higher order saddle points in non-convex optimization
A. Anandkumar and R. Ge · 2016
Earlier work this paper cites.
Dropping convexity for faster semi-definite optimization
S. Bhojanapalli, A. Kyrillidis, and S. Sanghavi · 2016
Earlier work this paper cites.
Foundations of data science
A. Blum, J. Hopcroft, and R. Kannan · 2016
Earlier work this paper cites.
Learning supervised pagerank with gradient-based and gradient-free optimization methods
Original
L. Bogolubsky, P. Dvurechensky, A. Gasnikov, G. Gusev, Y. Nesterov, A. M. Raigorodskii, A. Tikhonov, and M. Zhukovskii · 2016
Earlier work this paper cites.
Gradient descent efficiently finds the cubic-regularized non-convex newton step
Original
Y. Carmon and J. C. Duchi · 2016
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Earlier work this paper cites.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
Original
S. Ghadimi, G. Lan, and H. Zhang · 2016
Earlier work this paper cites.
Deep Learning
I. Goodfellow, Y. Bengio, and A. Courville · 2016
Earlier work this paper cites.
Deep learning
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
H. Karimi, J. Nutini, and M. Schmidt · 2016
Earlier work this paper cites.
Optimizing star-convex functions
J. C. H. Lee and P. Valiant · 2016
Earlier work this paper cites.
The power of normalization: Faster evasion of saddle points
Original
K. Y. Levy · 2016
Earlier work this paper cites.
Identifiability in blind deconvolution with subspace or sparsity constraints
Y. Li, K. Lee, and Y. Bresler · 2016
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Original
I. Loshchilov and F. Hutter · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2016
Earlier work this paper cites.
Proximal stochastic methods for nonsmooth nonconvex finite-sum optimization
S. J. Reddi, S. Sra, B. Poczos, and A. J. Smola · 2016
Earlier work this paper cites.
Algorithms and matching lower bounds for approximately-convex optimization
A. Risteski and Y. Li · 2016
Earlier work this paper cites.
Sdca without duality, regularization, and individual convexity
S. Shalev-Shwartz · 2016
Earlier work this paper cites.
Local minima in training of deep networks
G. Swirszcz, W. M. Czarnecki, and R. Pascanu · 2016
Earlier work this paper cites.
Finding approximate local minima faster than gradient descent
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2017
Earlier work this paper cites.