Fetching the paper…
Reading the bibliography…
Nonsmooth nonconvex optimization problems broadly emerge in machine learning and business decision making, whereas two core challenges impede the development of efficient solution methods with finite-time convergence guarantee: the lack of computationally tractable optimality criterion and the lack of computationally powerful oracles.
Stochastic optimization problems with nondifferentiable cost functionals
D. P. Bertsekas · 1973
Earlier work this paper cites.
Optimization of Lipschitz continuous functions
A. Goldstein · 1977
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
Some NP-complete problems in quadratic and nonlinear programming
K. G. Murty and S. N. Kabadi · 1987
Earlier work this paper cites.
Optimization and Nonsmooth Analysis
F. H. Clarke · 1990
Earlier work this paper cites.
On concepts of directional differentiability
A. Shapiro · 1990
Earlier work this paper cites.
Restricted step and Levenberg-Marquardt techniques in proximal bundle methods for nonconvex nondifferentiable optimization
K. C. Kiwiel · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
An Introduction to o-Minimal Geometry
M. Coste · 2000
Earlier work this paper cites.
Minimizing nonconvex nonsmooth functions via cutting planes and proximity control
A. Fuduli, M. Gaudioso, and G. Giallombardo · 2004
Earlier work this paper cites.
Stochastic approximations and differential inclusions
M. Benaïm, J. Hofbauer, and S. Sorin · 2005
Earlier work this paper cites.
A robust gradient sampling algorithm for nonsmooth, nonconvex optimization
J. V. Burke, A. S. Lewis, and M. L. Overton · 2005
Earlier work this paper cites.
Online convex optimization in the bandit setting: Gradient descent without a gradient
A. D. Flaxman, A. T. Kalai, and H. B. McMahan · 2005
Earlier work this paper cites.
Clarke subgradients of stratifiable functions
J. Bolte, A. Daniilidis, A. Lewis, and M. Shiota · 2007
Earlier work this paper cites.
Convergence of the gradient sampling algorithm for nonsmooth nonconvex optimization
K. C. Kiwiel · 2007
Earlier work this paper cites.
Large deviations of vector-valued martingales in 2-smooth normed spaces
A. Juditsky and A. S. Nemirovski · 2008
Earlier work this paper cites.
Supply chain management — an overview
H. Stadtler · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Variational Analysis , volume 317
R. T. Rockafellar and R. J-B. Wets · 2009
Earlier work this paper cites.
Optimal algorithms for online convex optimization with multi-point bandit feedback
A. Agarwal, O. Dekel, and L. Xiao · 2010
Earlier work this paper cites.
Dynamic Asset Pricing Theory
D. Duffie · 2010
Earlier work this paper cites.
Optimization via simulation over discrete decision variables
B. L. Nelson · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Randomized smoothing for stochastic optimization
J. C. Duchi, P. L. Bartlett, and M. J. Wainwright · 2012
Earlier work this paper cites.
On stochastic gradient and subgradient methods with adaptive steplength sequences
F. Yousefian, A. Nedić, and U. V. Shanbhag · 2012
Earlier work this paper cites.
Stochastic convex optimization with bandit feedback
A. Agarwal, D. P. Foster, D. Hsu, S. M. Kakade, and A. Rakhlin · 2013
Earlier work this paper cites.
Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward-backward splitting, and regularized Gauss-Seidel methods
H. Attouch, J. Bolte, and B. F. Svaiter · 2013
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Y. Nesterov · 2013
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov · 2014
Cited alongside, same era.
Complexity analysis of interior point algorithms for non-Lipschitz and nonconvex minimization
W. Bian, X. Chen, and Y. Ye · 2015
Cited alongside, same era.
The loss surfaces of multilayer networks
A. Choromanska, M. Henaff, M. Mathieu, G. B. Arous, and Y. LeCun · 2015
Cited alongside, same era.
Optimal rates for zero-order convex optimization: The power of two function evaluations
J. C. Duchi, M. I. Jordan, M. J. Wainwright, and A. Wibisono · 2015
Cited alongside, same era.
Escaping from saddle points—online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Analysis of nonsmooth stochastic approximation: the differential inclusion approach
S. Majewski, B. Miasojedow, and E. Moulines · 2018
Later among the works it cites.
Lectures on Convex Optimization , volume 137
Y. Nesterov · 2018
Later among the works it cites.
A geometric analysis of phase retrieval
J. Sun, Q. Qu, and J. Wright · 2018
Later among the works it cites.
Stochastic zeroth-order optimization in high dimensions
Y. Wang, S. Du, S. Balakrishnan, and A. Singh · 2018
Later among the works it cites.
ZO-AdaMM: zeroth-order adaptive momentum method for black-box optimization
X. Chen, S. Liu, K. Xu, X. Li, X. Lin, M. Hong, and D. Cox · 2019
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
D. Davis and D. Drusvyatskiy · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Discrete optimization via simulation
L. J. Hong, B. L. Nelson, and J. Xu · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Regularized M-estimators with nonconvexity: Statistical and algorithmic theory for local optima
P-L. Loh and M. J. Wainwright · 2015
Cited alongside, same era.
On the low-rank approach for semidefinite programs arising in synchronization and community detection
A. S. Bandeira, N. Boumal, and V. Voroninski · 2016
Cited alongside, same era.
Global optimality of local search for low rank matrix recovery
S. Bhojanapalli, B. Neyshabur, and N. Srebro · 2016
Cited alongside, same era.
The non-convex Burer-Monteiro approach works on smooth semidefinite programs
N. Boumal, V. Voroninski, and A. S. Bandeira · 2016
Cited alongside, same era.
Introduction to Optimization and Hadamard Semidifferential Calculus
M. C. Delfour · 2019
Later among the works it cites.
Efficiency of minimizing compositions of convex functions and smooth maps
D. Drusvyatskiy and C. Paquette · 2019
Later among the works it cites.
Improved zeroth-order variance reduced algorithms and analysis for nonconvex optimization
K. Ji, Z. Wang, Y. Zhou, and Y. Liang · 2019
Later among the works it cites.
Pytorch: an imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, and L. Antiga · 2019
Later among the works it cites.
High-Dimensional Statistics: A Non-asymptotic Viewpoint , volume 48
M. J. Wainwright · 2019
Later among the works it cites.
Second-order information in non-convex stochastic optimization: Power and limitations
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, A. Sekhari, and K. Sridharan · 2020
Later among the works it cites.
On the convergence to stationary points of deterministic and randomized feasible descent directions methods
A. Beck and N. Hallak · 2020
Later among the works it cites.
Gradient sampling methods for nonsmooth optimization
J. V. Burke, F. E. Curtis, A. S. Lewis, M. L. Overton, and L. E. A. Simões · 2020
Later among the works it cites.
Lower bounds for finding stationary points I
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2020
Later among the works it cites.
Pathological subgradient dynamics
A. Daniilidis and D. Drusvyatskiy · 2020
Later among the works it cites.
Stochastic subgradient method converges on tame functions
D. Davis, D. Drusvyatskiy, S. Kakade, and J. D. Lee · 2020
Later among the works it cites.
Implicit regularization in nonconvex statistical estimation: Gradient descent converges linearly for phase retrieval, matrix completion, and blind deconvolution
C. Ma, K. Wang, Y. Chi, and Y. Chen · 2020
Later among the works it cites.
Complexity of finding stationary points of nonconvex nonsmooth functions
J. Zhang, H. Lin, S. Jegelka, S. Sra, and A. Jadbabaie · 2020
Later among the works it cites.
Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning
J. Bolte and E. Pauwels · 2021
Later among the works it cites.
Lower bounds for finding stationary points II: First-order methods
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2021
Later among the works it cites.
On nonconvex optimization for machine learning: Gradients, stochasticity, and saddle points
C. Jin, P. Netrapalli, R. Ge, S. M. Kakade, and M. I. Jordan · 2021
Later among the works it cites.
Oracle complexity in nonsmooth nonconvex optimization
G. Kornowski and O. Shamir · 2021
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2022
Closest in time.
A gradient sampling method with complexity guarantees for Lipschitz functions in high and low dimensions
D. Davis, D. Drusvyatskiy, Y. T. Lee, S. Padmanabhan, and G. Ye · 2022
Closest in time.
Accelerated zeroth-order and first-order momentum methods from mini to minimax optimization
F. Huang, S. Gao, J. Pei, and H. Huang · 2022
Closest in time.
On the finite-time complexity and practical computation of approximate stationarity concepts of Lipschitz functions
L. Tian, K. Zhou, and A. M-C. So · 2022
Closest in time.