Fetching the paper…
Reading the bibliography…
This paper focuses on the problem of minimizing a locally Lipschitz continuous function.
USSR Comput. Math. Math. Phys. 4
Polyak, B.T.: Some methods of speeding up the convergence of iteration methods · 1964
Earlier work this paper cites.
McGraw-Hill, Inc., New York (1964)
Rudin, W.: Principles of mathematical analysis · 1964
Earlier work this paper cites.
USSR Comput. Math. Math. Phys. 7
Bregman, L.M.: The relaxation method of finding the common point of convex sets and its application to the solution of problems in convex programming · 1967
Earlier work this paper cites.
Wiley-Interscience (1983)
Nemirovskij, A.S., Yudin, D.B.: Problem complexity and method efficiency in optimization · 1983
Earlier work this paper cites.
SIAM (1990)
Clarke, F.H.: Optimization and Nonsmooth Analysis · 1990
Earlier work this paper cites.
McGraw-Hill, Inc., New York (1991)
Rudin, W.: Functional analysis 2nd ed · 1991
Earlier work this paper cites.
Marcus, M., Santorini, B., Marcinkiewicz, M.A.: Building a large annotated corpus of english: The penn treebank (1993)
1993
Earlier work this paper cites.
Duke Math. J. (1996)
Van den Dries, L., Miller, C.: Geometric categories and o-minimal structures · 1996
Earlier work this paper cites.
Neural Comput. 9
Hochreiter, S., Schmidhuber, J.: Long short-term memory · 1997
Earlier work this paper cites.
Princeton University Press (1997)
Rockafellar, R.T.: Convex Analysis, vol. 11 · 1997
Earlier work this paper cites.
Springer Verlag, Heidelberg, Berlin, New York (1998)
Rockafellar, R., Wets, R.J.B.: Variational Analysis · 1998
Earlier work this paper cites.
Istituti editoriali e poligrafici internazionali Pisa (2000)
Coste, M.: An introduction to o-minimal geometry · 2000
Earlier work this paper cites.
Oper. Res. Lett. 31
Beck, A., Teboulle, M.: Mirror descent and nonlinear projected subgradient methods for convex optimization · 2003
Earlier work this paper cites.
SIAM J. Control Optim. 42
Bolte, J., Teboulle, M.: Barrier operators and associated gradient-like dynamical systems for constrained minimization problems · 2003
Earlier work this paper cites.
SIAM J. Control Optim. 43
Alvarez, F., Bolte, J., Brahic, O.: Hessian Riemannian gradient flows in convex programming · 2004
Earlier work this paper cites.
SIAM J. Control Optim. 44
Benaïm, M., Hofbauer, J., Sorin, S.: Stochastic approximations and differential inclusions · 2005
Earlier work this paper cites.
SIAM J. Optim. 18
Bolte, J., Daniilidis, A., Lewis, A., Shiota, M.: Clarke subgradients of stratifiable functions · 2007
Earlier work this paper cites.
Springer (2009)
Borkar, V.S.: Stochastic approximation: a dynamical systems viewpoint, vol. 48 · 2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
2009
Earlier work this paper cites.
SIAM J. Sci. Comput. 37
Benamou, J.D., Carlier, G., Cuturi, M., Nenna, L., Peyré, G.: Iterative Bregman projections for regularized transportation problems · 2015
Earlier work this paper cites.
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
Math. Oper. Res. 42
Bauschke, H.H., Bolte, J., Teboulle, M.: A descent lemma beyond Lipschitz gradient continuity: first-order methods revisited and applications · 2017
Earlier work this paper cites.
arXiv preprint arXiv:1708.03888 (2017)
You, Y., Gitman, I., Ginsburg, B.: Large batch training of convolutional networks · 2017
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 560–569. PMLR (2018)
Bernstein, J., Wang, Y.X., Azizzadenesheli, K., Anandkumar, A.: signsgd: Compressed optimisation for non-convex problems · 2018
Cited alongside, same era.
SIAM J. Optim. 28
Bolte, J., Sabach, S., Teboulle, M., Vaisbourd, Y.: First order methods beyond convexity and lipschitz gradient continuity with applications to quadratic inverse problems · 2018
Cited alongside, same era.
SIAM J. Optim. 28
Duchi, J.C., Ruan, F.: Stochastic methods for composite and weakly convex optimization problems · 2018
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 1832–1841. PMLR (2018)
Gunasekar, S., Lee, J., Soudry, D., Srebro, N.: Characterizing implicit bias in terms of optimization geometry · 2018
Cited alongside, same era.
Comput. Optim. Appl. 79
Hanzely, F., Richtárik, P.: Fastest rates for stochastic mirror descent methods · 2021
Later among the works it cites.
arXiv preprint arXiv:2108.06808 (2021)
Li, Y., Ju, C., Fang, E.X., Zhao, T.: Implicit regularization of Bregman proximal point algorithm and mirror descent on separable data · 2021
Later among the works it cites.
SIAM J. Control Optim. 59
Ruszczynski, A.: A stochastic subgradient method for nonsmooth nonconvex multilevel composition optimization · 2021
Later among the works it cites.
Adv. Neural Inf. Process. Syst. 35
Ghai, U., Lu, Z., Hazan, E.: Non-convex online learning via algorithmic equivalence · 2022
Later among the works it cites.
J. Optim. Theory Appl. 194
Gürbüzbalaban, M., Ruszczyński, A., Zhu, L.: A stochastic subgradient method for distributionally robust non-convex and non-smooth learning · 2022
Later among the works it cites.
In: URL https://openreview.net/forum
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SIAM J. Optim. 28
Lu, H., Freund, R.M., Nesterov, Y.: Relatively smooth convex optimization by first-order methods, and applications · 2018
Cited alongside, same era.
arXiv preprint arXiv:1806.04781 (2018)
Zhang, S., He, N.: On the convergence rate of stochastic mirror descent for nonsmooth nonconvex optimization · 2018
Cited alongside, same era.
arXiv preprint arXiv:1904.00962 (2019)
You, Y., Li, J., Reddi, S., Hseu, J., Kumar, S., Bhojanapalli, S., Song, X., Demmel, J., Keutzer, K., Hsieh, C.J.: Large batch optimization for deep learning: Training bert in 76 minutes · 2019
Cited alongside, same era.
Adv. Neural Inf. Process. Syst. 33
Amid, E., Warmuth, M.K.: Reparameterizing mirror descent as gradient descent · 2020
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 2260–2268. PMLR (2020)
Cutkosky, A., Mehta, H.: Momentum improves normalized sgd · 2020
Cited alongside, same era.
Found. Comput. Math. 20
Davis, D., Drusvyatskiy, D., Kakade, S., Lee, J.D.: Stochastic subgradient method converges on tame functions · 2020
Cited alongside, same era.
arXiv preprint arXiv:2010.00406 (2020)
Defazio, A.: Understanding the role of momentum in non-convex optimization: Practical insights from a Lyapunov analysis · 2020
Cited alongside, same era.
Jelassi, S., Li, Y.: Towards understanding how momentum improves generalization in deep learning, 2022 · 2022
Later among the works it cites.
Adv. Neural Inf. Process. Syst. 35
Li, Z., Wang, T., Lee, J.D., Arora, S.: Implicit bias of gradient descent on reparametrized models: On equivalence to mirror descent · 2022
Later among the works it cites.
SIAM J. Optim. 32
Yang, L., Toh, K.C.: Bregman proximal point algorithm revisited: A new inexact version and its inertial variant · 2022
Later among the works it cites.
SIAM J. Optim. 33
Bolte, J., Le, T., Pauwels, E.: Subgradient sampling for nonsmooth nonconvex minimization · 2023
Later among the works it cites.
Comput. Optim. Appl. 85
Chu, H.T., Liang, L., Toh, K.C., Yang, L.: An efficient implementable inexact entropic proximal point algorithm for a class of linear programming problems · 2023
Later among the works it cites.
Math. Program. 198
Lan, G.: Policy mirror descent for reinforcement learning: Linear convergence, new sampling complexity, and generalized problem classes · 2023
Later among the works it cites.
J. Mach. Learn. Res. 24
Sun, H., Gatmiry, K., Ahn, K., Azizan, N.: A unified approach to controlling implicit regularization via mirror descent · 2023
Later among the works it cites.
arXiv preprint arXiv:2307.10053 (2023)
Xiao, N., Hu, X., Toh, K.C.: Convergence guarantees for stochastic subgradient methods in nonsmooth nonconvex optimization · 2023
Later among the works it cites.
SIAM J. Optim. 33
Zhan, W., Cen, S., Huang, B., Chen, Y., Lee, J.D., Chi, Y.: Policy mirror descent for regularized reinforcement learning: A generalized framework with linear convergence · 2023
Later among the works it cites.
J. Optim. Theory Appl. 201
Le, T.: Nonsmooth nonconvex stochastic heavy ball · 2024
Closest in time.
J. Mach. Learn. Res. 25
Ramezani-Kebrya, A., Antonakopoulos, K., Cevher, V., Khisti, A., Liang, B.: On the generalization of stochastic gradient descent with momentum · 2024
Closest in time.
arXiv preprint arXiv:2404.09438 (2024)
Xiao, N., Ding, K., Hu, X., Toh, K.C.: Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization · 2024
Closest in time.
J. Mach. Learn. Res. 25
Xiao, N., Hu, X., Liu, X., Toh, K.C.: Adam-family methods for nonsmooth optimization with convergence guarantees · 2024
Closest in time.
J. Mach. Learn. Res. 26
Ding, K., Li, J., Toh, K.C.: Nonconvex stochastic Bregman proximal gradient method with application to deep learning · 2025
Closest in time.
Trans. Mach. Learn. Res. (2025)
Ding, K., Xiao, N., Toh, K.C.: Adam-family methods with decoupled weight decay in deep learning · 2025
Closest in time.
Comput. Optim. Appl. 90
Takahashi, S., Takeda, A.: Approximate Bregman proximal gradient algorithm for relatively smooth nonconvex optimization · 2025
Closest in time.