Fetching the paper…
Reading the bibliography…
In this paper, we propose a generalized framework for developing learning-rate-free momentum stochastic gradient descent (SGD) methods in the minimization of nonsmooth nonconvex functions, especially in training nonsmooth neural networks.
Ussr computational mathematics and mathematical physics 4
Polyak, B.T.: Some methods of speeding up the convergence of iteration methods · 1964
Earlier work this paper cites.
Publications Mathématiques de l’IHÉS 67
Bierstone, E., Milman, P.D.: Semianalytic and subanalytic sets · 1988
Earlier work this paper cites.
SIAM (1990)
Clarke, F.H.: Optimization and nonsmooth analysis, vol. 5 · 1990
Earlier work this paper cites.
Journal of the American Mathematical Society 9
Wilkie, A.J.: Model completeness results for expansions of the ordered field of real numbers by restricted Pfaffian functions and the exponential function · 1996
Earlier work this paper cites.
SIAM Journal on Control and Optimization 44
Benaïm, M., Hofbauer, J., Sorin, S.: Stochastic approximations and differential inclusions · 2005
Earlier work this paper cites.
In: Seminaire de probabilites XXXIII, pp. 1–68. Springer (2006)
Benaïm, M.: Dynamics of stochastic approximation algorithms · 2006
Earlier work this paper cites.
SIAM Journal on Optimization 18
Bolte, J., Daniilidis, A., Lewis, A., Shiota, M.: Clarke subgradients of stratifiable functions · 2007
Earlier work this paper cites.
Springer (2009)
Borkar, V.S.: Stochastic approximation: a dynamical systems viewpoint, vol. 48 · 2009
Earlier work this paper cites.
In: 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255. Ieee (2009)
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database · 2009
Earlier work this paper cites.
Krizhevsky, A., Hinton, G., et al.: Learning multiple layers of features from tiny images (2009)
2009
Earlier work this paper cites.
Journal of machine learning research 12
Duchi, J., Hazan, E., Singer, Y.: Adaptive subgradient methods for online learning and stochastic optimization · 2011
Earlier work this paper cites.
Springer Science & Business Media (2012)
Aubin, J.P., Cellina, A.: Differential inclusions: set-valued maps and viability theory, vol. 264 · 2012
Earlier work this paper cites.
Cited on 14
Hinton, G., Srivastava, N., Swersky, K.: Neural networks for machine learning lecture 6a overview of mini-batch gradient descent · 2012
Earlier work this paper cites.
arXiv preprint arXiv:1211.2260 (2012)
Streeter, M., McMahan, H.B.: No-regret algorithms for unconstrained online convex optimization · 2012
Earlier work this paper cites.
Advances in Neural Information Processing Systems 27
Orabona, F.: Simultaneous model selection and optimization through parameter-free stochastic learning · 2014
Earlier work this paper cites.
In Proceedings of the 3rd International Conference for Learning Representations (2015)
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization · 2015
Earlier work this paper cites.
In: Conference on Learning Theory, pp. 1286–1304. PMLR (2015)
Luo, H., Schapire, R.E.: Achieving all with no parameters: Adanormalhedge · 2015
Earlier work this paper cites.
In: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778 (2016)
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition · 2016
Earlier work this paper cites.
Advances in Neural Information Processing Systems 29
Orabona, F., Pál, D.: Coin betting and parameter-free online learning · 2016
Earlier work this paper cites.
Advances in Neural Information Processing Systems 30
Orabona, F., Tommasi, T.: Training deep networks without learning rates through coin betting · 2017
Cited alongside, same era.
In: Conference On Learning Theory, pp. 1493–1529. PMLR (2018)
Cutkosky, A., Orabona, F.: Black-box reductions for parameter-free online learning in banach spaces · 2018
Cited alongside, same era.
In: International conference on machine learning, pp. 3321–3330. PMLR (2019)
Kempka, M., Kotlowski, W., Warmuth, M.K.: Adaptive scale-invariant online algorithms for learning linear models · 2019
Cited alongside, same era.
In: International Conference on Machine Learning, pp. 822–831. PMLR (2020)
Bhaskara, A., Cutkosky, A., Kumar, R., Purohit, M.: Online learning with imperfect hints · 2020
Cited alongside, same era.
Advances in Neural Information Processing Systems 33
Bolte, J., Pauwels, E.: A mathematical model for automatic differentiation in machine learning · 2020
Cited alongside, same era.
Information Sciences 587
Jin, H., Chen, J., Zheng, H., Wang, Z., Xiao, J., Yu, S., Ming, Z.: Roby: Evaluating the adversarial robustness of a deep model by its decision boundaries · 2022
Later among the works it cites.
In: International Conference on Machine Learning, pp. 26085–26115. PMLR (2022)
Zhang, Z., Cutkosky, A., Paschalidis, I.: PDE-based optimal strategy for unconstrained online learning · 2022
Later among the works it cites.
IEEE Transactions on Computational Social Systems (2022)
Zhou, B., Chen, L., Zhao, S., Li, S., Zheng, Z., Pan, G.: Unsupervised domain adaptation for crime risk prediction across cities · 2022
Later among the works it cites.
arXiv preprint arXiv:2302.06675 (2023)
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Liu, Y., Pham, H., Dong, X., Luong, T., Hsieh, C.J., et al.: Symbolic discovery of optimization algorithms · 2023
Later among the works it cites.
In: International Conference on Machine Learning, pp. 7449–7479. PMLR (2023)
Defazio, A., Mishchenko, K.: Learning-rate-free learning by D-adaptation · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Davis, D., Drusvyatskiy, D., Kakade, S., Lee, J.D.: Stochastic subgradient method converges on tame functions · 2020
Cited alongside, same era.
In: Conference on Learning Theory, pp. 2858–2887. PMLR (2020)
Mhammedi, Z., Koolen, W.M.: Lipschitz and comparator-norm adaptivity in online learning · 2020
Cited alongside, same era.
Optimization Letters 14
Ruszczyński, A.: Convergence of a stochastic subgradient method with averaging for nonsmooth nonconvex constrained optimization · 2020
Cited alongside, same era.
The Journal of Machine Learning Research 21
Ward, R., Wu, X., Bottou, L.: Adagrad stepsizes: Sharp convergence over nonconvex landscapes · 2020
Cited alongside, same era.
IEEE Transactions on Neural Networks and Learning Systems 32
Wei, L., Zhao, S., Bourahla, O.F., Li, X., Wu, F., Zhuang, Y., Han, J., Xu, M.: End-to-end video saliency detection via a deep contextual spatiotemporal network · 2020
Cited alongside, same era.
Advances in Neural Information Processing Systems 34
Bolte, J., Le, T., Pauwels, E., Silveti-Falls, T.: Nonsmooth implicit differentiation for machine-learning and optimization · 2021
Cited alongside, same era.
Mathematical Programming 188
Bolte, J., Pauwels, E.: Conservative set valued fields, automatic differentiation, stochastic gradient methods and deep learning · 2021
Cited alongside, same era.
IMA Journal of Numerical Analysis p. drad098 (2023)
Hu, X., Xiao, N., Liu, X., Toh, K.C.: A constraint dissolving approach for nonsmooth optimization over the Stiefel manifold · 2023
Later among the works it cites.
Ivgi, M., Hinder, O., Carmon, Y.: DoG is SGD’s best friend: A parameter-free dynamic step size schedule pp. 14465–14499 (2023)
2023
Later among the works it cites.
Mathematical Programming pp. 1–26 (2023)
Josz, C., Lai, L.: Global stability of first-order methods for coercive tame functions · 2023
Later among the works it cites.
Mathematical Programming pp. 1–10 (2023)
Josz, C., Lai, L.: Lyapunov stability of the subgradient method with constant step size · 2023
Later among the works it cites.
Advances in Neural Information Processing Systems 36
Khaled, A., Mishchenko, K., Jin, C.: DoWG unleashed: An efficient universal parameter-free gradient descent method · 2023
Later among the works it cites.
arXiv preprint arXiv:2306.06101 (2023)
Mishchenko, K., Defazio, A.: Prodigy: An expeditiously adaptive parameter-free learner · 2023
Later among the works it cites.
Set-Valued and Variational Analysis 31
Pauwels, E.: Conservative parametric optimality and the ridge method for tame min-max problems · 2023
Later among the works it cites.
IEEE Journal of Biomedical and Health Informatics (2023)
Wang, J., Zhou, C., Chen, S., Hu, J., Wu, M., Jiang, X., Xu, C., Qian, D.: Chromosome detection in metaphase cell images using morphological priors · 2023
Later among the works it cites.
arXiv preprint arXiv:2307.10053 (2023)
Xiao, N., Hu, X., Toh, K.C.: Convergence guarantees for stochastic subgradient methods in nonsmooth nonconvex optimization · 2023
Later among the works it cites.
Computers in Biology and Medicine 152
Zhou, B., Zhou, H., Zhang, X., Xu, X., Chai, Y., Zheng, Z., Kot, A.C., Zhou, Z.: Tempo: A transformer-based mutation prediction framework for sars-cov-2 evolution · 2023
Later among the works it cites.
arXiv preprint arXiv:2402.07793 (2024)
Khaled, A., Jin, C.: Tuning-free stochastic optimization · 2024
Closest in time.
arXiv preprint arXiv:2404.00666 (2024)
Kreisler, I., Ivgi, M., Hinder, O., Carmon, Y.: Accelerated parameter-free stochastic optimization · 2024
Closest in time.
Journal of Optimization Theory and Applications 201
Le, T.: Nonsmooth nonconvex stochastic heavy ball · 2024
Closest in time.
Journal of Machine Learning Research 25
Xiao, N., Hu, X., Liu, X., Toh, K.C.: Adam-family methods for nonsmooth optimization with convergence guarantees · 2024
Closest in time.