Fetching the paper…
Reading the bibliography…
In this work, we study optimization methods that leverage the linear minimization oracle (LMO) over a norm-ball.
Lower bounds for non-convex stochastic optimization, 2022
Arjevani, Y., Carmon, Y., Duchi, J. C., Foster, D. J., Srebro, N., and Woodworth, B · 1912
Earlier work this paper cites.
An algorithm for quadratic programming
Frank, M., Wolfe, P., et al · 1956
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
Trust region methods
Conn, A. R., Gould, N. I., and Toint, P. L · 2000
Earlier work this paper cites.
Numerical optimization, 2006
Wright, S. J · 2006
Earlier work this paper cites.
Optimal approximation for the submodular welfare problem in the value oracle model
Vondrák, J · 2008
Earlier work this paper cites.
Coresets, sparse greedy approximation, and the frank-wolfe algorithm
Clarkson, K. L · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization, 2010
McMahan, H. B. and Streeter, M · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
What is a Fenchel conjugate
Bauschke, H. and Lucet, Y · 2012
Earlier work this paper cites.
Projection-free online learning
Hazan, E. and Kale, S · 2012
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
Efficiency of coordinate descent methods on huge-scale optimization problems
Nesterov, Y · 2012
Earlier work this paper cites.
Revisiting Frank-Wolfe: Projection-free sparse convex optimization
Jaggi, M · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Saxe, A. M., McClelland, J. L., and Ganguli, S · 2013
Earlier work this paper cites.
An almost-linear-time algorithm for approximate max flow in undirected graphs, and its multicommodity generalizations
Kelner, J. A., Lee, Y. T., Orecchia, L., and Sidford, A · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P · 2014
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
Hazan, E., Levy, K., and Shalev-Shwartz, S · 2015
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K., Zhang, X., Ren, S., and Sun, J · 2015
Earlier work this paper cites.
Trust region policy optimization
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P · 2015
Earlier work this paper cites.
MM optimization algorithms
Lange, K · 2016
Cited alongside, same era.
Wasserstein generative adversarial networks
Arjovsky, M., Chintala, S., and Bottou, L · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Bartlett, P. L., Foster, D. J., and Telgarsky, M. J · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., and Usunier, N · 2017
Cited alongside, same era.
The duality structure gradient descent algorithm: analysis and applications to neural networks
Flynn, T · 2017
Cited alongside, same era.
A unified approach to adaptive regularization in online and stochastic optimization
Stochastic normalized gradient descent with momentum for large batch training
Zhao, S.-Y., Xie, Y.-P., and Li, W.-J · 2020
Later among the works it cites.
Training data-efficient image transformers & distillation through attention
Touvron, H., Cord, M., Douze, M., Massa, F., Sablayrolles, A., and Jégou, H · 2021
Later among the works it cites.
Tensor programs iv: Feature learning in infinite-width neural networks
Yang, G. and Hu, E. J · 2021
Later among the works it cites.
Learning pruning-friendly networks via Frank-Wolfe: One-shot, any-sparsity, and no retraining
Lu, M., Luo, X., Chen, T., Chen, W., Liu, D., and Wang, Z · 2022
Later among the works it cites.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gupta, V., Koren, T., and Singer, Y · 2017
Cited alongside, same era.
Preconditioned stochastic gradient descent
Li, X.-L · 2017
Cited alongside, same era.
Large batch training of convolutional networks
You, Y., Gitman, I., and Ginsburg, B · 2017
Cited alongside, same era.
Block-normalized gradient method: An empirical study for training deep neural network
Yu, A. W., Huang, L., Lin, Q., Salakhutdinov, R., and Carbonell, J · 2017
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Cited alongside, same era.
Learning with structured sparsity: From discrete to convex and back
El Halabi, M · 2018
Cited alongside, same era.
Spectral normalization for generative adversarial networks
Miyato, T., Kataoka, T., Koyama, M., and Yoshida, Y · 2018
Cited alongside, same era.
Later among the works it cites.
Lion secretly solves constrained optimization: As lyapunov predicts
Chen, L., Liu, B., Liang, K., and Liu, Q · 2023
Later among the works it cites.
Benchmarking neural network training algorithms
Dahl, G. E., Schneider, F., Nado, Z., Agarwal, N., Sastry, C. S., Hennig, P., Medapati, S., Eschenhagen, R., Kasimbeg, P., Suo, D., et al · 2023
Later among the works it cites.
Why do we need weight decay in modern deep learning?
D’Angelo, F., Andriushchenko, M., Varre, A., and Flammarion, N · 2023
Later among the works it cites.
Orabona, F · 2023
Later among the works it cites.
A spectral condition for feature learning
Yang, G., Simon, J. B., and Bernstein, J · 2023
Later among the works it cites.
Exact convergence rate of the last iterate in subgradient methods
Zamani, M. and Glineur, F · 2023
Later among the works it cites.
Defazio, A., Yang, X. A., Mehta, H., Mishchenko, K., Khaled, A., and Cutkosky, A · 2024
Later among the works it cites.
Cifar-10 airbench, 2024
Jordan, K · 2024
Later among the works it cites.
Analyzing and improving the training dynamics of diffusion models
Karras, T., Aittala, M., Lehtinen, J., Hellsten, J., Aila, T., and Laine, S · 2024
Later among the works it cites.
Scalable optimization in the modular norm
Large, T., Liu, Y., Huh, M., Bahng, H., Isola, P., and Bernstein, J · 2024
Later among the works it cites.
Curvature-informed sgd via general purpose lie-group preconditioners
Pooladzandi, O. and Li, X.-L · 2024
Later among the works it cites.
Implicit bias of AdamW: ℓ ∞ \ell_{\infty} norm constrained optimization
Xie, S. and Li, Z · 2024
Later among the works it cites.
Mars: Unleashing the power of variance reduction for training large models
Yuan, H., Liu, Y., Wu, S., Zhou, X., and Gu, Q · 2024
Later among the works it cites.
Muon is scalable for llm training
Liu, J., Su, J., Yao, X., Jiang, Z., Lai, G., Du, Y., Qin, Y., Xu, W., Lu, E., Yan, J., et al · 2025
Closest in time.
ν \nu sam: Memory-efficient sharpness-aware minimization via nuclear norm constraints
Pethick, T., Raman, P., Minorics, L., Hong, M., Sabach, S., and Cevher, V · 2025
Closest in time.