Fetching the paper…
Reading the bibliography…
Recent advancements in deep learning optimization have introduced new algorithms, such as Schedule-Free optimizers, AdEMAMix, MARS and Lion which modify traditional momentum mechanisms.
Laprop: Separating momentum and adaptivity in adam, 2021
Ziyin, L., Wang, Z. T., and Ueda, M · 2002
Earlier work this paper cites.
Defazio, A · 2010
Earlier work this paper cites.
Accelerating stochastic gradient descent for least squares regression
Jain, P., Kakade, S. M., Kidambi, R., Netrapalli, P., and Sidford, A · 2018
Earlier work this paper cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Kidambi, R., Netrapalli, P., Jain, P., and Kakade, S. M · 2018
Earlier work this paper cites.
Aggregated momentum: Stability through passive damping
Lucas, J., Sun, S., Zemel, R., and Grosse, R · 2019
Earlier work this paper cites.
Quasi-hyperbolic momentum and adam for deep learning
Ma, J. and Yarats, D · 2019
Cited alongside, same era.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M · 2019
Cited alongside, same era.
Accelerating sgd with momentum for over-parameterized learning
Liu, C. and Belkin, M · 2020
Cited alongside, same era.
Symbolic discovery of optimization algorithms
Chen, X., Liang, C., Huang, D., Real, E., Wang, K., Pham, H., Dong, X., Luong, T., Hsieh, C., Lu, Y., and Le, Q. V · 2023
Cited alongside, same era.
Achieving acceleration despite very noisy gradients, 2023
Gupta, K., Siegel, J., and Wojtowytsch, S · 2023
Cited alongside, same era.
Defazio, A., Yang, X., Mehta, H., Mishchenko, K., Khaled, A., and Cutkosky, A · 2024
Later among the works it cites.
The ademamix optimizer: Better, faster, older
Pagliardini, M., Ablin, P., and Grangier, D · 2024
Later among the works it cites.
Resolving discrepancies in compute-optimal scaling of language models
Porian, T., Wortsman, M., Jitsev, J., Schmidt, L., and Carmon, Y · 2024
Later among the works it cites.
Mars: Unleashing the power of variance reduction for training large models, 2024
Yuan, H., Liu, Y., Wu, S., Zhou, X., and Gu, Q · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…