Fetching the paper…
Reading the bibliography…
The momentum acceleration technique is widely adopted in many optimization algorithms.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
A method for solving a convex programming problem with convergence rate o (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Boosting: Foundations and algorithms
R. E. Schapire and Y. Freund · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Margins, shrinkage, and boosting
M. Telgarsky · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
A differential equation for modeling nesterov’s accelerated gradient method: theory and insights
W. Su, S. Boyd, and E. Candes · 2014
Earlier work this paper cites.
Global convergence of the heavy-ball method for convex optimization
E. Ghadimi, H. R. Feyzmahdavian, and M. Johansson · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
B. Neyshabur, R. Tomioka, and N. Srebro · 2015
Earlier work this paper cites.
Learning with incremental iterative regularization
L. Rosasco and S. Villa · 2015
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
N. S. Keskar, D. Mudigere, J. Nocedal, M. Smelyanskiy, and P. T. P. Tang · 2017
Earlier work this paper cites.
Optimal rates for multi-pass stochastic gradient methods
J. Lin and L. Rosasco · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Earlier work this paper cites.
Aggregated residual transformations for deep neural networks
S. Xie, R. Girshick, P. Dollár, Z. Tu, and K. He · 2017
Earlier work this paper cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2018
Cited alongside, same era.
S. De, A. Mukherjee, and E. Ullah · 2018
Cited alongside, same era.
Characterizing implicit bias in terms of optimization geometry
S. Gunasekar, J. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
Implicit bias of gradient descent on linear convolutional networks
S. Gunasekar, J. D. Lee, D. Soudry, and N. Srebro · 2018
Cited alongside, same era.
Gradient descent aligns the layers of deep linear networks
Z. Ji and M. Telgarsky · 2018
Cited alongside, same era.
Non-ergodic convergence analysis of heavy-ball algorithms
T. Sun, P. Yin, D. Li, C. Huang, L. Guan, and H. Jiang · 2019
Later among the works it cites.
Efficientnet: Rethinking model scaling for convolutional neural networks
M. Tan and Q. Le · 2019
Later among the works it cites.
On the linear speedup analysis of communication efficient momentum sgd for distributed non-convex optimization
H. Yu, R. Jin, and S. Yang · 2019
Later among the works it cites.
The implicit regularization of stochastic gradient flow for least squares
A. Ali, E. Dobriban, and R. Tibshirani · 2020
Later among the works it cites.
A simple convergence proof of adam and adagrad
A. Défossez, L. Bottou, F. Bach, and N. Usunier · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
D. Soudry, E. Hoffer, M. S. Nacson, S. Gunasekar, and N. Srebro · 2018
Cited alongside, same era.
A unified analysis of stochastic momentum methods for deep learning
Y. Yan, T. Yang, Z. Li, Q. Lin, and Y. Yang · 2018
Cited alongside, same era.
Fantastic generalization measures and where to find them
Y. Jiang, B. Neyshabur, H. Mobahi, D. Krishnan, and S. Bengio · 2019
Cited alongside, same era.
Gradient descent maximizes the margin of homogeneous neural networks
K. Lyu and J. Li · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
M. S. Nacson, J. Lee, S. Gunasekar, P. H. P. Savarese, N. Srebro, and D. Soudry · 2019
Cited alongside, same era.
Stochastic gradient descent on separable data: Exact convergence with a fixed learning rate
M. S. Nacson, N. Srebro, and D. Soudry · 2019
Cited alongside, same era.
Z. Ji and M. Telgarsky · 2020
Later among the works it cites.
An improved analysis of stochastic gradient descent with momentum
Y. Liu, Y. Gao, and W. Yin · 2020
Later among the works it cites.
A novel convergence analysis for algorithms of the adam family
Z. Guo, Y. Xu, W. Yin, R. Jin, and T. Yang · 2021
Closest in time.
Fast margin maximization via dual acceleration
Z. Ji, N. Srebro, and M. Telgarsky · 2021
Closest in time.
Characterizing the implicit bias via a primal-dual analysis
Z. Ji and M. Telgarsky · 2021
Closest in time.
Rmsprop converges with proper hyper-parameter
N. Shi, D. Li, M. Hong, and R. Sun · 2021
Closest in time.
The role of momentum parameters in the optimal convergence of adaptive polyak’s heavy-ball methods
W. Tao, S. Long, G. Wu, and Q. Tao · 2021
Closest in time.
The implicit bias for adaptive optimization algorithms on homogeneous neural networks
B. Wang, Q. Meng, W. Chen, and T.-Y. Liu · 2021
Closest in time.
A lyapunov analysis of accelerated methods in optimization
A. C. Wilson, B. Recht, and M. I. Jordan · 2021
Closest in time.