Fetching the paper…
Reading the bibliography…
Recent empirical advances show that training deep models with large learning rate often improves generalization performance.
Yuanzhi Li, Colin Wei, and Tengyu Ma · 1907
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Introduction to optimization. optimization software
Boris T Polyak · 1987
Earlier work this paper cites.
Neural networks and principal component analysis: Learning from examples without local minima
Pierre Baldi and Kurt Hornik · 1989
Earlier work this paper cites.
Chaos: an introduction to dynamical systems
Kathleen T Alligood, Tim D Sauer, and James A Yorke · 1996
Earlier work this paper cites.
The landscape of matrix factorization revisited
Hossein Valavi, Sulin Liu, and Peter J Ramadge · 2002
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2003
Earlier work this paper cites.
Convex optimization
Stephen Boyd, Stephen P Boyd, and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Geometric numerical integration
Ernst Hairer, Marlis Hochbruck, Arieh Iserles, and Christian Lubich · 2006
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Tensor decompositions and applications
Tamara G Kolda and Brett W Bader · 2009
Earlier work this paper cites.
Matrix completion from a few entries
Raghunandan H Keshavan, Andrea Montanari, and Sewoong Oh · 2010
Earlier work this paper cites.
On optimization methods for deep learning
Quoc V Le, Jiquan Ngiam, Adam Coates, Ahbik Lahiri, Bobby Prochnow, and Andrew Y Ng · 2011
Earlier work this paper cites.
Unifying nuclear norm and bilinear factorization approaches for low-rank matrix decomposition
Ricardo Cabral, Fernando De la Torre, João P Costeira, and Alexandre Bernardino · 2013
Earlier work this paper cites.
Geometric measure theory
Herbert Federer · 2014
Earlier work this paper cites.
Understanding alternating minimization for matrix completion
Moritz Hardt · 2014
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2014
Earlier work this paper cites.
Phase retrieval via wirtinger flow: Theory and algorithms
Emmanuel J Candes, Xiaodong Li, and Mahdi Soltanolkotabi · 2015
Cited alongside, same era.
Yudong Chen and Martin J Wainwright · 2015
Cited alongside, same era.
Dropping convexity for faster semi-definite optimization
Srinadh Bhojanapalli, Anastasios Kyrillidis, and Sujay Sanghavi · 2016
Cited alongside, same era.
Low-rank solutions of linear matrix equations via procrustes flow
Stephen Tu, Ross Boczar, Max Simchowitz, Mahdi Soltanolkotabi, and Ben Recht · 2016
Cited alongside, same era.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Cited alongside, same era.
An exponential learning rate schedule for deep learning
Zhiyuan Li and Sanjeev Arora · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Super-convergence: Very fast training of neural networks using large learning rates
Leslie N Smith and Nicholay Topin · 2019
Later among the works it cites.
Toward understanding the importance of noise in training neural networks
Mo Zhou, Tianyi Liu, Yan Li, Dachao Lin, Enlu Zhou, and Tuo Zhao · 2019
Later among the works it cites.
Stochasticity of deterministic gradient descent: Large learning rate for multiscale objective function
Lingkai Kong and Molei Tao · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2017
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
On landscape of lagrangian functions and stochastic search for constrained nonconvex optimization
Zhehui Chen, Xingguo Li, Lin F Yang, Jarvis Haupt, and Tuo Zhao · 2018
Cited alongside, same era.
Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced
Simon S Du, Wei Hu, and Jason D Lee · 2018
Cited alongside, same era.
A closer look at deep learning heuristics: Learning rate restarts, warmup and distillation
Akhilesh Gotmare, Nitish Shirish Keskar, Caiming Xiong, and Richard Socher · 2018
Cited alongside, same era.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nathan Srebro · 2018
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Implicit bias of gradient descent based adversarial training on separable data
Yan Li, Ethan X.Fang, Huan Xu, and Tuo Zhao · 2020
Later among the works it cites.
Learning rate annealing can provably help generalization, even for convex problems
Preetum Nakkiran · 2020
Later among the works it cites.
A comparison of optimization algorithms for deep learning
Derya Soydaner · 2020
Later among the works it cites.
Salr: Sharpness-aware learning rates for improved generalization
Xubo Yue, Maher Nouiehed, and Raed Al Kontar · 2020
Later among the works it cites.
Towards theoretically understanding why sgd generalizes better than adam in deep learning
Pan Zhou, Jiashi Feng, Chao Ma, Caiming Xiong, Steven Hoi, et al · 2020
Later among the works it cites.
Implicit regularization of bregman proximal point algorithm and mirror descent on separable data
Yan Li, Caleb Ju, Ethan X Fang, and Tuo Zhao · 2021
Closest in time.
Noisy gradient descent converges to flat minima for nonconvex matrix factorization
Tianyi Liu, Yan Li, Song Wei, Enlu Zhou, and Tuo Zhao · 2021
Closest in time.
Beyond procrustes: Balancing-free gradient descent for asymmetric low-rank matrix sensing
Cong Ma, Yuanxin Li, and Yuejie Chi · 2021
Closest in time.
Global convergence of gradient descent for asymmetric low-rank matrix factorization
Tian Ye and Simon S Du · 2021
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Closest in time.
Noise regularizes over-parameterized rank one matrix recovery, provably
Tianyi Liu, Yan Li, Enlu Zhou, and Tuo Zhao · 2022
Closest in time.