Fetching the paper…
Reading the bibliography…
Modern machine learning paradigms, such as deep learning, occur in or close to the interpolation regime, wherein the number of model parameters is much larger than the number of data samples.
“A note on analytic functions in the unit circle”
Raymond Paley and Antoni Zygmund · 1932
Earlier work this paper cites.
“Near-Optimal Methods for Minimizing Star-Convex Functions and Beyond”
Oliver Hinder, Aaron Sidford and Nimit Sohoni · 1938
Earlier work this paper cites.
“A topological property of real analytic subsets”
Stanislaw Lojasiewicz · 1963
Earlier work this paper cites.
“Gradient methods for minimizing functionals”
B.. Poljak · 1963
Earlier work this paper cites.
“Problem complexity and method efficiency in optimization”
Arkadijč Nemirovskij and David Yudin · 1983
Earlier work this paper cites.
“Error bounds and convergence analysis of feasible descent methods: a general approach”
Z.-Q. Luo and P. Tseng · 1993
Earlier work this paper cites.
“Metric regularity and subdifferential calculus”
Aleksandr Ioffe · 2000
Earlier work this paper cites.
“A nonlinear programming algorithm for solving semidefinite programs via low-rank factorization”
Samuel Burer and Renato Monteiro · 2003
Earlier work this paper cites.
“Convergence of descent methods for semi-algebraic and tame problems: proximal algorithms, forward–backward splitting, and regularized Gauss–Seidel methods”
Hedy Attouch, Jérôme Bolte and Benar Svaiter · 2013
Earlier work this paper cites.
“Curves of descent”
Dmitriy Drusvyatskiy, Alexander Ioffe and Adrian Lewis · 2015
Earlier work this paper cites.
“Learning without concentration”
Shahar Mendelson · 2015
Earlier work this paper cites.
“Geometric median and robust estimation in Banach spaces”
STANISLAV MINSKER · 2015
Earlier work this paper cites.
“Gradient descent learns linear dynamical systems”
Moritz Hardt, Tengyu Ma and Benjamin Recht · 2016
Earlier work this paper cites.
“Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition”
Hamed Karimi, Julie Nutini and Mark Schmidt · 2016
Cited alongside, same era.
“Error Bounds, Quadratic Growth, and Linear Convergence of Proximal Methods”
Dmitriy Drusvyatskiy and Adrian. Lewis · 2017
Cited alongside, same era.
“Three factors influencing minima in sgd”
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio and Amos Storkey · 2017
Cited alongside, same era.
“On exponential convergence of sgd in non-convex over-parametrized learning”
Raef Bassily, Mikhail Belkin and Siyuan Ma · 2018
Cited alongside, same era.
“Algorithmic regularization in learning deep homogeneous models: Layers are automatically balanced”
Simon Du, Wei Hu and Jason Lee · 2018
Cited alongside, same era.
“Efficientnet: Rethinking model scaling for convolutional neural networks”
Mingxing Tan and Quoc Le · 2019
Later among the works it cites.
“Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron”
Sharan Vaswani, Francis Bach and Mark Schmidt · 2019
Later among the works it cites.
“On the convergence of first order methods for quasar-convex optimization”
Jikai Jin · 2020
Later among the works it cites.
“Better theory for SGD in the nonconvex world”
Ahmed Khaled and Peter Richtárik · 2020
Later among the works it cites.
“Big transfer (bit): General visual representation learning”
Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly and Neil Houlsby · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Gradient Descent Provably Optimizes Over-parameterized Neural Networks”
Simon Du, Xiyu Zhai, Barnabas Poczos and Aarti Singh · 2018
Cited alongside, same era.
“Neural tangent kernel: Convergence and generalization in neural networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Cited alongside, same era.
“Reconciling modern machine-learning practice and the classical bias–variance trade-off”
Mikhail Belkin, Daniel Hsu, Siyuan Ma and Soumik Mandal · 2019
Cited alongside, same era.
“Gradient Descent Finds Global Minima of Deep Neural Networks”
Simon Du, Jason Lee, Haochuan Li, Liwei Wang and Xiyu Zhai · 2019
Cited alongside, same era.
“Gpipe: Efficient training of giant neural networks using pipeline parallelism”
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc Le and Yonghui Wu · 2019
Cited alongside, same era.
“Linear convergence of first order methods for non-strongly convex optimization”
Ion Necoara, Yu Nesterov and Francois Glineur · 2019
Cited alongside, same era.
“Overparameterized nonlinear learning: Gradient descent takes the shortest path?”
Samet Oymak and Mahdi Soltanolkotabi · 2019
Cited alongside, same era.
“On the linearity of large non-linear models: when and why the tangent kernel is constant”
Chaoyue Liu, Libin Zhu and Misha Belkin · 2020
Later among the works it cites.
“Gradient descent on neural networks typically occurs at the edge of stability”
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J Kolter and Ameet Talwalkar · 2021
Later among the works it cites.
“Sgd for structured nonconvex functions: Learning rates, minibatching and interpolation”
Robert Gower, Othmane Sebbouh and Nicolas Loizou · 2021
Later among the works it cites.
“Understanding deep learning (still) requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2021
Later among the works it cites.
“Loss landscapes and optimization in over-parameterized non-linear systems and neural networks”
Chaoyue Liu, Libin Zhu and Mikhail Belkin · 2022
Later among the works it cites.
“Accelerated Stochastic Optimization Methods under Quasar-convexity”
Qiang Fu, Dongchu Xu and Ashia Wilson · 2023
Closest in time.
“Handbook of convergence theorems for (stochastic) gradient methods”
Guillaume Garrigos and Robert Gower · 2023
Closest in time.