Fetching the paper…
Reading the bibliography…
This work analyzes the solution trajectory of gradient-based algorithms via a novel basis function decomposition.
Introductory lectures on convex programming, 1998
Yu Nesterov · 1998
Earlier work this paper cites.
On regularization algorithms in learning theory
Frank Bauer, Sergei Pereverzev, and Lorenzo Rosasco · 2007
Earlier work this paper cites.
NIST handbook of mathematical functions hardback and CD-ROM
Frank WJ Olver, Daniel W Lozier, Ronald F Boisvert, and Charles W Clark · 2010
Earlier work this paper cites.
Tensor completion for estimating missing values in visual data
Ji Liu, Przemyslaw Musialski, Peter Wonka, and Jieping Ye · 2012
Earlier work this paper cites.
Early stopping and non-parametric regression: an optimal data-dependent stopping rule
Garvesh Raskutti, Martin J Wainwright, and Bin Yu · 2014
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Matrix completion has no spurious local minimum
Rong Ge, Jason D Lee, and Tengyu Ma · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Earlier work this paper cites.
Complete dictionary recovery over the sphere i: Overview and the geometric picture
Ju Sun, Qing Qu, and John Wright · 2016
Earlier work this paper cites.
Analyzing tensor power method dynamics in overcomplete regime
Animashree Anandkumar, Rong Ge, and Majid Janzamin · 2017
Earlier work this paper cites.
No spurious local minima in nonconvex low rank problems: A unified geometric analysis
Rong Ge, Chi Jin, and Yi Zheng · 2017
Earlier work this paper cites.
How to escape saddle points efficiently
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M Kakade, and Michael I Jordan · 2017
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2017
Earlier work this paper cites.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2018
Cited alongside, same era.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Cited alongside, same era.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2018
Cited alongside, same era.
Zhiyuan Li, Yuping Luo, and Kaifeng Lyu · 2020
Later among the works it cites.
Beyond lazy training for over-parameterized tensor decomposition
Xiang Wang, Chenwei Wu, Jason D Lee, Tengyu Ma, and Rong Ge · 2020
Later among the works it cites.
Understanding deflation process in over-parametrized tensor decomposition
Rong Ge, Yunwei Ren, Xiang Wang, and Mo Zhou · 2021
Later among the works it cites.
On the random conjugate kernel and neural tangent kernel
Zhengmian Hu and Heng Huang · 2021
Later among the works it cites.
Properties of the after kernel
Philip M Long · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cong Fang, Zhouchen Lin, and Tong Zhang · 2019
Cited alongside, same era.
First-order methods almost always avoid strict saddle points
Jason D Lee, Ioannis Panageas, Georgios Piliouras, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2019
Cited alongside, same era.
First-order methods almost always avoid saddle points: The case of vanishing step-sizes
Ioannis Panageas, Georgios Piliouras, and Xiao Wang · 2019
Cited alongside, same era.
Implicit regularization for optimal sparse recovery
Tomas Vaskevicius, Varun Kanade, and Patrick Rebeschini · 2019
Cited alongside, same era.
Gradient descent for deep matrix factorization: Dynamics and implicit bias towards low rank
Hung-Hsu Chou, Carsten Gieshoff, Johannes Maly, and Holger Rauhut · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Cited alongside, same era.
Spectra of the conjugate kernel and neural tangent kernel for linear-width neural networks
Zhou Fan and Zhichao Wang · 2020
Cited alongside, same era.
Noam Razin, Asaf Maman, and Nadav Cohen · 2021
Later among the works it cites.
Small random initialization is akin to spectral learning: Optimization and generalization guarantees for overparameterized low-rank matrix reconstruction
Dominik Stöger and Mahdi Soltanolkotabi · 2021
Later among the works it cites.
Global convergence of gradient descent for asymmetric low-rank matrix factorization
Tian Ye and Simon S Du · 2021
Later among the works it cites.
Preconditioned gradient descent for over-parameterized nonconvex matrix factorization
Jialun Zhang, Salar Fattahi, and Richard Y Zhang · 2021
Later among the works it cites.
Sharp global guarantees for nonconvex low-rank matrix recovery in the overparameterized regime
Richard Y Zhang · 2021
Later among the works it cites.
On the computational and statistical complexity of over-parameterized matrix sensing
Jiacheng Zhuo, Jeongyeol Kwon, Nhat Ho, and Constantine Caramanis · 2021
Later among the works it cites.
Implicit regularization in hierarchical tensor factorization and deep convolutional neural networks
Noam Razin, Asaf Maman, and Nadav Cohen · 2022
Closest in time.
Scaling and scalability: Provable nonconvex low-rank tensor estimation from incomplete measurements
Tian Tong, Cong Ma, Ashley Prater-Bennette, Erin Tripp, and Yuejie Chi · 2022
Closest in time.
Limitations of the ntk for understanding generalization in deep learning
Nikhil Vyas, Yamini Bansal, and Preetum Nakkiran · 2022
Closest in time.