Fetching the paper…
Reading the bibliography…
We study the problem of training a two-layer neural network (NN) of arbitrary width using stochastic gradient descent (SGD) where the input $\boldsymbol{x}\in \mathbb{R}^d$ is Gaussian and the target $y \in \mathbb{R}$ follows a multiple-index model, i.e., $y=g(\langle\boldsymbol{u_1},\boldsymbol{x}\rangle,...,\langle\boldsymbol{u_k},\boldsymbol{x}\rangle)$ with a noisy link function $g$.
Ker-Chau Li and Naihua Duan, Regression Analysis Under Link Violation , The Annals of Statistics (1989)
1989
Earlier work this paper cites.
Ker-Chau Li, Sliced inverse regression for dimension reduction , Journal of the American Statistical Association (1991)
1991
Earlier work this paper cites.
Boris T Polyak and Anatoli B Juditsky, Acceleration of stochastic approximation by averaging , SIAM journal on control and optimization 30
1992
Earlier work this paper cites.
Larry Goldstein and Gesine Reinert, Stein’s method and the zero bias transformation with application to simple random sampling , The Annals of Applied Probability 7
1997
Earlier work this paper cites.
Olivier Bousquet and André Elisseeff, Stability and generalization , Journal of Machine Learning Research (2002)
2002
Earlier work this paper cites.
Ali Rahimi and Benjamin Recht, Random features for large-scale kernel machines , Advances in Neural Information Processing Systems, 2007
2007
Earlier work this paper cites.
Lawrence C. Evans, Partial differential equations , Graduate Studies in Mathematics, American Mathematical Society, 2010
2010
Earlier work this paper cites.
2014
Earlier work this paper cites.
Murat A Erdogdu, Newton-stein method: a second order method for glms via stein’s lemma , Proceedings of Advances in Neural Information Processing Systems, 2015, pp. 1216–1224
2015
Earlier work this paper cites.
Murat A Erdogdu, Lee H Dicker, and Mohsen Bayati, Scaled least squares estimator for glms in large-scale problems , Advances in Neural Information Processing Systems 29
2016
Earlier work this paper cites.
Moritz Hardt, Ben Recht, and Yoram Singer, Train faster, generalize better: Stability of stochastic gradient descent , International Conference on Machine Learning, 2016
2016
Earlier work this paper cites.
Alon Brutzkus and Amir Globerson, Globally Optimal Gradient Descent for a ConvNet with Gaussian Inputs , International Conference on Machine Learning, 2017
2017
Earlier work this paper cites.
Yuanzhi Li and Yang Yuan, Convergence Analysis of Two-layer Neural Networks with ReLU Activation , Advances in Neural Information Processing Systems, 2017
2017
Earlier work this paper cites.
Kai Zhong, Zhao Song, Prateek Jain, Peter L. Bartlett, and Inderjit S. Dhillon, Recovery Guarantees for One-hidden-layer Neural Networks , International Conference on Machine Learning, 2017
2017
Earlier work this paper cites.
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang, Stronger Generalization Bounds for Deep Nets via a Compression Approach , International Conference on Machine Learning, 2018
2018
Earlier work this paper cites.
Lenaic Chizat and Francis Bach, On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport , Advances in Neural Information Processing Systems, 2018
2018
Earlier work this paper cites.
Vitaly Feldman and Jan Vondrak, Generalization bounds for uniformly stable algorithms , Advances in Neural Information Processing Systems 31
2018
Earlier work this paper cites.
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro, Implicit Bias of Gradient Descent on Linear Convolutional Networks , Advances in Neural Information Processing Systems, 2018
2018
Earlier work this paper cites.
Arthur Jacot, Franck Gabriel, and Clement Hongler, Neural Tangent Kernel: Convergence and Generalization in Neural Networks , Advances in Neural Information Processing Systems, 2018
2018
Earlier work this paper cites.
Cosme Louart, Zhenyu Liao, and Romain Couillet, A random matrix approach to neural networks , The Annals of Applied Probability (2018)
2018
Earlier work this paper cites.
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang, Algorithmic Regularization in Over-parameterized Matrix Sensing and Neural Networks with Quadratic Activations , Conference on Learning Theory, 2018
2018
Earlier work this paper cites.
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar, Foundations of machine learning , MIT press, 2018
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro, The Implicit Bias of Gradient Descent on Separable Data , Journal of Machine Learning Research (2018)
2018
Earlier work this paper cites.
Roman Vershynin, High-dimensional probability: An introduction with applications in data science , Cambridge University Press, 2018
2018
Earlier work this paper cites.
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang, Fine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks , International Conference on Machine Learning, 2019
2019
Earlier work this paper cites.
Zeyuan Allen-Zhu and Yuanzhi Li, What Can ResNet Learn Efficiently, Going Beyond Kernels? , Advances in Neural Information Processing Systems, 2019
2019
Earlier work this paper cites.
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song, A Convergence Theory for Deep Learning via Over-Parameterization , International Conference on Machine Learning, 2019
2019
Earlier work this paper cites.
Benedikt Bauer and Michael Kohler, On deep learning as a remedy for the curse of dimensionality in nonparametric regression , The Annals of Statistics (2019)
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
Lenaic Chizat, Edouard Oyallon, and Francis Bach, On Lazy Training in Differentiable Programming , Advances in Neural Information Processing Systems, 2019
2019
Cited alongside, same era.
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh, Gradient Descent Provably Optimizes Over-parameterized Neural Networks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Murat A. Erdogdu, Mohsen Bayati, and Lee H. Dicker, Scalable Approximations to Generalized Linear Problems , Journal of Machine Learning Research (2019)
2019
Cited alongside, same era.
Sebastian Goldt, Madhu Advani, Andrew M Saxe, Florent Krzakala, and Lenka Zdeborová, Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup , Advances in Neural Information Processing Systems, 2019
2019
Cited alongside, same era.
Alexander Camuto, George Deligiannidis, Murat A Erdogdu, Mert Gurbuzbalaban, Umut Simsekli, and Lingjiong Zhu, Fractal structure and generalization properties of stochastic optimization algorithms , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Konstantin Donhauser, Mingqi Wu, and Fanny Yang, How rotational invariance of common kernels prevents generalization in high dimensions , International Conference on Machine Learning, 2021
2021
Later among the works it cites.
Tyler Farghly and Patrick Rebeschini, Time-independent generalization bounds for SGLD in non-convex settings , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Tianyang Hu, Wenjia Wang, Cong Lin, and Guang Cheng, Regularization matters: A nonparametric perspective on overparametrized neural network , International Conference on Artificial Intelligence and Statistics, 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gauthier Gidel, Francis Bach, and Simon Lacoste-Julien, Implicit regularization of discrete gradient dynamics in linear neural networks , Advances in Neural Information Processing Systems 32
2019
Cited alongside, same era.
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik, SGD: General analysis and improved rates , International Conference on Machine Learning, 2019
2019
Cited alongside, same era.
B. Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Limitations of Lazy Training of Two-layers Neural Networks , Advances in Neural Information Processing Systems, 2019
2019
Cited alongside, same era.
Larry Goldstein and Xiaohan Wei, Non-gaussian observations in nonlinear compressed sensing via stein discrepancies , Information and Inference: A Journal of the IMA 8
2019
Cited alongside, same era.
Ziwei Ji and Matus Telgarsky, The implicit bias of gradient descent on nonseparable data , Conference on Learning Theory, 2019
2019
Cited alongside, same era.
Song Mei, Theodor Misiakiewicz, and Andrea Montanari, Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit , Conference on Learning Theory, 2019
2019
Cited alongside, same era.
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro, The role of over-parametrization in generalization of neural networks , International Conference on Learning Representations, 2019
2019
Cited alongside, same era.
Martin J. Wainwright, High-dimensional statistics: A non-asymptotic viewpoint , Cambridge University Press, 2019
2019
Cited alongside, same era.
Ilja Kuzborskij and Csaba Szepesvári, Nonparametric regression with shallow overparameterized neural networks trained by gd with early stopping , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Eran Malach, Pritish Kamath, Emmanuel Abbe, and Nathan Srebro, Quantifying the benefit of using differentiable learning over tangent kernels , International Conference on Machine Learning, 2021
2021
Later among the works it cites.
Gergely Neu, Gintare Karolina Dziugaite, Mahdi Haghifam, and Daniel M. Roy, Information-theoretic generalization bounds for stochastic gradient descent , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Scott Pesme, Loucas Pillaud-Vivien, and Nicolas Flammarion, Implicit Bias of SGD for Diagonal Linear Networks: a Provable Benefit of Stochasticity , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová, Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed , International Conference on Machine Learning, 2021
2021
Later among the works it cites.
Itay M Safran, Gilad Yehudai, and Ohad Shamir, The Effects of Mild Over-parameterization on the Optimization Landscape of Shallow ReLU Neural Networks , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Tingran Wang, Sam Buchanan, Dar Gilboa, and John Wright, Deep networks provably classify data on curves , Advances in Neural Information Processing Systems 34
2021
Later among the works it cites.
Lu Yu, Krishna Balasubramanian, Stanislav Volgushev, and Murat A Erdogdu, An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and Bias , Advances in Neural Information Processing Systems, 2021
2021
Later among the works it cites.
Chulhee Yun, Shankar Krishnan, and Hossein Mobahi, A Unifying View on Implicit Bias in Training Linear Neural Networks , International Conference on Learning Representations, 2021
2021
Later among the works it cites.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, Understanding deep learning (still) requires rethinking generalization , Communications of the ACM (2021)
2021
Later among the works it cites.
Mo Zhou, Rong Ge, and Chi Jin, A Local Convergence Theory for Mildly Over-Parameterized Two-Layer Neural Network , Conference on Learning Theory, 2021
2021
Later among the works it cites.
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz, The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks , Conference on Learning Theory, 2022
2022
Closest in time.
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song, Learning single-index models with shallow neural networks , Advances in Neural Information Processing Systems, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi, Neural Networks can Learn Representations with Gradient Descent , Conference on Learning Theory, 2022
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
Song Mei and Andrea Montanari, The generalization error of random features regression: Precise asymptotics and the double descent curve , Communications on Pure and Applied Mathematics (2022)
2022
Closest in time.
Atsushi Nitanda, Denny Wu, and Taiji Suzuki, Convex analysis of the mean field langevin dynamics , International Conference on Artificial Intelligence and Statistics, PMLR, 2022, pp. 9741–9757
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
2022
Closest in time.
John Wright and Yi Ma, High-dimensional data analysis with low-dimensional models: Principles, computation, and applications , Cambridge University Press, 2022
2022
Closest in time.