Fetching the paper…
Reading the bibliography…
In this work, we consider the optimization process of minibatch stochastic gradient descent (SGD) on a 2-layer neural network with data separated by a quadratic ground truth function.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Earlier work this paper cites.
Gradient descent quantizes relu network features
Hartmut Maennel, Olivier Bousquet, and Sylvain Gelly · 2018
Earlier work this paper cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D Lee · 2019
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Gradient descent maximizes the margin of homogeneous neural networks
Kaifeng Lyu and Jian Li · 2019
Earlier work this paper cites.
Mean-field theory of two-layers neural networks: dimension-free bounds and kernel limit
Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
Yan Shuo Tan and Roman Vershynin · 2019
Earlier work this paper cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Colin Wei, Jason D Lee, Qiang Liu, and Tengyu Ma · 2019
Earlier work this paper cites.
On the universality of deep learning
Emmanuel Abbe and Colin Sandon · 2020
Earlier work this paper cites.
Backward feature correction: How deep learning performs deep learning
Zeyuan Allen-Zhu and Yuanzhi Li · 2020
Earlier work this paper cites.
Towards understanding hierarchical learning: Benefits of neural representations
Minshuo Chen, Yu Bai, Jason D Lee, Tuo Zhao, Huan Wang, Caiming Xiong, and Richard Socher · 2020
Earlier work this paper cites.
Learning polynomials in few relevant dimensions
Sitan Chen and Raghu Meka · 2020
Earlier work this paper cites.
Implicit bias of gradient descent for wide two-layer neural networks trained with the logistic loss
Lenaic Chizat and Francis Bach · 2020
Earlier work this paper cites.
Learning parities with neural networks
Amit Daniely and Eran Malach · 2020
Cited alongside, same era.
Learning over-parametrized two-layer relu neural networks beyond ntk
Yuanzhi Li, Tengyu Ma, and Hongyang R Zhang · 2020
Cited alongside, same era.
Benefit of deep learning with non-convex noisy gradient descent: Provable excess risk bound and superiority to kernel methods
Taiji Suzuki and Shunta Akiyama · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Online stochastic gradient descent on non-convex losses from high-dimensional inference
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2021
Cited alongside, same era.
High-dimensional limit theorems for sgd: Effective dynamics and critical scaling
Gerard Ben Arous, Reza Gheissari, and Aukosh Jagannath · 2022
Later among the works it cites.
Learning single-index models with shallow neural networks
Alberto Bietti, Joan Bruna, Clayton Sanford, and Min Jae Song · 2022
Later among the works it cites.
Neural networks can learn representations with gradient descent
Alexandru Damian, Jason Lee, and Mahdi Soltanolkotabi · 2022
Later among the works it cites.
Random feature amplification: Feature learning and generalization in neural networks
Spencer Frei, Niladri S Chatterji, and Peter L Bartlett · 2022
Later among the works it cites.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Linearized two-layer neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiaiewicz, and Andrea Montanari · 2021
Cited alongside, same era.
Dimension lower bounds for linear approaches to function approximation
Daniel Hsu · 2021
Cited alongside, same era.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2021
Cited alongside, same era.
Gradient descent on two-layer nets: Margin maximization and simplicity bias
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
Classifying high-dimensional gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Emmanuel Abbe, Enric Boix Adsera, and Theodor Misiakiewicz · 2022
Cited alongside, same era.
Hidden progress in deep learning: Sgd learns parities near the computational limit
Boaz Barak, Benjamin Edelman, Surbhi Goel, Sham Kakade, Eran Malach, and Cyril Zhang · 2022
Cited alongside, same era.
Later among the works it cites.
Eshaan Nichani, Yu Bai, and Jason D Lee · 2022
Later among the works it cites.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Matus Telgarsky · 2022
Later among the works it cites.
Sgd learning on neural networks: leap complexity and saddle-to-saddle dynamics
Emmanuel Abbe, Enric Boix-Adsera, and Theodor Misiakiewicz · 2023
Closest in time.
(s) gd over diagonal linear networks: Implicit regularisation, large stepsizes and edge of stability
Mathieu Even, Scott Pesme, Suriya Gunasekar, and Nicolas Flammarion · 2023
Closest in time.
Arvind Mahankali, Jeff Z Haochen, Kefan Dong, Margalit Glasgow, and Tengyu Ma · 2023
Closest in time.
Saddle-to-saddle dynamics in diagonal linear networks
Scott Pesme and Nicolas Flammarion · 2023
Closest in time.
Learning high-dimensional single-neuron relu networks with finite samples
Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman, Quanquan Gu, and Sham M Kakade · 2023
Closest in time.