Fetching the paper…
Reading the bibliography…
In this work, we provide a characterization of the feature-learning process in two-layer ReLU networks trained by gradient descent on the logistic loss following random initialization.
“Rademacher and Gaussian Complexities: Risk Bounds and Structural Results”
Peter Bartlett and Shahar Mendelson · 2003
Earlier work this paper cites.
“The Geometry of Logconcave Functions and Sampling Algorithms”
László Lovász and Santosh Vempala · 2007
Earlier work this paper cites.
“A Short Proof of Paouris’ Inequality”
Radoslaw Adamczak, Rafal Latala, Alexander. Litvak, Krzysztof Oleszkiewicz, Alain Pajor and Nicole Tomczak-Jaegermann · 2014
Earlier work this paper cites.
“Understanding machine learning: From theory to algorithms”
Shai Shalev-Shwartz and Shai Ben-David · 2014
Earlier work this paper cites.
“CS229T/STAT231: Statistical Learning Theory lecture notes (Fall 2017)”, https://web.archive.org/web/20200901203150/http://web.stanford.edu/class/cs229t/scribe_notes/10_17_final.pdf , 2017
Tengyu Ma · 2017
Earlier work this paper cites.
“Understanding deep learning requires rethinking generalization”
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht and Oriol Vinyals · 2017
Earlier work this paper cites.
“On the Global Convergence of Gradient Descent for Over-parameterized Models using Optimal Transport”
Lenaic Chizat and Francis Bach · 2018
Earlier work this paper cites.
“Neural Tangent Kernel: Convergence and Generalization in Neural Networks”
Arthur Jacot, Franck Gabriel and Clément Hongler · 2018
Earlier work this paper cites.
“A mean field view of the landscape of two-layers neural networks”
Song Mei, Andrea Montanari and Phan-Minh Nguyen · 2018
Earlier work this paper cites.
“What Can ResNet Learn Efficiently, Going Beyond Kernels?”
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Earlier work this paper cites.
“A Convergence Theory for Deep Learning via Over-Parameterization”
Zeyuan Allen-Zhu, Yuanzhi Li and Zhao Song · 2019
Earlier work this paper cites.
“On exact computation with an infinitely wide neural net”
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov and Ruosong Wang · 2019
Earlier work this paper cites.
“Gradient Descent Provably Optimizes Over-parameterized Neural Networks”
Simon Du, Xiyu Zhai, Barnabás Póczos and Aarti Singh · 2019
Earlier work this paper cites.
“Limitations of Lazy Training of Two-layers Neural Network”
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz and Andrea Montanari · 2019
Earlier work this paper cites.
“Gradient Descent with Early Stopping is Provably Robust to Label Noise for Overparameterized Neural Networks”
Mingchen Li, Mahdi Soltanolkotabi and Samet Oymak · 2019
Earlier work this paper cites.
“Theoretical Insights Into the Optimization Landscape of Over-Parameterized Shallow Neural Networks”
Mahdi Soltanolkotabi, Adel Javanmard and Jason. Lee · 2019
Earlier work this paper cites.
“High-dimensional statistics: A non-asymptotic viewpoint”
Martin Wainwright · 2019
Cited alongside, same era.
“Regularization Matters: Generalization and Optimization of Neural Nets vs their Induced Kernel”
Colin Wei, Jason Lee, Qiang Liu and Tengyu Ma · 2019
Cited alongside, same era.
“On the Power and Limitations of Random Features for Understanding Neural Networks”
Gilad Yehudai and Ohad Shamir · 2019
Cited alongside, same era.
“Gradient descent optimizes over-parameterized deep ReLU networks”
Difan Zou, Yuan Cao, Dongruo Zhou and Quanquan Gu · 2019
Cited alongside, same era.
“Beyond linearization: On quadratic and higher-order approximation of wide neural networks”
Yu Bai and Jason Lee · 2020
Cited alongside, same era.
“Benign Overfitting in Linear Regression”
Peter. Bartlett, Philip. Long, Gábor Lugosi and Alexander Tsigler · 2020
“On the Power of Differentiable Learning versus PAC and SQ Learning”
Emmanuel Abbe, Pritish Kamath, Eran Malach, Colin Sandon and Nathan Srebro · 2021
Later among the works it cites.
“Backward Feature Correction: How Deep Learning Performs Deep Learning”
Zeyuan Allen-Zhu and Yuanzhi Li · 2021
Later among the works it cites.
“Modeling from Features: A mean-field Framework for Over-parameterized Deep Neural Networks”
Cong Fang, Jason Lee, Pengkun Yang and Tong Zhang · 2021
Later among the works it cites.
“Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise”
Spencer Frei, Yuan Cao and Quanquan Gu · 2021
Later among the works it cites.
“Early-stopped neural networks are consistent”
Ziwei Ji, Justin. Li and Matus Telgarsky · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“A Generalized Neural Tangent Kernel Analysis for Two-layer Neural Networks”
Zixiang Chen, Yuan Cao, Quanquan Gu and Tong Zhang · 2020
Cited alongside, same era.
“Learning Parities with Neural Networks”
Amit Daniely and Eran Malach · 2020
Cited alongside, same era.
“Learning Halfspaces with Massart Noise Under Structured Distributions”
Ilias Diakonikolas, Vasilis Kontonis, Christos Tzamos and Nikos Zarifis · 2020
Cited alongside, same era.
“Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent Kernel”
Stanislav Fort, Gintare Dziugaite, Mansheej Paul, Sepideh Kharaghani, Daniel Roy and Surya Ganguli · 2020
Cited alongside, same era.
“Agnostic Learning of a Single Neuron with Gradient Descent”
Spencer Frei, Yuan Cao and Quanquan Gu · 2020
Cited alongside, same era.
“Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization Guarantee”
Wei Hu, Zhiyuan Li and Dingli Yu · 2020
Cited alongside, same era.
Philip Long · 2021
Later among the works it cites.
“Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity Bias”
Kaifeng Lyu, Zhiyuan Li, Runzhe Wang and Sanjeev Arora · 2021
Later among the works it cites.
“Quantifying the Benefit of Using Differentiable Learning over Tangent Kernels”
Eran Malach, Pritish Kamath, Emmanuel Abbe and Nathan Srebro · 2021
Later among the works it cites.
“The inductive bias of Re{LU} networks on orthogonally separable data”
Mary Phuong and Christoph Lampert · 2021
Later among the works it cites.
“Feature Learning in Infinite-Width Neural Networks”
Greg Yang and Edward Hu · 2021
Later among the works it cites.
“High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the Representation”
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Zhichao Wang, Denny Wu and Greg Yang · 2022
Closest in time.
“Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs”
Etienne Boursier, Loucas Pillaud-Vivien and Nicolas Flammarion · 2022
Closest in time.
“Neural Networks can Learn Representations with Gradient Descent”
Alex Damian, Jason. Lee and Mahdi Soltanolkotabi · 2022
Closest in time.
“Benign overfitting without linearity: Neural network classifiers trained by gradient descent for noisy linear data”
Spencer Frei, Niladri Chatterji and Peter Bartlett · 2022
Closest in time.
“Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data”
Spencer Frei, Gal Vardi, Peter Bartlett, Nathan Srebro and Wei Hu · 2023
Closest in time.