Fetching the paper…
Reading the bibliography…
Understanding the mechanisms through which neural networks extract statistics from input-label pairs through feature learning is one of the most important unsolved problems in supervised learning.
Structure adaptive approach for dimension reduction
Marian Hristache, Anatoli Juditsky, Jorg Polzehl, and Vladimir Spokoiny · 2001
Earlier work this paper cites.
Spectrum estimation for large dimensional covariance matrices using random matrix theory
Noureddine El Karoui · 2008
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
A consistent estimator of the expected gradient outerproduct
Shubhendu Trivedi, Jialei Wang, Samory Kpotufe, and Gregory Shakhnarovich · 2014
Earlier work this paper cites.
Layer-specific adaptive learning rates for deep networks
Bharat Singh, Soham De, Yangmuzi Zhang, Thomas Goldstein, and Gavin Taylor · 2015
Earlier work this paper cites.
Gradients weights improve regression and classification
Samory Kpotufe, Abdeslam Boularias, Thomas Schultz, and Kyoungok Kim · 2016
Earlier work this paper cites.
Free probability and random matrices , volume 35
James A Mingo and Roland Speicher · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Earlier work this paper cites.
Limitations of Lazy Training of Two-layers Neural Networks
B. Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2019
Earlier work this paper cites.
What Can ResNet Learn Efficiently, Going Beyond Kernels?
Zeyuan Allen-Zhu and Yuanzhi Li · 2019
Earlier work this paper cites.
On the Power and Limitations of Random Features for Understanding Neural Networks
Gilad Yehudai and Ohad Shamir · 2019
Earlier work this paper cites.
A random matrix perspective on mixtures of nonlinearities for deep learning
Ben Adlam, Jake Levinson, and Jeffrey Pennington · 2019
Earlier work this paper cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Earlier work this paper cites.
Generalization of two-layer neural networks: An asymptotic viewpoint
Jimmy Ba, Murat Erdogdu, Taiji Suzuki, Denny Wu, and Tianzong Zhang · 2019
Earlier work this paper cites.
A short note on concentration inequalities for random vectors with subgaussian norm
Chi Jin, Praneeth Netrapalli, Rong Ge, Sham M Kakade, and Michael I Jordan · 2019
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Cited alongside, same era.
Learning over-parametrized two-layer neural networks beyond NTK
Yuanzhi Li, Tengyu Ma, and Hongyang R Zhang · 2020
Cited alongside, same era.
Knowledge distillation in wide neural networks: Risk bound, data efficiency and imperfect teacher
Guangda Ji and Zhanxing Zhu · 2020
Cited alongside, same era.
When Do Neural Networks Outperform Kernel Methods?
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 2020
Cited alongside, same era.
Neural networks efficiently learn low-dimensional representations with sgd
Alireza Mousavi-Hosseini, Sejun Park, Manuela Girotti, Ioannis Mitliagkas, and Murat A Erdogdu · 2022
Later among the works it cites.
Theodor Misiakiewicz · 2022
Later among the works it cites.
Second-order regression models exhibit progressive sharpening to the edge of stability
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Later among the works it cites.
Quadratic models for understanding neural network dynamics
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, and Mikhail Belkin · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Greg Yang and Edward J Hu · 2020
Cited alongside, same era.
The neural tangent kernel in high dimensions: Triple descent and a multi-scale theory of generalization
Ben Adlam and Jeffrey Pennington · 2020
Cited alongside, same era.
Kernel and rich regimes in overparametrized models
Blake Woodworth, Suriya Gunasekar, Jason D Lee, Edward Moroshko, Pedro Savarese, Itay Golan, Daniel Soudry, and Nathan Srebro · 2020
Cited alongside, same era.
Temperature check: theory and practice for training models with softmax-cross-entropy losses
Atish Agarwala, Jeffrey Pennington, Yann Dauphin, and Sam Schoenholz · 2020
Cited alongside, same era.
Nerf: Representing scenes as neural radiance fields for view synthesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng · 2021
Cited alongside, same era.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Cited alongside, same era.
Classifying high-dimensional Gaussian mixtures: Where kernel methods fail and neural networks succeed
Maria Refinetti, Sebastian Goldt, Florent Krzakala, and Lenka Zdeborová · 2021
Cited alongside, same era.
Hong Hu and Yue M Lu · 2022
Later among the works it cites.
Provable guarantees for nonlinear feature learning in three-layer neural networks
Eshaan Nichani, Alex Damian, and Jason D Lee · 2023
Later among the works it cites.
A theory of non-linear feature learning with one gradient step in two-layer neural networks
Behrad Moniri, Donghwan Lee, Hamed Hassani, and Edgar Dobriban · 2023
Later among the works it cites.
Linear neural network layers promote learning single-and multiple-index models
Suzanna Parkinson, Greg Ongie, and Rebecca Willett · 2023
Later among the works it cites.
Mechanism of feature learning in convolutional neural networks
Daniel Beaglehole, Adityanarayanan Radhakrishnan, Parthe Pandit, and Mikhail Belkin · 2023
Later among the works it cites.
Efficient estimation of the central mean subspace via smoothed gradient outer products
Gan Yuan, Mingyue Xu, Samory Kpotufe, and Daniel Hsu · 2023
Later among the works it cites.
Dichotomy of early and late phase implicit biases can provably induce grokking
Kaifeng Lyu, Jikai Jin, Zhiyuan Li, Simon S Du, Jason D Lee, and Wei Hu · 2023
Later among the works it cites.
Hessian inertia in neural networks
Xuchan Bao, Alberto Bietti, Aaron Defazio, and Vivien Cabannes · 2023
Later among the works it cites.
High-dimensional sgd aligns with emerging outlier eigenspaces
Gerard Ben Arous, Reza Gheissari, Jiaoyang Huang, and Aukosh Jagannath · 2023
Later among the works it cites.
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, and Mikhail Belkin · 2023
Later among the works it cites.
The fast committor machine: Interpretable prediction with kernels
D Aristoff, M Johnson, G Simpson, and RJ Webber · 2024
Closest in time.
Average gradient outer product as a mechanism for deep neural collapse, 2024
Daniel Beaglehole, Peter Súkeník, Marco Mondelli, and Mikhail Belkin · 2024
Closest in time.
Nonlinear spiked covariance matrices and signal propagation in deep neural networks
Zhichao Wang, Denny Wu, and Zhou Fan · 2024
Closest in time.
Lénaïc Chizat and Praneeth Netrapalli · 2024
Closest in time.