Fetching the paper…
Reading the bibliography…
Wide neural networks are biased towards learning certain functions, influencing both the rate of convergence of gradient descent (GD) and the functions that are reachable with GD in finite training time.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 1904
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen J Wright · 1999
Earlier work this paper cites.
A generalized representer theorem
Bernhard Schölkopf, Ralf Herbrich, and Alex J Smola · 2001
Earlier work this paper cites.
Spectral properties of the kernel matrix and their relation to kernel methods in machine learning
Mikio Ludwig Braun · 2005
Earlier work this paper cites.
Beyond the point cloud: from transductive to semi-supervised learning
Vikas Sindhwani, Partha Niyogi, and Mikhail Belkin · 2005
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence Saul · 2009
Earlier work this paper cites.
On learning with integral operators
Lorenzo Rosasco, Mikhail Belkin, and Ernesto De Vito · 2010
Earlier work this paper cites.
Bayesian learning for neural networks , volume 118
Radford M Neal · 2012
Earlier work this paper cites.
Mercer’s theorem on general domains: On the interaction between measures, kernels, and rkhss
Ingo Steinwart and Clint Scovel · 2012
Earlier work this paper cites.
Classification and construction of closed-form kernels for signal representation on the 2-sphere
Rodney A Kennedy, Parastoo Sadeghi, Zubair Khalid, and Jason D McEwen · 2013
Earlier work this paper cites.
Preconditioned spectral descent for deep learning
David E Carlson, Edo Collins, Ya-Ping Hsieh, Lawrence Carin, and Volkan Cevher · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Theory of reproducing kernels and applications
Saburou Saitoh and Yoshihiro Sawano · 2016
Earlier work this paper cites.
Breaking the curse of dimensionality with convex neural networks
Francis Bach · 2017
Earlier work this paper cites.
Large-scale data-dependent kernel approximation
Catalin Ionescu, Alin Popa, and Cristian Sminchisescu · 2017
Earlier work this paper cites.
Diving into the shallows: a computational perspective on large-scale shallow learning
Siyuan Ma and Mikhail Belkin · 2017
Earlier work this paper cites.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Cited alongside, same era.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Spectral bias in practice: The role of function frequency in generalization
Sara Fridovich-Keil, Raphael Gontijo-Lopes, and Rebecca Roelofs · 2021
Later among the works it cites.
Learning with convolution and pooling operations in kernel methods
Theodor Misiakiewicz and Song Mei · 2021
Later among the works it cites.
Tight bounds on the smallest eigenvalue of the neural tangent kernel for deep relu networks
Quynh Nguyen, Marco Mondelli, and Guido F Montufar · 2021
Later among the works it cites.
Zhichao Wang and Yizhe Zhu · 2021
Later among the works it cites.
A kernel perspective of skip connections in convolutional networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 2019
Cited alongside, same era.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Cited alongside, same era.
Frequency principle: Fourier analysis sheds light on deep neural networks
Zhi-Qin John Xu, Yaoyu Zhang, Tao Luo, Yanyang Xiao, and Zheng Ma · 2019
Cited alongside, same era.
Greg Yang · 2019
Cited alongside, same era.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Cited alongside, same era.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2019
Cited alongside, same era.
Daniel Barzilai, Amnon Geifman, Meirav Galun, and Ronen Basri · 2022
Later among the works it cites.
Spectral bias outside the training set for deep networks in the kernel regime
Benjamin Bowman and Guido Montufar · 2022
Later among the works it cites.
How wide convolutional neural networks learn hierarchical tasks
Francesco Cagnetta, Alessandro Favero, and Matthieu Wyart · 2022
Later among the works it cites.
On the spectral bias of convolutional neural tangent and gaussian process kernels
Amnon Geifman, Meirav Galun, David Jacobs, and Ronen Basri · 2022
Later among the works it cites.
Benign, tempered, or catastrophic: A taxonomy of overfitting
Neil Mallinar, James B Simon, Amirhesam Abedsoltan, Parthe Pandit, Mikhail Belkin, and Preetum Nakkiran · 2022
Later among the works it cites.
The interpolation phase transition in neural networks: Memorization and generalization under lazy training
Andrea Montanari and Yiqiao Zhong · 2022
Later among the works it cites.
Characterizing the spectrum of the ntk via a power series expansion
Michael Murray, Hui Jin, Benjamin Bowman, and Guido Montufar · 2022
Later among the works it cites.
Fast finite width neural tangent kernel
Roman Novak, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2022
Later among the works it cites.
On kernel regression with data-dependent kernels
James B Simon · 2022
Later among the works it cites.
Kernel-based smoothness analysis of residual networks
Tom Tirer, Joan Bruna, and Raja Giryes · 2022
Later among the works it cites.
When and why pinns fail to train: A neural tangent kernel perspective
Sifan Wang, Xinling Yu, and Paris Perdikaris · 2022
Later among the works it cites.
Eigenspace restructuring: a principle of space and frequency in neural networks
Lechao Xiao · 2022
Later among the works it cites.
Overview frequency principle/spectral bias in deep learning
Zhi-Qin John Xu, Yaoyu Zhang, and Tao Luo · 2022
Later among the works it cites.
Generalization in kernel regression under realistic assumptions
Daniel Barzilai and Ohad Shamir · 2023
Closest in time.
A fast, well-founded approximation to the empirical neural tangent kernel
Mohamad Amin Mohamadi, Wonho Bae, and Danica J Sutherland · 2023
Closest in time.
Benign overfitting in ridge regression
Alexander Tsigler and Peter L Bartlett · 2023
Closest in time.