Fetching the paper…
Reading the bibliography…
We provide quantitative bounds measuring the $L^2$ difference in function space between the trajectory of a finite-width network trained on finitely many samples from the idealized kernel dynamics of infinite width and infinite data.
A fine-grained spectral perspective on neural networks, 2019
Greg Yang and Hadi Salman · 1907
Earlier work this paper cites.
Geometric and probabilistic estimates for entropy and approximation numbers of operators
Y. Gordon, Hermann König, and Carsten Schütt · 1987
Earlier work this paper cites.
The Volume of Convex Bodies and Banach Space Geometry
Gilles Pisier · 1989
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Rademacher and Gaussian complexities: Risk bounds and structural results
Peter L. Bartlett and Shahar Mendelson · 2003
Earlier work this paper cites.
Covering an ellipsoid with equal balls
Ilya Dumer · 2006
Earlier work this paper cites.
Implicit bias of gradient descent for mean squared error regression with wide neural networks, 2021
Hui Jin and Guido Montúfar · 2006
Earlier work this paper cites.
Tensor programs II: Neural tangent kernel for any architecture, 2020
Greg Yang · 2006
Earlier work this paper cites.
Matplotlib: A 2d graphics environment
J. D. Hunter · 2007
Earlier work this paper cites.
IPython: a system for interactive scientific computing
Fernando Pérez and Brian E. Granger · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2010
Earlier work this paper cites.
On the exact computation of linear frequency principle dynamics and its generalization, 2020
Tao Luo, Zheng Ma, Zhi-Qin John Xu, and Yaoyu Zhang · 2010
Earlier work this paper cites.
On learning with integral operators
Lorenzo Rosasco, Mikhail Belkin, and Ernesto De Vito · 2010
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2012
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Jupyter notebooks – a publishing format for reproducible computational workflows
Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Jason Grout, Sylvain Corlay, Paul Ivanov, Damián Avila, Safia Abdalla, and Carol Willing · 2016
Earlier work this paper cites.
Implicit regularization in matrix factorization
Suriya Gunasekar, Blake E Woodworth, Srinadh Bhojanapalli, Behnam Neyshabur, and Nati Srebro · 2017
Earlier work this paper cites.
Geometry of optimization and implicit regularization in deep learning, 2017
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Earlier work this paper cites.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Earlier work this paper cites.
Implicit bias of gradient descent on linear convolutional networks
Suriya Gunasekar, Jason D Lee, Daniel Soudry, and Nati Srebro · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
Algorithmic regularization in over-parameterized matrix sensing and neural networks with quadratic activations
Yuanzhi Li, Tengyu Ma, and Hongyang Zhang · 2018
Cited alongside, same era.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Later among the works it cites.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
Ziwei Ji and Matus Telgarsky · 2020
Later among the works it cites.
Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks
Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak · 2020
Later among the works it cites.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Misha Belkin · 2020
Later among the works it cites.
Global convergence of deep networks with one wide layer followed by pyramidal topology
Quynh Nguyen and Marco Mondelli · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
The convergence rate of neural networks for learned functions of different frequencies
Ronen Basri, David Jacobs, Yoni Kasten, and Shira Kritchman · 2019
Cited alongside, same era.
The implicit bias of gradient descent on nonseparable data
Ziwei Ji and Matus Telgarsky · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Convergence of gradient descent on separable data
Mor Shpigel Nacson, J. Lee, Suriya Gunasekar, Nathan Srebro, and Daniel Soudry · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
On the spectral bias of neural networks
Nasim Rahaman, Aristide Baratin, Devansh Arpit, Felix Draxler, Min Lin, Fred Hamprecht, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Toward moderate overparameterization: Global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Generalization guarantees for neural networks via harnessing the low-rank structure of the Jacobian, 2020
Samet Oymak, Zalan Fabian, Mingchen Li, and Mahdi Soltanolkotabi · 2020
Later among the works it cites.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Later among the works it cites.
SciPy 1.0: Fundamental Algorithms for Scientific Computing in Python
Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, Stéfan J. van der Walt, Matthew Brett, Joshua Wilson, K. Jarrod Millman, Nikolay Mayorov, Andrew R. J. Nelson, Eric Jones, Robert Kern, Eric Larson, C J Carey, İlhan Polat, Yu Feng, Eric W. Moore, Jake VanderPlas, Denis Laxalde, Josef Perktold, Robert Cimrman, Ian Henriksen, E. A. Quintero, Charles R. Harris, Anne M. Archibald, Antônio H. Ribeiro, Fabian Pedregosa, Paul van Mulbregt, and SciPy 1.0 Contributors · 2020
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep ReLU networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Implicit regularization via neural feature alignment
Aristide Baratin, Thomas George, César Laurent, R Devon Hjelm, Guillaume Lajoie, Pascal Vincent, and Simon Lacoste-Julien · 2021
Later among the works it cites.
Deep networks and the multiple manifold problem
Sam Buchanan, Dar Gilboa, and John Wright · 2021
Later among the works it cites.
Towards understanding the spectral bias of deep learning
Yuan Cao, Zhiying Fang, Yue Wu, Ding-Xuan Zhou, and Quanquan Gu · 2021
Later among the works it cites.
Pathological Spectra of the Fisher Information Metric and Its Variants in Deep Neural Networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2021
Later among the works it cites.
On the proof of global convergence of gradient descent for deep relu networks with linear widths
Quynh Nguyen · 2021
Later among the works it cites.
Deep learning theory lecture notes
Matus Telgarsky · 2021
Later among the works it cites.
Explicit loss asymptotics in the gradient descent training of neural networks
Maksim Velikanov and Dmitry Yarotsky · 2021
Later among the works it cites.
Does the data induce capacity control in deep learning?
Rubing Yang, Jialin Mao, and Pratik Chaudhari · 2021
Later among the works it cites.
Neural networks as kernel learners: The silent alignment effect
Alexander Atanasov, Blake Bordelon, and Cengiz Pehlevan · 2022
Closest in time.
Implicit bias of MSE gradient optimization in underparameterized neural networks
Benjamin Bowman and Guido Montúfar · 2022
Closest in time.
Overcoming the spectral bias of neural value approximation
Ge Yang, Anurag Ajay, and Pulkit Agrawal · 2022
Closest in time.