Fetching the paper…
Reading the bibliography…
We perform a careful, thorough, and large scale empirical study of the correspondence between wide neural networks and kernel methods.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Greg Yang · 1902
Earlier work this paper cites.
Linearized two-layers neural networks in high dimension
Behrooz Ghorbani, Song Mei, Theodor Misiakiewicz, and Andrea Montanari · 1904
Earlier work this paper cites.
Randaugment: Practical data augmentation with no separate search
Ekin D Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V Le · 1909
Earlier work this paper cites.
Enhanced convolutional neural tangent kernels
Zhiyuan Li, Ruosong Wang, Dingli Yu, Simon S Du, Wei Hu, Ruslan Salakhutdinov, and Sanjeev Arora · 1911
Earlier work this paper cites.
What size net gives valid generalization?
Eric B. Baum and David Haussler · 1989
Earlier work this paper cites.
Generalization and network design strategies
Yann Lecun · 1989
Earlier work this paper cites.
Combining forecasts: A review and annotated bibliography
Robert T Clemen · 1989
Earlier work this paper cites.
On the ability of the optimal perceptron to generalise
M Opper, W Kinzel, J Kleinz, and R Nehl · 1990
Earlier work this paper cites.
Decision theoretic generalizations of the pac model for neural net and other learning applications
David Haussler · 1992
Earlier work this paper cites.
Priors for infinite networks (tech. rep. no. crg-tr-94-1)
Radford M. Neal · 1994
Earlier work this paper cites.
Probable networks and plausible predictions—a review of practical bayesian methods for supervised neural networks
David JC MacKay · 1995
Earlier work this paper cites.
Python reference manual
Guido Van Rossum and Fred L Drake Jr · 1995
Earlier work this paper cites.
A desicion-theoretic generalization of on-line learning and an application to boosting
Yoav Freund and Robert E Schapire · 1995
Earlier work this paper cites.
Bagging predictors
Leo Breiman · 1996
Earlier work this paper cites.
Generating accurate and diverse members of a neural-network ensemble
David W Opitz and Jude W Shavlik · 1996
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
The “independent components” of natural scenes are edge filters
Anthony J Bell and Terrence J Sejnowski · 1997
Earlier work this paper cites.
What size neural network gives optimal generalization? convergence properties of backpropagation
Steve Lawrence, C Lee Giles, and Ah Chung Tsoi · 1998
Earlier work this paper cites.
The sample complexity of pattern classification with neural networks: the size of the weights is more important than the size of the network
Peter L Bartlett · 1998
Earlier work this paper cites.
Statistical learning theory
Vladimir Vapnik · 1998
Earlier work this paper cites.
Popular ensemble methods: An empirical study
David Opitz and Richard Maclin · 1999
Earlier work this paper cites.
Ensemble methods in machine learning
Thomas G Dietterich · 2000
Earlier work this paper cites.
Using the nyström method to speed up kernel machines
Christopher KI Williams and Matthias Seeger · 2001
Earlier work this paper cites.
Random forests
Leo Breiman · 2001
Earlier work this paper cites.
Infinitely wide graph convolutional networks: semi-supervised learning via gaussian processes
Jilin Hu, Jianbing Shen, Bin Yang, and Ling Shao · 2002
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
Stability and generalization
Olivier Bousquet and André Elisseeff · 2002
Earlier work this paper cites.
Statistical learning: Stability is sufficient for generalization and necessary and sufficient for consistency of empirical risk minimization
Sayan Mukherjee, Partha Niyogi, Tomaso Poggio, and Ryan Rifkin · 2004
Earlier work this paper cites.
General conditions for predictivity in learning theory
Tomaso Poggio, Ryan Rifkin, Sayan Mukherjee, and Partha Niyogi · 2004
Earlier work this paper cites.
Gaussian processes for machine learning , volume 1
Carl Edward Rasmussen and Christopher KI Williams · 2006
Earlier work this paper cites.
Matplotlib: A 2D Graphics Environment
J. D. Hunter · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K Saul · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Data Structures for Statistical Computing in Python
Wes McKinney · 2010
Earlier work this paper cites.
Ensemble-based classifiers
Lior Rokach · 2010
Earlier work this paper cites.
Machine learning for the new york city power grid
Cynthia Rudin, David Waltz, Roger N Anderson, Albert Boulanger, Ansaf Salleb-Aouissi, Maggie Chow, Haimonti Dutta, Philip N Gross, Bert Huang, Steve Ierome, et al · 2011
Earlier work this paper cites.
The NumPy Array: A Structure for Efficient Numerical Computation
S. van der Walt, S. C. Colbert, and G. Varoquaux · 2011
Earlier work this paper cites.
Steps toward deep kernel methods from infinite neural networks
Tamir Hazan and Tommi Jaakkola · 2015
Earlier work this paper cites.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Machine learning methods for attack detection in the smart grid
Mete Ozay, Inaki Esnaola, Fatos Tunay Yarman Vural, Sanjeev R Kulkarni, and H Vincent Poor · 2015
Earlier work this paper cites.
Probabilistic backpropagation for scalable learning of bayesian neural networks
José Miguel Hernández-Lobato and Ryan Adams · 2015
Earlier work this paper cites.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
An analysis of deep neural network models for practical applications
Alfredo Canziani, Adam Paszke, and Eugenio Culurciello · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Equality of opportunity in supervised learning
Moritz Hardt, Eric Price, and Nati Srebro · 2016
Cited alongside, same era.
Big data’s disparate impact
Solon Barocas and Andrew D Selbst · 2016
Cited alongside, same era.
The role of a layer in deep neural networks: a gaussian process perspective
Oded Ben-David and Zohar Ringel · 2019
Later among the works it cites.
A fine-grained spectral perspective on neural networks
Greg Yang and Hadi Salman · 2019
Later among the works it cites.
Function space particle optimization for bayesian neural networks
Ziyu Wang, Tongzheng Ren, Jun Zhu, and Bo Zhang · 2019
Later among the works it cites.
A bayesian perspective on the deep image prior
Zezhou Cheng, Matheus Gadelha, Subhransu Maji, and Daniel Sheldon · 2019
Later among the works it cites.
Deeper connections between neural networks and gaussian processes speed-up active learning
Evgenii Tsymbalov, Sergei Makarychev, Alexander Shapeev, and Maxim Panov · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, et al · 2016
Cited alongside, same era.
Jupyter Notebooks – a publishing format for reproducible computational workflows
Thomas Kluyver, Benjamin Ragan-Kelley, Fernando Pérez, Brian Granger, Matthias Bussonnier, Jonathan Frederic, Kyle Kelley, Jessica Hamrick, Jason Grout, Sylvain Corlay, Paul Ivanov, Damián Avila, Safia Abdalla, and Carol Willing · 2016
Cited alongside, same era.
Dropout as a bayesian approximation: Representing model uncertainty in deep learning
Yarin Gal and Zoubin Ghahramani · 2016
Cited alongside, same era.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Cited alongside, same era.
SGD learns the conjugate kernel class of the network
Amit Daniely · 2017
Cited alongside, same era.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
On random deep weight-tied autoencoders: Exact asymptotic analysis, phase transitions, and implications to training
Ping Li and Phan-Minh Nguyen · 2019
Later among the works it cites.
A mean field theory of quantized deep networks: The quantization-depth trade-off
Yaniv Blumenfeld, Dar Gilboa, and Daniel Soudry · 2019
Later among the works it cites.
Mean-field behaviour of neural tangent kernel for deep neural networks, 2019
Soufiane Hayou, Arnaud Doucet, and Judith Rousseau · 2019
Later among the works it cites.
Disentangling feature and lazy learning in deep neural networks: an empirical study
Mario Geiger, Stefano Spigler, Arthur Jacot, and Matthieu Wyart · 2019
Later among the works it cites.
Finite size corrections for neural network gaussian processes
Joseph M Antognini · 2019
Later among the works it cites.
A type of generalization error induced by initialization in deep neural networks
Yaoyu Zhang, Zhi-Qin John Xu, Tao Luo, and Zheng Ma · 2019
Later among the works it cites.
Understanding generalization of deep neural networks trained with noisy labels
Wei Hu, Zhiyuan Li, and Dingli Yu · 2019
Later among the works it cites.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel S. Park, Jascha Sohl-Dickstein, Quoc V. Le, and Samuel L. Smith · 2019
Later among the works it cites.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Zeyuan Allen-Zhu, Yuanzhi Li, and Yingyu Liang · 2019
Later among the works it cites.
Why do larger models generalize better? A theoretical perspective via the XOR problem
Alon Brutzkus and Amir Globerson · 2019
Later among the works it cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
Preetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang, Boaz Barak, and Ilya Sutskever · 2019
Later among the works it cites.
A continuous-time view of early stopping for least squares regression
Alnur Ali, J Zico Kolter, and Ryan J Tibshirani · 2019
Later among the works it cites.
Fairness and Machine Learning
Solon Barocas, Moritz Hardt, and Arvind Narayanan · 2019
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2019
Later among the works it cites.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.
Infinite attention: NNGP and NTK for deep attention networks
Jiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, and Roman Novak · 2020
Closest in time.
On the infinite width limit of neural networks with a standard parameterization
Jascha Sohl-Dickstein, Roman Novak, Samuel S Schoenholz, and Jaehoon Lee · 2020
Closest in time.
On the neural tangent kernel of deep networks with orthogonal initialization
Wei Huang, Weitao Du, and Richard Yi Da Xu · 2020
Closest in time.
Disentangling trainability and generalization in deep learning
Lechao Xiao, Jeffrey Pennington, and Samuel S Schoenholz · 2020
Closest in time.
Sebastian W Ober and Laurence Aitchison · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
On the training dynamics of deep networks with
Aitor Lewkowycz and Guy Gur-Ari · 2020
Closest in time.
Harnessing the power of infinitely wide deep nets on small-data tasks
Sanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov, Ruosong Wang, and Dingli Yu · 2020
Closest in time.
Neural kernels without tangents
Vaishaal Shankar, Alex Chengyu Fang, Wenshuo Guo, Sara Fridovich-Keil, Ludwig Schmidt, Jonathan Ragan-Kelley, and Benjamin Recht · 2020
Closest in time.
Scalable uncertainty for computer vision with functional variational inference
Eduardo D C Carvalho, Ronald Clark, Andrea Nicastro, and Paul H. J. Kelly · 2020
Closest in time.
Scaling description of generalization with number of parameters in deep learning
Mario Geiger, Arthur Jacot, Stefano Spigler, Franck Gabriel, Levent Sagun, Stéphane d’Ascoli, Giulio Biroli, Clément Hongler, and Matthieu Wyart · 2020
Closest in time.
Asymptotics of wide networks from feynman diagrams
Ethan Dyer and Guy Gur-Ari · 2020
Closest in time.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Closest in time.
Non-Gaussian processes and neural networks at finite widths
Sho Yaida · 2020
Closest in time.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Yu Bai and Jason D. Lee · 2020
Closest in time.
Bayesian deep learning and a probabilistic perspective of generalization
Andrew Gordon Wilson and Pavel Izmailov · 2020
Closest in time.
Asymptotics of wide convolutional neural networks
Anders Andreassen and Ethan Dyer · 2020
Closest in time.
Why bigger is not always better: on finite and infinite neural networks
Laurence Aitchison · 2020
Closest in time.
Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, and Jascha Sohl-Dickstein · 2020
Closest in time.
Truth or backpropaganda? an empirical investigation of deep learning theory
Micah Goldblum, Jonas Geiping, Avi Schwarzschild, Michael Moeller, and Tom Goldstein · 2020
Closest in time.
Scipy 1.0: fundamental algorithms for scientific computing in python
Pauli Virtanen, Ralf Gommers, Travis E Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al · 2020
Closest in time.