Fetching the paper…
Reading the bibliography…
Large width limits have been a recent focus of deep learning research: modulo computational practicalities, do wider networks outperform narrower ones? Answering this question has been challenging, as conventional networks gain representational power with width, potentially masking any negative effects.
Metric spaces and completely monotone functions
I. J. Schoenberg · 1938
Earlier work this paper cites.
Correlation theory of stationary and related random functions
A. Yaglom · 1987
Earlier work this paper cites.
Bayesian learning for neural networks
R. M. Neal · 1995
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner · 1998
Earlier work this paper cites.
Classes of kernels for machine learning: a statistics perspective
M. G. Genton · 2001
Earlier work this paper cites.
The curse of dimensionality for local kernel machines
Y. Bengio, O. Delalleau, and N. Le Roux · 2005
Earlier work this paper cites.
Pattern recognition and machine learning
C. M. Bishop · 2006
Earlier work this paper cites.
Universal kernels
C. A. Micchelli, Y. Xu, and H. Zhang · 2006
Earlier work this paper cites.
Gaussian processes for machine learning , volume 1
C. E. Rasmussen and C. Williams · 2006
Earlier work this paper cites.
UCI machine learning repository, 2007
A. Asuncion and D. Newman · 2007
Earlier work this paper cites.
Random features for large-scale kernel machines
A. Rahimi, B. Recht, et al · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Y. Cho and L. Saul · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Convergence of probability measures
P. Billingsley · 2013
Earlier work this paper cites.
Deep Gaussian processes
A. Damianou and N. Lawrence · 2013
Earlier work this paper cites.
Avoiding pathologies in very deep networks
D. Duvenaud, O. Rippel, R. Adams, and Z. Ghahramani · 2014
Earlier work this paper cites.
The no-u-turn sampler: adaptively setting path lengths in Hamiltonian Monte Carlo
M. D. Hoffman and A. Gelman · 2014
Earlier work this paper cites.
On the number of linear regions of deep neural networks
G. Montúfar, R. Pascanu, K. Cho, and Y. Bengio · 2014
Earlier work this paper cites.
Weight uncertainty in neural network
C. Blundell, J. Cornebise, K. Kavukcuoglu, and D. Wierstra · 2015
Earlier work this paper cites.
Deep Gaussian processes and variational propagation of uncertainty
A. Damianou · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2015
Earlier work this paper cites.
Deep Gaussian processes for regression using approximate expectation propagation
T. Bui, D. Hernández-Lobato, J. Hernandez-Lobato, Y. Li, and R. Turner · 2016
Earlier work this paper cites.
Variational auto-encoded deep Gaussian processes
Z. Dai, A. C. Damianou, J. González, and N. D. Lawrence · 2016
Earlier work this paper cites.
Dropout as a Bayesian approximation: Representing model uncertainty in deep learning
Y. Gal and Z. Ghahramani · 2016
Earlier work this paper cites.
Deep learning
I. Goodfellow, Y. Bengio, A. Courville, and Y. Bengio · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Structured and efficient variational deep learning with matrix Gaussian posteriors
C. Louizos and M. Welling · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
B. Poole, S. Lahiri, M. Raghu, J. Sohl-Dickstein, and S. Ganguli · 2016
Earlier work this paper cites.
Benefits of depth in neural networks
M. Telgarsky · 2016
Cited alongside, same era.
Sequential inference for deep Gaussian process
Y. Wang, M. Brubaker, B. Chaib-Draa, and R. Urtasun · 2016
Cited alongside, same era.
Wide residual networks
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Random feature expansions for deep Gaussian processes
K. Cutajar, E. V. Bonilla, P. Michiardi, and M. Filippone · 2017
Cited alongside, same era.
The expressive power of neural networks: A view from the width
Z. Lu, H. Pu, F. Wang, Z. Hu, and L. Wang · 2017
Cited alongside, same era.
The loss surface of deep and wide neural networks
Q. Nguyen and M. Hein · 2017
Cited alongside, same era.
On the expressive power of deep neural networks
Understanding priors in Bayesian neural networks at the unit level
M. Vladimirova, J. Verbeek, P. Mesejo, and J. Arbel · 2019
Later among the works it cites.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
C. Wei, J. Lee, Q. Liu, and T. Ma · 2019
Later among the works it cites.
Tensor programs I: Wide feedforward or recurrent neural networks of any architecture are gaussian processes
G. Yang · 2019
Later among the works it cites.
On the power and limitations of random features for understanding neural networks
G. Yehudai and O. Shamir · 2019
Later among the works it cites.
Wide neural networks with bottlenecks are deep Gaussian processes
D. Agrawal, T. Papamarkou, and J. Hinkle · 2020
Later among the works it cites.
Why bigger is not always better: on finite and infinite neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Raghu, B. Poole, J. Kleinberg, S. Ganguli, and J. Sohl-Dickstein · 2017
Cited alongside, same era.
Doubly stochastic variational inference for deep Gaussian processes
H. Salimbeni and M. Deisenroth · 2017
Cited alongside, same era.
On the global convergence of gradient descent for over-parameterized models using optimal transport
L. Chizat and F. Bach · 2018
Cited alongside, same era.
How deep are deep Gaussian processes?
M. M. Dunlop, M. A. Girolami, A. M. Stuart, and A. L. Teckentrup · 2018
Cited alongside, same era.
GPyTorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration
J. R. Gardner, G. Pleiss, K. Q. Weinberger, D. Bindel, and A. G. Wilson · 2018
Cited alongside, same era.
Inference in deep Gaussian processes using stochastic gradient Hamiltonian Monte Carlo
M. Havasi, J. M. Hernández-Lobato, and J. J. Murillo-Fuentes · 2018
Cited alongside, same era.
L. Aitchison · 2020
Later among the works it cites.
Backward feature correction: How deep learning performs deep learning
Z. Allen-Zhu and Y. Li · 2020
Later among the works it cites.
Harnessing the power of infinitely wide deep nets on small-data tasks
S. Arora, S. S. Du, Z. Li, R. Salakhutdinov, R. Wang, and D. Yu · 2020
Later among the works it cites.
Beyond linearization: On quadratic and higher-order approximation of wide neural networks
Y. Bai and J. D. Lee · 2020
Later among the works it cites.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Y. Cao and Q. Gu · 2020
Later among the works it cites.
Towards understanding hierarchical learning: Benefits of neural representations
M. Chen, Y. Bai, J. D. Lee, T. Zhao, H. Wang, C. Xiong, and R. Socher · 2020
Later among the works it cites.
Bayesian image classification with deep convolutional Gaussian processes
V. Dutordoir, M. Wilk, A. Artemev, and J. Hensman · 2020
Later among the works it cites.
On the expressiveness of approximate inference in bayesian neural networks
A. Y. Foong, D. R. Burt, Y. Li, and R. E. Turner · 2020
Later among the works it cites.
Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the neural tangent kernel
S. Fort, G. K. Dziugaite, M. Paul, S. Kharaghani, D. M. Roy, and S. Ganguli · 2020
Later among the works it cites.
Scaling description of generalization with number of parameters in deep learning
M. Geiger, A. Jacot, S. Spigler, F. Gabriel, L. Sagun, S. d’Ascoli, G. Biroli, C. Hongler, and M. Wyart · 2020
Later among the works it cites.
When do neural networks outperform kernel methods?
B. Ghorbani, S. Mei, T. Misiakiewicz, and A. Montanari · 2020
Later among the works it cites.
Towards a general theory of infinite-width limits of neural classifiers
E. Golikov · 2020
Later among the works it cites.
Finite versus infinite neural networks: an empirical study
J. Lee, S. Schoenholz, J. Pennington, B. Adlam, L. Xiao, R. Novak, and J. Sohl-Dickstein · 2020
Later among the works it cites.
Learning over-parametrized two-layer neural networks beyond NTK
Y. Li, T. Ma, and H. R. Zhang · 2020
Later among the works it cites.
Interpretable deep Gaussian processes with moments
C.-K. Lu, S. C.-H. Yang, X. Hao, and P. Shafto · 2020
Later among the works it cites.
Deep double descent: Where bigger models and more data hurt
P. Nakkiran, G. Kaplun, Y. Bansal, T. Yang, B. Barak, and I. Sutskever · 2020
Later among the works it cites.
Neural kernels without tangents
V. Shankar, A. Fang, W. Guo, S. Fridovich-Keil, J. Ragan-Kelley, L. Schmidt, and B. Recht · 2020
Later among the works it cites.
Tensor programs II: Neural tangent kernel for any architecture
G. Yang · 2020
Later among the works it cites.
Deep kernel processes
L. Aitchison, A. Yang, and S. W. Ober · 2021
Closest in time.
Neural networks and quantum field theory
J. Halverson, A. Maiti, and K. Stoner · 2021
Closest in time.
What are bayesian neural network posteriors really like?
P. Izmailov, S. Vikram, M. D. Hoffman, and A. G. Wilson · 2021
Closest in time.
Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processes
S. W. Ober and L. Aitchison · 2021
Closest in time.
Tensor programs IV: Feature learning in infinite-width neural networks
G. Yang and E. J. Hu · 2021
Closest in time.
Exact priors of finite neural networks
J. A. Zavatone-Veth and C. Pehlevan · 2021
Closest in time.