Fetching the paper…
Reading the bibliography…
Analysing and computing with Gaussian processes arising from infinitely wide neural networks has recently seen a resurgence in popularity.
Yang, G. 2019a · 1902
Earlier work this paper cites.
A fine-grained spectral perspective on neural networks
Yang, G.; and Salman, H. 2019 · 1907
Earlier work this paper cites.
Richer priors for infinitely wide multi-layer perceptrons
Tsuchida, R.; Roosta, F.; and Gallagher, M. 2019b · 1911
Earlier work this paper cites.
Exact expressions for double descent and implicit regularization via surrogate random design
Dereziński, M.; Liang, F.; and Mahoney, M. W. 2019 · 1912
Earlier work this paper cites.
Moments of a truncated bivariate normal distribution
Rosenbaum, S. 1961 · 1961
Earlier work this paper cites.
The theory of generalised functions
Jones, D. S. 1982 · 1982
Earlier work this paper cites.
Probability and measure
Billingsley, P. 1995 · 1995
Earlier work this paper cites.
Bayesian learning for neural networks
Neal, R. M. 1995 · 1995
Earlier work this paper cites.
Computing with infinite networks
Williams, C. K. 1997 · 1997
Earlier work this paper cites.
Fixed point theory and applications , volume 141
Agarwal, R. P.; Meehan, M.; and O’regan, D. 2001 · 2001
Earlier work this paper cites.
Stable behaviour of infinitely wide deep neural networks
Favaro, S.; Fortini, S.; and Peluchetti, S. 2020 · 2003
Earlier work this paper cites.
A first course in dynamics: with a panorama of recent developments
Hasselblatt, B.; and Katok, A. 2003 · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms , 547
MacKay, D. J. 2003 · 2003
Earlier work this paper cites.
Beyond Gaussian processes: On the distributions of infinite networks
Der, R.; and Lee, D. D. 2006 · 2006
Earlier work this paper cites.
Continuous neural networks
Le Roux, N.; and Bengio, Y. 2007 · 2007
Earlier work this paper cites.
Kernel Methods for Deep Learning
Cho, Y.; and Saul, L. K. 2009 · 2009
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Probabilistic Backpropagation for Scalable Learning of Bayesian Neural Networks
Hernández-Lobato, J. M.; and Adams, R. P. 2015 · 2015
Cited alongside, same era.
In search of the real inductive bias: On the role of implicit regularization in deep learning
Neyshabur, B.; Tomioka, R.; and Srebro, N. 2015 · 2015
Cited alongside, same era.
Striving for Simplicity: The All Convolutional Net
Springenberg, J.; Dosovitskiy, A.; Brox, T.; and Riedmiller, M. 2015 · 2015
Cited alongside, same era.
Deep Gaussian processes for regression using approximate expectation propagation
Bui, T.; Hernández-Lobato, D.; Hernandez-Lobato, J.; Li, Y.; and Turner, R. 2016 · 2016
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (ELUs)
Clevert, D.-A.; Unterthiner, T.; and Hochreiter, S. 2016 · 2016
Cited alongside, same era.
Improving language understanding by generative pre-training
Radford, A.; Narasimhan, K.; Salimans, T.; and Sutskever, I. 2018 · 2018
Later among the works it cites.
Invariance of Weight Distributions in Rectified MLPs
Tsuchida, R.; Roosta, F.; and Gallagher, M. 2018 · 2018
Later among the works it cites.
Learning and generalization in overparameterized neural networks, going beyond two layers
Allen-Zhu, Z.; Li, Y.; and Liang, Y. 2019 · 2019
Later among the works it cites.
On exact computation with an infinitely wide neural net
Arora, S.; Du, S. S.; Hu, W.; Li, Z.; Salakhutdinov, R. R.; and Wang, R. 2019 · 2019
Later among the works it cites.
On Lazy Training in Differentiable Programming
Chizat, L.; Oyallon, E.; and Bach, F. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.; Lee, K.; and Toutanova, K. 2019 · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hendrycks, D.; and Gimpel, K. 2016 · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Poole, B.; Lahiri, S.; Raghu, M.; Sohl-Dickstein, J.; and Ganguli, S. 2016 · 2016
Cited alongside, same era.
Self-normalizing neural networks
Klambauer, G.; Unterthiner, T.; Mayr, A.; and Hochreiter, S. 2017 · 2017
Cited alongside, same era.
Searching for activation functions
Ramachandran, P.; Zoph, B.; and Le, Q. V. 2017 · 2017
Cited alongside, same era.
Deep information propagation
Schoenholz, S. S.; Gilmer, J.; Ganguli, S.; and Sohl-Dickstein, J. 2017 · 2017
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Zhang, C.; Bengio, S.; Hardt, M.; Recht, B.; and Vinyals, O. 2017 · 2017
Cited alongside, same era.
Deep reinforcement learning in a handful of trials using probabilistic dynamics models
Chua, K.; Calandra, R.; McAllister, R.; and Levine, S. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J.; Xiao, L.; Schoenholz, S.; Bahri, Y.; Novak, R.; Sohl-Dickstein, J.; and Pennington, J. 2019 · 2019
Later among the works it cites.
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
Novak, R.; Xiao, L.; Lee, J.; Bahri, Y.; Yang, G.; Hron, J.; Abolafia, D. A.; Pennington, J.; and Sohl-Dickstein, J. 2019 · 2019
Later among the works it cites.
Expressive Priors in Bayesian Neural Networks: Kernel Combinations and Periodic Functions
Pearce, T.; Tsuchida, R.; Zaki, M.; Brintrup, A.; and Neely, A. 2019 · 2019
Later among the works it cites.
Deep learning generalizes because the parameter-function map is biased towards simple functions
Valle-Pérez, G.; Camargo, C. Q.; and Louis, A. A. 2019 · 2019
Later among the works it cites.
Why bigger is not always better: on finite and infinite neural networks
Aitchison, L. 2020 · 2020
Closest in time.
Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
Arora, S.; Du, S. S.; Li, Z.; Salakhutdinov, R.; Wang, R.; and Yu, D. 2020 · 2020
Closest in time.
Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks?–A Neural Tangent Kernel Perspective
Huang, K.; Wang, Y.; Tao, M.; and Zhao, T. 2020 · 2020
Closest in time.
Finite versus infinite neural networks: an empirical study
Lee, J.; Schoenholz, S.; Pennington, J.; Adlam, B.; Xiao, L.; Novak, R.; and Sohl-Dickstein, J. 2020 · 2020
Closest in time.
Stationary Activations for Uncertainty Calibration in Deep Learning
Meronen, L.; Irwanto, C.; and Solin, A. 2020 · 2020
Closest in time.
Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Novak, R.; Xiao, L.; Hron, J.; Lee, J.; Alemi, A. A.; Sohl-Dickstein, J.; and Schoenholz, S. S. 2020 · 2020
Closest in time.