Fetching the paper…
Reading the bibliography…
It is well-known that the distribution over functions induced through a zero-mean iid prior distribution over the parameters of a multi-layer perceptron (MLP) converges to a Gaussian process (GP), under mild conditions.
Greg Yang · 1902
Earlier work this paper cites.
Representations for partially exchangeable arrays of random variables
D.J. Aldous · 1981
Earlier work this paper cites.
Bayesian learning for neural networks
Radford M Neal · 1995
Earlier work this paper cites.
Computing with infinite networks
Christopher KI Williams · 1997
Earlier work this paper cites.
Information theory, inference and learning algorithms , page 547
David J.C. MacKay · 2003
Earlier work this paper cites.
Numerical computation of rectangular bivariate and trivariate normal and t probabilities
Alan Genz · 2004
Earlier work this paper cites.
Probabilistic symmetries and invariance principles , page 325
Olav Kallenberg · 2006
Earlier work this paper cites.
Gaussian Processes for Machine Learning
CE. Rasmussen and CKI. Williams · 2006
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
Edward Snelson and Zoubin Ghahramani · 2006
Earlier work this paper cites.
Continuous neural networks
Nicolas Le Roux and Yoshua Bengio · 2007
Earlier work this paper cites.
Kernel methods for deep learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
F2py: a tool for connecting fortran and python programs
Pearu Peterson · 2009
Earlier work this paper cites.
Rectified linear units improve restricted boltzmann machines
Vinod Nair and Geoffrey E Hinton · 2010
Cited alongside, same era.
GPy: A Gaussian process framework in Python
GPy · 2012
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
Analysis of boolean functions , page 330
Ryan O’Donnell · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
When and why are deep networks better than shallow ones?
Hrushikesh Mhaskar, Qianli Liao, and Tomaso Poggio · 2017
Later among the works it cites.
Swish: a self-gated activation function
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Later among the works it cites.
Analytic expressions for probabilistic moments of pl-dnn with gaussian input
Adel Bibi, Modar Alfadly, and Bernard Ghanem · 2018
Later among the works it cites.
Deep convolutional networks as shallow gaussian processes
Adrià Garriga-Alonso, Carl Edward Rasmussen, and Laurence Aitchison · 2018
Later among the works it cites.
Deep neural networks as Gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Going deeper with convolutions
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich · 2015
Cited alongside, same era.
Fast and accurate deep network learning by exponential linear units (elus)
Djork-Arné Clevert, Thomas Unterthiner, and Sepp Hochreiter · 2016
Cited alongside, same era.
Toward deeper understanding of neural networks: The power of initialization and a dual view on expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Cited alongside, same era.
On moments of folded and truncated multivariate normal distributions
Raymond Kan and Cesare Robotti · 2017
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Later among the works it cites.
Invariance of weight distributions in rectified MLPs
Russell Tsuchida, Farbod Roosta-Khorasani, and Marcus Gallagher · 2018
Later among the works it cites.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Closest in time.
Bayesian deep convolutional networks with many channels are gaussian processes
Roman Novak, Lechao Xiao, Jaehoon Lee, Yasaman Bahri, Greg Yang, Jiri Hron, Daniel A Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2019
Closest in time.
Expressive priors in bayesian neural networks: Kernel combinations and periodic functions
Tim Pearce, Mohamed Zaki, Alexandra Brintrup, and Andy Neely · 2019
Closest in time.
Exchangeability and kernel invariance in trained MLPs
Russell Tsuchida, Farbod Roosta-Khorasani, and Marcus Gallagher · 2019
Closest in time.