Fetching the paper…
Reading the bibliography…
Whilst deep neural networks have shown great empirical success, there is still much work to be done to understand their theoretical properties.
Some theorems on distribution functions
H. Cramér and H. Wold · 1936
Earlier work this paper cites.
Central limit theorems for interchangeable processes
J. R. Blum, H. Chernoff, M. Rosenblatt, and H. Teicher · 1958
Earlier work this paper cites.
Stochastic Variational Deep Kernel Learning
Andrew G. Wilson, Zhiting Hu, Ruslan R. Salakhutdinov, and Eric P. Xing · 1958
Earlier work this paper cites.
Probability and Measure
Patrick Billingsley · 1986
Earlier work this paper cites.
Bayesian Learning for Neural Networks
Radford M. Neal · 1996
Earlier work this paper cites.
Asymptotic Statistics
A. W. van der Vaart · 1998
Earlier work this paper cites.
Computing with Infinite Networks
Christopher K. I. Williams · 1998
Earlier work this paper cites.
Convergence of Probability Measures
Patrick Billingsley · 1999
Earlier work this paper cites.
Information Theory, Inference & Learning Algorithms
David J. C. MacKay · 2002
Earlier work this paper cites.
The Curse of Dimensionality for Local Kernel Machines
Yoshua Bengio, Olivier Delalleau, and Nicolas Le Roux · 2005
Earlier work this paper cites.
Sparse Gaussian processes using pseudo-inputs
Edward Snelson and Zoubin Ghahramani · 2005
Earlier work this paper cites.
Gaussian Processes for Machine Learning
Carl E. Rasmussen and Christopher K. I. Williams · 2006
Earlier work this paper cites.
Kernel Methods for Deep Learning
Youngmin Cho and Lawrence K. Saul · 2009
Earlier work this paper cites.
Elliptical Slice Sampling
Iain Murray, Ryan P. Adams, and David J. C. MacKay · 2010
Cited alongside, same era.
MCMC using Hamiltonian Dynamics
Radford M. Neal · 2010
Cited alongside, same era.
Practical Variational Inference for Neural Networks
Alex Graves · 2011
Cited alongside, same era.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee Whye Teh · 2011
Cited alongside, same era.
A Kernel Two-sample test
Arthur Gretton, Karsten M. Borgwardt, Malte J. Rasch, Bernhard Schölkopf, and Alexander Smola · 2012
Cited alongside, same era.
Hamiltonian Annealed Importance Sampling for partition function estimation
Jascha Sohl-Dickstein and Benjamin J. Culpepper · 2012
Cited alongside, same era.
Toward Deeper Understanding of Neural Networks: The Power of Initialization and a Dual View on Expressivity
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Later among the works it cites.
Black-box alpha divergence minimization
José M. Hernández-Lobato, Yingzhen Li, Mark Rowland, Thang Bui, Daniel Hernández-Lobato, and Richard E. Turner · 2016
Later among the works it cites.
Exponential expressivity in Deep Neural Networks through Transient Chaos
Ben Poole, Subhaneil Lahiri, Maithreyi Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Later among the works it cites.
Learning Scalable Deep Kernels with Recurrent Structure
Maruan Al-Shedivat, Andrew G. Wilson, Yunus Saatchi, Zhiting Hu, and Eric P. Xing · 2017
Later among the works it cites.
Self-normalizing neural networks
Günter Klambauer, Thomas Unterthiner, Andreas Mayr, and Sepp Hochreiter · 2017
Later among the works it cites.
AutoGP: Exploring the capabilities and limitations of Gaussian Process models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Deep Gaussian Processes
Andreas C. Damianou and Neil D. Lawrence · 2013
Cited alongside, same era.
The Bayesian Approach To Inverse Problems
M. Dashti and A. M. Stuart · 2013
Cited alongside, same era.
Avoiding Pathologies in very Deep Networks
David Duvenaud, Oren Rippel, Ryan P. Adams, and Zoubin Ghahramani · 2014
Cited alongside, same era.
Weight Uncertainty in Neural Networks
Charles Blundell, Julien Cornebise, Koray Kavukcuoglu, and Daan Wierstra · 2015
Cited alongside, same era.
Sandwiching the marginal likelihood using bidirectional Monte Carlo
Roger B. Grosse, Zoubin Ghahramani, and Ryan P. Adams · 2015
Cited alongside, same era.
Steps Toward Deep Kernel Methods from Infinite Neural Networks
Tamir Hazan and Tommi Jaakkola · 2015
Cited alongside, same era.
Karl Krauth, Edwin V. Bonilla, Kurt Cutajar, and Maurizio Filippone · 2017
Later among the works it cites.
Stochastic Gradient Descent as Approximate Bayesian Inference
S. Mandt, M. D. Hoffman, and D. M. Blei · 2017
Later among the works it cites.
GPflow: A Gaussian Process Library using TensorFlow
Alexander G. de G. Matthews, Mark van der Wilk, Tom Nickson, Keisuke Fujii, Alexis Boukouvalas, Pablo León-Villagrá, Zoubin Ghahramani, and James Hensman · 2017
Later among the works it cites.
Deep Kernel Machines via the Kernel Reparametrization Trick
Jovana Mitrovic, Dino Sejdinovic, and Yee Whye Teh · 2017
Later among the works it cites.
Deep Information Propagation
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Deep Neural Networks as Gaussian Processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S. Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Closest in time.
Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G. de G. Matthews, Jiri Hron, Mark Rowland, Richard E. Turner, and Zoubin Ghahramani · 2018
Closest in time.
A Bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Closest in time.