Why and when can deep-but not shallow-networks avoid the curse of dimensionality: A review
Tomaso Poggio, Hrushikesh Mhaskar, Lorenzo Rosasco, Brando Miranda, and Qianli Liao · 2017
Later among the works it cites.
Deep information propagation
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2017
Later among the works it cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Later among the works it cites.
Convolutional gaussian processes
Mark van der Wilk, Carl Edward Rasmussen, and James Hensman · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
Mean field residual networks: On the edge of chaos
Greg Yang and Samuel Schoenholz · 2017
Later among the works it cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Later among the works it cites.
Assessing the scalability of biologically-motivated deep learning algorithms and architectures
Sergey Bartunov, Adam Santoro, Blake Richards, Luke Marris, Geoffrey E Hinton, and Timothy Lillicrap · 2018
Closest in time.
Deep convolutional gaussian processes
Original
Kenneth Blomqvist, Samuel Kaski, and Markus Heinonen · 2018
Closest in time.
A gaussian process perspective on convolutional neural networks
Original
Anastasia Borovykh · 2018
Closest in time.
Dynamical isometry and a mean field theory of RNNs: Gating enables signal propagation in recurrent neural networks
Minmin Chen, Jeffrey Pennington, and Samuel Schoenholz · 2018
Closest in time.
Deep convolutional networks as shallow Gaussian processes
Original
Adrià Garriga-Alonso, Laurence Aitchison, and Carl Edward Rasmussen · 2018
Closest in time.
How to start training: The effect of initialization and architecture
Original
Boris Hanin and David Rolnick · 2018
Closest in time.
Deep gaussian processes with convolutional kernels
Original
Vinayak Kumar, Vaibhav Singh, PK Srijith, and Andreas Damianou · 2018
Closest in time.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Sam Schoenholz, Jeffrey Pennington, and Jascha Sohl-dickstein · 2018
Closest in time.
Sensitivity and generalization in neural networks: an empirical study
Roman Novak, Yasaman Bahri, Daniel A. Abolafia, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2018
Closest in time.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Closest in time.
Deep Mean Field Theory: Layerwise Variance and Width Variation as Methods to Control Gradient Explosion
Greg Yang and Sam S. Schoenholz · 2018
Closest in time.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Sam S. Schoenholz · 2018
Closest in time.
On the effect of the activation function on the distribution of hidden nodes in a deep network
Original
Philip M Long and Hanie Sedghi · 2019
Closest in time.
Scaling limits of wide neural networks with weight sharing: Gaussian process behavior, gradient independence, and neural tangent kernel derivation
Original
Greg Yang · 2019
Closest in time.
Neural tangents: Fast and easy infinite neural networks in python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2020
Closest in time.