Fetching the paper…
Reading the bibliography…
Recently mean field theory has been successfully used to analyze properties of wide, random neural networks.
Bayesian Learning for Neural Networks , volume 118 of Lecture Notes in Statistics
Radford M. Neal · 1996
Earlier work this paper cites.
The Geometry of Algorithms with Orthogonality Constraints
Alan Edelman, Tomás A. Arias, and Steven T. Smith · 1998
Earlier work this paper cites.
Convex Optimization, With Corrections 2008
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Optimization Algorithms on Matrix Manifolds
P.-A. Absil, R. Mahony, and R. Sepulchre · 2007
Earlier work this paper cites.
Spectral Theory of Block Operator Matrices and Applications
Christiane Tretter · 2008
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Unitary Evolution Recurrent Neural Networks
Martin Arjovsky, Amar Shah, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Optimization Methods for Large-Scale Machine Learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Earlier work this paper cites.
Generalized BackPropagation, \’{E}tude De Cas: Orthogonality
Mehrtash Harandi and Basura Fernando · 2016
Earlier work this paper cites.
Recurrent Orthogonal Networks and Long-Memory Tasks
Mikael Henaff, Arthur Szlam, and Yann LeCun · 2016
Cited alongside, same era.
Optimization on Submanifolds of Convolution Kernels in CNNs
Mete Ozay and Takayuki Okatani · 2016
Cited alongside, same era.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
Tim Salimans and Diederik P. Kingma · 2016
Cited alongside, same era.
Samuel S. Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Cited alongside, same era.
All you need is beyond a good init: Exploring better solution for training extremely deep convolutional neural networks with orthonormality and modulation
Di Xie, Jiang Xiong, and Shiliang Pu · 2017
Later among the works it cites.
Fisher Information and Natural Gradient Learning of Random Deep Networks
Shun-ichi Amari, Ryo Karakida, and Masafumi Oizumi · 2018
Closest in time.
Dynamical Isometry and a Mean Field Theory of RNNs: Gating Enables Signal Propagation in Recurrent Neural Networks
Minmin Chen, Jeffrey Pennington, and Samuel Schoenholz · 2018
Closest in time.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Closest in time.
Universal Statistics of Fisher Information in Deep Neural Networks: Mean Field Approach
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scott Wisdom, Thomas Powers, John R. Hershey, Jonathan Le Roux, and Les Atlas · 2016
Cited alongside, same era.
Riemannian approach to batch normalization
Minhyung Cho and Jaehyung Lee · 2017
Cited alongside, same era.
Parseval networks: Improving robustness to adversarial examples
Moustapha Cisse, Piotr Bojanowski, Edouard Grave, Yann Dauphin, and Nicolas Usunier · 2017
Cited alongside, same era.
Deep Neural Networks as Gaussian Processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, et al · 2017
Cited alongside, same era.
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Sam Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
On orthogonality and learning recurrent networks with long term dependencies
Eugene Vorontsov, Chiheb Trabelsi, Samuel Kadoury, and Chris Pal · 2017
Cited alongside, same era.
Dynamical Isometry and a Mean Field Theory of CNNs: How to Train 10,000-Layer Vanilla Convolutional Neural Networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington
Cited in the paper.
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 2018
Closest in time.
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi, et al · 2018
Closest in time.
Gaussian Process Behaviour in Wide Deep Neural Networks
Alexander G. de G. Matthews, Mark Rowland, Jiri Hron, Richard E. Turner, and Zoubin Ghahramani · 2018
Closest in time.
The Emergence of Spectral Universality in Deep Networks
Jeffrey Pennington, Samuel S. Schoenholz, and Surya Ganguli · 2018
Closest in time.
How Does Batch Normalization Help Optimization? (No, It Is Not About Internal Covariate Shift)
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Madry · 2018
Closest in time.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, et al · 2019
Closest in time.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Closest in time.