Fetching the paper…
Reading the bibliography…
The prevailing thinking is that orthogonal weights are crucial to enforcing dynamical isometry and speeding up training.
Multivariate normal approximation using exchangeable pairs
Sourav Chatterjee and Elizabeth Meckes · 2007
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Earlier work this paper cites.
Human-level control through deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Andrei A Rusu, Joel Veness, Marc G Bellemare, Alex Graves, Martin Riedmiller, Andreas K Fidjeland, Georg Ostrovski, et al · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Exponential expressivity in deep neural networks through transient chaos
Ben Poole, Subhaneil Lahiri, Maithra Raghu, Jascha Sohl-Dickstein, and Surya Ganguli · 2016
Earlier work this paper cites.
Samuel S Schoenholz, Justin Gilmer, Surya Ganguli, and Jascha Sohl-Dickstein · 2016
Earlier work this paper cites.
Mean field residual networks: On the edge of chaos
Ge Yang and Samuel Schoenholz · 2017
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Earlier work this paper cites.
Deep neural networks as gaussian processes
Jaehoon Lee, Yasaman Bahri, Roman Novak, Samuel S Schoenholz, Jeffrey Pennington, and Jascha Sohl-Dickstein · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel S Schoenholz, and Jeffrey Pennington · 2018
Cited alongside, same era.
Minmin Chen, Jeffrey Pennington, and Samuel S Schoenholz · 2018
Cited alongside, same era.
The emergence of spectral universality in deep networks
Jeffrey Pennington, Samuel S Schoenholz, and Surya Ganguli · 2018
Cited alongside, same era.
Dynamical isometry is achieved in residual networks in a universal way for any activation function
Wojciech Tarnowski, Piotr Warchoł, Stanisław Jastrzębski, Jacek Tabor, and Maciej A Nowak · 2018
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Later among the works it cites.
Greg Yang · 2019
Later among the works it cites.
Dynamics of deep neural networks and neural tangent hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2019
Later among the works it cites.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Russ R Salakhutdinov, and Ruosong Wang · 2019
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Information geometry of orthogonal initializations and training
Piotr A Sokol and Il Memming Park · 2018
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Cited alongside, same era.
Gaussian process behaviour in wide deep neural networks
Alexander G de G Matthews, Mark Rowland, Jiri Hron, Richard E Turner, and Zoubin Ghahramani · 2018
Cited alongside, same era.
Feature extraction and image processing for computer vision
Mark Nixon and Alberto Aguado · 2019
Cited alongside, same era.
Mean field theory for deep dropout networks: digging up gradient backpropagation deeply
Wei Huang, Richard Yi Da Xu, Weitao Du, Yutian Zeng, and Yunce Zhao · 2019
Cited alongside, same era.
A mean field theory of batch normalization
Greg Yang, Jeffrey Pennington, Vinay Rao, Jascha Sohl-Dickstein, and Samuel S Schoenholz · 2019
Cited alongside, same era.
Spectrum concentration in deep residual learning: a free probability approach
Zenan Ling and Robert C Qiu · 2019
Cited alongside, same era.
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
Provable benefit of orthogonal initialization in optimizing deep linear networks
Wei Hu, Lechao Xiao, and Jeffrey Pennington · 2020
Closest in time.
Tensor programs ii: Neural tangent kernel for any architecture
Greg Yang · 2020
Closest in time.
On the infinite width limit of neural networks with a standard parameterization
Jascha Sohl-Dickstein, Roman Novak, Samuel S Schoenholz, and Jaehoon Lee · 2020
Closest in time.
On the linearity of large non-linear models: when and why the tangent kernel is constant
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Constructing exchangeable pairs by diffusion on manifolds and its application
Weitao Du · 2020
Closest in time.