Fetching the paper…
Reading the bibliography…
Natural gradient descent has proven effective at mitigating the effects of pathological curvature in neural network optimization, but little is known theoretically about its convergence properties, especially for \emph{nonlinear} networks.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 1901
Earlier work this paper cites.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
Jssai Schur · 1911
Earlier work this paper cites.
Improving the convergence of backpropagation learning with second order methods
S. Becker and Y. LeCun · 1989
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim Blum and Ronald L. Rivest · 1992
Earlier work this paper cites.
Neural learning in structured parameter spaces-natural riemannian gradient
Shun-ichi Amari · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, Patrick Haffner, et al · 1998
Earlier work this paper cites.
Rademacher and gaussian complexities: Risk bounds and structural results
Peter L Bartlett and Shahar Mendelson · 2002
Earlier work this paper cites.
A linear lower bound on the unbounded error probabilistic communication complexity
Jürgen Forster · 2002
Earlier work this paper cites.
Stochastic gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2002
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M Saxe, James L McClelland, and Surya Ganguli · 2013
Earlier work this paper cites.
New insights and perspectives on the natural gradient method
James Martens · 2014
Earlier work this paper cites.
Escaping from saddle points—online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Norm-based capacity control in neural networks
Behnam Neyshabur, Ryota Tomioka, and Nathan Srebro · 2015
Earlier work this paper cites.
Riemannian metrics for neural networks I: Feedforward networks
Yann Ollivier · 2015
Earlier work this paper cites.
An introduction to matrix concentration inequalities
Joel A Tropp et al · 2015
Cited alongside, same era.
A Kronecker-factored approximate Fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Cited alongside, same era.
Identity matters in deep learning
Moritz Hardt and Tengyu Ma · 2016
Cited alongside, same era.
Deep learning without poor local minima
Kenji Kawaguchi · 2016
Cited alongside, same era.
Gradient descent converges to minimizers
Jason D Lee, Max Simchowitz, Michael I Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
No bad local minima: Data independent training error guarantees for multilayer neural networks
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Shun Liao, Roger Grosse, and Jimmy Ba · 2017
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2018
Later among the works it cites.
Stronger generalization bounds for deep nets via a compression approach
Sanjeev Arora, Rong Ge, Behnam Neyshabur, and Yi Zhang · 2018
Later among the works it cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Alberto Bernacchia, Mate Lengyel, and Guillaume Hennequin · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Daniel Soudry and Yair Carmon · 2016
Cited alongside, same era.
Diverse neural network learns true target functions
Bo Xie, Yingyu Liang, and Le Song · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Cited alongside, same era.
High-dimensional dynamics of generalization error in neural networks
Madhu S Advani and Andrew M Saxe · 2017
Cited alongside, same era.
Spectrally-normalized margin bounds for neural networks
Peter L Bartlett, Dylan J Foster, and Matus J Telgarsky · 2017
Cited alongside, same era.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Cited alongside, same era.
Later among the works it cites.
Data-dependent pac-bayes priors via differential privacy
Gintare Karolina Dziugaite and Daniel M Roy · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
Depth with nonlinearity creates no bad local minima in resnets
Kenji Kawaguchi and Yoshua Bengio · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
A coordinate-free construction of scalable natural gradient
Kevin Luk and Roger Grosse · 2018
Later among the works it cites.
Kronecker-factored curvature approximations for recurrent neural networks
James Martens, Jimmy Ba, and Matt Johnson · 2018
Later among the works it cites.
Foundations of machine learning
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar · 2018
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Closest in time.
Learning one-hidden-layer neural networks under general input distributions
Weihao Gao, Ashok Makkuva, Sewoong Oh, and Pramod Viswanath · 2019
Closest in time.
Wide neural network of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S Schoenholz, Yasaman Bahri, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Closest in time.
The role of over-parametrization in generalization of neural networks
Behnam Neyshabur, Zhiyuan Li, Srinadh Bhojanapalli, Yann LeCun, and Nathan Srebro · 2019
Closest in time.
Samet Oymak and Mahdi Soltanolkotabi · 2019
Closest in time.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Closest in time.