Fetching the paper…
Reading the bibliography…
This paper relates parameter distance to gradient breakdown for a broad class of nonlinear compositional functions.
Problem complexity and method efficiency in optimization
Arkady S. Nemirovsky and David B. Yudin · 1983
Earlier work this paper cites.
Learning representations by back-propagating errors
David E. Rumelhart, Geoffrey E. Hinton, and Ronald J. Williams · 1986
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Smoothed analysis of the condition numbers and growth factors of matrices
Arvind Sankar, Daniel A. Spielman, and Shang-Hua Teng · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Smoothed analysis of Moore-Penrose inversion
Peter Bürgisser and Felipe Cucker · 2010
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
Andrew M. Saxe, James L. McClelland, and Surya Ganguli · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Path-SGD: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan Salakhutdinov, and Nathan Srebro · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Rethinking the Inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2015
Cited alongside, same era.
ImageNet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Natasha 2: Faster non-convex optimization than SGD
Zeyuan Allen-Zhu · 2018
Later among the works it cites.
signSGD: Compressed optimisation for non-convex problems
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Later among the works it cites.
Dynamical isometry and a mean field theory of CNNs: How to train 10,000-layer vanilla convolutional neural networks
Lechao Xiao, Yasaman Bahri, Jascha Sohl-Dickstein, Samuel Schoenholz, and Jeffrey Pennington · 2018
Later among the works it cites.
Spectral normalization for generative adversarial networks
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida · 2018
Later among the works it cites.
A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks
Behnam Neyshabur, Srinadh Bhojanapalli, and Nathan Srebro · 2018
Later among the works it cites.
cGANs with projection discriminator
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Information geometry and its applications
Shunichi Amari · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
GANs trained by a two time-scale update rule converge to a local Nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter · 2017
Cited alongside, same era.
Scaling SGD batch size to 32K for Imagenet training
Yang You, Igor Gitman, and Boris Ginsburg · 2017
Cited alongside, same era.
Resurrecting the sigmoid in deep learning through dynamical isometry: theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Gradient descent can take exponential time to escape saddle points
Simon S. Du, Chi Jin, Jason D. Lee, Michael I. Jordan, Aarti Singh, and Barnabas Poczos · 2017
Cited alongside, same era.
Takeru Miyato and Masanori Koyama · 2018
Later among the works it cites.
Progressive growing of GANs for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen · 2018
Later among the works it cites.
Large scale GAN training for high fidelity natural image synthesis
Andrew Brock, Jeff Donahue, and Karen Simonyan · 2019
Later among the works it cites.
Competitive gradient descent
Florian Schaefer and Anima Anandkumar · 2019
Later among the works it cites.
Measuring and regularizing networks in function space
Ari Benjamin, David Rolnick, and Konrad Kording · 2019
Later among the works it cites.
Optimization for deep learning: theory and algorithms
Ruoyu Sun · 2019
Later among the works it cites.
Stochastic gradient/mirror descent: Minimax optimality and implicit regularization
Navid Azizan and Babak Hassibi · 2019
Later among the works it cites.
Navid Azizan, Sahin Lale, and Babak Hassibi · 2019
Later among the works it cites.
Sorting out Lipschitz function approximation
Cem Anil, James Lucas, and Roger Grosse · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training BERT in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Closest in time.