Fetching the paper…
Reading the bibliography…
While stochastic gradient descent (SGD) and variants have been surprisingly successful for training deep nets, several aspects of the optimization dynamics and generalization are still not well understood.
Information and the accuracy attainable in the estimation of statistical parameters
C. Radhakrishna Rao · 1945
Earlier work this paper cites.
An iteration method for the solution of the eigenvalue problem of linear differential and integral operators
Cornelius Lanczos · 1950
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Naive Set Theory
Paul Halmos · 1960
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation. iii
C. Davis and W. Kahan · 1970
Earlier work this paper cites.
Principles of mathematical analysis
Walter Rudin · 1976
Earlier work this paper cites.
Statistical Inference
G. Casella and R.L. Berger · 1990
Earlier work this paper cites.
Probability in Banach Spaces: isoperimetry and processes
Michel Ledoux and Michel Talagrand · 1991
Earlier work this paper cites.
Probability with martingales
David Williams · 1991
Earlier work this paper cites.
Fast exact multiplication by the hessian
Barak A. Pearlmutter · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Theory of Point Estimation
E.L. Lehmann and G. Casella · 1998
Earlier work this paper cites.
Nonlinear Programming
D.P. Bertsekas · 1999
Earlier work this paper cites.
Pac-bayesian model averaging
David A McAllester · 1999
Earlier work this paper cites.
Methods of Information Geometry
Shun-ichi Amari and Hiroshi Nagaoka · 2000
Earlier work this paper cites.
Introduction to the gamma function
Pascal Sebah and Xavier Gourdon · 2002
Earlier work this paper cites.
Pac-bayes & margins
John Langford and John Shawe-Taylor · 2003
Earlier work this paper cites.
Elements of Information Theory (Wiley Series in Telecommunications and Signal Processing)
Thomas M. Cover and Joy A. Thomas · 2006
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
All of Statistics: A Concise Course in Statistical Inference
Larry Wasserman · 2010
Earlier work this paper cites.
Information theory and dynamical system predictability
Richard Kleeman · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
A tail inequality for quadratic forms of subgaussian random vectors
Daniel Hsu, Sham Kakade, and Tong Zhang · 2012
Earlier work this paper cites.
Probability essentials
Jean Jacod and Philip Protter · 2012
Earlier work this paper cites.
Introduction to the non-asymptotic analysis of random matrices
Roman Vershynin · 2012
Earlier work this paper cites.
Concentration inequalities: A nonasymptotic theory of independence
Stéphane Boucheron, Gábor Lugosi, and Pascal Massart · 2013
Cited alongside, same era.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Matrix Computations
Gene H. Golub and Charles F. van Loan · 2013
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Later among the works it cites.
Gradient descent learns one-hidden-layer CNN: Don’t be afraid of spurious local minima
Simon S. Du, Jason D. Lee, Yuandong Tian, Barnabás Póczos, and Aarti Singh · 2018
Later among the works it cites.
An overview on the evolution and adoption of deep learning applications used in the industry
Sourav Dutta · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Later among the works it cites.
Three factors influencing minima in SGD
Stanislaw Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos J. Storkey · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Autograd: Effortless gradients in numpy
Dougal Maclaurin, David Duvenaud, and Ryan P. Adams · 2015
Cited alongside, same era.
Local smoothness in variance reduced optimization
Daniel Vainsencher, Han Liu, and Tong Zhang · 2015
Cited alongside, same era.
A useful variant of the davis–kahan theorem for statisticians
Y. Yu, T. Wang, and R. J. Samworth · 2015
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
Brownian Motion, Martingales, and Stochastic Calculus
Jean-Francois Le Gall · 2016
Cited alongside, same era.
Stanislaw Jastrzebski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos J. Storkey · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Recent advances in deep learning: An overview
Matiur Rahman Minar and Jibon Naher · 2018
Later among the works it cites.
The full spectrum of deep net hessians at scale: Dynamics with sample size
Vardan Papyan · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
Spurious local minima are common in two-layer relu neural networks
Itay Safran and Ohad Shamir · 2018
Later among the works it cites.
Exponential convergence time of gradient descent for one-dimensional deep linear neural networks
Ohad Shamir · 2018
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Later among the works it cites.
High-dimensional probability: An introduction with applications in data science
Roman Vershynin · 2018
Later among the works it cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Y Bengio · 2018
Later among the works it cites.
Small nonlinearities in activation functions create bad local minima in neural networks
Chulhee Yun, Suvrit Sra, and Ali Jadbabaie · 2018
Later among the works it cites.
Stochastic Gradient Descent Optimizes Over-parameterized Deep ReLU Networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2018
Later among the works it cites.
A convergence analysis of gradient descent for deep linear neural networks
Sanjeev Arora, Nadav Cohen, Noah Golowich, and Wei Hu · 2019
Closest in time.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Closest in time.
Gradient descent provably optimizes over-parameterized neural networks
Simon S. Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Closest in time.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Closest in time.
Deterministic PAC-bayesian generalization bounds for deep networks via generalizing noise-resilience
Vaishnavh Nagarajan and Zico Kolter · 2019
Closest in time.
Vardan Papyan · 2019
Closest in time.
A scale invariant flatness measure for deep network minima
Akshay Rangamani, Nam H Nguyen, Abhishek Kumar, Dzung Phan, Sang H Chin, and Trac D Tran · 2019
Closest in time.
Escaping saddle points with adaptive gradient methods
Matthew Staib, Sashank Reddi, Satyen Kale, Sanjiv Kumar, and Suvrit Sra · 2019
Closest in time.
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2019
Closest in time.
Positively scale-invariant flatness of relu neural networks
Mingyang Yi, Qi Meng, Wei Chen, Zhi-ming Ma, and Tie-Yan Liu · 2019
Closest in time.