Fetching the paper…
Reading the bibliography…
We study the effect of mini-batching on the loss landscape of deep neural networks using spiked, field-dependent random matrix theory.
Summationsmethoden und momentfolgen. i
Felix Hausdorff · 1921
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
Vladimir A Marčenko and Leonid Andreevich Pastur · 1967
Earlier work this paper cites.
A bound for the error in the normal approximation to the distribution of a sum of dependent random variables
Charles Stein · 1972
Earlier work this paper cites.
A stochastic estimator of the trace of the influence matrix for Laplacian smoothing splines
Michael F Hutchinson · 1990
Earlier work this paper cites.
Free random variables
Dan V Voiculescu, Ken J Dykema, and Alexandru Nica · 1992
Earlier work this paper cites.
Matrices, moments and quadrature
Gene H Golub and Gérard Meurant · 1994
Earlier work this paper cites.
Fast exact multiplication by the Hessian
Barak A Pearlmutter · 1994
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay · 2003
Earlier work this paper cites.
Eigenvalues of large sample covariance matrices of spiked population models
Jinho Baik and Jack W Silverstein · 2006
Earlier work this paper cites.
The Lanczos and conjugate gradient algorithms in finite precision arithmetic
Gérard Meurant and Zdeněk Strakoš · 2006
Earlier work this paper cites.
Numerical optimization
Jorge Nocedal and Stephen Wright · 2006
Earlier work this paper cites.
Convergence rate of expected spectral distributions of large random matrices part i: Wigner matrices
Zhi Dong Bai · 2008
Earlier work this paper cites.
Convex optimization
Stephen P. Boyd and Lieven Vandenberghe · 2009
Earlier work this paper cites.
Deep learning via Hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
On the marchenko-pastur and circular laws for some classes of random matrices with dependent entries
Radoslaw Adamczak · 2011
Earlier work this paper cites.
The eigenvalues and eigenvectors of finite, low rank perturbations of large random matrices
Florent Benaych-Georges and Raj Rao Nadakuditi · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Elements of information theory
Thomas M Cover and Joy A Thomas · 2012
Earlier work this paper cites.
Methods of proof in random matrix theory
Adina Roxana Feier · 2012
Earlier work this paper cites.
Matrix computations , volume 3
Gene H Golub and Charles F Van Loan · 2012
Earlier work this paper cites.
Semicircle law for a class of random matrices with dependent entries
Friedrich Götze, A Naumov, and A Tikhomirov · 2012
Earlier work this paper cites.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Earlier work this paper cites.
Training deep and recurrent networks with Hessian-free optimization
James Martens and Ilya Sutskever · 2012
Earlier work this paper cites.
A note on the marchenko-pastur law for a class of random matrices with dependent entries
Sean O’Rourke et al · 2012
Earlier work this paper cites.
Topics in random matrix theory , volume 132
Terence Tao · 2012
Earlier work this paper cites.
Distributions of angles in random packing on spheres
Tony Cai, Jianqing Fan, and Tiefeng Jiang · 2013
Earlier work this paper cites.
On bilinear forms based on the resolvent of large random matrices
Walid Hachem, Philippe Loubaton, Jamal Najim, and Pascal Vallet · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Cited alongside, same era.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Geometry of neural network loss surfaces via random matrix theory
Jeffrey Pennington and Yasaman Bahri · 2017
Later among the works it cites.
Empirical analysis of the Hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2017
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2017
Later among the works it cites.
Fast estimation of tr(f(a)) via stochastic Lanczos quadrature
Shashanka Ubaru, Jie Chen, and Yousef Saad · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Cited alongside, same era.
Limit theorems for two classes of random matrices with dependent entries
F Gotze, AA Naumov, and AN Tikhomirov · 2015
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Optimizing neural networks with Kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Improved bounds on sample size for implicit matrix trace estimators
Farbod Roosta-Khorasani and Uri Ascher · 2015
Cited alongside, same era.
Rotational invariant estimator for general noisy matrices
Joël Bun, Romain Allez, Jean-Philippe Bouchaud, Marc Potters, et al · 2016
Cited alongside, same era.
Entropy-SGD: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Later among the works it cites.
Deep Frank-Wolfe for neural network optimization
Leonard Berrada, Andrew Zisserman, and M Pawan Kumar · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
Gpytorch: Blackbox matrix-matrix Gaussian process inference with GPU acceleration
Jacob Gardner, Geoff Pleiss, Kilian Q Weinberger, David Bindel, and Andrew G Wilson · 2018
Later among the works it cites.
On the computational inefficiency of large batch sizes for stochastic gradient descent
Noah Golmant, Nikita Vemuri, Zhewei Yao, Vladimir Feinberg, Amir Gholami, Kai Rothauge, Michael W Mahoney, and Joseph Gonzalez · 2018
Later among the works it cites.
The deep learning limit: are negative neural network eigenvalues just noise?
Diego Granziol, Timur Garipov, Stefan Zohren, Dmitry Vetrov, Stephen Roberts, and Andrew Gordon Wilson · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
On the relation between the sharpest directions of DNN loss and the SGD step length
Stanisław Jastrzębski, Zachary Kenton, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Later among the works it cites.
The full spectrum of deep net Hessians at scale: Dynamics with sample size
Vardan Papyan · 2018
Later among the works it cites.
The spectrum of the fisher information matrix of a single-hidden-layer neural network
Jeffrey Pennington and Pratik Worah · 2018
Later among the works it cites.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Later among the works it cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, et al · 2018
Later among the works it cites.
Exact and inexact subsampled newton methods for optimization
Raghu Bollapragada, Richard H Byrd, and Jorge Nocedal · 2019
Later among the works it cites.
An investigation into neural net optimization via Hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
Diego Granziol, Xingchen Wan, Timur Garipov, Dmitry Vetrov, and Stephen Roberts · 2019
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Later among the works it cites.
Limitations of the empirical Fisher approximation
Frederik Kunstner, Lukas Balles, and Philipp Hennig · 2019
Later among the works it cites.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Closest in time.
Traces of class/cross-class structure pervade deep learning spectra
Vardan Papyan · 2020
Closest in time.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2020
Closest in time.
Understanding deep learning (still) requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2021
Closest in time.