Fetching the paper…
Reading the bibliography…
The slow convergence rate and pathological curvature issues of first-order gradient methods for training deep neural networks, initiated an ongoing effort for developing faster $\mathit{second}$-$\mathit{order}$ optimization algorithms beyond SGD, without compromising the generalization error.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff · 1952
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Sue Becker and Yann Le Cun · 1988
Earlier work this paper cites.
A new algorithm for minimizing convex functions over convex sets
Pravin M Vaidya · 1989
Earlier work this paper cites.
Speeding-up linear programming using fast matrix multiplication
Pravin M Vaidya · 1989
Earlier work this paper cites.
Nearly-linear time algorithms for graph partitioning, graph sparsification, and solving linear systems
Daniel A. Spielman and Shang-Hua Teng · 2004
Earlier work this paper cites.
Group-theoretic algorithms for matrix multiplication
Henry Cohn, Robert Kleinberg, Balazs Szegedy, and Christopher Umans · 2005
Earlier work this paper cites.
Approximate nearest neighbors and the fast johnson-lindenstrauss transform
Nir Ailon and Bernard Chazelle · 2006
Earlier work this paper cites.
Sampling algorithms for l2 regression and applications
Petros Drineas, Michael W Mahoney, and Shan Muthukrishnan · 2006
Earlier work this paper cites.
Improved approximation algorithms for large matrices via random projections
Tamás Sarlós · 2006
Earlier work this paper cites.
Fast linear algebra is stable
James Demmel, Ioana Dumitriu, and Olga Holtz · 2007
Earlier work this paper cites.
A fast randomized algorithm for overdetermined linear least-squares regression
Vladimir Rokhlin and Mark Tygert · 2008
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Improved analysis of the subsampled randomized hadamard transform
Joel A Tropp · 2011
Earlier work this paper cites.
Fast approximation of matrix coherence and statistical leverage
Petros Drineas, Malik Magdon-Ismail, Michael W Mahoney, and David P Woodruff · 2012
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
A simple, combinatorial algorithm for solving sdd systems in nearly-linear time
Jonathan A Kelner, Lorenzo Orecchia, Aaron Sidford, and Zeyuan Allen Zhu · 2013
Earlier work this paper cites.
Faster ridge regression via the subsampled randomized hadamard transform
Yichao Lu, Paramveer Dhillon, Dean P Foster, and Lyle Ungar · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Optimal cur matrix decompositions
Christos Boutsidis and David P Woodruff · 2014
Earlier work this paper cites.
Powers of tensors and fast matrix multiplication
François Le Gall · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
David P. Woodruff · 2014
Earlier work this paper cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Christos Boutsidis, David P Woodruff, and Peilin Zhong · 2016
Earlier work this paper cites.
A kronecker-factored approximate fisher matrix for convolution layers
Roger Grosse and James Martens · 2016
Earlier work this paper cites.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Benjamin Recht, and Yoram Singer · 2016
Cited alongside, same era.
A linearly-convergent stochastic l-bfgs algorithm
Philipp Moritz, Robert Nishihara, and Michael Jordan · 2016
Cited alongside, same era.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Cited alongside, same era.
Distributed low rank approximation of implicit functions of a matrix
David P Woodruff and Peilin Zhong · 2016
Cited alongside, same era.
Sub-sampled newton methods with non-uniform sampling
Peng Xu, Jiyan Yang, Fred Roosta, Christopher Ré, and Michael W Mahoney · 2016
Cited alongside, same era.
Second-order stochastic optimization for machine learning in linear time
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
Exact and inexact subsampled newton methods for optimization
Raghu Bollapragada, Richard H Byrd, and Jorge Nocedal · 2019
Later among the works it cites.
Robust and sample optimal algorithms for psd low-rank approximation
Ainesh Bakshi, Nadiia Chepurko, and David P Woodruff · 2019
Later among the works it cites.
Learning two layer rectified neural networks in polynomial time
Ainesh Bakshi, Rajesh Jayaram, and David P Woodruff · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Cited alongside, same era.
Random fourier features for kernel ridge regression: Approximation bounds and statistical guarantees
Haim Avron, Michael Kapralov, Cameron Musco, Christopher Musco, Ameya Velingker, and Amir Zandieh · 2017
Cited alongside, same era.
Practical gauss-newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Cited alongside, same era.
Optimality of the johnson-lindenstrauss lemma
Kasper Green Larsen and Jelani Nelson · 2017
Cited alongside, same era.
Convergence analysis of two-layer neural networks with ReLU activation
Yuanzhi Li and Yang Yuan · 2017
Cited alongside, same era.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J Wainwright · 2017
Cited alongside, same era.
Scalable trust-region method for deep reinforcement learning using kronecker-factored approximation
Yuhuai Wu, Elman Mansimov, Roger B Grosse, Shun Liao, and Jimmy Ba · 2017
Cited alongside, same era.
Later among the works it cites.
A gram-gauss-newton method learning overparameterized deep neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Later among the works it cites.
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Later among the works it cites.
Optimal sketching for kronecker product regression and low rank approximation
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2019
Later among the works it cites.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Later among the works it cites.
Towards moderate overparameterization: global convergence guarantees for training shallow neural networks
Samet Oymak and Mahdi Soltanolkotabi · 2019
Later among the works it cites.
Matrix Theory : Optimization, Concentration and Algorithms
Zhao Song · 2019
Later among the works it cites.
Efficient symmetric norm regression via linear sketching
Zhao Song, Ruosong Wang, Lin Yang, Hongyang Zhang, and Peilin Zhong · 2019
Later among the works it cites.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Zhao Song and Xin Yang · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Later among the works it cites.
Fast convergence of natural gradient descent for over-parameterized neural networks
Guodong Zhang, James Martens, and Roger B Grosse · 2019
Later among the works it cites.
Second order optimization made practical
Rohan Anil, Vineet Gupta, Tomer Koren, Kevin Regan, and Yoram Singer · 2020
Closest in time.
Network size and weights size for memorization with two-layers neural networks
Sébastien Bubeck, Ronen Eldan, Yin Tat Lee, and Dan Mikulincer · 2020
Closest in time.
Regularized weighted low rank approximation
Frank Ban, David P. Woodruff, and Richard Zhang · 2020
Closest in time.
Memorizing gaussians with no over-parameterizaion via gradient decent on neural networks
Amit Daniely · 2020
Closest in time.
A faster interior point method for semidefinite programming
Haotian Jiang, Tarun Kathuria, Yin Tat Lee, Swati Padmanabhan, and Zhao Song · 2020
Closest in time.
An improved cutting plane method for convex optimization, convex-concave games and its applications
Haotian Jiang, Yin Tat Lee, Zhao Song, and Sam Chiu-wai Wong · 2020
Closest in time.
Faster dynamic matrix inverse for faster lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2020
Closest in time.
Ziwei Ji and Matus Telgarsky · 2020
Closest in time.
Automatic differentiation of sketched regression
Hang Liao, Barak A. Pearlmutter, Vamsi K. Potluru, and David P. Woodruff · 2020
Closest in time.
Near input sparsity time kernel embeddings via adaptive sampling
David P. Woodruff and Amir Zandieh · 2020
Closest in time.