Fetching the paper…
Reading the bibliography…
We consider the problem of training a multi-layer over-parametrized neural network to minimize the empirical risk induced by a loss function.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
J. Schur · 1911
Earlier work this paper cites.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
A measure of asymptotic efficiency for tests of a hypothesis based on the sum of observations
Herman Chernoff · 1952
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Speeding-up linear programming using fast matrix multiplication
Pravin M Vaidya · 1989
Earlier work this paper cites.
Improved approximation algorithms for large matrices via random projections
Tamas Sarlos · 2006
Earlier work this paper cites.
Faster approximate lossy generalized flow via interior point algorithms
Samuel I Daitch and Daniel A Spielman · 2008
Earlier work this paper cites.
Smallest singular value of a random rectangular matrix
Mark Rudelson and Roman Vershynin · 2009
Earlier work this paper cites.
Restricted isometries for partial random circulant matrices
Holger Rauhut, Justin Romberg, and Joel A Tropp · 2012
Earlier work this paper cites.
Multiplying matrices faster than coppersmith-winograd
Virginia Vassilevska Williams · 2012
Earlier work this paper cites.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
Sparsity lower bounds for dimensionality reducing maps
Jelani Nelson and Huy L Nguyen · 2013
Earlier work this paper cites.
Subspace embeddings for the polynomial kernel
Haim Avron, Huy L. Nguyen, and David P. Woodruff · 2014
Earlier work this paper cites.
Suprema of chaos processes and the restricted isometry property
Felix Krahmer, Shahar Mendelson, and Holger Rauhut · 2014
Earlier work this paper cites.
Path finding methods for linear programming: Solving linear programs in õ(sqrt(rank)) iterations and faster algorithms for maximum flow
Yin Tat Lee and Aaron Sidford · 2014
Earlier work this paper cites.
Sketching as a tool for numerical linear algebra
David P. Woodruff · 2014
Earlier work this paper cites.
A faster cutting plane method and its implications for combinatorial and convex optimization
Yin Tat Lee, Aaron Sidford, and Sam Chiu-wai Wong · 2015
Earlier work this paper cites.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Christos Boutsidis, David P. Woodruff, and Peilin Zhong · 2016
Earlier work this paper cites.
Second-order stochastic optimization for machine learning in linear time
Naman Agarwal, Brian Bullins, and Elad Hazan · 2017
Earlier work this paper cites.
Practical Gauss-Newton optimisation for deep learning
Aleksandar Botev, Hippolyt Ritter, and David Barber · 2017
Earlier work this paper cites.
Newton sketch: A near linear-time optimization algorithm with linear-quadratic convergence
Mert Pilanci and Martin J. Wainwright · 2017
Earlier work this paper cites.
Low rank approximation with entrywise l1-norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Earlier work this paper cites.
Exact natural gradient in deep linear networks and its application to the nonlinear case
Alberto Bernacchia, Mate Lengyel, and Guillaume Hennequin · 2018
Earlier work this paper cites.
Improved rectangular matrix multiplication using powers of the coppersmith-winograd tensor
François Le Gall and Florent Urrutia · 2018
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Yuan Cao and Quanquan Gu · 2019
Cited alongside, same era.
Gram-gauss-newton method: Learning overparameterized neural networks for regression problems
Tianle Cai, Ruiqi Gao, Jikai Hou, Siyu Chen, Dong Wang, Di He, Zhihua Zhang, and Liwei Wang · 2019
Cited alongside, same era.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Closest in time.
How much over-parameterization is sufficient to learn deep ReLU networks?
Zixiang Chen, Yuan Cao, Difan Zou, and Quanquan Gu · 2021
Closest in time.
{MONGOOSE}: A learnable {lsh} framework for efficient neural network training
Beidi Chen, Zichang Liu, Binghui Peng, Zhaozhuo Xu, Jonathan Lingjie Li, Tri Dao, Zhao Song, Anshumali Shrivastava, and Christopher Re · 2021
Closest in time.
Newton-less: Sparsification without trade-offs for the sketched newton update, 2021
Michał Dereziński, Jonathan Lacotte, Mert Pilanci, and Michael W. Mahoney · 2021
Closest in time.
Fl-ntk: A neural tangent kernel-based framework for federated learning convergence analysis
Baihe Huang, Xiaoxiao Li, Zhao Song, and Xin Yang · 2021
Closest in time.
Faster dynamic matrix inverse for faster lps
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Solving linear programs in the current matrix multiplication time
Michael B Cohen, Yin Tat Lee, and Zhao Song · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon S Du, Jason D Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Jaehoon Lee, Lechao Xiao, Samuel S. Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Efficient symmetric norm regression via linear sketching
Zhao Song, Ruosong Wang, Lin Yang, Hongyang Zhang, and Peilin Zhong · 2019
Cited alongside, same era.
Relative error tensor low rank approximation
Zhao Song, David P Woodruff, and Peilin Zhong · 2019
Cited alongside, same era.
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Closest in time.
Fast sketching of polynomial kernels of polynomial degree
Zhao Song, David P. Woodruff, Zheng Yu, and Lichen Zhang · 2021
Closest in time.
Oblivious sketching-based central path method for linear programming
Zhao Song and Zheng Yu · 2021
Closest in time.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Closest in time.
Adahessian: An adaptive second order optimizer for machine learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, and Michael Mahoney · 2021
Closest in time.
Fast distance oracles for any symmetric norm
Yichuan Deng, Zhao Song, Omri Weinstein, and Ruizhe Zhang · 2022
Closest in time.
Improved sliding window algorithms for clustering and coverage via bucketing-based sketches
Alessandro Epasto, Mohammad Mahdian, Vahab Mirrokni, and Peilin Zhong · 2022
Closest in time.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Closest in time.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Closest in time.
Solving sdp faster: A robust ipm framework and efficient implementation
Baihe Huang, Shunhua Jiang, Zhao Song, Runzhou Tao, and Ruizhe Zhang · 2022
Closest in time.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Closest in time.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Closest in time.
A faster interior-point method for sum-of-squares optimization
Shunhua Jiang, Bento Natura, and Omri Weinstein · 2022
Closest in time.
Last layer re-training is sufficient for robustness to spurious correlations, 2022
Polina Kirichenko, Pavel Izmailov, and Andrew Gordon Wilson · 2022
Closest in time.
Sparse fourier transform over lattices: A unified approach to signal reconstruction
Zhao Song, Baocheng Sun, Omri Weinstein, and Ruizhe Zhang · 2022
Closest in time.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2023
Closest in time.
Space-efficient interior point method, with applications to linear programming and maximum weight bipartite matching
S. Cliff Liu, Zhao Song, Hengjie Zhang, Lichen Zhang, and Tianyi Zhou · 2023
Closest in time.
Relu strikes back: Exploiting activation sparsity in large language models, 2023
Iman Mirzadeh, Keivan Alizadeh, Sachin Mehta, Carlo C Del Mundo, Oncel Tuzel, Golnoosh Samei, Mohammad Rastegari, and Mehrdad Farajtabar · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
Convergence and generalization of wide neural networks with large bias, 2023
Hongru Yang, Ziyu Jiang, Ruizhe Zhang, Zhangyang Wang, and Yingbin Liang · 2023
Closest in time.
Faster rectangular matrix multiplication by combination loss analysis
Francois Le Gall · 2024
Closest in time.
New bounds for matrix multiplication: from alpha to omega
Virginia Vassilevska Williams, Yinzhan Xu, Zixuan Xu, and Renfei Zhou · 2024
Closest in time.