Fetching the paper…
Reading the bibliography…
A rising trend in theoretical deep learning is to understand why deep learning works through Neural Tangent Kernel (NTK) [jgh18], a kernel method that is equivalent to using gradient descent to train a multi-layer infinitely-wide neural network.
Bemerkungen zur theorie der beschränkten bilinearformen mit unendlich vielen veränderlichen
Jssai Schur · 1911
Earlier work this paper cites.
On a modification of chebyshev’s inequality and of the error formula of laplace
Sergei Bernstein · 1924
Earlier work this paper cites.
Probability inequalities for sums of bounded random variables
Wassily Hoeffding · 1963
Earlier work this paper cites.
Gaussian processes: inequalities, small ball probabilities and applications
Wenbo V Li and Q-M Shao · 2001
Earlier work this paper cites.
Low rank approximation and regression in input sparsity time
Kenneth L. Clarkson and David P. Woodruff · 2013
Earlier work this paper cites.
Low-distortion subspace embeddings in input-sparsity time and applications to robust linear regression
Xiangrui Meng and Michael W Mahoney · 2013
Earlier work this paper cites.
Osnap: Faster numerical linear algebra algorithms via sparser subspace embeddings
Jelani Nelson and Huy L Nguyên · 2013
Earlier work this paper cites.
Delving deep into rectifiers: Surpassing human-level performance on imagenet classification
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Optimal principal component analysis in distributed and streaming models
Christos Boutsidis, David P Woodruff, and Peilin Zhong · 2016
Earlier work this paper cites.
Amit Daniely, Roy Frostig, and Yoram Singer · 2016
Earlier work this paper cites.
Weighted low rank approximations with provable guarantees
Ilya Razenshteyn, Zhao Song, and David P Woodruff · 2016
Earlier work this paper cites.
Distributed low rank approximation of implicit functions of a matrix
David P Woodruff and Peilin Zhong · 2016
Earlier work this paper cites.
Faster kernel ridge regression using sketching and preconditioning
Haim Avron, Kenneth L Clarkson, and David P Woodruff · 2017
Earlier work this paper cites.
Modular multitask reinforcement learning with policy sketches
Jacob Andreas, Dan Klein, and Sergey Levine · 2017
Earlier work this paper cites.
Globally optimal gradient descent for a convnet with gaussian inputs
Alon Brutzkus and Amir Globerson · 2017
Earlier work this paper cites.
Sgd learns the conjugate kernel class of the network
Amit Daniely · 2017
Earlier work this paper cites.
Protein interface prediction using graph convolutional networks
Alex M Fout · 2017
Earlier work this paper cites.
Learning one-hidden-layer neural networks with landscape design
Rong Ge, Jason D Lee, and Tengyu Ma · 2017
Earlier work this paper cites.
Near optimal sketching of low-rank tensor regression
Jarvis Haupt, Xingguo Li, and David P Woodruff · 2017
Earlier work this paper cites.
Knowledge transfer for out-of-knowledge-base entities: A graph neural network approach
Takuo Hamaguchi, Hidekazu Oiwa, Masashi Shimbo, and Yuji Matsumoto · 2017
Earlier work this paper cites.
Convergence analysis of two-layer neural networks with relu activation
Yuanzhi Li and Yang Yuan · 2017
Earlier work this paper cites.
Learning relus via gradient descent
Mahdi Soltanolkotabi · 2017
Earlier work this paper cites.
Low rank approximation with entrywise ℓ 1 \ell_{1} -norm error
Zhao Song, David P Woodruff, and Peilin Zhong · 2017
Earlier work this paper cites.
An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis
Yuandong Tian · 2017
Earlier work this paper cites.
Predicting multicellular function through multi-layer tissue networks
Marinka Zitnik and Jure Leskovec · 2017
Earlier work this paper cites.
Learning non-overlapping convolutional neural networks with multiple kernels
Kai Zhong, Zhao Song, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Recovery guarantees for one-hidden-layer neural networks
Kai Zhong, Zhao Song, Prateek Jain, Peter L Bartlett, and Inderjit S Dhillon · 2017
Earlier work this paper cites.
Subspace embedding and linear regression with orlicz norm
Alexandr Andoni, Chengyu Lin, Ying Sheng, Peilin Zhong, and Ruiqi Zhong · 2018
Earlier work this paper cites.
Gradient descent learns one-hidden-layer cnn: Don’t be afraid of spurious local minima
Simon Du, Jason Lee, Yuandong Tian, Aarti Singh, and Barnabas Poczos · 2018
Earlier work this paper cites.
Sega: Variance reduction via gradient sketching
Filip Hanzely, Konstantin Mishchenko, and Peter Richtárik · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Earlier work this paper cites.
Bourgan: generative networks with metric embeddings
Chang Xiao, Peilin Zhong, and Changxi Zheng · 2018
Earlier work this paper cites.
Graph convolutional neural networks for web-scale recommender systems
Rex Ying, Ruining He, Kaifeng Chen, Pong Eksombatchai, William L Hamilton, and Jure Leskovec · 2018
Earlier work this paper cites.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
On exact computation with an infinitely wide neural net
Sanjeev Arora, Simon S Du, Wei Hu, Zhiyuan Li, Ruslan Salakhutdinov, and Ruosong Wang · 2019
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
Learning two layer rectified neural networks in polynomial time
Ainesh Bakshi, Rajesh Jayaram, and David P Woodruff · 2019
Cited alongside, same era.
Almost linear time density level set estimation via dbscan
Hossein Esfandiari, Vahab Mirrokni, and Peilin Zhong · 2021
Later among the works it cites.
Streaming and distributed algorithms for robust column subset selection
Shuli Jiang, Dennis Li, Irene Mengze Li, Arvind V Mahankali, and David Woodruff · 2021
Later among the works it cites.
Optimal sketching for trace estimation
Shuli Jiang, Hai Pham, David Woodruff, and Richard Zhang · 2021
Later among the works it cites.
Faster dynamic matrix inverse for faster lps
Shunhua Jiang, Zhao Song, Omri Weinstein, and Hengjie Zhang · 2021
Later among the works it cites.
Fast sketching of polynomial kernels of polynomial degree
Zhao Song, David Woodruff, Zheng Yu, and Lichen Zhang · 2021
Later among the works it cites.
Oblivious sketching-based central path method for solving linear programming problems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lenaic Chizat and Francis Bach · 2019
Cited alongside, same era.
A near-optimal algorithm for approximating the john ellipsoid
Michael B Cohen, Ben Cousins, Yin Tat Lee, and Xin Yang · 2019
Cited alongside, same era.
Graph neural tangent kernel: Fusing graph neural networks with graph kernels
Simon S Du, Kangcheng Hou, Russ R Salakhutdinov, Barnabas Poczos, Ruosong Wang, and Keyulu Xu · 2019
Cited alongside, same era.
Optimal sketching for kronecker product regression and low rank approximation
Huaian Diao, Rajesh Jayaram, Zhao Song, Wen Sun, and David Woodruff · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Cited alongside, same era.
Solving empirical risk minimization in the current matrix multiplication time
Yin Tat Lee, Zhao Song, and Qiuyi Zhang · 2019
Cited alongside, same era.
Zhao Song and Zheng Yu · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2021
Later among the works it cites.
Breaking the linear iteration cost barrier for some well-known conditional gradient methods using maxip data-structures
Zhaozhuo Xu, Zhao Song, and Anshumali Shrivastava · 2021
Later among the works it cites.
Memrein: Rein the domain shift for cross-domain few-shot learning
Yi Xu, Lichen Wang, Yizhou Wang, Can Qin, Yulun Zhang, and Yun Fu · 2021
Later among the works it cites.
Scaling neural tangent kernels via sketching and random features
Amir Zandieh, Insu Han, Haim Avron, Neta Shoham, Chaewon Kim, and Jinwoo Shin · 2021
Later among the works it cites.
Josh Alman, Jiehao Liang, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2022
Later among the works it cites.
A sublinear adversarial training algorithm
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Later among the works it cites.
A faster small treewidth sdp solver
Yuzhou Gu and Zhao Song · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Later among the works it cites.
Adore: Differentially oblivious relational database operators
Lianke Qin, Rajesh Jayaram, Elaine Shi, Zhao Song, Danyang Zhuo, and Shumo Chu · 2022
Later among the works it cites.
Adaptive and dynamic multi-resolution hashing for pairwise summations
Lianke Qin, Aravind Reddy, Zhao Song, Zhaozhuo Xu, and Danyang Zhuo · 2022
Later among the works it cites.
Robust semi-supervised domain adaptation against noisy labels
Can Qin, Yizhou Wang, and Yun Fu · 2022
Later among the works it cites.
Faster algorithm for structured john ellipsoid computation
Zhao Song, Xin Yang, Yuanyuan Yang, and Tianyi Zhou · 2022
Later among the works it cites.
Making reconstruction-based method great again for video anomaly detection
Yizhou Wang, Can Qin, Yue Bai, Yi Xu, Xu Ma, and Yun Fu · 2022
Later among the works it cites.
Adaptive trajectory prediction via transferable gnn
Yi Xu, Lichen Wang, Yizhou Wang, and Yun Fu · 2022
Later among the works it cites.
Speeding up optimizations via data structures: Faster search, sample and maintenance
Lichen Zhang · 2022
Later among the works it cites.
Fast attention requires bounded entries
Josh Alman and Zhao Song · 2023
Closest in time.
Attention scheme inspired softmax regression
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
An improved sample complexity for rank-1 matrix sensing
Yichuan Deng, Zhihang Li, and Zhao Song · 2023
Closest in time.
An iterative algorithm for rescaled hyperbolic functions regression
Yeqi Gao, Zhao Song, and Junze Yin · 2023
Closest in time.
Low rank matrix completion via robust alternating minimization in nearly linear time
Yuzhou Gu, Zhao Song, Junze Yin, and Lichen Zhang · 2023
Closest in time.
A nearly-linear time algorithm for structured support vector machines
Yuzhou Gu, Zhao Song, and Lichen Zhang · 2023
Closest in time.
Solving regularized exp, cosh and sinh regression problems
Zhihang Li, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
A general algorithm for solving rank-one matrix sensing
Lianke Qin, Zhao Song, and Ruizhe Zhang · 2023
Closest in time.
An online and unified algorithm for projection matrix vector multiplication with application to empirical risk minimization
Lianke Qin, Zhao Song, Lichen Zhang, and Danyang Zhuo · 2023
Closest in time.
A tale of two efficient value iteration algorithms for solving linear mdps with large action space
Anshumali Shrivastava, Zhao Song, and Zhaozhuo Xu · 2023
Closest in time.
Ritwik Sinha, Zhao Song, and Tianyi Zhou · 2023
Closest in time.
Towards explainable visual anomaly detection
Yizhou Wang, Dongliang Guo, Sheng Li, and Yun Fu · 2023
Closest in time.