To understand deep learning we need to understand kernel learning
Mikhail Belkin, Siyuan Ma, and Soumik Mandal · 2018
Later among the works it cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Later among the works it cites.
A faster subquadratic algorithm for finding outlier correlations
Matti Karppa, Petteri Kaski, and Jukka Kohonen · 2018
Later among the works it cites.
Quantum machine learning in chemical compound space
O Anatole Von Lilienfeld · 2018
Later among the works it cites.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Later among the works it cites.
Oracle inequalities for sparse additive quantile regression in reproducing kernel hilbert space
Shaogao Lv, Huazhen Lin, Heng Lian, and Jian Huang · 2018
Later among the works it cites.
Faster all-pairs shortest paths via circuit complexity
R Ryan Williams · 2018
Later among the works it cites.
An illuminating algorithm for the light bulb problem
Original
Josh Alman · 2019
Later among the works it cites.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
On the convergence rate of training recurrent neural networks
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Later among the works it cites.
Matrix rigidity and the croot-lev-pach lemma
Zeev Dvir and Benjamin L Edelman · 2019
Later among the works it cites.
Fourier and circulant matrices are not rigid
Zeev Dvir and Allen Liu · 2019
Later among the works it cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2019
Later among the works it cites.
A log-sobolev inequality for the multislice, with applications
Yuval Filmus, Ryan O’Donnell, and Xinyu Wu · 2019
Later among the works it cites.
Reformer: The efficient transformer
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya · 2019
Later among the works it cites.
(Nearly) Sample-optimal sparse Fourier transform in any dimension; RIPless and Filterless
Original
Vasileios Nakos, Zhao Song, and Zhengyu Wang · 2019
Later among the works it cites.
Spectral graph theory and its applications, 2018
Daniel Spielman · 2019
Later among the works it cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Original
Zhao Song and Xin Yang · 2019
Later among the works it cites.
Algorithms and hardness for linear algebra on geometric graphs
Josh Alman, Timothy Chu, Aaron Schild, and Zhao Song · 2020
Closest in time.
Rethinking attention with performers
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamas Sarlos, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, et al · 2020
Closest in time.
Exact computation of a manifold metric, via lipschitz embeddings and shortest paths on a graph
Timothy Chu, Gary L. Miller, and Donald Sheehy · 2020
Closest in time.
A robust multi-dimensional sparse fourier transform in the continuous setting
Original
Yaonan Jin, Daogao Liu, and Zhao Song · 2020
Closest in time.
Transformers
Ria Kulshrestha · 2020
Closest in time.
Generalized leverage score sampling for neural networks
Jason D Lee, Ruoqi Shen, Zhao Song, Mengdi Wang, and Zheng Yu · 2020
Closest in time.
Linformer: Self-attention with linear complexity
Original
Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma · 2020
Closest in time.
Training (overparametrized) neural networks in near-linear time
Jan van den Brand, Binghui Peng, Zhao Song, and Omri Weinstein · 2021
Closest in time.
Algorithmic foundations for the diffraction limit
Original
Sitan Chen and Ankur Moitra · 2021
Closest in time.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
Original
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Closest in time.