Pruning neural networks without any data by iteratively conserving synaptic flow
Hidenori Tanaka, Daniel Kunin, Daniel L Yamins, and Surya Ganguli · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Difan Zou, Yuan Cao, Dongruo Zhou, and Quanquan Gu · 2020
Later among the works it cites.
Only train once: A one-shot neural network training and pruning framework
Tianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang, Guanyi Wang, Zhihui Zhu, Luming Liang, Yixin Shi, Sheng Yi, and Xiao Tu · 2021
Later among the works it cites.
Linearized two-layers neural networks in high dimension
BEHROOZ GHORBANI, SONG MEI, THEODOR MISIAKIEWICZ, and ANDREA MONTANARI · 2021
Later among the works it cites.
Sparse is enough in scaling transformers
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Later among the works it cites.
The unreasonable effectiveness of random pruning: Return of the most naive baseline for sparse training
Shiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen, Decebal Constantin Mocanu, Zhangyang Wang, and Mykola Pechenizkiy · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Zhao Song, Shuo Yang, and Ruizhe Zhang · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li, Chong You, Jeremias Sulam, and Qing Qu · 2021
Later among the works it cites.
A sublinear adversarial training algorithm
Original
Yeqi Gao, Lianke Qin, Zhao Song, and Yitan Wang · 2022
Later among the works it cites.
Training overparametrized neural networks in sublinear time
Original
Hang Hu, Zhao Song, Omri Weinstein, and Danyang Zhuo · 2022
Later among the works it cites.
On the convergence of shallow neural network training with randomly masked neurons
Fangshuo Liao and Anastasios Kyrillidis · 2022
Later among the works it cites.
The lazy neuron phenomenon: On emergence of activation sparsity in transformers
Zonglin Li, Chong You, Srinadh Bhojanapalli, Daliang Li, Ankit Singh Rawat, Sashank J Reddi, Ke Ye, Felix Chern, Felix Yu, Ruiqi Guo, et al · 2022
Later among the works it cites.
Bounding the width of neural networks via coupled initialization a worst case analysis
Alexander Munteanu, Simon Omlor, Zhao Song, and David Woodruff · 2022
Later among the works it cites.
Bypass exponential time preprocessing: Fast neural network training via weight-data correlation preprocessing
Josh Alman, Zhao Song, Ruizhe Zhang, and Danyang Zhuo · 2024
Closest in time.
Training multi-layer over-parametrized neural network in subquadratic time
Zhao Song, Lichen Zhang, and Ruizhe Zhang · 2024
Closest in time.