Fetching the paper…
Reading the bibliography…
It has been observed in practice that applying pruning-at-initialization methods to neural networks and training the sparsified networks can not only retain the testing performance of the original dense models, but also sometimes even slightly boost the generalization performance.
The state of sparsity in deep neural networks
Gale, T · 1902
Earlier work this paper cites.
Quadratic suffices for over-parametrization via matrix chernoff bound
Song, Z · 1906
Earlier work this paper cites.
Optimal brain damage
LeCun, Y · 1989
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Towards understanding ensemble, knowledge distillation and self-distillation in deep learning
Allen-Zhu, Z · 2012
Earlier work this paper cites.
The mnist database of handwritten digit images for machine learning research [best of the web]
Deng, L · 2012
Earlier work this paper cites.
Diversity networks: Neural network compression using determinantal point processes
Mariet, Z · 2015
Earlier work this paper cites.
Deep residual learning for image recognition
He, K · 2016
Earlier work this paper cites.
A topological insight into restricted boltzmann machines
Mocanu, D. C · 2016
Earlier work this paper cites.
Wide residual networks
Zagoruyko, S · 2016
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
He, Y · 2017
Earlier work this paper cites.
An entropy-based pruning method for cnn compression
Luo, J.-H · 2017
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Molchanov, P · 2017
Earlier work this paper cites.
On the global convergence of gradient descent for over-parameterized models using optimal transport
Chizat, L · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Jacot, A · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C · 2018
Earlier work this paper cites.
Deep expander networks: Efficient deep networks from graph theory
Prabhu, A · 2018
Earlier work this paper cites.
Neural networks as interacting particle systems: Asymptotic convexity of the loss landscape and universal scaling of the approximation error
Rotskoff, G. M · 2018
Earlier work this paper cites.
A mean field view of the landscape of two-layers neural networks
Song, M · 2018
Earlier work this paper cites.
Network compression using correlation analysis of layer responses
Suau, X · 2018
Earlier work this paper cites.
A convergence theory for deep learning via over-parameterization
Allen-Zhu, Z · 2019
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Arora, S · 2019
Cited alongside, same era.
Generalization bounds of stochastic gradient descent for wide and deep neural networks
Cao, Y · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Chizat, L · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Du, S · 2019
Cited alongside, same era.
Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow relu networks
Ji, Z · 2019
Cited alongside, same era.
Sanity-checking pruning methods: Random tickets can win the jackpot
Su, J · 2020
Later among the works it cites.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H · 2020
Later among the works it cites.
Deephoyer: Learning sparser neural network with differentiable scale-invariant sparsity measures
Yang, H · 2020
Later among the works it cites.
Good subnetworks provably exist: Pruning via greedy forward selection
Ye, M · 2020
Later among the works it cites.
Gradient descent optimizes over-parameterized deep relu networks
Zou, D · 2020
Later among the works it cites.
Feature purification: How adversarial training performs robust deep learning
Allen-Zhu, Z · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Radix-net: Structured sparse matrices for deep neural networks
Kepner, J · 2019
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
Lee, J · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H · 2019
Cited alongside, same era.
Picking winning tickets before training by preserving gradient flow
Wang, C · 2019
Cited alongside, same era.
Regularization matters: Generalization and optimization of neural nets vs their induced kernel
Wei, C · 2019
Cited alongside, same era.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H · 2019
Cited alongside, same era.
Audio lottery: Speech recognition made ultra-lightweight, noise-robust, and transferable
Ding, S · 2021
Later among the works it cites.
Modeling from features: a mean-field framework for over-parameterized deep neural networks
Fang, C · 2021
Later among the works it cites.
Ac/dc: Alternating compressed/decompressed training of deep neural networks
Peste, A · 2021
Later among the works it cites.
A theoretical analysis on feature learning in neural networks: Emergence from inputs and advantage over fixed features
Shi, Z · 2021
Later among the works it cites.
Does preprocessing help training over-parameterized neural networks?
Song, Z · 2021
Later among the works it cites.
Finding everything within random binary networks
Sreenivasan, K · 2021
Later among the works it cites.
Dynamic regularization on activation sparsity for neural network efficiency improvement
Yang, Q · 2021
Later among the works it cites.
Why lottery ticket wins? a theoretical perspective of sample complexity on sparse neural networks
Zhang, S · 2021
Later among the works it cites.
A geometric analysis of neural collapse with unconstrained features
Zhu, Z · 2021
Later among the works it cites.
Understanding the generalization of adam in learning neural networks with proper regularization
Zou, D · 2021
Later among the works it cites.
Benign overfitting in two-layer convolutional neural networks
Cao, Y · 2022
Later among the works it cites.
Sparse double descent: Where network pruning aggravates overfitting
He, Z · 2022
Later among the works it cites.
Rare gems: Finding lottery tickets at initialization
Sreenivasan, K · 2022
Later among the works it cites.
Feature selection with gradient descent on two-layer networks in low-rotation regimes
Telgarsky, M · 2022
Later among the works it cites.
Zhou, J · 2022
Later among the works it cites.