Fetching the paper…
Reading the bibliography…
Sparse Neural Networks (NNs) can match the generalization of dense NNs using a fraction of the compute/storage for inference, and also have the potential to enable efficient training.
Multidimensional scaling by optimizing goodness of fit to a nonmetric hypothesis
Kruskal, J. 1964 · 1964
Earlier work this paper cites.
Backpropagation applied to handwritten zip code recognition
LeCun, Y.; Boser, B.; Denker, J. S.; Henderson, D.; Howard, R. E.; Hubbard, W.; and Jackel, L. D. 1989 · 1989
Earlier work this paper cites.
Skeletonization: A Technique for Trimming the Fat from a Network via Relevance Assessment
Mozer, M. C.; and Smolensky, P. 1989 · 1989
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Glorot, X.; and Bengio, Y. 2010 · 2010
Earlier work this paper cites.
Are wider nets better given the same number of parameters?
Golubeva, A.; Neyshabur, B.; and Gur-Ari, G. 2021 · 2010
Earlier work this paper cites.
The NumPy Array: A Structure for Efficient Numerical Computation
van der Walt, S.; Colbert, S. C.; and Varoquaux, G. 2011 · 2011
Earlier work this paper cites.
Qualitatively characterizing neural network optimization problems
Goodfellow, I. J.; Vinyals, O.; and Saxe, A. M. 2015 · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
Han, S.; Pool, J.; Tran, J.; and Dally, W. 2015 · 2015
Earlier work this paper cites.
Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Ioffe, S.; and Szegedy, C. 2015 · 2015
Earlier work this paper cites.
ImageNet Large Scale Visual Recognition Challenge
Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; Berg, A. C.; and Fei-Fei, L. 2015 · 2015
Earlier work this paper cites.
Bayesian compression for deep learning
Louizos, C.; Ullrich, K.; and Welling, M. 2017 · 2017
Earlier work this paper cites.
Variational Dropout Sparsifies Deep Neural Networks
Molchanov, D.; Ashukha, A.; and Vetrov, D. 2017 · 2017
Earlier work this paper cites.
Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
Sagun, L.; Evci, U.; Güney, V. U.; Dauphin, Y.; and Bottou, L. 2017 · 2017
Earlier work this paper cites.
Essentially No Barriers in Neural Network Energy Landscape
Draxler, F.; Veschgini, K.; Salmhofer, M.; and Hamprecht, F. A. 2018 · 2018
Earlier work this paper cites.
Loss Surfaces, M Connectivity, and Fast Ensembling of DNNs
Garipov, T.; Izmailov, P.; Podoprikhin, D.; Vetrov, D. P.; and Wilson, A. G. 2018 · 2018
Earlier work this paper cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C.; Mocanu, E.; Stone, P.; Nguyen, P. H.; Gibescu, M.; and Liotta, A. 2018 · 2018
Earlier work this paper cites.
Deep residual learning for image steganalysis
Wu, S.; Zhong, S.; and Liu, Y. 2018 · 2018
Earlier work this paper cites.
Dynamical isometry and a mean field theory of cnns: How to train 10,000-layer vanilla convolutional neural networks
Xiao, L.; Bahri, Y.; Sohl-Dickstein, J.; Schoenholz, S. S.; and Pennington, J. 2018 · 2018
Earlier work this paper cites.
To Prune, or Not to Prune: Exploring the Efficacy of Pruning for Model Compression
Zhu, M.; and Gupta, S. 2018 · 2018
Cited alongside, same era.
Sparse Networks from Scratch: Faster Training without Losing Performance
Dettmers, T.; and Zettlemoyer, L. 2019 · 2019
Cited alongside, same era.
The Difficulty of Training Sparse Neural Networks
Evci, U.; Pedregosa, F.; Gomez, A. N.; and Elsen, E. 2019 · 2019
Cited alongside, same era.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J.; and Carbin, M. 2019 · 2019
Cited alongside, same era.
Stabilizing the Lottery Ticket Hypothesis
Frankle, J.; Dziugaite, G. K.; Roy, D. M.; and Carbin, M. 2019 · 2019
Cited alongside, same era.
The State of Sparsity in Deep Neural Networks
Gale, T.; Elsen, E.; and Hooker, S. 2019 · 2019
The Early Phase of Neural Network Training
Frankle, J.; Schwab, D. J.; and Morcos, A. S. 2020 · 2020
Closest in time.
Soft Threshold Weight Reparameterization for Learnable Sparsity
Kusupati, A.; Ramanujan, V.; Somani, R.; Wortsman, M.; Jain, P.; Kakade, S.; and Farhadi, A. 2020 · 2020
Closest in time.
A Signal Propagation Perspective for Pruning Neural Networks at Initialization
Lee, N.; Ajanthan, T.; Gould, S.; and Torr, P. H. S. 2020 · 2020
Closest in time.
The large learning rate phase of deep learning: the catapult mechanism
Lewkowycz, A.; Bahri, Y.; Dyer, E.; Sohl-Dickstein, J.; and Gur-Ari, G. 2020 · 2020
Closest in time.
Towards Practical Lottery Ticket Hypothesis for Adversarial Training
Li, B.; Wang, S.; Jia, Y.; Lu, Y.; Zhong, Z.; Carin, L.; and Jana, S. 2020 · 2020
Closest in time.
Comparing Rewinding and Fine-tuning in Neural Network Pruning
Renda, A.; Frankle, J.; and Carbin, M. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
Ghorbani, B.; Krishnan, S.; and Xiao, Y. 2019 · 2019
Cited alongside, same era.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Lee, N.; Ajanthan, T.; and Torr, P. H. S. 2019 · 2019
Cited alongside, same era.
Rethinking the Value of Network Pruning
Liu, Z.; Sun, M.; Zhou, T.; Huang, G.; and Darrell, T. 2019 · 2019
Cited alongside, same era.
One ticket to win them all: generalizing lottery ticket initializations across datasets and optimizers
Morcos, A.; Yu, H.; Paganini, M.; and Tian, Y. 2019 · 2019
Cited alongside, same era.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H.; and Wang, X. 2019 · 2019
Cited alongside, same era.
Measurements of Three-Level Hierarchical Structure in the Outliers in the Spectrum of Deepnet Hessians
Papyan, V. 2019 · 2019
Cited alongside, same era.
Closest in time.
On the Transferability of Winning Tickets in Non-Natural Image Datasets
Sabatelli, M.; Kestemont, M.; and Geurts, P. 2020 · 2020
Closest in time.
Pruning neural networks without any data by iteratively conserving synaptic flow
Tanaka, H.; Kunin, D.; Yamins, D. L. K.; and Ganguli, S. 2020 · 2020
Closest in time.
Calibrate and Prune: Improving Reliability of Lottery Tickets Through Prediction Calibration
Venkatesh, B.; Thiagarajan, J. J.; Thopalli, K.; and Sattigeri, P. 2020 · 2020
Closest in time.
Picking Winning Tickets Before Training by Preserving Gradient Flow
Wang, C.; Zhang, G.; and Grosse, R. 2020 · 2020
Closest in time.
Chasing Sparsity in Vision Transformers: An End-to-End Exploration
Chen, T.; Cheng, Y.; Gan, Z.; Yuan, L.; Zhang, L.; and Wang, Z. 2021 · 2021
Closest in time.
Truly Sparse Neural Networks at Scale
Curci, S.; Mocanu, D. C.; and Pechenizkiy, M. 2021 · 2021
Closest in time.
Towards Structured Dynamic Sparse Pre-Training of BERT
Dietrich, A.; Gressmann, F.; Orr, D.; Chelombiev, I.; Justus, D.; and Luschi, C. 2021 · 2021
Closest in time.
A Gradient Flow Framework For Analyzing Network Pruning
Lubana, E. S.; and Dick, R. 2021 · 2021
Closest in time.
Towards Understanding Iterative Magnitude Pruning: Why Lottery Tickets Win
Maene, J.; Li, M.; and Moens, M. 2021 · 2021
Closest in time.
Price, I.; and Tanner, J. 2021 · 2021
Closest in time.
Dynamic Sparse Training for Deep Reinforcement Learning
Sokar, G.; Mocanu, E.; Mocanu, D. C.; Pechenizkiy, M.; and Stone, P. 2021 · 2021
Closest in time.
Keep the Gradients Flowing: Using Gradient Flow to Study Sparse Network Optimization
Tessera, K.; Hooker, S.; and Rosman, B. 2021 · 2021
Closest in time.