Fetching the paper…
Reading the bibliography…
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels.
The state of sparsity in deep neural networks
Gale, T., Elsen, E., and Hooker, S. (2019) · 1902
Earlier work this paper cites.
The lottery ticket hypothesis at scale
Frankle, J., Dziugaite, G. K., Roy, D. M., and Carbin, M. (2019) · 1903
Earlier work this paper cites.
Generating long sequences with sparse transformers
Child, R., Gray, S., Radford, A., and Sutskever, I. (2019) · 1904
Earlier work this paper cites.
Deconstructing lottery tickets: Zeros, signs, and the supermask
Zhou, H., Lan, J., Liu, R., and Yosinski, J. (2019) · 1905
Earlier work this paper cites.
A back-propagation algorithm with optimal use of hidden units
Chauvin, Y. (1988) · 1988
Earlier work this paper cites.
Skeletonization: A technique for trimming the fat from a network via relevance assessment
Mozer, M. C. and Smolensky, P. (1988) · 1988
Earlier work this paper cites.
Optimal brain damage
LeCun, Y., Denker, J. S., and Solla, S. A. (1989) · 1989
Earlier work this paper cites.
A simple procedure for pruning back-propagation trained neural networks
Karnin, E. D. (1990) · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Hassibi, B. and Stork, D. G. (1992) · 1992
Earlier work this paper cites.
Structural learning with forgetting
Ishikawa, M. (1996) · 1996
Earlier work this paper cites.
Gradient-based learning applied to document recognition
LeCun, Y. (1998) · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Qian, N. (1999) · 1999
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L. (2009) · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A. and Hinton, G. (2009) · 2009
Earlier work this paper cites.
Connectivity-driven white matter scaling and folding in primate cerebral cortex
Herculano-Houzel, S., Mota, B., Wong, P., and Kaas, J. H. (2010) · 2010
Cited alongside, same era.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Simonyan, K., Vedaldi, A., and Zisserman, A. (2013) · 2013
Cited alongside, same era.
Striving for simplicity: The all convolutional net
Springenberg, J. T., Dosovitskiy, A., Brox, T., and Riedmiller, M. A. (2014) · 2014
Cited alongside, same era.
Visualizing and understanding convolutional networks
Zeiler, M. D. and Fergus, R. (2014) · 2014
Cited alongside, same era.
Learning both weights and connections for efficient neural network
Han, S., Pool, J., Tran, J., and Dally, W. (2015) · 2015
Cited alongside, same era.
Automatic differentiation in pytorch
Paszke, A., Gross, S., Chintala, S., Chanan, G., Yang, E., DeVito, Z., Lin, Z., Desmaison, A., Antiga, L., and Lerer, A. (2017) · 2017
Later among the works it cites.
Soft weight-sharing for neural network compression
Ullrich, K., Meeds, E., and Welling, M. (2017) · 2017
Later among the works it cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Later among the works it cites.
Deep rewiring: Training very sparse deep networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R. A. (2018) · 2018
Later among the works it cites.
“learning-compression” algorithms for neural net pruning
Carreira-Perpinán, M. A. and Idelbayev, Y. (2018) · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic network surgery for efficient dnns
Guo, Y., Yao, A., and Chen, Y. (2016) · 2016
Cited alongside, same era.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J. (2016) · 2016
Cited alongside, same era.
Nest: A neural network synthesis tool based on a grow-and-prune paradigm
Dai, X., Yin, H., and Jha, N. K. (2017) · 2017
Cited alongside, same era.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Dong, X., Chen, S., and Pan, S. J. (2017) · 2017
Cited alongside, same era.
Gpu kernels for block-sparse weights
Gray, S., Radford, A., and Kingma, D. P. (2017) · 2017
Cited alongside, same era.
Bayesian compression for deep learning
Louizos, C., Ullrich, K., and Welling, M. (2017) · 2017
Cited alongside, same era.
Variational dropout sparsifies deep neural networks
Molchanov, D., Ashukha, A., and Vetrov, D. P. (2017) · 2017
Cited alongside, same era.
Dai, X., Yin, H., and Jha, N. K. (2018) · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Later among the works it cites.
Learning sparse neural networks through l 0 l_{0} regularization
Louizos, C., Welling, M., and Kingma, D. P. (2018) · 2018
Later among the works it cites.
Scalable training of artificial neural networks with adaptive sparse connectivity inspired by network science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A. (2018) · 2018
Later among the works it cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Frankle, J. and Carbin, M. (2019) · 2019
Closest in time.
Snip: Single-shot network pruning based on connection sensitivity
Lee, N., Ajanthan, T., and Torr, P. H. S. (2019) · 2019
Closest in time.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
Mostafa, H. and Wang, X. (2019) · 2019
Closest in time.
To prune, or not to prune: Exploring the efficacy of pruning for model compression
Zhu, M. and Gupta, S. (2018) · 2019
Closest in time.