Fetching the paper…
Reading the bibliography…
Sparse training has received an upsurging interest in machine learning due to its tantalizing saving potential for the entire training process as well as inference.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Diversity networks: Neural network compression using determinantal point processes
Z. Mariet and S. Sra · 2016
Earlier work this paper cites.
Gpu kernels for block-sparse weights
S. Gray, A. Radford, and D. P. Kingma · 2017
Earlier work this paper cites.
Channel pruning for accelerating very deep neural networks
Y. He, X. Zhang, and J. Sun · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. Diamos, H. Jun, H. Kianinejad, M. Patwary, M. Ali, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Learning efficient convolutional networks through network slimming
Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang · 2017
Earlier work this paper cites.
Exploring sparsity in recurrent neural networks
S. Narang, E. Elsen, G. Diamos, and S. Sengupta · 2017
Earlier work this paper cites.
Deep rewiring: Training very sparse deep networks
G. Bellec, D. Kappel, W. Maass, and R. Legenstein · 2018
Earlier work this paper cites.
Snip: Single-shot network pruning based on connection sensitivity
N. Lee, T. Ajanthan, and P. H. Torr · 2018
Earlier work this paper cites.
D. C. Mocanu, E. Mocanu, P. Stone, P. H. Nguyen, M. Gibescu, and A. Liotta · 2018
Earlier work this paper cites.
To prune, or not to prune: Exploring the efficacy of pruning for model compression, 2018
M. H. Zhu and S. Gupta · 2018
Earlier work this paper cites.
Legr: Filter pruning via learned global ranking
T.-W. Chin, R. Ding, C. Zhang, and D. Marculescu · 2019
Earlier work this paper cites.
Sparse networks from scratch: Faster training without losing performance
T. Dettmers and L. Zettlemoyer · 2019
Earlier work this paper cites.
Network pruning via transformable architecture search
X. Dong and Y. Yang · 2019
Earlier work this paper cites.
The state of sparsity in deep neural networks
T. Gale, E. Elsen, and S. Hooker · 2019
Earlier work this paper cites.
Estimation of energy consumption in machine learning
E. García-Martín, C. F. Rodrigues, G. Riley, and H. Grahn · 2019
Earlier work this paper cites.
Filter pruning via geometric median for deep convolutional neural networks acceleration
Y. He, P. Liu, Z. Wang, Z. Hu, and Y. Yang · 2019
Earlier work this paper cites.
Dissecting the graphcore ipu architecture via microbenchmarking
Z. Jia, B. Tillman, M. Maggioni, and D. P. Scarpazza · 2019
Earlier work this paper cites.
Parameter efficient training of deep convolutional neural networks by dynamic sparse reparameterization
H. Mostafa and X. Wang · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
E. Strubell, A. Ganesh, and A. McCallum · 2019
Earlier work this paper cites.
Gate decorator: Global filter pruning method for accelerating deep convolutional neural networks
Z. You, K. Yan, J. Ye, M. Ma, and P. Wang · 2019
Cited alongside, same era.
Autoslim: Towards one-shot architecture search for channel numbers
J. Yu and T. Huang · 2019
Cited alongside, same era.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
Fast sparse convnets
E. Elsen, M. Dukhan, T. Gale, and K. Simonyan · 2020
Sparse training via boosting pruning plasticity with neuroregeneration
S. Liu, T. Chen, X. Chen, Z. Atashgahi, L. Yin, H. Kou, L. Shen, M. Pechenizkiy, Z. Wang, and D. C. Mocanu · 2021
Later among the works it cites.
Sparse evolutionary deep learning with over one million artificial neurons on commodity hardware
S. Liu, D. C. Mocanu, A. R. R. Matavalam, Y. Pei, and M. Pechenizkiy · 2021
Later among the works it cites.
Do we actually need dense over-parameterization? in-time over-parameterization in sparse training
S. Liu, L. Yin, D. C. Mocanu, and M. Pechenizkiy · 2021
Later among the works it cites.
Carbon emissions and large neural network training
D. Patterson, J. Gonzalez, Q. Le, C. Liang, L.-M. Munguia, D. Rothchild, D. So, M. Texier, and J. Dean · 2021
Later among the works it cites.
Ac/dc: Alternating compressed/decompressed training of deep neural networks
A. Peste, E. Iofinova, A. Vladu, and D. Alistarh · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
U. Evci, T. Gale, J. Menick, P. S. Castro, and E. Elsen · 2020
Cited alongside, same era.
Sparse gpu kernels for deep learning
T. Gale, M. Zaharia, C. Young, and E. Elsen · 2020
Cited alongside, same era.
Top-kast: Top-k always sparse training
S. Jayakumar, R. Pascanu, J. Rae, S. Osindero, and E. Elsen · 2020
Cited alongside, same era.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
M. Kurtz, J. Kopinsky, R. Gelashvili, A. Matveev, J. Carr, M. Goin, W. Leiserson, S. Moore, B. Nell, N. Shavit, and D. Alistarh · 2020
Cited alongside, same era.
Hrank: Filter pruning using high-rank feature map
M. Lin, R. Ji, Y. Wang, Y. Zhang, B. Zhang, Y. Tian, and L. Shao · 2020
Cited alongside, same era.
Nvidia a100 tensor core gpu architecture
Nvidia · 2020
Cited alongside, same era.
Channel permutations for n:m sparsity
J. Pool and C. Yu · 2021
Later among the works it cites.
Powerpropagation: A sparsity inducing weight reparameterisation
J. Schwarz, S. Jayakumar, R. Pascanu, P. E. Latham, and Y. Teh · 2021
Later among the works it cites.
Locally free weight sharing for network width search
X. Su, S. You, T. Huang, F. Wang, C. Qian, C. Zhang, and C. Xu · 2021
Later among the works it cites.
Chip: Channel independence-based pruning for compact neural networks
Y. Sui, M. Yin, Y. Xie, H. Phan, S. Aliari Zonouz, and B. Yuan · 2021
Later among the works it cites.
Mest: Accurate and fast memory-economic sparse training framework on the edge
G. Yuan, X. Ma, W. Niu, Z. Li, Z. Kong, N. Liu, Y. Gong, Z. Zhan, C. He, Q. Jin, et al · 2021
Later among the works it cites.
Learning n:m fine-grained structured sparse neural networks from scratch
A. Zhou, Y. Ma, J. Zhu, J. Liu, Z. Zhang, K. Yuan, W. Sun, and H. Li · 2021
Later among the works it cites.
Coarsening the granularity: Towards structurally sparse lottery tickets
T. Chen, X. Chen, X. Ma, Y. Wang, and Z. Wang · 2022
Later among the works it cites.
Chex: Channel exploration for cnn model compression
Z. Hou, M. Qin, F. Sun, X. Ma, K. Yuan, Y. Xu, Y.-K. Chen, R. Jin, Y. Xie, and S.-Y. Kung · 2022
Later among the works it cites.
How well do sparse imagenet models transfer?
E. Iofinova, A. Peste, M. Kurtz, and D. Alistarh · 2022
Later among the works it cites.
Training your sparse neural network better with any mask
A. K. Jaiswal, H. Ma, T. Chen, Y. Ding, and Z. Wang · 2022
Later among the works it cites.
Exposing and exploiting fine-grained block structures for fast and accurate sparse training
P. Jiang, L. Hu, and S. Song · 2022
Later among the works it cites.
More convnets in the 2020s: Scaling up kernels beyond 51x51 using sparsity
S. Liu, T. Chen, X. Chen, X. Chen, Q. Xiao, B. Wu, M. Pechenizkiy, D. Mocanu, and Z. Wang · 2022
Later among the works it cites.
S. Liu, T. Chen, X. Chen, L. Shen, D. C. Mocanu, Z. Wang, and M. Pechenizkiy · 2022
Later among the works it cites.
Swin transformer v2: Scaling up capacity and resolution
Z. Liu, H. Hu, Y. Lin, Z. Yao, Z. Xie, Y. Wei, J. Ning, Y. Cao, Z. Zhang, L. Dong, et al · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al · 2022
Later among the works it cites.