Fetching the paper…
Reading the bibliography…
N:M Structured sparsity has garnered significant interest as a result of relatively modest overhead and improved efficiency.
Optimal brain damage
LeCun, Y., Denker, J., and Solla, S · 1989
Earlier work this paper cites.
Occam’s Razor
Rasmussen, C. and Ghahramani, Z · 2001
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., and Fei-Fei, L · 2009
Earlier work this paper cites.
Learning Multiple Layers of Features from Tiny Images
Krizhevsky, A., Hinton, G., et al · 2009
Earlier work this paper cites.
Accelerating Stochastic Gradient Descent Using Predictive Variance Reduction
Johnson, R. and Zhang, T · 2013
Earlier work this paper cites.
Variance Reduction for Stochastic Gradient Optimization
Wang, C., Chen, X., Smola, A. J., and Xing, E. P · 2013
Earlier work this paper cites.
Dynamic Network Surgery for Efficient DNNs
Guo, Y., Yao, A., and Chen, Y · 2016
Earlier work this paper cites.
Pruning Filters for Efficient ConvNets
Li, H., Kadav, A., Durdanovic, I., Samet, H., and Graf, H. P · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Molchanov, P., Tyree, S., Karras, T., Aila, T., and Kautz, J · 2016
Earlier work this paper cites.
Learning Structured Sparsity in Deep Neural Networks
Wen, W., Wu, C., Wang, Y., Chen, Y., and Li, H · 2016
Earlier work this paper cites.
Findings of the 2017 Conference on Machine Translation (WMT17)
Bojar, O., Chatterjee, R., Federmann, C., Graham, Y., Haddow, B., Huang, S., Huck, M., Koehn, P., Liu, Q., Logacheva, V., Monz, C., Negri, M., Post, M., Rubino, R., Specia, L., and Turchi, M · 2017
Earlier work this paper cites.
GPU Kernels for Block-Sparse Weights
Gray, S., Radford, A., and Kingma, D. P · 2017
Earlier work this paper cites.
Channel Pruning for Accelerating Very Deep Neural Networks
He, Y., Zhang, X., and Sun, J · 2017
Earlier work this paper cites.
Scnn: An accelerator for compressed-sparse convolutional neural networks
Parashar, A., Rhu, M., Mukkara, A., Puglielli, A., Venkatesan, R., Khailany, B., Emer, J., Keckler, S. W., and Dally, W. J · 2017
Earlier work this paper cites.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
To Prune, or not to Prune: Exploring the Efficacy of Pruning for Model Compression
Zhu, M. and Gupta, S · 2017
Earlier work this paper cites.
Snapea: Predictive Early Activation for Reducing Computation in Deep Convolutional Neural Networks
Akhlaghi, V., Yazdanbakhsh, A., Samadi, K., Gupta, R. K., and Esmaeilzadeh, H · 2018
Earlier work this paper cites.
Deep Rewiring: Training very Sparse Deep Networks
Bellec, G., Kappel, D., Maass, W., and Legenstein, R · 2018
Earlier work this paper cites.
Rethinking the Value of Network Pruning
Liu, Z., Sun, M., Zhou, T., Huang, G., and Darrell, T · 2018
Earlier work this paper cites.
Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity Inspired by Network Science
Mocanu, D. C., Mocanu, E., Stone, P., Nguyen, P. H., Gibescu, M., and Liotta, A · 2018
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers
Child, R., Gray, S., Radford, A., and Sutskever, I · 2019
Earlier work this paper cites.
Sparse Networks from Scratch: Faster Training without Losing Performance
Dettmers, T. and Zettlemoyer, L · 2019
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Earlier work this paper cites.
The Difficulty of Training Sparse Neural Networks
Evci, U., Pedregosa, F., Gomez, A., and Elsen, E · 2019
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Frankle, J. and Carbin, M · 2019
Earlier work this paper cites.
The State of Sparsity in Deep Neural Networks
Gale, T., Elsen, E., and Hooker, S · 2019
Earlier work this paper cites.
Accelerator-Aware Pruning for Convolutional Neural Networks
Kang, H.-J · 2019
Cited alongside, same era.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Lee, N., Ajanthan, T., and Torr, P. H · 2019
Cited alongside, same era.
Importance Estimation for Neural Network Pruning
Molchanov, P., Mallya, A., Tyree, S., Frosio, I., and Kautz, J · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2019
Cited alongside, same era.
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Filter Sketch for Network Pruning
Lin, M., Cao, L., Li, S., Ye, Q., Tian, Y., Liu, J., Tian, Q., and Ji, R · 2021
Later among the works it cites.
Non-Structured DNN Weight Pruning – Is It Beneficial in Any Platform?
Ma, X., Lin, S., Ye, S., He, Z., Zhang, L., Yuan, G., Tan, S. H., Li, Z., Fan, D., Qian, X., Lin, X., Ma, K., and Wang, Y · 2021
Later among the works it cites.
Accelerating Sparse Deep Neural Networks
Mishra, A., Latorre, J. A., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., and Micikevicius, P · 2021
Later among the works it cites.
Channel Permutations for N: M Sparsity
Pool, J. and Yu, C · 2021
Later among the works it cites.
Extending Sparse Tensor Accelerators to Support Multiple Compression Formats
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2019
Cited alongside, same era.
PyTorch Image Models
Wightman, R · 2019
Cited alongside, same era.
Wortsman, M., Farhadi, A., and Rastegari, M · 2019
Cited alongside, same era.
Balanced Sparsity for Efficient DNN Inference on GPU
Yao, Z., Cao, S., Xiao, W., Zhang, C., and Nie, L · 2019
Cited alongside, same era.
Zafrir, O., Boudoukh, G., Izsak, P., and Wasserblat, M · 2019
Cited alongside, same era.
Zhu, M., Zhang, T., Gu, Z., and Xie, Y · 2019
Cited alongside, same era.
Longformer: The Long-Document Transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Cited alongside, same era.
Elsen, E., Dukhan, M., Gale, T., and Simonyan, K · 2020
Cited alongside, same era.
Qin, E., Jeong, G., Won, W., Kao, S.-C., Kwon, H., Srinivasan, S., Das, D., Moon, G. E., Rajamanickam, S., and Krishna, T · 2021
Later among the works it cites.
Efficient Content-based Sparse Attention with Routing Transformers
Roy, A., Saffar, M., Vaswani, A., and Grangier, D · 2021
Later among the works it cites.
1-bit Adam: Communication Efficient Large-scale Training with Adam’s Convergence Speed
Tang, H., Gan, S., Awan, A. A., Rajbhandari, S., Li, C., Lian, X., Liu, J., Zhang, C., and He, Y · 2021
Later among the works it cites.
Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
Yeom, S.-K., Seegerer, P., Lapuschkin, S., Binder, A., Wiedemann, S., Müller, K.-R., and Samek, W · 2021
Later among the works it cites.
Learning N:M Fine-grained Structured Sparse Neural Networks from Scratch
Zhou, A., Ma, Y., Zhu, J., Liu, J., Zhang, Z., Yuan, K., Sun, W., and Li, H · 2021
Later among the works it cites.
An algorithm–hardware co-optimized framework for accelerating n:m sparse transformers
Fang, C., Zhou, A., and Wang, Z · 2022
Later among the works it cites.
1-bit LAMB: Communication Efficient Large-scale Large-batch Training with LAMB’s Convergence Speed
Li, C., Awan, A. A., Tang, H., Rajbhandari, S., and He, Y · 2022
Later among the works it cites.
S2ta: Exploiting structured sparsity for energy-efficient mobile cnn acceleration
Liu, Z., Whatmough, P. N., Zhu, Y., and Mattina, M · 2022
Later among the works it cites.
Maximizing communication Efficiency for Large-scale Training via 0/1 Adam
Lu, Y., Li, C., Zhang, M., De Sa, C., and He, Y · 2022
Later among the works it cites.
Enabling flexibility for sparse tensor acceleration via heterogeneity, 2022
Qin, E., Garg, R., Bambhaniya, A., Pellauer, M., Parashar, A., Rajamanickam, S., Hao, C., and Krishna, T · 2022
Later among the works it cites.
Efficient Transformers: A Survey
Tay, Y., Dehghani, M., Bahri, D., and Metzler, D · 2022
Later among the works it cites.
Rethinking Weight Decay for Efficient Neural Network Pruning
Tessier, H., Gripon, V., Léonardon, M., Arzel, M., Hannagan, T., and Bertrand, D · 2022
Later among the works it cites.
Emergent Abilities of Large Language Models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W · 2022
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models
Anil, R., Borgeaud, S., Wu, Y., Alayrac, J.-B., Yu, J., Soricut, R., Schalkwyk, J., Dai, A. M., Hauth, A., et al · 2023
Later among the works it cites.
Accelerating attention based models via hw-sw co-design using fine-grained sparsification
Bambhaniya, A. R., Yazdanbakhsh, A., Subramanian, S., and Krishna, T · 2023
Later among the works it cites.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot , 2023
Frantar, E. and Alistarh, D · 2023
Later among the works it cites.
Vegeta: Vertically-integrated extensions for sparse/dense gemm tile acceleration on cpus
Jeong, G., Damani, S., Bambhaniya, A. R., Qin, E., Hughes, C. J., Subramoney, S., Kim, H., and Krishna, T · 2023
Later among the works it cites.
STEP: Learning N:M Structured Sparsity Masks from Scratch with Precondition
Lu, Y., Agrawal, S., Subramanian, S., Rybakov, O., Sa, C. D., and Yazdanbakhsh, A · 2023
Later among the works it cites.
GPT-4 Technical Report, 2023
OpenAI · 2023
Later among the works it cites.
BitSET: Bit-Serial Early Termination for Computation Reduction in Convolutional Neural Networks
Pan, Y., Yu, J., Lukefahr, A., Das, R., and Mahlke, S · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.