Fetching the paper…
Reading the bibliography…
In deep learning, fine-grained N:M sparsity reduces the data footprint and bandwidth of a General Matrix multiply (GEMM) up to x2, and doubles throughput by skipping computation of zero values.
Accelerating cnn training by pruning activation gradients
Xucheng Ye, P. Dai, J. Luo, X. Guo, Y. Qi, Jianlei Yang, and Yiran Chen · 1908
Earlier work this paper cites.
Loss aware post-training quantization
Yury Nahshan, Brian Chmiel, Chaim Baskin, Evgenii Zheltonozhskii, Ron Banner, Alex M. Bronstein, and Avi Mendelson · 1911
Earlier work this paper cites.
Pruning versus clipping in neural networks
S. A. Janowsky · 1989
Earlier work this paper cites.
Sparse weight activation training
Md Aamir Raihan and Tor M. Aamodt · 2001
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and W. Lin · 2017
Earlier work this paper cites.
meprop: Sparsified back propagation for accelerated deep learning with reduced overfitting
Xu Sun, Xuancheng Ren, Shuming Ma, and Houfeng Wang · 2017
Earlier work this paper cites.
Scalable methods for 8-bit training of neural networks
Ron Banner, Itay Hubara, Elad Hoffer, and Daniel Soudry · 2018
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
The lottery ticket hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2018
Cited alongside, same era.
Rethinking the value of network pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2018
Cited alongside, same era.
Learned step size quantization
Steven K Esser, Jeffrey L McKinstry, Deepika Bablani, Rathinakumar Appuswamy, and Dharmendra S Modha · 2019
Cited alongside, same era.
Learning to quantize deep networks by optimizing quantization intervals with task loss
Sangil Jung, Changyong Son, Seohyung Lee, Jinwoo Son, Jae-Joon Han, Youngjun Kwak, Sung Ju Hwang, and Changkyu Choi · 2019
Cited alongside, same era.
Hybrid 8-bit floating point (hfp8) training and inference for deep neural networks
Xiao Sun, Jungwook Choi, Chia-Yu Chen, Naigang Wang, Swagath Venkataramani, Vijayalakshmi Srinivasan, Xiaodong Cui, Wei Zhang, and Kailash Gopalakrishnan · 2019
Nxmtransformer: Semi-structured sparsification for natural language understanding via admm
Connor Holmes, Minjia Zhang, Yuxiong He, and Bo Wu · 2021
Later among the works it cites.
Accelerated sparse neural training: A provable and efficient method to find n: M transposable masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Seffi Naor, and Daniel Soudry · 2021
Later among the works it cites.
Sparse is enough in scaling transformers
Sebastian Jaszczur, Aakanksha Chowdhery, Afroz Mohiuddin, Lukasz Kaiser, Wojciech Gajewski, Henryk Michalewski, and Jonni Kanerva · 2021
Later among the works it cites.
Accelerating sparse deep neural networks
Asit K. Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Later among the works it cites.
Channel permutations for n:m sparsity
Jeff Pool and Chong Yu · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Rigging the lottery: Making all tickets winners
Utku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro, and Erich Elsen · 2020
Cited alongside, same era.
Inducing and exploiting activation sparsity for fast inference on deep neural networks
Mark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev, John Carr, Michael Goin, William Leiserson, Sage Moore, Nir Shavit, and Dan Alistarh · 2020
Cited alongside, same era.
Fpnew: An open-source multiformat floating-point unit architecture for energy-proportional transprecision computing
Stefan Mach, Fabian Schuiki, Florian Zaruba, and Luca Benini · 2020
Cited alongside, same era.
a100 tensor core gpu architecture
Nvidia · 2020
Cited alongside, same era.
Comparing rewinding and fine-tuning in neural network pruning
A. Renda, Jonathan Frankle, and Michael Carbin · 2020
Cited alongside, same era.
Differentiable joint pruning and quantization for hardware efficiency
Ying Wang, Yadong Lu, and Tijmen Blankevoort · 2020
Cited alongside, same era.
URL https://docs.graphcore.ai/projects/tensorflow1-user-guide/en/latest/tensorflow/rand_and_fp.html
Greaphcore ipu
Cited in the paper.
Search spaces for neural model training
Darko Stosic and Dusan Stosic · 2021
Later among the works it cites.
Dominosearch: Find layer-wise fine-grained n:m sparse schemes from dense neural networks
Wei Sun, Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven, Andrew Nelson, Hongsheng Li, and Henk Corporaal · 2021
Later among the works it cites.
Learning n:m fine-grained structures sparse neural networks from scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li · 2021
Later among the works it cites.
Stochastic rounding: implementation, error analysis and applications
Matteo Croci, Massimiliano Fasi, Nicholas Higham, Theo Mary, and Mantas Mikaitis · 2022
Closest in time.
Accelerating dnn training with structured data gradient pruning
Bradley McDanel, Helia Dinh, and J. R. Magallanes · 2022
Closest in time.
Towards fully sparse training: Information restoration with spatial similarity
Xu Weixiang, Xiangyu He, Ke Cheng, Peisong Wang, and Jian Cheng · 2022
Closest in time.