Fetching the paper…
Reading the bibliography…
Sparsity has become one of the promising methods to compress and accelerate Deep Neural Networks (DNNs).
Optimal Brain Damage
Yann LeCun, John Denker, and Sara Solla · 1989
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Learning both Weights and Connections for Efficient Neural Network
Song Han, Jeff Pool, John Tran, and William Dally · 2015
Earlier work this paper cites.
Dynamic Network Surgery for Efficient DNNs
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Earlier work this paper cites.
Deep Residual Learning for Image Recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Pruning Filters for Efficient Convnets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Earlier work this paper cites.
Pruning Convolutional Neural Networks for Resource Efficient Inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2016
Earlier work this paper cites.
Learning Structured Sparsity in Deep Neural Networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
https://www.statmt.org/wmt17/ , 2017
EMNLP 2017 Second Conference on Machine Translation (WMT17) · 2017
Earlier work this paper cites.
Deep Rewiring: Training Very Sparse Deep Networks
Guillaume Bellec, David Kappel, Wolfgang Maass, and Robert Legenstein · 2017
Earlier work this paper cites.
GPU Kernels for Block-sparse Weights
Scott Gray, Alec Radford, and Diederik P Kingma · 2017
Earlier work this paper cites.
Channel Pruning for Accelerating Very Deep Neural Networks
Yihui He, Xiangyu Zhang, and Jian Sun · 2017
Earlier work this paper cites.
Exploring Sparsity in Recurrent Neural Networks
Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta · 2017
Earlier work this paper cites.
Block-sparse Recurrent Neural Networks
Sharan Narang, Eric Undersander, and Gregory Diamos · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
To Prune, or not to Prune: Exploring the Efficacy of Pruning for Model Compression
Michael Zhu and Suyog Gupta · 2017
Earlier work this paper cites.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle and Michael Carbin · 2018
Earlier work this paper cites.
Efficient Neural Audio Synthesis
Nal Kalchbrenner, Erich Elsen, Karen Simonyan, Seb Noury, Norman Casagrande, Edward Lockhart, Florian Stimberg, Aaron Oord, Sander Dieleman, and Koray Kavukcuoglu · 2018
Earlier work this paper cites.
SNIP: Single-shot Network Pruning based on Connection Sensitivity
Namhoon Lee, Thalaiyasingam Ajanthan, and Philip HS Torr · 2018
Earlier work this paper cites.
Rethinking the Value of Network Pruning
Zhuang Liu, Mingjie Sun, Tinghui Zhou, Gao Huang, and Trevor Darrell · 2018
Earlier work this paper cites.
Scalable Training of Artificial Neural Networks with Adaptive Sparse Connectivity Inspired by Network Science
Decebal Constantin Mocanu, Elena Mocanu, Peter Stone, Phuong H Nguyen, Madeleine Gibescu, and Antonio Liotta · 2018
Cited alongside, same era.
SQuantizer: Simultaneous Learning for both Sparse and Low-precision Neural Networks
Mi Sun Park, Xiaofan Xu, and Cormac Brick · 2018
Cited alongside, same era.
ReleQ: An Automatic Reinforcement Learning Approach for Deep Quantization of Neural Networks
Amir Yazdanbakhsh, Ahmed T Elthakeb, Prannoy Pilligundla, F Mireshghallah, and Hadi Esmaeilzadeh · 2018
Cited alongside, same era.
Sparse Networks from Scratch: Faster Training without Losing Performance
Tim Dettmers and Luke Zettlemoyer · 2019
Cited alongside, same era.
Pruning Neural Networks at Initialization: Why are we Missing the Mark?
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M Roy, and Michael Carbin · 2020
Later among the works it cites.
Campfire: Compressible, Regularization-free, Structured Sparse Training for Hardware Accelerators
Noah Gamboa, Kais Kudrolli, Anand Dhoot, and Ardavan Pedram · 2020
Later among the works it cites.
Soft Threshold Weight Reparameterization for Learnable Sparsity
Aditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman, Prateek Jain, Sham Kakade, and Ali Farhadi · 2020
Later among the works it cites.
Comparing Rewinding and Fine-tuning in Neural Network Pruning
Alex Renda, Jonathan Frankle, and Michael Carbin · 2020
Later among the works it cites.
Q-BERT: Hessian based Ultra Low Precision Quantization of BERT
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Utku Evci, Fabian Pedregosa, Aidan Gomez, and Erich Elsen · 2019
Cited alongside, same era.
The State of Sparsity in Deep Neural Networks
Trevor Gale, Erich Elsen, and Sara Hooker · 2019
Cited alongside, same era.
TinyBERT: Distilling BERT for Natural Language Understanding
Xiaoqi Jiao, Yichun Yin, Lifeng Shang, Xin Jiang, Xiao Chen, Linlin Li, Fang Wang, and Qun Liu · 2019
Cited alongside, same era.
Accelerator-aware Pruning for Convolutional Neural Networks
Hyeong-Ju Kang · 2019
Cited alongside, same era.
Importance Estimation for Neural Network Pruning
Pavlo Molchanov, Arun Mallya, Stephen Tyree, Iuri Frosio, and Jan Kautz · 2019
Cited alongside, same era.
Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2019
Cited alongside, same era.
DistilBERT, a Distilled version of BERT: Smaller, Faster, Cheaper and Lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Discovering Neural Wirings
Mitchell Wortsman, Ali Farhadi, and Mohammad Rastegari · 2019
Cited alongside, same era.
Later among the works it cites.
MobileBERT: A Compact Task-agnostic BERT for Resource-Limited Devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Later among the works it cites.
PCNN: Pattern-based Fine-grained Regular Pruning Towards Optimizing CNN Accelerators
Zhanhong Tan, Jiebo Song, Xiaolong Ma, Sia-Huat Tan, Hongyang Chen, Yuanqing Miao, Yifu Wu, Shaokai Ye, Yanzhi Wang, Dehui Li, and Kaisheng Ma · 2020
Later among the works it cites.
MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
Wenhui Wang, Furu Wei, Li Dong, Hangbo Bao, Nan Yang, and Ming Zhou · 2020
Later among the works it cites.
TernaryBERT: Distillation-aware Ultra-low Bit BERT
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
https://images.nvidia.com/aem-dam/en-zz/Solutions/data-center/nvidia-ampere-architecture-whitepaper.pdf , 2021
NVIDIA Ampere Architecture Whitepaper · 2021
Later among the works it cites.
https://github.com/NVIDIA/apex/tree/master/apex/contrib/sparsity , 2021
NVIDIA ASP (Automatic Sparsity) · 2021
Later among the works it cites.
I-BERT: Integer-only BERT Quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2021
Later among the works it cites.
Filter Sketch for Network Pruning
Mingbao Lin, Liujuan Cao, Shaojie Li, Qixiang Ye, Yonghong Tian, Jianzhuang Liu, Qi Tian, and Rongrong Ji · 2021
Later among the works it cites.
Accelerating Sparse Deep Neural Networks
Asit Mishra, Jorge Albericio Latorre, Jeff Pool, Darko Stosic, Dusan Stosic, Ganesh Venkatesh, Chong Yu, and Paulius Micikevicius · 2021
Later among the works it cites.
Channel Permutations for N:M Sparsity
Jeff Pool and Chong Yu · 2021
Later among the works it cites.
Extending Sparse Tensor Accelerators to Support Multiple Compression Formats
Eric Qin, Geonhwa Jeong, William Won, Sheng-Chun Kao, Hyoukjun Kwon, Sudarshan Srinivasan, Dipankar Das, Gordon E Moon, Sivasankaran Rajamanickam, and Tushar Krishna · 2021
Later among the works it cites.
Pruning by Explaining: A Novel Criterion for Deep Neural Network Pruning
Seul-Ki Yeom, Philipp Seegerer, Sebastian Lapuschkin, Alexander Binder, Simon Wiedemann, Klaus-Robert Müller, and Wojciech Samek · 2021
Later among the works it cites.
Learning N:M Fine-grained Structured Sparse Neural Networks from Scratch
Aojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu, Zhijie Zhang, Kun Yuan, Wenxiu Sun, and Hongsheng Li · 2021
Later among the works it cites.
Non-Structured DNN Weight Pruning – Is It Beneficial in Any Platform?
Xiaolong Ma, Sheng Lin, Shaokai Ye, Zhezhi He, Linfeng Zhang, Geng Yuan, Sia Huat Tan, Zhengang Li, Deliang Fan, Xuehai Qian, Xue Lin, Kaisheng Ma, and Yanzhi Wang · 2022
Closest in time.
OPT: Open Pre-trained Transformer Language Models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Closest in time.