Fetching the paper…
Reading the bibliography…
This paper presents the philosophy, design and feature-set of Neural Network Distiller, an open-source Python package for DNN compression research.
Model selection and estimation in regression with grouped variables
Ming Yuan and Yi Lin · 2006
Earlier work this paper cites.
Restructuring of deep neural network acoustic models with singular value decomposition
J. Xue, J. Li, and Y. Gong · 2013
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey E. Hinton, Oriol Vinyals, and Jeffrey Dean · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural networks
Song Han, Jeff Pool, John Tran, and William J. Dally · 2015
Earlier work this paper cites.
MXNet: A flexible and efficient machine learning library for heterogeneous distributed systems
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
EIE: Efficient inference engine on compressed deep neural network
Song Han, Xingyu Liu, Huizi Mao, Jing Pu, Ardavan Pedram, Mark A. Horowitz, and William J. Dally · 2016
Earlier work this paper cites.
BranchyNet: Fast inference via early exiting from deep neural networks
Surat Teerapittayanon, Bradley McDanel, and H.T. Kung · 2016
Earlier work this paper cites.
Google’s Neural Machine Translation System: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Gregory S. Corrado, Macduff Hughes, and Jeffrey Dean · 2016
Earlier work this paper cites.
Learning structured sparsity in deep neural networks
Wei Wen, Chunpeng Wu, Yandan Wang, Yiran Chen, and Hai Li · 2016
Earlier work this paper cites.
Network Trimming: A data-driven neuron pruning approach towards efficient deep architectures
Hengyuan Hu, Rui Peng, Yu-Wing Tai, and Chi-Keung Tang · 2016
Earlier work this paper cites.
Pruning convolutional neural networks for resource efficient inference
Pavlo Molchanov, Stephen Tyree, Tero Karras, Timo Aila, and Jan Kautz · 2016
Earlier work this paper cites.
Dynamic network surgery for efficient DNNs
Yiwen Guo, Anbang Yao, and Yurong Chen · 2016
Earlier work this paper cites.
Pruning Filters for Efficient ConvNets
Hao Li, Asim Kadav, Igor Durdanovic, Hanan Samet, and Hans Peter Graf · 2016
Earlier work this paper cites.
DoReFa-Net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Zekun Ni, Xinyu Zhou, He Wen, Yuxin Wu, and Yuheng Zou · 2016
Earlier work this paper cites.
Speed/accuracy trade-offs for modern convolutional object detectors
Jonathan Huang, Vivek Rathod, Chen Sun, Menglong Zhu, Anoop Korattikara, Alireza Fathi, Ian Fischer, Zbigniew Wojna, Yang Song, Sergio Guadarrama, and Kevin Murphy · 2017
Cited alongside, same era.
Automatic differentiation in PyTorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adem Lerer · 2017
Cited alongside, same era.
Faster R-CNN: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2017
Cited alongside, same era.
Structural compression of convolutional neural networks based on greedy filter pruning
Reza Abbasi-Asl and Bin Yu · 2017
Cited alongside, same era.
Channel pruning for accelerating very deep neural networks
Yihui He, Xiangyu Zhang, and Jian Sun · 2017
Cited alongside, same era.
PACT: Parameterized clipping activation for quantized neural networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Later among the works it cites.
Glow: Graph lowering compiler techniques for neural networks
Nadav Rotem, Jordan Fix, Saleem Abdulrasool, Summer Deng, Roman Dzhabarov, James Hegeman, Roman Levenstein, Bert Maher, Nadathur Satish, Jakob Olesen, Jongsoo Park, Artem Rakhov, and Misha Smelyanskiy · 2018
Later among the works it cites.
TVM: An automated end-to-end optimizing compiler for deep learning
Tianqi Chen, Thierry Moreau, Ziheng Jiang, Lianmin Zheng, Eddie Yan, Haichen Shen, Meghan Cowan, Leyuan Wang, Yuwei Hu, Luis Ceze, Carlos Guestrin, and Arvind Krishnamurthy · 2018
Later among the works it cites.
PocketFlow: An automated framework for compressing and accelerating deep neural networks
Jiaxiang Wu, Yao Zhang, Haoli Bai, Huasong Zhong, Jinlong Hou, Wei Liu, and Junzhou Huang · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Michael Zhu and Suyog Gupta · 2017
Cited alongside, same era.
Exploring sparsity in recurrent neural networks
Sharan Narang, Erich Elsen, Gregory Diamos, and Shubho Sengupta · 2017
Cited alongside, same era.
Accelerating convolutional neural network with FFT on embedded hardware
T. Abtahi, C. Shea, A. Kulkarni, and T. Mohsenin · 2018
Cited alongside, same era.
AMC: AutoML for Model Compression and Acceleration on Mobile Devices
Yihui He, Ji Lin, Zhijian Liu, Hanrui Wang, Li-Jia Li, and Song Han · 2018
Cited alongside, same era.
Attention-based guided structured sparsity of deep neural networks
Amirsina Torfi, Rouzbeh A. Shirvani, Sobhan Soleymani, and Nasser M. Nasrabadi · 2018
Cited alongside, same era.
Smallify: Learning network size while training
Guillaume Leclerc, Manasi Vartak, Raul Castro Fernandez, Tim Kraska, and Samuel Madden · 2018
Cited alongside, same era.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Benoit Jacob, Skirmantas Kligys, Bo Chen, Menglong Zhu, Matthew Tang, Andrew G. Howard, Hartwig Adam, and Dmitry Kalenichenko · 2018
Cited alongside, same era.
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Chris De Sa, and Zhiru Zhang · 2019
Closest in time.
SinReQ: Generalized sinusoidal regularization for low-bitwidth deep quantized training
Ahmed T. Elthakeb, Prannoy Pilligundla, and Hadi Esmaeilzadeh · 2019
Closest in time.
SMT-SA: Simultaneous multithreading in systolic arrays
Gil Shomron, Tal Horowitz, and Uri C. Weiser · 2019
Closest in time.
FAKTA: An automatic end-to-end fact checking system
Moin Nadeem, Wei Fang, Brian Xu, Mitra Mohtarami, and James Glass · 2019
Closest in time.
Cross domain model compression by structurally weight sharing
Shangqian Gao, Cheng Deng, and Heng Huang · 2019
Closest in time.
Analog/mixed-signal hardware error modeling for deep learning inference
Angad S. Rekhi, Brian Zimmer, Nikola Nedovic, Ningxi Liu, Rangharajan Venkatesan, Miaorong Wang, Brucek Khailany, William J. Dally, and C. Thomas Gray · 2019
Closest in time.
The Lottery Ticket Hypothesis: Finding sparse, trainable neural networks
Jonathan Frankle and Michael Carbin · 2019
Closest in time.
Thinning of convolutional neural network with mixed pruning
W. Yang, L. Jin, S. Wang, Z. Cu, X. Chen, and L. Chen · 2019
Closest in time.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Ron Bannner, Yury Nahshan, and Daniel Soudry · 2019
Closest in time.