Fetching the paper…
Reading the bibliography…
Deep neural network (DNN) model compression for efficient on-device inference is becoming increasingly important to reduce memory requirements and keep user data on-device.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdana, Kyunghyun Cho, and Yoshua Bengio · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural network with pruning, trained quantization and huffman coding
Song Han, Huizi Mao, and William J. Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Earlier work this paper cites.
Towards the limit of network quantization
Yoojin Choi, Mostafa El-Khamy, and Jungwon Lee · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2017
Earlier work this paper cites.
Towards accurate binary convolutional neural network
Xiaofan Lin, Cong Zhao, and Wei Pan · 2017
Earlier work this paper cites.
Weighted-entropy-based quantization for deep neural networks
Eunhyeok Park, Junwhan Ahn, and Sungjoo Yoo · 2017
Earlier work this paper cites.
Soft weight-sharing for neural network compression
Karen Ullrich, Edward Meeds, and Max Welling · 2017
Earlier work this paper cites.
Value-aware quantization for training and inference of neural networks
Eunhyeok Park, Sungjoo Yoo, and Peter Vajda · 2018
Earlier work this paper cites.
Model compression via distillation and quantization
Antonio Polino, Razvan Pascanu, and Dan-Adrian Alistarh · 2018
Earlier work this paper cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zhmoginov, and Liang-Chieh Chen · 2018
Earlier work this paper cites.
Learning discrete weights using the local reparameterization trick
Oran Shayer, Dan Levi, and Ethan Fetaya · 2018
Earlier work this paper cites.
A general reinforcement learning algorithm that masters chess, shogi, and go through self-play
David Silver, Thomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Matthew Lai, Arthur Guez, Marc Lanctot, Laurent Sifre, Dharshan Kumaran, Thore Graepel, et al · 2018
Earlier work this paper cites.
Deep k-means: Re-training and parameter sharing with harder cluster assignments for compressing deep convolutions
Junru Wu, Yue Wang, Zhenyu Wu, Zhangyang Wang, Ashok Veeraraghavan, and Yingyan Lin · 2018
Cited alongside, same era.
Nisp: Pruning networks using neuron importance score propagation
Ruichi Yu, Ang Li, Chun-Fu Chen, Jui-Hsin Lai, Vlad I Morariu, Xintong Han, Mingfei Gao, Ching-Yung Lin, and Larry S Davis · 2018
Cited alongside, same era.
Lq-nets: Learned quantization for highly accurate and compact deep neural networks
Dongqing Zhang, Jiaolong Yang, Dongqiangzi Ye, and Gang Hua · 2018
Cited alongside, same era.
Tensorflow eager: A multi-stage, python-embedded dsl for machine learning
Akshay Agrawal, Akshay Naresh Modi, Alexandre Passos, Allen Lavoie, Ashish Agarwal, Asim Shankar, Igor Ganichev, Josh Levenberg, Mingsheng Hong, Rajat Monga, et al · 2019
Cited alongside, same era.
Metaquant: Learning to quantize by learning to penetrate non-differentiable quantization
Shangyu Chen, Wenya Wang, and Sinno Jialin Pan · 2019
Well-read students learn better: The impact of student initialization on knowledge distillation
Iulia Turc, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Later among the works it cites.
Understanding straight-through estimator in training activation quantized neural nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang, Stanley Osher, Yingyong Qi, and Jack Xin · 2019
Later among the works it cites.
Linear symmetric quantization of neural networks for low-precision integer hardware
Xiandong Zhao, Ying Wang, Xuyi Cai, Cheng Liu, and Lei Zhang · 2019
Later among the works it cites.
Neural epitome search for architecture-agnostic network compression
Daquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang, Jianchao Yang, and Jiashi Feng · 2019
Later among the works it cites.
Universal deep neural network compression
Yoojin Choi, Mostafa El-Khamy, and Jungwon Lee · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ALBERT: A lite BERT for self-supervised learning of language representations
Zhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel, Piyush Sharma, and Radu Soricut · 2019
Cited alongside, same era.
Additive powers-of-two quantization: An efficient non-uniform discretization for neural networks
Yuhang Li, Xin Dong, and Wei Wang · 2019
Cited alongside, same era.
Darts: Differentiable architecture search
Hanxiao Liu, Karen Simonyan, and Yiming Yang · 2019
Cited alongside, same era.
Espnetv2: A light-weight, power efficient, and general purpose convolutional neural network
Sachin Mehta, Mohammad Rastegari, Linda Shapiro, and Hannaneh Hajishirzi · 2019
Cited alongside, same era.
Lookahead: A far-sighted alternative of magnitude-based pruning
Sejun Park, Jaeho Lee, Sangwoo Mo, and Jinwoo Shin · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf · 2019
Cited alongside, same era.
Hawq-v2: Hessian aware trace-weighted quantization of neural networks
Zhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Up or down? Adaptive rounding for post-training quantization
Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort · 2020
Later among the works it cites.
Profit: A novel training method for sub-4-bit mobilenet models
Eunhyeok Park and Sungjoo Yoo · 2020
Later among the works it cites.
And the bit goes down: Revisiting the quantization of neural networks
Pierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham, and Hervé Jégou · 2020
Later among the works it cites.
Mobilebert: a compact task-agnostic bert for resource-limited devices
Zhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu, Yiming Yang, and Denny Zhou · 2020
Later among the works it cites.
Gobo: Quantizing attention-based nlp models for low latency and energy efficient inference
Ali Hadi Zadeh, Isak Edo, Omar Mohamed Awad, and Andreas Moshovos · 2020
Later among the works it cites.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin · 2021
Closest in time.
Network quantization with element-wise gradient scaling
B. Ham J. Lee, D. Kim · 2021
Closest in time.
Efficientnetv2: Smaller models and faster training
Mingxing Tan and Quoc V Le · 2021
Closest in time.