Fetching the paper…
Reading the bibliography…
As soon as abstract mathematical computations were adapted to computation on digital computers, the problem of efficient representation, manipulation, and communication of the numerical values in those computations arose.
A logical calculus of the ideas immanent in nervous activity
Warren S McCulloch and Walter Pitts · 1943
Earlier work this paper cites.
Spectra of quantized signals
William Ralph Bennett · 1948
Earlier work this paper cites.
The philosophy of pcm
BM Oliver, JR Pierce, and Claude E Shannon · 1948
Earlier work this paper cites.
A mathematical theory of communication
Claude E Shannon · 1948
Earlier work this paper cites.
A method for the construction of minimum-redundancy codes
David A Huffman · 1952
Earlier work this paper cites.
The perceptron, a perceiving and recognizing automaton Project Para
Frank Rosenblatt · 1957
Earlier work this paper cites.
Coding theorems for a discrete source with a fidelity criterion
Claude E Shannon · 1959
Earlier work this paper cites.
Principles of neurodynamics. perceptrons and the theory of brain mechanisms
Frank Rosenblatt · 1961
Earlier work this paper cites.
The performance of a class of n dimensional quantizers for a gaussian source
JG Dunn · 1965
Earlier work this paper cites.
The History of Statistics: The Measurement of Uncertainty before 1900
S. M. Stigler · 1986
Earlier work this paper cites.
Using simulated annealing to design good codes
AE Gamal, L Hemachandra, Itzhak Shperling, and V Wei · 1987
Earlier work this paper cites.
A new vector quantization clustering algorithm
William H Equitz · 1989
Earlier work this paper cites.
Optimal brain damage
Yann LeCun, John S Denker, and Sara A Solla · 1990
Earlier work this paper cites.
A deterministic annealing approach to clustering
Kenneth Rose, Eitan Gurewitz, and Geoffrey Fox · 1990
Earlier work this paper cites.
Second order derivatives for network pruning: Optimal brain surgeon
Babak Hassibi and David G Stork · 1993
Earlier work this paper cites.
Reduced-size neural networks through singular value decomposition and subset selection
PP Kanjilal, PK Dey, and DN Banerjee · 1993
Earlier work this paper cites.
Numerical Linear Algebra
L.N. Trefethen and D. Bau III · 1997
Earlier work this paper cites.
Quantization
Robert M. Gray and David L. Neuhoff · 1998
Earlier work this paper cites.
An introduction to natural computation
Dana Harry Ballard · 1999
Earlier work this paper cites.
Is perception discrete or continuous?
Rufin VanRullen and Christof Koch · 2003
Earlier work this paper cites.
Optimal information storage in noisy synapses under resource constraints
Lav R Varshney, Per Jesper Sjöström, and Dmitri B Chklovskii · 2006
Earlier work this paper cites.
Challenges and advances in parallel sparse matrix-matrix multiplication
Aydin Buluc and John R Gilbert · 2008
Earlier work this paper cites.
Noise in the nervous system
A Aldo Faisal, Luc PJ Selen, and Daniel M Wolpert · 2008
Earlier work this paper cites.
Product quantization for nearest neighbor search
Herve Jegou, Matthijs Douze, and Cordelia Schmid · 2010
Earlier work this paper cites.
Simplifying convnets for fast learning
Franck Mamalet and Christophe Garcia · 2012
Earlier work this paper cites.
A framework for bayesian optimality of psychophysical laws
John Z Sun, Grace I Wang, Vivek K Goyal, and Lav R Varshney · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Low-rank matrix factorization for deep neural network training with high-dimensional output targets
Tara N Sainath, Brian Kingsbury, Vikas Sindhwani, Ebru Arisoy, and Bhuvana Ramabhadran · 2013
Earlier work this paper cites.
Training deep neural networks with low precision multiplications
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2014
Earlier work this paper cites.
Compressing deep convolutional networks using vector quantization
Yunchao Gong, Liu Liu, Ming Yang, and Lubomir Bourdev · 2014
Earlier work this paper cites.
Generative adversarial networks
Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Mark Horowitz · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
Adriana Romero, Nicolas Ballas, Samira Ebrahimi Kahou, Antoine Chassang, Carlo Gatta, and Yoshua Bengio · 2014
Earlier work this paper cites.
BinaryConnect: Training deep neural networks with binary weights during propagations
Matthieu Courbariaux, Yoshua Bengio, and Jean-Pierre David · 2015
Earlier work this paper cites.
Deep learning with limited numerical precision
Suyog Gupta, Ankur Agrawal, Kailash Gopalakrishnan, and Pritish Narayanan · 2015
Earlier work this paper cites.
Song Han, Huizi Mao, and William J Dally · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Single-trial spike trains in parietal cortex reveal discrete steps during decision-making
Kenneth W Latimer, Jacob L Yates, Miriam LR Meister, Alexander C Huk, and Jonathan W Pillow · 2015
Earlier work this paper cites.
Difference target propagation
Dong-Hyun Lee, Saizheng Zhang, Asja Fischer, and Yoshua Bengio · 2015
Earlier work this paper cites.
Neural networks with few multiplications
Zhouhan Lin, Matthieu Courbariaux, Roland Memisevic, and Yoshua Bengio · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Resiliency of deep neural networks under quantization
Wonyong Sung, Sungho Shin, and Kyuyeon Hwang · 2015
Earlier work this paper cites.
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton · 2016
Earlier work this paper cites.
Computational principles of memory
Rishidev Chaudhuri and Ila Fiete · 2016
Earlier work this paper cites.
Towards the limit of network quantization
Yoojin Choi, Mostafa El-Khamy, and Jungwon Lee · 2016
Earlier work this paper cites.
Hardware-oriented approximation of convolutional neural networks
Philipp Gysel, Mohammad Motamedi, and Soheil Ghiasi · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Gaussian error linear units (GELUs)
Dan Hendrycks and Kevin Gimpel · 2016
Earlier work this paper cites.
Loss-aware binarization of deep networks
Lu Hou, Quanming Yao, and James T Kwok · 2016
Earlier work this paper cites.
Binarized neural networks
Itay Hubara, Matthieu Courbariaux, Daniel Soudry, Ran El-Yaniv, and Yoshua Bengio · 2016
Earlier work this paper cites.
SqueezeNet: Alexnet-level accuracy with 50x fewer parameters and< 0.5 mb model size
Forrest N Iandola, Song Han, Matthew W Moskewicz, Khalid Ashraf, William J Dally, and Kurt Keutzer · 2016
Earlier work this paper cites.
Categorical reparameterization with gumbel-softmax
Eric Jang, Shixiang Gu, and Ben Poole · 2016
Earlier work this paper cites.
Minje Kim and Paris Smaragdis · 2016
Earlier work this paper cites.
Fengfu Li, Bo Zhang, and Bin Liu · 2016
Earlier work this paper cites.
Fixed point quantization of deep convolutional networks
Darryl Lin, Sachin Talathi, and Sreekanth Annapureddy · 2016
Earlier work this paper cites.
Convolutional neural networks using logarithmic data representation
Daisuke Miyashita, Edward H Lee, and Boris Murmann · 2016
Earlier work this paper cites.
Xnor-net: Imagenet classification using binary convolutional neural networks
Mohammad Rastegari, Vicente Ordonez, Joseph Redmon, and Ali Farhadi · 2016
Earlier work this paper cites.
Fixed-point performance analysis of recurrent neural networks
Sungho Shin, Kyuyeon Hwang, and Wonyong Sung · 2016
Earlier work this paper cites.
Rethinking the Inception architecture for computer vision
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna · 2016
Earlier work this paper cites.
Decision making with quantized priors leads to discrimination
Lav R Varshney and Kush R Varshney · 2016
Earlier work this paper cites.
Quantized convolutional neural networks for mobile devices
Jiaxiang Wu, Cong Leng, Yuhang Wang, Qinghao Hu, and Jian Cheng · 2016
Earlier work this paper cites.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
Shuchang Zhou, Yuxin Wu, Zekun Ni, Xinyu Zhou, He Wen, and Yuheng Zou · 2016
Earlier work this paper cites.
Chenzhuo Zhu, Song Han, Huizi Mao, and William J Dally · 2016
Earlier work this paper cites.
Neural architecture search with reinforcement learning
Barret Zoph and Quoc V Le · 2016
Earlier work this paper cites.
Soft-to-hard vector quantization for end-to-end learning compressible representations
Eirikur Agustsson, Fabian Mentzer, Michael Tschannen, Lukas Cavigelli, Radu Timofte, Luca Benini, and Luc Van Gool · 2017
Earlier work this paper cites.
Deep learning with low precision by half-wave gaussian quantization
Zhaowei Cai, Xiaodong He, Jian Sun, and Nuno Vasconcelos · 2017
Earlier work this paper cites.
Learning to prune deep neural networks via layer-wise optimal brain surgeon
Xin Dong, Shangyu Chen, and Sinno Jialin Pan · 2017
Earlier work this paper cites.
Learning deep binary descriptor with multi-quantization
Yueqi Duan, Jiwen Lu, Ziwei Wang, Jianjiang Feng, and Jie Zhou · 2017
Earlier work this paper cites.
Deep learning as a mixed convex-combinatorial optimization problem
Abram L Friesen and Pedro Domingos · 2017
Earlier work this paper cites.
Tensor processing using low precision format, December 28 2017
Boris Ginsburg, Sergei Nikolaev, Ahmad Kiswani, Hao Wu, Amir Gholaminejad, Slawomir Kierat, Michael Houston, and Alex Fit-Florea · 2017
Earlier work this paper cites.
Network sketching: Exploiting binary structure in deep cnns
Yiwen Guo, Anbang Yao, Hao Zhao, and Yurong Chen · 2017
Earlier work this paper cites.
MobileNets: Efficient convolutional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam · 2017
Earlier work this paper cites.
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens Van Der Maaten, and Kilian Q Weinberger · 2017
Earlier work this paper cites.
Deep roots: Improving cnn efficiency with hierarchical filter groups
Yani Ioannou, Duncan Robertson, Roberto Cipolla, and Antonio Criminisi · 2017
Earlier work this paper cites.
Local binary convolutional neural networks
Felix Juefei-Xu, Vishnu Naresh Boddeti, and Marios Savvides · 2017
Earlier work this paper cites.
Discrete adjustment to a changing environment: Experimental evidence
Mel Win Khaw, Luminita Stevens, and Michael Woodford · 2017
Earlier work this paper cites.
Learning from noisy labels with distillation
Yuncheng Li, Jianchao Yang, Yale Song, Liangliang Cao, Jiebo Luo, and Li-Jia Li · 2017
Earlier work this paper cites.
Performance guaranteed network acceleration via high-order residual quantization
Zefan Li, Bingbing Ni, Wenjun Zhang, Xiaokang Yang, and Wen Gao · 2017
Earlier work this paper cites.
Towards accurate binary convolutional neural network
Xiaofan Lin, Cong Zhao, and Wei Pan · 2017
Earlier work this paper cites.
Thinet: A filter level pruning method for deep neural network compression
Jian-Hao Luo, Jianxin Wu, and Weiyao Lin · 2017
Earlier work this paper cites.
Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al · 2017
Earlier work this paper cites.
Nvidia 8-bit inference with tensorrt
Szymon Migacz · 2017
Earlier work this paper cites.
Apprentice: Using knowledge distillation techniques to improve low-precision network accuracy
Asit Mishra and Debbie Marr · 2017
Earlier work this paper cites.
Wrpn: Wide reduced-precision networks
Asit Mishra, Eriko Nurvitadhi, Jeffrey J Cook, and Debbie Marr · 2017
Earlier work this paper cites.
Weighted-entropy-based quantization for deep neural networks
Eunhyeok Park, Junwhan Ahn, and Sungjoo Yoo · 2017
Earlier work this paper cites.
Searching for activation functions
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
Swish: a self-gated activation function
Prajit Ramachandran, Barret Zoph, and Quoc V Le · 2017
Earlier work this paper cites.
How to train a compact binary neural network with high accuracy?
Wei Tang, Gang Hua, and Liang Wang · 2017
Earlier work this paper cites.
Antti Tarvainen and Harri Valpola · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
A gift from knowledge distillation: Fast optimization, network minimization and transfer learning
Junho Yim, Donggyu Joo, Jihoon Bae, and Junmo Kim · 2017
Earlier work this paper cites.
Learning from multiple teacher networks
Shan You, Chang Xu, Chao Xu, and Dacheng Tao · 2017
Earlier work this paper cites.
Incremental network quantization: Towards lossless cnns with low-precision weights
Aojun Zhou, Anbang Yao, Yiwen Guo, Lin Xu, and Yurong Chen · 2017
Earlier work this paper cites.
Adaptive quantization for deep neural network
Yiren Zhou, Seyed-Mohsen Moosavi-Dezfooli, Ngai-Man Cheung, and Pascal Frossard · 2017
Earlier work this paper cites.
An empirical study of binary neural networks’ optimisation
Milad Alizadeh, Javier Fernández-Marqués, Nicholas D Lane, and Yarin Gal · 2018
Earlier work this paper cites.
Proxquant: Quantized neural networks via proximal operators
Yu Bai, Yu-Xiang Wang, and Edo Liberty · 2018
Cited alongside, same era.
Scalable methods for 8-bit training of neural networks
Ron Banner, Itay Hubara, Elad Hoffer, and Daniel Soudry · 2018
Cited alongside, same era.
Post-training 4-bit quantization of convolution networks for rapid-deployment
Ron Banner, Yury Nahshan, Elad Hoffer, and Daniel Soudry · 2018
Cited alongside, same era.
Uniq: Uniform noise injection for non-uniform quantization of neural networks
Chaim Baskin, Eli Schwartz, Evgenii Zheltonozhskii, Natan Liss, Raja Giryes, Alex M Bronstein, and Avi Mendelson · 2018
Cited alongside, same era.
Proxylessnas: Direct neural architecture search on target task and hardware
Quantization networks
Jiwei Yang, Xu Shen, Jun Xing, Xinmei Tian, Houqiang Li, Bing Deng, Jianqiang Huang, and Xian-sheng Hua · 2019
Later among the works it cites.
Understanding straight-through estimator in training activation quantized neural nets
Penghang Yin, Jiancheng Lyu, Shuai Zhang, Stanley Osher, Yingyong Qi, and Jack Xin · 2019
Later among the works it cites.
Blended coarse gradient descent for full quantization of deep neural networks
Penghang Yin, Shuai Zhang, Jiancheng Lyu, Stanley Osher, Yingyong Qi, and Jack Xin · 2019
Later among the works it cites.
Be your own teacher: Improve the performance of convolutional neural networks via self distillation
Linfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen, Chenglong Bao, and Kaisheng Ma · 2019
Later among the works it cites.
Variational convolutional neural network pruning
Chenglong Zhao, Bingbing Ni, Jian Zhang, Qiwei Zhao, Wenjun Zhang, and Qi Tian · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Han Cai, Ligeng Zhu, and Song Han · 2018
Cited alongside, same era.
Distilled binary neural network for monaural speech separation
Xiuyi Chen, Guangcan Liu, Jing Shi, Jiaming Xu, and Bo Xu · 2018
Cited alongside, same era.
Pact: Parameterized clipping activation for quantized neural networks
Jungwook Choi, Zhuo Wang, Swagath Venkataramani, Pierce I-Jen Chuang, Vijayalakshmi Srinivasan, and Kailash Gopalakrishnan · 2018
Cited alongside, same era.
Learning low precision deep neural networks through regularization
Yoojin Choi, Mostafa El-Khamy, and Jungwon Lee · 2018
Cited alongside, same era.
Moonshine: Distilling with cheap convolutions
Elliot J Crowley, Gavin Gray, and Amos J Storkey · 2018
Cited alongside, same era.
Bnn+: Improved binary network training
Sajad Darabi, Mouloud Belbahri, Matthieu Courbariaux, and Vahid Partovi Nia · 2018
Cited alongside, same era.
Gxnor-net: Training deep neural networks with ternary weights and activations without full-precision memory under a unified discretization framework
Lei Deng, Peng Jiao, Jing Pei, Zhenzhi Wu, and Guoqi Li · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Cited alongside, same era.
Learning efficient tensor representations with ring-structured networks
Qibin Zhao, Masashi Sugiyama, Longhao Yuan, and Andrzej Cichocki · 2019
Later among the works it cites.
Improving neural network quantization without retraining using outlier channel splitting
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Christopher De Sa, and Zhiru Zhang · 2019
Later among the works it cites.
Binary ensemble neural network: More bits per network or more networks per bit?
Shilin Zhu, Xin Dong, and Hao Su · 2019
Later among the works it cites.
Structured binary neural networks for accurate image classification and semantic segmentation
Bohan Zhuang, Chunhua Shen, Mingkui Tan, Lingqiao Liu, and Ian Reid · 2019
Later among the works it cites.
Universally quantized neural compression
Eirikur Agustsson and Lucas Theis · 2020
Later among the works it cites.
Gradient l1 regularization for quantization robustness
Milad Alizadeh, Arash Behboodi, Mart van Baalen, Christos Louizos, Tijmen Blankevoort, and Max Welling · 2020
Later among the works it cites.
Binarybert: Pushing the limit of bert quantization
Haoli Bai, Wei Zhang, Lu Hou, Lifeng Shang, Jing Jin, Xin Jiang, Qun Liu, Michael Lyu, and Irwin King · 2020
Later among the works it cites.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Yash Bhalgat, Jinwon Lee, Markus Nagel, Tijmen Blankevoort, and Nojun Kwak · 2020
Later among the works it cites.
What is the state of neural network pruning?
Davis Blalock, Jose Javier Gonzalez Ortiz, Jonathan Frankle, and John Guttag · 2020
Later among the works it cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Later among the works it cites.
Zeroq: A novel zero shot quantization framework
Yaohui Cai, Zhewei Yao, Zhen Dong, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Shifted and squeezed 8-bit floating point format for low-precision training of deep neural networks
Léopold Cambier, Anahita Bhiwandiwalla, Ting Gong, Mehran Nekuii, Oguz H Elibol, and Hanlin Tang · 2020
Later among the works it cites.
A statistical framework for low-bitwidth training of deep neural networks
Jianfei Chen, Yu Gai, Zhewei Yao, Michael W Mahoney, and Joseph E Gonzalez · 2020
Later among the works it cites.
One weight bitwidth to rule them all
Ting-Wu Chin, Pierce I-Jen Chuang, Vikas Chandra, and Diana Marculescu · 2020
Later among the works it cites.
Data-free network quantization with adversarial knowledge distillation
Yoojin Choi, Jihwan Choi, Mostafa El-Khamy, and Jungwon Lee · 2020
Later among the works it cites.
HAWQ-V2: Hessian aware trace-weighted quantization of neural networks
Zhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami, Michael W. Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Adaptive gradient quantization for data-parallel sgd
Fartash Faghri, Iman Tabrizian, Ilia Markov, Dan Alistarh, Daniel Roy, and Ali Ramezani-Kebrya · 2020
Later among the works it cites.
Training with quantization noise for extreme model compression
Angela Fan, Pierre Stock, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, and Armand Joulin · 2020
Later among the works it cites.
Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz, David Thorsley, Georgios Georgiadis, and Joseph Hassoun · 2020
Later among the works it cites.
Post-training piecewise linear quantization for deep neural networks
Jun Fang, Ali Shafiee, Hamzah Abdel-Aziz, David Thorsley, Georgios Georgiadis, and Joseph H Hassoun · 2020
Later among the works it cites.
An integrated approach to neural network design, training, and inference
Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Hmq: Hardware friendly mixed precision quantization block for cnns
Hai Victor Habi, Roy H Jennings, and Arnon Netzer · 2020
Later among the works it cites.
Training binary neural networks through learning with noisy supervision
Kai Han, Yunhe Wang, Yixing Xu, Chunjing Xu, Enhua Wu, and Chang Xu · 2020
Later among the works it cites.
The knowledge within: Methods for data-free model compression
Matan Haroush, Itay Hubara, Elad Hoffer, and Daniel Soudry · 2020
Later among the works it cites.
Improving post training neural quantization: Layer-wise calibration and integer programming
Itay Hubara, Yury Nahshan, Yair Hanani, Ron Banner, and Daniel Soudry · 2020
Later among the works it cites.
Efficient execution of quantized deep learning models: A compiler approach
Animesh Jain, Shoubhik Bhattacharya, Masahiro Masuda, Vin Sharma, and Yida Wang · 2020
Later among the works it cites.
Biqgemm: matrix multiplication with lookup table for binary-coding-based quantized dnns
Yongkweon Jeon, Baeseong Park, Se Jung Kwon, Byeongwook Kim, Jeongin Yun, and Dongsoo Lee · 2020
Later among the works it cites.
Efficient exact verification of binarized neural networks
Kai Jia and Martin Rinard · 2020
Later among the works it cites.
Adabits: Neural network quantization with adaptive bit-widths
Qing Jin, Linjie Yang, and Zhenyu Liao · 2020
Later among the works it cites.
Comparing fisher information regularization with distillation for dnn quantization
Prad Kadambi, Karthikeyan Natesan Ramamurthy, and Visar Berisha · 2020
Later among the works it cites.
Binaryduo: Reducing gradient mismatch in binary activation network by coupling binary activations
Hyungjun Kim, Kyungsu Kim, Jinseok Kim, and Jae-Joon Kim · 2020
Later among the works it cites.
Position-based scaled gradient for model quantization and sparse training
Jangho Kim, KiYoon Yoo, and Nojun Kwak · 2020
Later among the works it cites.
Structured compression by weight encryption for unstructured pruning and quantization
Se Jung Kwon, Dongsoo Lee, Byeongwook Kim, Parichay Kapoor, Baeseong Park, and Gu-Yeon Wei · 2020
Later among the works it cites.
Flexor: Trainable fractional quantization
Dongsoo Lee, Se Jung Kwon, Byeongwook Kim, Yongkweon Jeon, Baeseong Park, and Jeongin Yun · 2020
Later among the works it cites.
Dms: Differentiable dimension search for binary neural networks
Yuhang Li, Ruihao Gong, Fengwei Yu, Xin Dong, and Xianglong Liu · 2020
Later among the works it cites.
Rotated binary neural network
Mingbao Lin, Rongrong Ji, Zihan Xu, Baochang Zhang, Yan Wang, Yongjian Wu, Feiyue Huang, and Chia-Wen Lin · 2020
Later among the works it cites.
Training binary neural networks with real-to-binary convolutions
Brais Martinez, Jing Yang, Adrian Bulat, and Georgios Tzimiropoulos · 2020
Later among the works it cites.
Up or down? adaptive rounding for post-training quantization
Markus Nagel, Rana Ali Amjad, Mart Van Baalen, Christos Louizos, and Tijmen Blankevoort · 2020
Later among the works it cites.
Wrapnet: Neural net inference with ultra-low-resolution arithmetic
Renkun Ni, Hong-min Chu, Oscar Castañeda, Ping-yeh Chiang, Christoph Studer, and Tom Goldstein · 2020
Later among the works it cites.
Lookahead: a far-sighted alternative of magnitude-based pruning
Sejun Park, Jaeho Lee, Sangwoo Mo, and Jinwoo Shin · 2020
Later among the works it cites.
Binary neural networks: A survey
Haotong Qin, Ruihao Gong, Xianglong Liu, Xiao Bai, Jingkuan Song, and Nicu Sebe · 2020
Later among the works it cites.
Forward and backward information retention for accurate binary neural networks
Haotong Qin, Ruihao Gong, Xianglong Liu, Mingzhu Shen, Ziran Wei, Fengwei Yu, and Jingkuan Song · 2020
Later among the works it cites.
Adaptive loss-aware quantization for multi-bit networks
Zhongnan Qu, Zimu Zhou, Yun Cheng, and Lothar Thiele · 2020
Later among the works it cites.
Leveraging automated mixed-low-precision quantization for tiny edge microcontrollers
Manuele Rusci, Marco Fariselli, Alessandro Capotondi, and Luca Benini · 2020
Later among the works it cites.
Path sample-analytic gradient estimators for stochastic binary networks
Alexander Shekhovtsov, Viktor Yanush, and Boris Flach · 2020
Later among the works it cites.
Balanced binary neural networks with gated residual
Mingzhu Shen, Xianglong Liu, Ruihao Gong, and Kai Han · 2020
Later among the works it cites.
Q-BERT: Hessian based ultra low precision quantization of bert
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gholami, Michael W Mahoney, and Kurt Keutzer · 2020
Later among the works it cites.
Robust quantization: One model to rule them all
Moran Shkolnik, Brian Chmiel, Ron Banner, Gil Shomron, Yuri Nahshan, Alex Bronstein, and Uri Weiser · 2020
Later among the works it cites.
Is information in the brain represented in continuous or discrete form?
James Tee and Desmond P Taylor · 2020
Later among the works it cites.
Bayesian bits: Unifying quantization and pruning
Mart van Baalen, Christos Louizos, Markus Nagel, Rana Ali Amjad, Ying Wang, Tijmen Blankevoort, and Max Welling · 2020
Later among the works it cites.
Attentivenas: Improving neural architecture search via attentive sampling
Dilin Wang, Meng Li, Chengyue Gong, and Vikas Chandra · 2020
Later among the works it cites.
Apq: Joint search for network architecture, pruning and quantization policy
Tianzhe Wang, Kuan Wang, Han Cai, Ji Lin, Zhijian Liu, Hanrui Wang, Yujun Lin, and Song Han · 2020
Later among the works it cites.
Differentiable joint pruning and quantization for hardware efficiency
Ying Wang, Yadong Lu, and Tijmen Blankevoort · 2020
Later among the works it cites.
Integer quantization for deep learning inference: Principles and empirical evaluation
Hao Wu, Patrick Judd, Xiaojie Zhang, Mikhail Isaev, and Paulius Micikevicius · 2020
Later among the works it cites.
Generative low-bitwidth data free quantization
Shoukai Xu, Haokun Li, Bohan Zhuang, Jing Liu, Jiezhang Cao, Chuangrun Liang, and Mingkui Tan · 2020
Later among the works it cites.
Automatic neural network compression by sparsity-quantization joint learning: A constrained optimization-based approach
Haichuan Yang, Shupeng Gui, Yuhao Zhu, and Ji Liu · 2020
Later among the works it cites.
Searching for low-bit weights in quantized neural networks
Zhaohui Yang, Yunhe Wang, Kai Han, Chunjing Xu, Chao Xu, Dacheng Tao, and Chang Xu · 2020
Later among the works it cites.
Hawqv3: Dyadic neural network quantization
Zhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami, Jiali Yu, Eric Tan, Leyuan Wang, Qijing Huang, Yida Wang, Michael W Mahoney, et al · 2020
Later among the works it cites.
Distillation guided residual learning for binary convolutional neural networks
Jianming Ye, Shiliang Zhang, and Jingdong Wang · 2020
Later among the works it cites.
Dreaming to distill: Data-free knowledge transfer via deepinversion
Hongxu Yin, Pavlo Molchanov, Jose M Alvarez, Zhizhong Li, Arun Mallya, Derek Hoiem, Niraj K Jha, and Jan Kautz · 2020
Later among the works it cites.
Ternarybert: Distillation-aware ultra-low bit bert
Wei Zhang, Lu Hou, Yichun Yin, Lifeng Shang, Xiao Chen, Xin Jiang, and Qun Liu · 2020
Later among the works it cites.
High-capacity expert binary networks
Adrian Bulat, Brais Martinez, and Georgios Tzimiropoulos · 2021
Closest in time.
Incremental few-shot learning via vector quantization in deep embedded space
Kuilin Chen and Chi-Guhn Lee · 2021
Closest in time.
Neural gradients are near-lognormal: improved quantized and sparse training
Brian Chmiel, Liad Ben-Uri, Moran Shkolnik, Elad Hoffer, Ron Banner, and Daniel Soudry · 2021
Closest in time.
Multi-prize lottery ticket hypothesis: Finding accurate binary neural networks by pruning a randomly weighted network
James Diffenderfer and Bhavya Kailkhura · 2021
Closest in time.
Confounding tradeoffs for neural network quantization
Sahaj Garg, Anirudh Jain, Joe Lou, and Mitchell Nahmias · 2021
Closest in time.
Dynamic precision analog computing for neural networks
Sahaj Garg, Joe Lou, Anirudh Jain, and Mitchell Nahmias · 2021
Closest in time.
Boolnet: Minimizing the energy consumption of binary neural networks
Nianhui Guo, Joseph Bethge, Haojin Yang, Kai Zhong, Xuefei Ning, Christoph Meinel, and Yu Wang · 2021
Closest in time.
Ps and qs: Quantization-aware pruning for efficient low latency neural network inference
Benjamin Hawks, Javier Duarte, Nicholas J Fraser, Alessandro Pappalardo, Nhan Tran, and Yaman Umuroglu · 2021
Closest in time.
Generative zero-shot network quantization
Xiangyu He, Qinghao Hu, Peisong Wang, and Jian Cheng · 2021
Closest in time.
Torsten Hoefler, Dan Alistarh, Tal Ben-Nun, Nikoli Dryden, and Alexandra Peste · 2021
Closest in time.
Opq: Compressing deep neural networks with one-shot pruning-quantization
Peng Hu, Xi Peng, Hongyuan Zhu, Mohamed M Sabry Aly, and Jie Lin · 2021
Closest in time.
Codenet: Efficient deployment of input-adaptive object detection on embedded fpgas
Qijing Huang, Dequan Wang, Zhen Dong, Yizhao Gao, Yaohui Cai, Tian Li, Bichen Wu, Kurt Keutzer, and John Wawrzynek · 2021
Closest in time.
On the distribution, sparsity, and inference-time quantization of attention values in transformers
Tianchu Ji, Shraddhan Jain, Michael Ferdman, Peter Milder, H Andrew Schwartz, and Niranjan Balasubramanian · 2021
Closest in time.
Kdlsq-bert: A quantized bert combining knowledge distillation with learned step size quantization
Jing Jin, Cai Liang, Tiancheng Wu, Liqin Zou, and Zhiliang Gan · 2021
Closest in time.
I-bert: Integer-only bert quantization
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W Mahoney, and Kurt Keutzer · 2021
Closest in time.
Brecq: Pushing the limit of post-training quantization by block reconstruction
Yuhang Li, Ruihao Gong, Xu Tan, Yang Yang, Peng Hu, Qi Zhang, Fengwei Yu, Wei Wang, and Shi Gu · 2021
Closest in time.
Pruning and quantization for deep neural network acceleration: A survey
Tailin Liang, John Glossner, Lei Wang, and Shaobo Shi · 2021
Closest in time.
Sparse quantized spectral clustering
Zhenyu Liao, Romain Couillet, and Michael W Mahoney · 2021
Closest in time.
Layer importance estimation with imprinting for neural network quantization
Hongyang Liu, Sara Elkerdawy, Nilanjan Ray, and Mostafa Elhoushi · 2021
Closest in time.
A white paper on neural network quantization
Markus Nagel, Marios Fournarakis, Rana Ali Amjad, Yelysei Bondarenko, Mart van Baalen, and Tijmen Blankevoort · 2021
Closest in time.
Simple augmentation goes a long way: {ADRL} for {dnn} quantization
Lin Ning, Guoyang Chen, Weifeng Zhang, and Xipeng Shen · 2021
Closest in time.
Fully integer-based quantization for mobile convolutional neural network inference
Peng Peng, Mingyu You, Weisheng Xu, and Jiaxin Li · 2021
Closest in time.
Bipointnet: Binary neural network for point clouds
Haotong Qin, Zhongang Cai, Mingyuan Zhang, Yifu Ding, Haiyu Zhao, Shuai Yi, Xianglong Liu, and Hao Su · 2021
Closest in time.
Adaptive binary-ternary quantization
Ryan Razani, Gregoire Morin, Eyyub Sari, and Vahid Partovi Nia · 2021
Closest in time.
Post-training sparsity-aware quantization
Gil Shomron, Freddy Gabbay, Samer Kurzum, and Uri Weiser · 2021
Closest in time.
Training with quantization noise for extreme model compression
Pierre Stock, Angela Fan, Benjamin Graham, Edouard Grave, Rémi Gribonval, Herve Jegou, and Armand Joulin · 2021
Closest in time.
Degree-quant: Quantization-aware training for graph neural networks
Shyam A Tailor, Javier Fernandez-Marques, and Nicholas D Lane · 2021
Closest in time.
Bsq: Exploring bit-level sparsity for mixed-precision neural network quantization
Huanrui Yang, Lin Duan, Yiran Chen, and Hai Li · 2021
Closest in time.
Hessian-aware pruning and optimal neural implant
Shixing Yu, Zhewei Yao, Amir Gholami, Zhen Dong, Michael W Mahoney, and Kurt Keutzer · 2021
Closest in time.
Distribution-aware adaptive multi-bit quantization
Sijie Zhao, Tao Yue, and Xuemei Hu · 2021
Closest in time.