Fetching the paper…
Reading the bibliography…
Model size and inference speed/power have become a major challenge in the deployment of Neural Networks for many applications.
Modeling by shortest data description
J. Rissanen · 1978
Earlier work this paper cites.
Experimental determination of precision requirements for back-propagation training of artificial neural networks
K. Asanovic and N. Morgan · 1991
Earlier work this paper cites.
Flat minima
S. Hochreiter and J. Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, P. Haffner, et al · 1998
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Exploiting linear structure within convolutional networks for efficient evaluation
E. L. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus · 2014
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2014
Earlier work this paper cites.
Fitnets: Hints for thin deep nets
A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio · 2014
Earlier work this paper cites.
Binaryconnect: Training deep neural networks with binary weights during propagations
M. Courbariaux, Y. Bengio, and J.-P. David · 2015
Earlier work this paper cites.
Learning both weights and connections for efficient neural network
S. Han, J. Pool, J. Tran, and W. Dally · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Earlier work this paper cites.
Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
S. Han, H. Mao, and W. J. Dally · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Squeezenet: Alexnet-level accuracy with 50x fewer parameters and¡ 0.5 mb model size
F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer · 2016
Earlier work this paper cites.
Pruning filters for efficient convnets
H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf · 2016
Cited alongside, same era.
Pruning convolutional neural networks for resource efficient inference
P. Molchanov, S. Tyree, T. Karras, T. Aila, and J. Kautz · 2016
Cited alongside, same era.
Xnor-net: Imagenet classification using binary convolutional neural networks
M. Rastegari, V. Ordonez, J. Redmon, and A. Farhadi · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Cited alongside, same era.
Dorefa-net: Training low bitwidth convolutional neural networks with low bitwidth gradients
S. Zhou, Y. Wu, Z. Ni, X. Zhou, H. Wen, and Y. Zou · 2016
Quantization and training of neural networks for efficient integer-arithmetic-only inference
B. Jacob, S. Kligys, B. Chen, M. Zhu, M. Tang, A. Howard, H. Adam, and D. Kalenichenko · 2018
Later among the works it cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
R. Krishnamoorthi · 2018
Later among the works it cites.
Shufflenet v2: Practical guidelines for efficient cnn architecture design
N. Ma, X. Zhang, H.-T. Zheng, and J. Sun · 2018
Later among the works it cites.
Value-aware quantization for training and inference of neural networks
E. Park, S. Yoo, and P. Vajda · 2018
Later among the works it cites.
Mobilenetv2: Inverted residuals and linear bottlenecks
M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Xception: Deep learning with depthwise separable convolutions
F. Chollet · 2017
Cited alongside, same era.
Learning accurate low-bit deep neural networks with stochastic quantization
Y. Dong, R. Ni, J. Li, Y. Chen, J. Zhu, and H. Su · 2017
Cited alongside, same era.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
A. G. Howard, M. Zhu, B. Chen, D. Kalenichenko, W. Wang, T. Weyand, M. Andreetto, and H. Adam · 2017
Cited alongside, same era.
Quantized neural networks: Training neural networks with low precision weights and activations
I. Hubara, M. Courbariaux, D. Soudry, R. El-Yaniv, and Y. Bengio · 2017
Cited alongside, same era.
Exploring the regularity of sparse structure in convolutional neural networks
H. Mao, S. Han, J. Pool, W. Li, X. Liu, Y. Wang, and W. J. Dally · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Cited alongside, same era.
Incremental network quantization: Towards lossless cnns with low-precision weights
A. Zhou, A. Yao, Y. Guo, L. Xu, and Y. Chen · 2017
Cited alongside, same era.
Bit fusion: Bit-level dynamically composable architecture for accelerating deep neural networks
H. Sharma, J. Park, N. Suda, L. Lai, B. Chau, V. Chandra, and H. Esmaeilzadeh · 2018
Later among the works it cites.
Mixed precision quantization of convnets via differentiable neural architecture search
B. Wu, Y. Wang, P. Zhang, Y. Tian, P. Vajda, and K. Keutzer · 2018
Later among the works it cites.
Large batch size training of neural networks with adversarial training and second-order information
Z. Yao, A. Gholami, K. Keutzer, and M. Mahoney · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Z. Yao, A. Gholami, Q. Lei, K. Keutzer, and M. W. Mahoney · 2018
Later among the works it cites.
LQ-Nets: Learned quantization for highly accurate and compact deep neural networks
D. Zhang, J. Yang, D. Ye, and G. Hua · 2018
Later among the works it cites.
Shufflenet: An extremely efficient convolutional neural network for mobile devices
X. Zhang, X. Zhou, M. Lin, and J. Sun · 2018
Later among the works it cites.
Adaptive quantization for deep neural network
Y. Zhou, S.-M. Moosavi-Dezfooli, N.-M. Cheung, and P. Frossard · 2018
Later among the works it cites.
HAQ: Hardware-aware automated quantization
K. Wang, Z. Liu, Y. Lin, J. Lin, and S. Han · 2019
Closest in time.
Synetgy: Algorithm-hardware co-design for convnet accelerators on embedded fpgas
Y. Yang, Q. Huang, B. Wu, T. Zhang, L. Ma, G. Gambardella, M. Blott, L. Lavagno, K. Vissers, J. Wawrzynek, et al · 2019
Closest in time.