Fetching the paper…
Reading the bibliography…
In this paper, we investigate lossy compression of deep neural networks (DNNs) by weight quantization and lossless source coding for memory-efficient deployment.
J. Ziv and A. Lempel, “A universal algorithm for sequential data compression,” IEEE Transactions on Information Theory , vol. 23, no. 3, pp. 337–343, 1977
1977
Earlier work this paper cites.
——, “Compression of individual sequences via variable-rate coding,” IEEE Transactions on Information Theory , vol. 24, no. 5, pp. 530–536, 1978
1978
Earlier work this paper cites.
A. Gersho, “Asymptotically optimal block quantization,” IEEE Transactions on Information Theory , vol. 25, no. 4, pp. 373–380, 1979
1979
Earlier work this paper cites.
T. A. Welch, “A technique for high-performance data compression,” Computer , vol. 6, no. 17, pp. 8–19, 1984
1984
Earlier work this paper cites.
J. Ziv, “On universal quantization,” IEEE Transactions on Information Theory , vol. 31, no. 3, pp. 344–347, 1985
1985
Earlier work this paper cites.
R. Zamir and M. Feder, “On universal quantization by randomized uniform/lattice quantizers,” IEEE Transactions on Information Theory , vol. 38, no. 2, pp. 428–436, 1992
1992
Earlier work this paper cites.
R. M. Gray and D. L. Neuhoff, “Quantization,” IEEE Transactions on Information Theory , vol. 44, no. 6, pp. 2325–2383, 1998
1998
Earlier work this paper cites.
J. Seward, “bzip2,” 1998. [Online]. Available: www.bzip.org
1998
Earlier work this paper cites.
M. Effros, K. Visweswariah, S. R. Kulkarni, and S. Verdú, “Universal lossless source coding with the Burrows Wheeler transform,” IEEE Transactions on Information Theory , vol. 48, no. 5, pp. 1061–1081, 2002
2002
Cited alongside, same era.
A. Krizhevsky, “Learning multiple layers of features from tiny images,” 2009
2009
Cited alongside, same era.
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems , 2012, pp. 1097–1105
2012
Cited alongside, same era.
S. Han, J. Pool, J. Tran, and W. Dally, “Learning both weights and connections for efficient neural network,” in Advances in Neural Information Processing Systems , 2015, pp. 1135–1143
2015
Cited alongside, same era.
S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and Huffman coding,” in International Conference on Learning Representations , 2016
Y. Choi, M. El-Khamy, and J. Lee, “Towards the limit of network quantization,” in International Conference on Learning Representations , 2017
2017
Later among the works it cites.
K. Ullrich, E. Meeds, and M. Welling, “Soft weight-sharing for neural network compression,” in International Conference on Learning Representations , 2017
2017
Later among the works it cites.
E. Agustsson, F. Mentzer, M. Tschannen, L. Cavigelli, R. Timofte, L. Benini, and L. V. Gool, “Soft-to-hard vector quantization for end-to-end learning compressible representations,” in Advances in Neural Information Processing Systems , 2017, pp. 1141–1151
2017
Later among the works it cites.
D. Molchanov, A. Ashukha, and D. Vetrov, “Variational dropout sparsifies deep neural networks,” in International Conference on Machine Learning , 2017, pp. 2498–2507
2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
Y. Guo, A. Yao, and Y. Chen, “Dynamic network surgery for efficient DNNs,” in Advances In Neural Information Processing Systems , 2016, pp. 1379–1387
2016
Cited alongside, same era.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 770–778
2016
Cited alongside, same era.
C. Louizos, K. Ullrich, and M. Welling, “Bayesian compression for deep learning,” in Advances in Neural Information Processing Systems , 2017, pp. 3290–3300
2017
Later among the works it cites.
J. Lin, Y. Rao, J. Lu, and J. Zhou, “Runtime neural pruning,” in Advances in Neural Information Processing Systems , 2017, pp. 2178–2188
2017
Later among the works it cites.
B. Dai, C. Zhu, and D. Wipf, “Compressing neural networks using the variational information bottleneck,” in International Conference on Machine Learning , 2018
2018
Closest in time.