Fetching the paper…
Reading the bibliography…
While neural networks have advanced the frontiers in many applications, they often come at a high computational cost.
Jain, S. R., Gural, A., Wu, M., and Dick, C · 1903
Earlier work this paper cites.
And the bit goes down: Revisiting the quantization of neural networks
Stock, P., Joulin, A., Gribonval, R., Graham, B., and Jégou, H · 1907
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Bengio, Y., Léonard, N., and Courville, A · 2013
Earlier work this paper cites.
1.1 computing’s energy problem (and what we can do about it)
Horowitz, M · 2014
Earlier work this paper cites.
Deep learning with limited numerical precision
Gupta, S., Agrawal, A., Gopalakrishnan, K., and Narayanan, P · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Ioffe, S. and Szegedy, C · 2015
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation, 2017
Chen, L.-C., Papandreou, G., Schroff, F., and Adam, H · 2017
Earlier work this paper cites.
Mobilenets: Efficient convolutional neural networks for mobile vision applications
Howard, A. G., Zhu, M., Chen, B., Kalenichenko, D., Wang, W., Weyand, T., Andreetto, M., and Adam, H · 2017
Earlier work this paper cites.
Searching for activation functions
Ramachandran, P., Zoph, B., and Le, Q. V · 2017
Earlier work this paper cites.
Quantization and training of neural networks for efficient integer-arithmetic-only inference
Jacob, B., Kligys, S., Chen, B., Zhu, M., Tang, M., Howard, A., Adam, H., and Kalenichenko, D · 2018
Earlier work this paper cites.
Quantizing deep convolutional networks for efficient inference: A whitepaper
Krishnamoorthi, R · 2018
Cited alongside, same era.
Mobilenetv2: Inverted residuals and linear bottlenecks
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C · 2018
Cited alongside, same era.
A quantization-friendly separable convolution for mobilenets
Sheng, T., Feng, C., Zhuo, S., Zhang, X., Shen, L., and Aleksic, M · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S · 2018
Cited alongside, same era.
Post training 4-bit quantization of convolutional networks for rapid-deployment
Banner, R., Nahshan, Y., and Soudry, D · 2019
Cited alongside, same era.
Lsq+: Improving low-bit quantization through learnable offsets and better initialization
Bhalgat, Y., Lee, J., Nagel, M., Blankevoort, T., and Kwak, N · 2020
Later among the works it cites.
One weight bitwidth to rule them all
Chin, T.-W., Chuang, P. I.-J., Chandra, V., and Marculescu, D · 2020
Later among the works it cites.
Learned step size quantization
Esser, S. K., McKinstry, J. L., Bablani, D., Appuswamy, R., and Modha, D. S · 2020
Later among the works it cites.
Distilling optimal neural networks: Rapid search in diverse spaces
Moons, B., Noorzad, P., Skliar, A., Mariani, G., Mehta, D., Lott, C., and Blankevoort, T · 2020
Later among the works it cites.
Up or down? Adaptive rounding for post-training quantization
Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., and Blankevoort, T · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
HAWQ: hessian aware quantization of neural networks with mixed-precision
Dong, Z., Yao, Z., Gholami, A., Mahoney, M. W., and Keutzer, K · 2019
Cited alongside, same era.
Fighting quantization bias with bias
Finkelstein, A., Almog, U., and Grobman, M · 2019
Cited alongside, same era.
Same, same but different: Recovering neural network quantization error through weight factorization
Meller, E., Finkelstein, A., Almog, U., and Grobman, M · 2019
Cited alongside, same era.
Data-free quantization through weight equalization and bias correction
Nagel, M., van Baalen, M., Blankevoort, T., and Welling, M · 2019
Cited alongside, same era.
Dsconv: Efficient convolution operator
Nascimento, M. G. d., Fawcett, R., and Prisacariu, V. A · 2019
Cited alongside, same era.
A quantization-friendly separable convolution for mobilenets
Sheng, T., Feng, C., Zhuo, S., Zhang, X., Shen, L., and Aleksic, M
Cited in the paper.
Rouhani, B., Lo, D., Zhao, R., Liu, M., Fowers, J., Ovtcharov, K., Vinogradsky, A., Massengill, S., Yang, L., Bittner, R., Forin, A., Zhu, H., Na, T., Patel, P., Che, S., Koppaka, L. C., Song, X., Som, S., Das, K., Tiwary, S., Reinhardt, S., Lanka, S., Chung, E., and Burger, D · 2020
Later among the works it cites.
Efficientdet: Scalable and efficient object detection, 2020
Tan, M., Pang, R., and Le, Q. V · 2020
Later among the works it cites.
Mixed precision dnns: All you need is a good parametrization
Uhlich, S., Mauch, L., Cardinaux, F., Yoshiyama, K., Garcia, J. A., Tiedemann, S., Kemp, T., and Nakamura, A · 2020
Later among the works it cites.
Bayesian bits: Unifying quantization and pruning
van Baalen, M., Louizos, C., Nagel, M., Amjad, R. A., Wang, Y., Blankevoort, T., and Welling, M · 2020
Later among the works it cites.