2021

BSQ: Exploring Bit-Level Sparsity for Mixed-Precision Neural Network Quantization

Yang, Huanrui, Duan, Lin, Chen, Yiran et al.

Understand

Mixed-precision quantization can potentially achieve the optimal tradeoff between performance and compression rate of deep neural networks, and thus, have been widely investigated.

  • However, it lacks a systematic method to determine the exact quantization scheme.
  • Previous methods either examine only a small manually-designed search space or utilize a cumbersome neural architecture search to explore the vast search space.
  • These approaches cannot lead to an optimal quantization scheme efficiently.

Reading the bibliography…