Understand
To facilitate efficient embedded and hardware implementations of deep neural networks (DNNs), two important categories of DNN model compression techniques: weight pruning and weight quantization are investigated.
- The former leverages the redundancy in the number of weights, whereas the latter leverages the redundancy in bit representation of weights.
- However, there lacks a systematic framework of joint weight pruning and quantization of DNNs, thereby limiting the available model compression ratio.
- Moreover, the computation reduction, energy efficiency improvement, and hardware performance overhead need to be accounted for besides simply model size reduction.
Built on
Nothing clear enough to list yet.
Similar
Nothing clear enough to list yet.
Then
Nothing clear enough to list yet.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…