Pith. sign in

REVIEW 2 cited by

Learning to Quantize Deep Networks by Optimizing Quantization Intervals with Task Loss

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1808.05779 v3 pith:XIJWF43S submitted 2018-08-17 cs.CV

classification cs.CV
keywords networksaccuracyquantizationquantizequantizeractivationsbit-widthbit-widths
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reducing bit-widths of activations and weights of deep networks makes it efficient to compute and store them in memory, which is crucial in their deployments to resource-limited devices, such as mobile phones. However, decreasing bit-widths with quantization generally yields drastically degraded accuracy. To tackle this problem, we propose to learn to quantize activations and weights via a trainable quantizer that transforms and discretizes them. Specifically, we parameterize the quantization intervals and obtain their optimal values by directly minimizing the task loss of the network. This quantization-interval-learning (QIL) allows the quantized networks to maintain the accuracy of the full-precision (32-bit) networks with bit-width as low as 4-bit and minimize the accuracy degeneration with further bit-width reduction (i.e., 3 and 2-bit). Moreover, our quantizer can be trained on a heterogeneous dataset, and thus can be used to quantize pretrained networks without access to their training data. We demonstrate the effectiveness of our trainable quantizer on ImageNet dataset with various network architectures such as ResNet-18, -34 and AlexNet, on which it outperforms existing methods to achieve the state-of-the-art accuracy.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A trainable soft-tanh quantizer with learned steepness and clipping improves 1-4 bit network accuracy and yields fast ARM kernels.

  2. Efficient Deep Neural Networks

    cs.CV 2019-08 conditional novelty 4.0 of 10

    A dissertation showing that deep learning can be made practical on edge devices through four complementary routes: model, data, hardware, and design efficiency.

Pith tools