Pith. sign in

REVIEW 7 cited by

Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1711.05852 v1 pith:ZQKJNDAJ submitted 2017-11-15 cs.LG cs.CVcs.NE

classification cs.LGcs.CVcs.NE
keywords techniquesdistillationknowledgemodelsnetworkslow-precisionaccuraciesapprentice
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep learning networks have achieved state-of-the-art accuracies on computer vision workloads like image classification and object detection. The performant systems, however, typically involve big models with numerous parameters. Once trained, a challenging aspect for such top performing models is deployment on resource constrained inference systems - the models (often deep networks or wide networks or both) are compute and memory intensive. Low-precision numerics and model compression using knowledge distillation are popular techniques to lower both the compute requirements and memory footprint of these deployed models. In this paper, we study the combination of these two techniques and show that the performance of low-precision networks can be significantly improved by using knowledge distillation techniques. Our approach, Apprentice, achieves state-of-the-art accuracies using ternary precision and 4-bit precision for variants of ResNet architecture on ImageNet dataset. We present three schemes using which one can apply knowledge distillation techniques to various stages of the train-and-deploy pipeline.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A GAN-guided rectified training scheme narrows the accuracy gap between binary and full-precision convolutional networks on classification and tracking.

  2. Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks

    cs.CV 2019-08 conditional novelty 6.0 of 10

    A trainable soft-tanh quantizer with learned steepness and clipping improves 1-4 bit network accuracy and yields fast ARM kernels.

  3. GDRQ: Group-based Distribution Reshaping for Quantization

    cs.CV 2019-08 conditional novelty 5.0 of 10

    A training-time method that reshapes weights and activations toward uniform distributions and splits filters into groups with separate scales, yielding strong low-bit quantization results on classification, detection,...

  4. Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Quantization-aware distillation of an INT4 Qwen3.5-4B target plus a two-stage-trained, GPTQ-quantized, SWA-equipped DFlash drafter yields 6.978× average speedup on A10G while meeting quality thresholds.

  5. Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Using a relative-entropy coreset and layer-wise feature alignment, QuaRC improves 2-bit quantization-aware training accuracy by up to 5.72 points on ImageNet-1K when training on 1% of the data.

  6. Pan-infection Foundation Framework Enables Multiple Pathogen Prediction

    cs.LG 2024-12 reject novelty 4.0 of 10

    A teacher-student knowledge distillation framework trained on 11,247 blood transcriptomes reports high AUCs for pan-infection, four pathogens, and sepsis diagnosis.

  7. Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix

    cs.AI 2021-12 unverdicted novelty 4.0 of 10

    Proposes a modality relation distillation method that transfers teacher modality relationships via the modality-level Gram Matrix.

Pith tools