REVIEW 7 cited by
Apprentice: Using Knowledge Distillation Techniques To Improve Low-Precision Network Accuracy
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deep learning networks have achieved state-of-the-art accuracies on computer vision workloads like image classification and object detection. The performant systems, however, typically involve big models with numerous parameters. Once trained, a challenging aspect for such top performing models is deployment on resource constrained inference systems - the models (often deep networks or wide networks or both) are compute and memory intensive. Low-precision numerics and model compression using knowledge distillation are popular techniques to lower both the compute requirements and memory footprint of these deployed models. In this paper, we study the combination of these two techniques and show that the performance of low-precision networks can be significantly improved by using knowledge distillation techniques. Our approach, Apprentice, achieves state-of-the-art accuracies using ternary precision and 4-bit precision for variants of ResNet architecture on ImageNet dataset. We present three schemes using which one can apply knowledge distillation techniques to various stages of the train-and-deploy pipeline.
Forward citations
Cited by 7 Pith papers
-
RBCN: Rectified Binary Convolutional Networks for Enhancing the Performance of 1-bit DCNNs
A GAN-guided rectified training scheme narrows the accuracy gap between binary and full-precision convolutional networks on classification and tracking.
-
Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural Networks
A trainable soft-tanh quantizer with learned steepness and clipping improves 1-4 bit network accuracy and yields fast ARM kernels.
-
GDRQ: Group-based Distribution Reshaping for Quantization
A training-time method that reshapes weights and activations toward uniform distributions and splits filters into groups with separate scales, yielding strong low-bit quantization results on classification, detection,...
-
Quantize the Target, Quantize the Drafter: Efficient Inference with Qwen3.5-4B
Quantization-aware distillation of an INT4 Qwen3.5-4B target plus a two-stage-trained, GPTQ-quantized, SWA-equipped DFlash drafter yields 6.978× average speedup on A10G while meeting quality thresholds.
-
Enhancing Quantization-Aware Training on Edge Devices via Relative Entropy Coreset Selection and Cascaded Layer Correction
Using a relative-entropy coreset and layer-wise feature alignment, QuaRC improves 2-bit quantization-aware training accuracy by up to 5.72 points on ImageNet-1K when training on 1% of the data.
-
Pan-infection Foundation Framework Enables Multiple Pathogen Prediction
A teacher-student knowledge distillation framework trained on 11,247 blood transcriptomes reports high AUCs for pan-infection, four pathogens, and sepsis diagnosis.
-
Multi-Modality Distillation via Learning the teacher's modality-level Gram Matrix
Proposes a modality relation distillation method that transfers teacher modality relationships via the modality-level Gram Matrix.
Discussion (0). Continue with ORCID to comment.