Pith. sign in

REVIEW 3 cited by

Ensemble Knowledge Distillation for Learning Improved and Efficient Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.08097 v3 pith:5VJHHRAA submitted 2019-09-17 cs.CV cs.LG

classification cs.CVcs.LG
keywords networkensembleknowledgelearningmodelnetworksstudentbranches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensemble models comprising of deep Convolutional Neural Networks (CNN) have shown significant improvements in model generalization but at the cost of large computation and memory requirements. In this paper, we present a framework for learning compact CNN models with improved classification performance and model generalization. For this, we propose a CNN architecture of a compact student model with parallel branches which are trained using ground truth labels and information from high capacity teacher networks in an ensemble learning fashion. Our framework provides two main benefits: i) Distilling knowledge from different teachers into the student network promotes heterogeneity in feature learning at different branches of the student network and enables the network to learn diverse solutions to the target problem. ii) Coupling the branches of the student network through ensembling encourages collaboration and improves the quality of the final predictions by reducing variance in the network outputs. Experiments on the well established CIFAR-10 and CIFAR-100 datasets show that our Ensemble Knowledge Distillation (EKD) improves classification accuracy and model generalization especially in situations with limited training data. Experiments also show that our EKD based compact networks outperform in terms of mean accuracy on the test datasets compared to state-of-the-art knowledge distillation based methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Clustering of Seafloor Imagery for Interpretation during Long-Term AUV Operations

    cs.CV 2025-09 conditional novelty 6.0 of 10

    An online clustering framework with dynamic cluster splitting/merging and fixed-size representative sampling achieves about 0.68 average F1 on three seafloor image datasets with bounded runtime.

  2. Right Time to Learn:Promoting Generalization via Bio-inspired Spacing Effect in Knowledge Distillation

    cs.LG 2025-02 conditional novelty 5.0 of 10

    Training the teacher a small number of steps ahead of the student and freezing it during distillation improves student generalization by up to 3.4% on image benchmarks.

  3. A Comprehensive Survey on Imbalanced Data Learning

    cs.LG 2025-02 conditional novelty 3.0 of 10

    A structured survey and benchmark that groups imbalanced data learning methods into data re-balancing, feature representation, training strategy, and ensemble learning.

Pith tools