Pith. sign in

REVIEW 3 cited by

FEED: Feature-level Ensemble for Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.10754 v1 pith:TJY2AGOB submitted 2019-09-24 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords knowledgedistillationensemblestudentteacherfeedmultiplenetwork
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Knowledge Distillation (KD) aims to transfer knowledge in a teacher-student framework, by providing the predictions of the teacher network to the student network in the training stage to help the student network generalize better. It can use either a teacher with high capacity or {an} ensemble of multiple teachers. However, the latter is not convenient when one wants to use feature-map-based distillation methods. For a solution, this paper proposes a versatile and powerful training algorithm named FEature-level Ensemble for knowledge Distillation (FEED), which aims to transfer the ensemble knowledge using multiple teacher networks. We introduce a couple of training algorithms that transfer ensemble knowledge to the student at the feature map level. Among the feature-map-based distillation methods, using several non-linear transformations in parallel for transferring the knowledge of the multiple teacher{s} helps the student find more generalized solutions. We name this method as parallel FEED, andexperimental results on CIFAR-100 and ImageNet show that our method has clear performance enhancements, without introducing any additional parameters or computations at test time. We also show the experimental results of sequentially feeding teacher's information to the student, hence the name sequential FEED, and discuss the lessons obtained. Additionally, the empirical results on measuring the reconstruction errors at the feature map give hints for the enhancements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. All-in-One: Transferring Vision Foundation Models into Stereo Matching

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AIO-Stereo transfers and selectively fuses knowledge from DINOv2, SAM, and Depth Anything into a stereo matching network, achieving top results on Middlebury and ETH3D.

  2. Distillation of Diffusion Features for Semantic Correspondence

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A DINOv2 student trained with LoRA to imitate DINOv2-plus-SDXL-Turbo similarity maps, then fine-tuned on 3D-derived correspondences, sets new state-of-the-art on three semantic correspondence benchmarks.

  3. SFedKD: Sequential Federated Learning with Discrepancy-Aware Multi-Teacher Knowledge Distillation

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SFedKD uses discrepancy-weighted multi-teacher knowledge distillation and greedy teacher selection to reduce catastrophic forgetting in sequential federated learning.

Pith tools