Pith. sign in

REVIEW 1 cited by

How to Train the Teacher Model for Effective Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.18041 v1 pith:WLIREWVI submitted 2024-07-25 cs.LG

classification cs.LG
keywords teacherbcpdstudentlossoutputtraineddistillationerror
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, it was shown that the role of the teacher in knowledge distillation (KD) is to provide the student with an estimate of the true Bayes conditional probability density (BCPD). Notably, the new findings propose that the student's error rate can be upper-bounded by the mean squared error (MSE) between the teacher's output and BCPD. Consequently, to enhance KD efficacy, the teacher should be trained such that its output is close to BCPD in MSE sense. This paper elucidates that training the teacher model with MSE loss equates to minimizing the MSE between its output and BCPD, aligning with its core responsibility of providing the student with a BCPD estimate closely resembling it in MSE terms. In this respect, through a comprehensive set of experiments, we demonstrate that substituting the conventional teacher trained with cross-entropy loss with one trained using MSE loss in state-of-the-art KD methods consistently boosts the student's accuracy, resulting in improvements of up to 2.6\%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning an Adaptive and View-Invariant Vision Transformer for Real-Time UAV Tracking

    cs.CV 2024-12 conditional novelty 5.0 of 10

    An adaptive block-activation ViT and a mutual-information multi-teacher distillation variant achieve state-of-the-art speed/accuracy trade-offs on six UAV tracking benchmarks.

Pith tools