Pith. sign in

REVIEW 1 cited by

The Modality Focusing Hypothesis: Towards Understanding Crossmodal Knowledge Distillation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2206.06487 v3 pith:WIUWXGUT submitted 2022-06-13 cs.CV cs.LG

classification cs.CVcs.LG
keywords crossmodalknowledgemodalitydistillationhypothesistransferfocusinglearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Crossmodal knowledge distillation (KD) extends traditional knowledge distillation to the area of multimodal learning and demonstrates great success in various applications. To achieve knowledge transfer across modalities, a pretrained network from one modality is adopted as the teacher to provide supervision signals to a student network learning from another modality. In contrast to the empirical success reported in prior works, the working mechanism of crossmodal KD remains a mystery. In this paper, we present a thorough understanding of crossmodal KD. We begin with two case studies and demonstrate that KD is not a universal cure in crossmodal knowledge transfer. We then present the modality Venn diagram to understand modality relationships and the modality focusing hypothesis revealing the decisive factor in the efficacy of crossmodal KD. Experimental results on 6 multimodal datasets help justify our hypothesis, diagnose failure cases, and point directions to improve crossmodal knowledge transfer in the future.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Domain Adaptation-Based Crossmodal Knowledge Distillation for 3D Semantic Segmentation

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A 3D self-calibrated convolution module plus feature and semantic distillation lets a LiDAR network learn from 2D image teachers without 3D labels.

Pith tools