Pith. sign in

REVIEW 2 cited by

VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.16930 v1 pith:DYYBCIIY submitted 2024-08-29 cs.CV

classification cs.CV
keywords knowledgemodeldistillationsupervisiontextvisuallong-tailnovel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For visual recognition, knowledge distillation typically involves transferring knowledge from a large, well-trained teacher model to a smaller student model. In this paper, we introduce an effective method to distill knowledge from an off-the-shelf vision-language model (VLM), demonstrating that it provides novel supervision in addition to those from a conventional vision-only teacher model. Our key technical contribution is the development of a framework that generates novel text supervision and distills free-form text into a vision encoder. We showcase the effectiveness of our approach, termed VLM-KD, across various benchmark datasets, showing that it surpasses several state-of-the-art long-tail visual classifiers. To our knowledge, this work is the first to utilize knowledge distillation with text supervision generated by an off-the-shelf VLM and apply it to vanilla randomly initialized vision encoders.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Answer-conditioned CoT distillation lets a 3B VLM outperform direct LoRA and sometimes GPT-4.1 on industrial few-shot classification using 18–30 labeled images.

  2. HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A progressive two-stage knowledge distillation framework (HKD4VLM) reports first-place F1 scores of 98.2% and 98.4% on multimodal hallucination and factuality detection, but its ablation lacks a directly fine-tuned baseline.

Pith tools