Pith. sign in

REVIEW 3 cited by

FROSTER: Frozen CLIP Is A Strong Teacher for Open-Vocabulary Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.03241 v1 pith:3HYT5FB2 submitted 2024-02-05 cs.CV cs.LG

classification cs.CVcs.LG
keywords clipactionfrosterrecognitionfeatureopen-vocabularydistillationresidual
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce FROSTER, an effective framework for open-vocabulary action recognition. The CLIP model has achieved remarkable success in a range of image-based tasks, benefiting from its strong generalization capability stemming from pretaining on massive image-text pairs. However, applying CLIP directly to the open-vocabulary action recognition task is challenging due to the absence of temporal information in CLIP's pretraining. Further, fine-tuning CLIP on action recognition datasets may lead to overfitting and hinder its generalizability, resulting in unsatisfactory results when dealing with unseen actions. To address these issues, FROSTER employs a residual feature distillation approach to ensure that CLIP retains its generalization capability while effectively adapting to the action recognition task. Specifically, the residual feature distillation treats the frozen CLIP model as a teacher to maintain the generalizability exhibited by the original CLIP and supervises the feature learning for the extraction of video-specific features to bridge the gap between images and videos. Meanwhile, it uses a residual sub-network for feature distillation to reach a balance between the two distinct objectives of learning generalizable and video-specific features. We extensively evaluate FROSTER on open-vocabulary action recognition benchmarks under both base-to-novel and cross-dataset settings. FROSTER consistently achieves state-of-the-art performance on all datasets across the board. Project page: https://visual-ai.github.io/froster.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Richer text prompts from LLM synonyms and cleaner image regions from activation maps improve zero-shot vision-language classification.

  2. Single Domain Generalization for Few-Shot Counting via Universal Representation Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    URM distills CLIP vision-language representations into learnable prototypes for few-shot counting, improving single-domain generalization on unseen datasets.

  3. MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Combining joint, limb, RGB, Taylor-video, optical-flow, and depth streams with two video backbones and a validation-tuned weighted ensemble reaches 73.213% top-1 accuracy on iMiGUE, the best MiGA challenge result to date.

Pith tools