Pith. sign in

REVIEW 1 cited by

M2-CLIP: A Multimodal, Multi-task Adapting Framework for Video Action Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.11649 v1 pith:RCEMPZ3C submitted 2024-01-22 cs.CV

classification cs.CV
keywords multimodalperformancesupervisedframeworkgeneralizationmulti-taskstrongtemporal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, the rise of large-scale vision-language pretrained models like CLIP, coupled with the technology of Parameter-Efficient FineTuning (PEFT), has captured substantial attraction in video action recognition. Nevertheless, prevailing approaches tend to prioritize strong supervised performance at the expense of compromising the models' generalization capabilities during transfer. In this paper, we introduce a novel Multimodal, Multi-task CLIP adapting framework named \name to address these challenges, preserving both high supervised performance and robust transferability. Firstly, to enhance the individual modality architectures, we introduce multimodal adapters to both the visual and text branches. Specifically, we design a novel visual TED-Adapter, that performs global Temporal Enhancement and local temporal Difference modeling to improve the temporal representation capabilities of the visual encoder. Moreover, we adopt text encoder adapters to strengthen the learning of semantic label information. Secondly, we design a multi-task decoder with a rich set of supervisory signals to adeptly satisfy the need for strong supervised performance and generalization within a multimodal framework. Experimental results validate the efficacy of our approach, demonstrating exceptional performance in supervised learning while maintaining strong generalization in zero-shot scenarios.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. T-MASK: Temporal Masking for Probing Foundation Models across Camera Views in Driver Monitoring

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    T-MASK uses temporal token masking to improve cross-view driver activity recognition with foundation models, claiming gains of +1.23% over probing and +8.0% over PEFT on Drive&Act.

Pith tools