Pith. sign in

REVIEW 13 cited by

CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08860 v2 pith:NLXDZ2HY submitted 2021-04-18 cs.CV

classification cs.CV
keywords clipmodelretrievalvideo-textclip4clipdatasetsempiricalpre-training
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training), an image-language pre-training model, has demonstrated the power of visual concepts learning from web collected image-text datasets. In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner. Several questions are investigated via empirical studies: 1) Whether image feature is enough for video-text retrieval? 2) How a post-pretraining on a large-scale video-text dataset based on the CLIP affect the performance? 3) What is the practical mechanism to model temporal dependency between video frames? And 4) The Hyper-parameters sensitivity of the model on video-text retrieval task. Extensive experimental results present that the CLIP4Clip model transferred from the CLIP can achieve SOTA results on various video-text retrieval datasets, including MSR-VTT, MSVC, LSMDC, ActivityNet, and DiDeMo. We release our code at https://github.com/ArrowLuo/CLIP4Clip.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Trajectory-aware Cross-view Geo-localization with Sequential Observations

    cs.CV 2026-07 conditional novelty 7.0 of 10

    TrajLoc combines video and text route descriptions with trajectory geometry to retrieve satellite images, introducing the SeqGeo-VL benchmark and improving state-of-the-art on both video and text geo-localization.

  2. Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data

    cs.CV 2026-07 accept novelty 7.0 of 10

    A camera-trap TVR benchmark of 135 ethology queries plus an interpretable SALMA-to-JSON plus constrained-LLM-parser pipeline yields 34% set F1, beating zero-shot VLMs at 18%.

  3. LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Latent-action modeling (inverse dynamics plus forward world model) with a patch-level anti-collapse regularizer improves surgical action-triplet recognition and makes encoder change features land more on instrument-ti...

  4. CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CLaMR jointly encodes four video modalities in a vision-language model and uses token-level, per-modality matching to retrieve the right video and the right modality for a query.

  5. Knowledge-guided Disentanglement with Atomic Actions for Action Recognition

    cs.CV 2026-07 conditional novelty 5.0 of 10

    LLM-generated atomic action descriptions injected into scene-graph node features improve prompt-guided disentanglement for multi-label action recognition, with reported gains on Charades-oracle and SportsHHI but not o...

  6. Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.

  7. MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval

    cs.CV 2025-10 conditional novelty 5.0 of 10

    MSAM introduces two drone-video/text datasets and a CLIP-based multi-semantic pooling model that reports 0.6–3.8 point R@1 gains over earlier video-text retrieval methods.

  8. Regularizing Subspace Redundancy of Low-Rank Adaptation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    ReSoRA adds a penalty that reduces redundancy among rank-1 subspaces of LoRA-style adapters, producing modest accuracy improvements on vision-language retrieval and visual classification.

  9. Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking

    cs.AI 2025-07 conditional novelty 5.0 of 10

    Recipe-derived object status phrases added to vision-language matching improve recipe-step prediction by 20 to 26 accuracy points on instructional and real-world non-visual cooking videos.

  10. Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing

    cs.CV 2026-07 conditional novelty 4.5 of 10

    Multi-scale temporal convolutions plus capsule routing improve video-text moment localization to 42.9% R@0.5 and 41.1% mIoU on ActivityNet Captions.

  11. Video Understanding by Design: How Datasets Shape Video Models

    cs.CV 2025-09 reject novelty 4.0 of 10

    A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.

  12. ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing

    cs.MM 2025-06 reject novelty 4.0 of 10

    ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.

  13. Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.

Pith tools