REVIEW 13 cited by
CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Video-text retrieval plays an essential role in multi-modal research and has been widely used in many real-world web applications. The CLIP (Contrastive Language-Image Pre-training), an image-language pre-training model, has demonstrated the power of visual concepts learning from web collected image-text datasets. In this paper, we propose a CLIP4Clip model to transfer the knowledge of the CLIP model to video-language retrieval in an end-to-end manner. Several questions are investigated via empirical studies: 1) Whether image feature is enough for video-text retrieval? 2) How a post-pretraining on a large-scale video-text dataset based on the CLIP affect the performance? 3) What is the practical mechanism to model temporal dependency between video frames? And 4) The Hyper-parameters sensitivity of the model on video-text retrieval task. Extensive experimental results present that the CLIP4Clip model transferred from the CLIP can achieve SOTA results on various video-text retrieval datasets, including MSR-VTT, MSVC, LSMDC, ActivityNet, and DiDeMo. We release our code at https://github.com/ArrowLuo/CLIP4Clip.
Forward citations
Cited by 13 Pith papers
-
Trajectory-aware Cross-view Geo-localization with Sequential Observations
TrajLoc combines video and text route descriptions with trajectory geometry to retrieve satellite images, introducing the SeqGeo-VL benchmark and improving state-of-the-art on both video and text geo-localization.
-
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
A camera-trap TVR benchmark of 135 ethology queries plus an interpretable SALMA-to-JSON plus constrained-LLM-parser pipeline yields 34% set F1, beating zero-shot VLMs at 18%.
-
LAVIFT: Latent-Action-Guided Vision Fine-Tuning for Surgical Interaction Recognition
Latent-action modeling (inverse dynamics plus forward world model) with a patch-level anti-collapse regularizer improves surgical action-triplet recognition and makes encoder change features land more on instrument-ti...
-
CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval
CLaMR jointly encodes four video modalities in a vision-language model and uses token-level, per-modality matching to retrieve the right video and the right modality for a query.
-
Knowledge-guided Disentanglement with Atomic Actions for Action Recognition
LLM-generated atomic action descriptions injected into scene-graph node features improve prompt-guided disentanglement for multi-label action recognition, with reported gains on Charades-oracle and SportsHHI but not o...
-
Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification
A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.
-
MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval
MSAM introduces two drone-video/text datasets and a CLIP-based multi-semantic pooling model that reports 0.6–3.8 point R@1 gains over earlier video-text retrieval methods.
-
Regularizing Subspace Redundancy of Low-Rank Adaptation
ReSoRA adds a penalty that reduces redundancy among rank-1 subspaces of LoRA-style adapters, producing modest accuracy improvements on vision-language retrieval and visual classification.
-
Exploring Object Status Recognition for Recipe Progress Tracking in Non-Visual Cooking
Recipe-derived object status phrases added to vision-language matching improve recipe-step prediction by 20 to 26 accuracy points on instructional and real-world non-visual cooking videos.
-
Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing
Multi-scale temporal convolutions plus capsule routing improve video-text moment localization to 42.9% R@0.5 and 41.1% mIoU on ActivityNet Captions.
-
Video Understanding by Design: How Datasets Shape Video Models
A dataset-centric framework that explains video architectures as responses to structural properties of benchmark datasets.
-
ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing
ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.
-
Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
A structured survey of 2023-2025 LLM and VLM methods for crash detection in video, with notable internal inconsistencies in reported numbers.
Discussion (0). Continue with ORCID to comment.