Pith. sign in

REVIEW 3 cited by

GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.13824 v1 pith:EX6ZE7Z6 submitted 2024-05-22 cs.CV

classification cs.CV
keywords gmmformertext-clipmatchingmomentsprvruncertainty-awarerelevantvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given a text query, partially relevant video retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments. Due to the lack of moment annotations, the uncertainty lying in clip modeling and text-clip correspondence leads to major challenges. Despite the great progress, existing solutions either sacrifice efficiency or efficacy to capture varying and uncertain video moments. What's worse, few methods have paid attention to the text-clip matching pattern under such uncertainty, exposing the risk of semantic collapse. To address these issues, we present GMMFormer v2, an uncertainty-aware framework for PRVR. For clip modeling, we improve a strong baseline GMMFormer with a novel temporal consolidation module upon multi-scale contextual features, which maintains efficiency and improves the perception for varying moments. To achieve uncertainty-aware text-clip matching, we upgrade the query diverse loss in GMMFormer to facilitate fine-grained uniformity and propose a novel optimal matching loss for fine-grained text-clip alignment. Their collaboration alleviates the semantic collapse phenomenon and neatly promotes accurate correspondence between texts and moments. We conduct extensive experiments and ablation studies on three PRVR benchmarks, demonstrating remarkable improvement of GMMFormer v2 compared to the past SOTA competitor and the versatility of uncertainty-aware text-clip matching for PRVR. Code is available at \url{https://github.com/huangmozhi9527/GMMFormer_v2}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning

    cs.CV 2025-09 conditional novelty 6.0 of 10

    RAL models PRVR with probabilistic Gaussian alignment plus confidence-weighted word matching, improving SumR by 9.7 over prior best on TVR.

  2. ProPy: Building Interactive Prompt Pyramids upon CLIP for Partially Relevant Video Retrieval

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A hierarchical prompt pyramid over CLIP with ancestor-descendant attention improves partially relevant video retrieval.

  3. Uneven Event Modeling for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    UEM retrieves partially relevant videos by adaptively segmenting frames into uneven events and refining the best-matching event with text-conditioned attention.

Pith tools