Pith. sign in

REVIEW 2 cited by

A Hierarchical Multi-Modal Encoder for Moment Localization in Video Corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.09046 v2 pith:JSTND3TF submitted 2020-11-18 cs.CV cs.CL

classification cs.CVcs.CL
keywords videolocalizationsegmentchallengingcorpusdifferentencoderhierarchical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Identifying a short segment in a long video that semantically matches a text query is a challenging task that has important application potentials in language-based video search, browsing, and navigation. Typical retrieval systems respond to a query with either a whole video or a pre-defined video segment, but it is challenging to localize undefined segments in untrimmed and unsegmented videos where exhaustively searching over all possible segments is intractable. The outstanding challenge is that the representation of a video must account for different levels of granularity in the temporal domain. To tackle this problem, we propose the HierArchical Multi-Modal EncodeR (HAMMER) that encodes a video at both the coarse-grained clip level and the fine-grained frame level to extract information at different scales based on multiple subtasks, namely, video retrieval, segment temporal localization, and masked language modeling. We conduct extensive experiments to evaluate our model on moment localization in video corpus on ActivityNet Captions and TVR datasets. Our approach outperforms the previous methods as well as strong baselines, establishing new state-of-the-art for this task.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HLFormer: Enhancing Partially Relevant Video Retrieval with Hyperbolic Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    HLFormer adds hybrid Euclidean and Lorentz attention plus a partial-order cone loss to partially relevant video retrieval and reports the best total recall on ActivityNet Captions, Charades-STA, and TVR.

  2. Ambiguity-Restrained Text-Video Representation Learning for Partially Relevant Video Retrieval

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ARL detects ambiguous text-video pairs using uncertainty and similarity, then trains retrieval models with ambiguity-aware contrastive and triplet losses, achieving state-of-the-art on TVR and ActivityNet Captions.

Pith tools