Pith. sign in

REVIEW 5 cited by

QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.09609 v2 pith:4ZXCQWSE submitted 2021-07-20 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords videoshighlightsmomentsvideomomentqueriesqueryactivities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Detecting customized moments and highlights from videos given natural language (NL) user queries is an important but under-studied topic. One of the challenges in pursuing this direction is the lack of annotated data. To address this issue, we present the Query-based Video Highlights (QVHIGHLIGHTS) dataset. It consists of over 10,000 YouTube videos, covering a wide range of topics, from everyday activities and travel in lifestyle vlog videos to social and political activities in news videos. Each video in the dataset is annotated with: (1) a human-written free-form NL query, (2) relevant moments in the video w.r.t. the query, and (3) five-point scale saliency scores for all query-relevant clips. This comprehensive annotation enables us to develop and evaluate systems that detect relevant moments as well as salient highlights for diverse, flexible user queries. We also present a strong baseline for this task, Moment-DETR, a transformer encoder-decoder model that views moment retrieval as a direct set prediction problem, taking extracted video and query representations as inputs and predicting moment coordinates and saliency scores end-to-end. While our model does not utilize any human prior, we show that it performs competitively when compared to well-engineered architectures. With weakly supervised pretraining using ASR captions, MomentDETR substantially outperforms previous methods. Lastly, we present several ablations and visualizations of Moment-DETR. Data and code is publicly available at https://github.com/jayleicn/moment_detr

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MMEB-V3: Measuring the Performance Gaps of Omni-Modality Embedding Models

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    On 190 tasks and a 12-direction cross-modal diagnostic, seven embedding models frequently fail to honor explicit target-modality instructions: retrieval is biased toward the query modality and instruction-induced shif...

  2. SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A new benchmark links full soccer match broadcasts from SoccerNet with official league highlight summaries, plus a baseline model and a summary-length-constrained metric.

  3. The Repeated-Stimulus Confound in Electroencephalography

    q-bio.NC 2025-08 unverdicted novelty 5.0 of 10

    The repeated-stimulus confound, where models are trained and tested on repeated presentations of identical stimuli, inflates reported EEG decoding accuracies by an estimated 4.46 to 7.42 percent.

  4. Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A video-language model that combines adaptive frame sampling, explicit timestamps, and a reinforcement-learning reward for refusing irrelevant queries, beating prior methods on QVHighlights by about 3.5%.

  5. MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering

    cs.CV 2025-06

Pith tools