Pith. sign in

REVIEW 5 cited by

Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.10300 v3 pith:MGAOBBNQ submitted 2023-12-16 cs.CV

classification cs.CV
keywords videomulti-shotunderstandingcomprehensivesummariesvideosbenchmarkcaptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A short clip of video may contain progression of multiple events and an interesting story line. A human need to capture both the event in every shot and associate them together to understand the story behind it. In this work, we present a new multi-shot video understanding benchmark Shot2Story with detailed shot-level captions, comprehensive video summaries and question-answering pairs. To facilitate better semantic understanding of videos, we provide captions for both visual signals and human narrations. We design several distinct tasks including single-shot video captioning, multi-shot video summarization, and multi-shot video question answering. Preliminary experiments show some challenges to generate a long and comprehensive video summary for multi-shot videos. Nevertheless, the generated imperfect summaries can already achieve competitive performance on existing video understanding tasks such as video question-answering, promoting an under-explored setting of video understanding with detailed summaries.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting

    cs.CV 2025-07 conditional novelty 7.0 of 10

    HIPPO-Video contributes 2,040 LLM-simulated watch-history and saliency-score pairs, and the HiPHer model uses these histories to beat generic and query-based baselines on the new benchmark.

  2. AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 2B-parameter video-language model using an RWKV linear-RNN backbone and sorted token merging achieves competitive long-video QA accuracy with far lower memory cost than transformer-based models.

  3. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  4. NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding

    cs.HC 2025-08 conditional novelty 5.0 of 10

    NoteIt converts instructional videos into interactive notes that preserve chapter and step structure and key visual and verbal information, and users significantly preferred it over a commercial baseline.

  5. Comparing Learning Paradigms for Egocentric Video Summarization

    cs.CV 2025-06 reject novelty 4.0 of 10

    A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.

Pith tools