Pith. sign in

REVIEW 1 cited by

Zero-Shot Long-Form Video Understanding through Screenplay

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.17309 v1 pith:OVW74IGH submitted 2024-06-25 cs.CV

classification cs.CV
keywords videolong-formunderstandingaccuracybreakpointcontentinformationmm-screenplayer
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The Long-form Video Question-Answering task requires the comprehension and analysis of extended video content to respond accurately to questions by utilizing both temporal and contextual information. In this paper, we present MM-Screenplayer, an advanced video understanding system with multi-modal perception capabilities that can convert any video into textual screenplay representations. Unlike previous storytelling methods, we organize video content into scenes as the basic unit, rather than just visually continuous shots. Additionally, we developed a ``Look Back'' strategy to reassess and validate uncertain information, particularly targeting breakpoint mode. MM-Screenplayer achieved highest score in the CVPR'2024 LOng-form VidEo Understanding (LOVEU) Track 1 Challenge, with a global accuracy of 87.5% and a breakpoint accuracy of 68.8%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VEU-Bench: Towards Comprehensive Understanding of Video Editing

    cs.CV 2025-04 conditional novelty 6.0 of 10

    A new 19-task video editing benchmark shows that current video LLMs struggle to understand editing concepts, and a model fine-tuned on the benchmark improves both editing and general video reasoning.

Pith tools