Pith. sign in

REVIEW 3 cited by

Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.14040 v2 pith:A4HH3ZBE submitted 2024-05-22 cs.MM

classification cs.MM
keywords videonarrationsstorylinegenerategenerationintroducestorytellingsynchronized
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video captioning and video story generation have made some progress. However, in practical applications, we typically require synchronized narrations for ongoing visual scenes. In this work, we introduce a new task of Synchronized Video Storytelling, which aims to generate synchronous and informative narrations for videos. These narrations, associated with each video clip, should relate to the visual content, integrate relevant knowledge, and have an appropriate word count corresponding to the clip's duration. Specifically, a structured storyline is beneficial to guide the generation process, ensuring coherence and integrity. To support the exploration of this task, we introduce a new benchmark dataset E-SyncVidStory with rich annotations. Since existing Multimodal LLMs are not effective in addressing this task in one-shot or few-shot settings, we propose a framework named VideoNarrator that can generate a storyline for input videos and simultaneously generate narrations with the guidance of the generated or predefined storyline. We further introduce a set of evaluation metrics to thoroughly assess the generation. Both automatic and human evaluations validate the effectiveness of our approach. Our dataset, codes, and evaluations will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AdsQA: Towards Advertisement Video Understanding

    cs.CV 2025-09 conditional novelty 6.0 of 10

    AdsQA adds an ad-video question-answering benchmark and ReAd-R, a GRPO-trained model that beats 7B baselines but not larger closed models.

  2. Movie2Story: A framework for understanding videos and telling stories in the form of novel text

    cs.CV 2024-12 reject novelty 4.0 of 10

    MSBench evaluates video-plus-audio to novel-style story generation; the M2S pipeline combines existing video, speech, emotion, and speaker tools with an LLM and reportedly beats video-only baselines.

  3. Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

    cs.CV 2025-01

Pith tools