Pith. sign in

REVIEW 7 cited by

VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15091 v2 pith:KO2EJIUP submitted 2023-09-26 cs.CV cs.AIcs.CLcs.LG

classification cs.CVcs.AIcs.CLcs.LG
keywords videogenerationcontrolmulti-sceneconsistencyconsistententitiesframework
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent text-to-video (T2V) generation methods have seen significant advancements. However, the majority of these works focus on producing short video clips of a single event (i.e., single-scene videos). Meanwhile, recent large language models (LLMs) have demonstrated their capability in generating layouts and programs to control downstream visual modules. This prompts an important question: can we leverage the knowledge embedded in these LLMs for temporally consistent long video generation? In this paper, we propose VideoDirectorGPT, a novel framework for consistent multi-scene video generation that uses the knowledge of LLMs for video content planning and grounded video generation. Specifically, given a single text prompt, we first ask our video planner LLM (GPT-4) to expand it into a 'video plan', which includes the scene descriptions, the entities with their respective layouts, the background for each scene, and consistency groupings of the entities. Next, guided by this video plan, our video generator, named Layout2Vid, has explicit control over spatial layouts and can maintain temporal consistency of entities across multiple scenes, while being trained only with image-level annotations. Our experiments demonstrate that our proposed VideoDirectorGPT framework substantially improves layout and movement control in both single- and multi-scene video generation and can generate multi-scene videos with consistency, while achieving competitive performance with SOTAs in open-domain single-scene T2V generation. Detailed ablation studies, including dynamic adjustment of layout control strength with an LLM and video generation with user-provided images, confirm the effectiveness of each component of our framework and its future potential.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Authoring for Living Worlds: Tool-Constrained LLM Agents for Executable Multi-Actor Scenarios

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Tool-validated LLM agents author formal event-graph story specifications that a deterministic game engine executes into multi-actor videos with perfect annotations, reaching 80% seeded end-to-end success versus 0% for...

  2. Crayotter: Traceable Multi-Agent Workflows for Long-Form Video Editing

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Crayotter introduces a traceable three-phase multi-agent workflow for long-form video editing that scores 3.40/5 in human evaluations, outperforming two baselines on 23 themes.

  3. The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

    cs.CV 2026-01 reject novelty 5.0 of 10

    An agentic dialogue-to-video pipeline (ScripterAgent, DirectorAgent, CriticAgent) claims to improve long-horizon cinematic coherence, but its supporting evaluation is partly self-referential.

  4. A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data

    cs.AI 2026-01 conditional novelty 5.0 of 10

    A metric-oriented survey that classifies intrinsic quality and trustworthiness metrics for LLM-generated data across six modalities and documents systematic evaluation gaps in the current literature.

  5. Enhancing Scene Transition Awareness in Video Generation via Post-Training

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A post-training dataset of transition-centered clips increases the number of scenes an open-source video generator produces for multi-scene prompts, with mixed effects on quality.

  6. Test-time Prompt Refinement for Text-to-Image Models

    cs.LG 2025-07 conditional novelty 5.0 of 10

    A training-free closed loop, in which a multimodal LLM rewrites a text prompt after inspecting the generated image, improves overall text-to-image alignment but degrades some attribute and spatial categories.

  7. A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality

    cs.CV 2025-07 conditional novelty 3.0 of 10

    A survey of 32 long-video generation papers, presenting a taxonomy and component recommendations for backbones, text encoders, objectives, and positional encodings.

Pith tools