Pith. sign in

REVIEW 2 cited by

FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.08189 v4 pith:63DFS4C3 submitted 2024-08-15 cs.CV

classification cs.CV
keywords textualcross-framefancyvideoguidancevideosconditionsconsistentframe-specific
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Synthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations without frame-specific textual guidance. Thus, the model's capacity to comprehend the temporal logic conveyed in prompts and generate videos with coherent motion is restricted. To tackle this limitation, we introduce FancyVideo, an innovative video generator that improves the existing text-control mechanism with the well-designed Cross-frame Textual Guidance Module (CTGM). Specifically, CTGM incorporates the Temporal Information Injector (TII) and Temporal Affinity Refiner (TAR) at the beginning and end of cross-attention, respectively, to achieve frame-specific textual guidance. Firstly, TII injects frame-specific information from latent features into text conditions, thereby obtaining cross-frame textual conditions. Then, TAR refines the correlation matrix between cross-frame textual conditions and latent features along the time dimension. Extensive experiments comprising both quantitative and qualitative evaluations demonstrate the effectiveness of FancyVideo. Our approach achieves state-of-the-art T2V generation results on the EvalCrafter benchmark and facilitates the synthesis of dynamic and consistent videos. Note that the T2V process of FancyVideo essentially involves a text-to-image step followed by T+I2V. This means it also supports the generation of videos from user images, i.e., the image-to-video (I2V) task. A significant number of experiments have shown that its performance is also outstanding.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    Layout-Togglable storytelling is introduced: diffusion transformers conditioned on layout enable precise control over character position and appearance, supported by a new large-scale dataset and benchmark.

  2. Numerical Study of Oblique Detonation Initiation Assisted by Local Energy Deposition

    physics.flu-dyn 2025-08 unverdicted novelty 5.0 of 10

    Pulsatile local energy deposition can initiate sustainable oblique detonation on a finite wedge with less than 10% of the average power needed by continuous deposition.

Pith tools