Pith. sign in

REVIEW 10 cited by

ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04324 v2 pith:QX77AURY submitted 2024-02-06 cs.CV

classification cs.CV
keywords consistencygenerationconsisti2vframevideofirstvisualapproaches
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle to preserve the integrity of the subject, background, and style from the first frame, as well as ensure a fluid and logical progression within the video narrative. To mitigate these issues, we propose ConsistI2V, a diffusion-based method to enhance visual consistency for I2V generation. Specifically, we introduce (1) spatiotemporal attention over the first frame to maintain spatial and motion consistency, (2) noise initialization from the low-frequency band of the first frame to enhance layout consistency. These two approaches enable ConsistI2V to generate highly consistent videos. We also extend the proposed approaches to show their potential to improve consistency in auto-regressive long video generation and camera motion control. To verify the effectiveness of our method, we propose I2V-Bench, a comprehensive evaluation benchmark for I2V generation. Our automatic and human evaluation results demonstrate the superiority of ConsistI2V over existing methods.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LongAnimation: Long Animation Generation with Dynamic Global-Local Memory

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LongAnimation uses a dynamic global-local memory, built from a long-video-understanding model's KV cache, to colorize animation sequences of about 500 frames with stable color consistency.

  2. Populate-A-Scene: Affordance-Aware Human Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A fine-tuned text-to-video model inserts a person into a scene and generates an interaction video without bounding boxes or pose input, and its attention maps reveal a latent sense of affordance.

  3. CDP: Towards Robust Autoregressive Visuomotor Policy Learning via Causal Diffusion

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Causal Diffusion Policy adds historical action conditioning and attention cache sharing to diffusion-based robot policies, improving success rates on most tested manipulation tasks under degraded observations.

  4. VideoMAR: Autoregressive Video Generatio with Continuous Tokens

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A decoder-only autoregressive video model with continuous tokens, frame-wise causal attention, and a next-frame diffusion loss reports a higher VBench-I2V score than Cosmos I2V with a much smaller model and dataset.

  5. SG2VID: Scene Graphs Enable Fine-Grained Control for Video Synthesis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SG2VID conditions a latent video diffusion model on scene graphs with temporal features to generate controllable surgical videos across cataract and cholecystectomy datasets.

  6. FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FlowMo reduces temporal artifacts in video generation by guiding the denoising process to lower the maximum patch-wise variance of consecutive-frame differences in the latent space.

  7. ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProphetDWM is a one-stage diffusion world model that jointly predicts future driving video and low-level actions from a current frame and a short action sequence.

  8. HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A closed-loop image-to-video quality-control system reports 87–97% expert agreement on internal clips, with no released data, code, or thresholds.

  9. Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.

  10. LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LiON-LoRA adds a learned scaling token to video-diffusion LoRA adapters, enabling linear and independent control of camera trajectory and object motion strength.

Pith tools