Pith. sign in

REVIEW 3 cited by

DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03018 v4 pith:FCJS6WUR submitted 2023-12-05 cs.CV

classification cs.CV
keywords imagegenerationdreamvideoimage-to-videomodelreferencevideodiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation models. Nevertheless, these methods often result in either low fidelity or flickering over time due to their limitation to shallow image guidance and poor temporal consistency. To tackle these problems, we propose a high-fidelity image-to-video generation method by devising a frame retention branch based on a pre-trained video diffusion model, named DreamVideo. Instead of integrating the reference image into the diffusion process at a semantic level, our DreamVideo perceives the reference image via convolution layers and concatenates the features with the noisy latents as model input. By this means, the details of the reference image can be preserved to the greatest extent. In addition, by incorporating double-condition classifier-free guidance, a single image can be directed to videos of different actions by providing varying prompt texts. This has significant implications for controllable video generation and holds broad application prospects. We conduct comprehensive experiments on the public dataset, and both quantitative and qualitative results indicate that our method outperforms the state-of-the-art method. Especially for fidelity, our model has a powerful image retention ability and delivers the best results in UCF101 compared to other image-to-video models to our best knowledge. Also, precise control can be achieved by giving different text prompts. Further details and comprehensive results of our model will be presented in https://anonymous0769.github.io/DreamVideo/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers

    cs.CV 2025-08 conditional novelty 6.0 of 10

    MiraMo turns a static image into a video using a linear-attention transformer that learns inter-frame motion residuals, with DCT-based noise refinement and a user-controllable dynamics knob.

  2. InteRecon: Towards Reconstructing Interactivity of Personal Memorable Items in Mixed Reality

    cs.HC 2025-02 conditional novelty 6.0 of 10

    The paper demonstrates a prototype that lets people turn cherished objects into interactive AR versions that preserve their original motions, buttons, and embedded media.

  3. HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A Diffusion Transformer with a prefix-latent reference strategy and a Keypoint-DiT pose generator produces long-form, pose-accurate human videos at variable resolution, outperforming prior U-Net-based animation method...

Pith tools