Pith. sign in

REVIEW 5 cited by

LivePhoto: Real Image Animation with Text-guided Motion Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02928 v1 pith:BJGNNASU submitted 2023-12-05 cs.CV

classification cs.CV
keywords textmotioncontrolimageintensitymodulemotionscontents
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this work presents a practical system, named LivePhoto, which allows users to animate an image of their interest with text descriptions. We first establish a strong baseline that helps a well-learned text-to-image generator (i.e., Stable Diffusion) take an image as a further input. We then equip the improved generator with a motion module for temporal modeling and propose a carefully designed training pipeline to better link texts and motions. In particular, considering the facts that (1) text can only describe motions roughly (e.g., regardless of the moving speed) and (2) text may include both content and motion descriptions, we introduce a motion intensity estimation module as well as a text re-weighting module to reduce the ambiguity of text-to-motion mapping. Empirical evidence suggests that our approach is capable of well decoding motion-related textual instructions into videos, such as actions, camera movements, or even conjuring new contents from thin air (e.g., pouring water into an empty glass). Interestingly, thanks to the proposed intensity learning mechanism, our system offers users an additional control signal (i.e., the motion intensity) besides text for video customization.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation

    cs.CV 2025-01 conditional novelty 6.0 of 10

    OmniPhysGS lets each Gaussian in a 3D scene select from 12 expert material models, supervised by a text-to-video diffusion model, to generate dynamics for multiple materials.

  2. Generative Physical AI in Vision: A Survey

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.

  3. Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Adversarial noise added by a trained encoder or PGD degrades generated images across seven customization methods, including tuning-free reference-based approaches, and shows qualitative transfer to commercial APIs.

  4. DiffSim: Taming Diffusion Models for Evaluating Visual Similarity

    cs.CV 2024-12 reject novelty 5.0 of 10

    A diffusion U-Net's attention features, aligned with a bidirectional attention score, can rank visual similarity competitively with CLIP, DINO, and LPIPS.

  5. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.

Pith tools