REVIEW 5 cited by
LivePhoto: Real Image Animation with Text-guided Motion Control
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Despite the recent progress in text-to-video generation, existing studies usually overlook the issue that only spatial contents but not temporal motions in synthesized videos are under the control of text. Towards such a challenge, this work presents a practical system, named LivePhoto, which allows users to animate an image of their interest with text descriptions. We first establish a strong baseline that helps a well-learned text-to-image generator (i.e., Stable Diffusion) take an image as a further input. We then equip the improved generator with a motion module for temporal modeling and propose a carefully designed training pipeline to better link texts and motions. In particular, considering the facts that (1) text can only describe motions roughly (e.g., regardless of the moving speed) and (2) text may include both content and motion descriptions, we introduce a motion intensity estimation module as well as a text re-weighting module to reduce the ambiguity of text-to-motion mapping. Empirical evidence suggests that our approach is capable of well decoding motion-related textual instructions into videos, such as actions, camera movements, or even conjuring new contents from thin air (e.g., pouring water into an empty glass). Interestingly, thanks to the proposed intensity learning mechanism, our system offers users an additional control signal (i.e., the motion intensity) besides text for video customization.
Forward citations
Cited by 5 Pith papers
-
OmniPhysGS: 3D Constitutive Gaussians for General Physics-Based Dynamics Generation
OmniPhysGS lets each Gaussian in a 3D scene select from 12 expert material models, supervised by a text-to-video diffusion model, to generate dynamics for multiple materials.
-
Generative Physical AI in Vision: A Survey
A structured review that categorizes physics-aware generative models in vision into explicit-simulation and implicit-learning families and proposes six integration paradigms.
-
Anti-Reference: Universal and Immediate Defense Against Reference-Based Generation
Adversarial noise added by a trained encoder or PGD degrades generated images across seven customization methods, including tuning-free reference-based approaches, and shows qualitative transfer to commercial APIs.
-
DiffSim: Taming Diffusion Models for Evaluating Visual Similarity
A diffusion U-Net's attention features, aligned with a bidirectional attention score, can rank visual similarity competitively with CLIP, DINO, and LPIPS.
-
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
VBench++ is a benchmark that scores text-to-video and image-to-video models on 16 quality dimensions plus trustworthiness, reporting human-alignment correlations for each.
Discussion (0). Continue with ORCID to comment.