REVIEW 3 cited by
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work, we propose a new T2V generation setting$\unicode{x2014}$One-Shot Video Tuning, where only one text-video pair is presented. Our model is built on state-of-the-art T2I diffusion models pre-trained on massive image data. We make two key observations: 1) T2I models can generate still images that represent verb terms; 2) extending T2I models to generate multiple images concurrently exhibits surprisingly good content consistency. To further learn continuous motion, we introduce Tune-A-Video, which involves a tailored spatio-temporal attention mechanism and an efficient one-shot tuning strategy. At inference, we employ DDIM inversion to provide structure guidance for sampling. Extensive qualitative and numerical experiments demonstrate the remarkable ability of our method across various applications.
Forward citations
Cited by 3 Pith papers
-
OSVE: One Step Video Editing with One Step Diffusion Models
OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...
-
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.
-
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
VGTeam uses communicating LLM agents plus commercial APIs to turn a text prompt into a slideshow video for about $0.10 per clip, with a self-reported 75.7% quality rate.
Discussion (0). Continue with ORCID to comment.