Pith. sign in

REVIEW 3 cited by

Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.11565 v2 pith:QX2M2SGY submitted 2022-12-22 cs.CV

classification cs.CV
keywords modelsgenerationone-shottuningdiffusionemploygenerateimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

To replicate the success of text-to-image (T2I) generation, recent works employ large-scale video datasets to train a text-to-video (T2V) generator. Despite their promising results, such paradigm is computationally expensive. In this work, we propose a new T2V generation setting$\unicode{x2014}$One-Shot Video Tuning, where only one text-video pair is presented. Our model is built on state-of-the-art T2I diffusion models pre-trained on massive image data. We make two key observations: 1) T2I models can generate still images that represent verb terms; 2) extending T2I models to generate multiple images concurrently exhibits surprisingly good content consistency. To further learn continuous motion, we introduce Tune-A-Video, which involves a tailored spatio-temporal attention mechanism and an efficient one-shot tuning strategy. At inference, we employ DDIM inversion to provide structure guidance for sampling. Extensive qualitative and numerical experiments demonstrate the remarkable ability of our method across various applications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OSVE: One Step Video Editing with One Step Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...

  2. Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.

  3. Communicative Agents for Slideshow Storytelling Video Generation based on LLMs

    cs.AI 2025-09 conditional novelty 4.0 of 10

    VGTeam uses communicating LLM agents plus commercial APIs to turn a text prompt into a slideshow video for about $0.10 per clip, with a self-reported 75.7% quality rate.

Pith tools