Pith. sign in

REVIEW 3 cited by

NewMove: Customizing text-to-video models with novel motions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04966 v2 pith:IYBIL2L2 submitted 2023-12-07 cs.CV

classification cs.CV
keywords motionmotionsapproachmethodcustomcustomizationinputintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce an approach for augmenting text-to-video generation models with customized motions, extending their capabilities beyond the motions depicted in the original training data. By leveraging a few video samples demonstrating specific movements as input, our method learns and generalizes the input motion patterns for diverse, text-specified scenarios. Our contributions are threefold. First, to achieve our results, we finetune an existing text-to-video model to learn a novel mapping between the depicted motion in the input examples to a new unique token. To avoid overfitting to the new custom motion, we introduce an approach for regularization over videos. Second, by leveraging the motion priors in a pretrained model, our method can produce novel videos featuring multiple people doing the custom motion, and can invoke the motion in combination with other motions. Furthermore, our approach extends to the multimodal customization of motion and appearance of individualized subjects, enabling the generation of videos featuring unique characters and distinct motions. Third, to validate our method, we introduce an approach for quantitatively evaluating the learned custom motion and perform a systematic ablation study. We show that our method significantly outperforms prior appearance-based customization approaches when extended to the motion customization task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Personalized Search with Regularized Low-Rank Parameter Updates

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Regularized rank-one LoRA updates to the final value transform of CLIP's text encoder beat textual inversion for personalized retrieval while preserving general knowledge.

  2. Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    ATOP personalizes a pre-trained multi-view diffusion model with a few reference videos to generate part motion from text and masks, then lifts that motion to a 3D articulation axis via score distillation.

  3. VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer

    cs.CV 2025-02 conditional novelty 5.0 of 10

    A controllable diffusion transformer generates VFX videos from a reference image, text, and mask and timestamp conditions, with a new 675-video dataset and a temporal accuracy metric.

Pith tools