Pith. sign in

REVIEW 3 cited by

MotionCraft: Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21136 v3 pith:6S77GYXD submitted 2024-07-30 cs.CV

classification cs.CV
keywords motiongenerationmultimodaldifferenttaskswhole-bodyacrossmotioncraft
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Whole-body multimodal motion generation, controlled by text, speech, or music, has numerous applications including video generation and character animation. However, employing a unified model to achieve various generation tasks with different condition modalities presents two main challenges: motion distribution drifts across different tasks (e.g., co-speech gestures and text-driven daily actions) and the complex optimization of mixed conditions with varying granularities (e.g., text and audio). Additionally, inconsistent motion formats across different tasks and datasets hinder effective training toward multimodal motion generation. In this paper, we propose MotionCraft, a unified diffusion transformer that crafts whole-body motion with plug-and-play multimodal control. Our framework employs a coarse-to-fine training strategy, starting with the first stage of text-to-motion semantic pre-training, followed by the second stage of multimodal low-level control adaptation to handle conditions of varying granularities. To effectively learn and transfer motion knowledge across different distributions, we design MC-Attn for parallel modeling of static and dynamic human topology graphs. To overcome the motion format inconsistency of existing benchmarks, we introduce MC-Bench, the first available multimodal whole-body motion generation benchmark based on the unified SMPL-X format. Extensive experiments show that MotionCraft achieves state-of-the-art performance on various standard motion generation tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A single autoregressive diffusion model, trained on a new 286-hour SMPL-X dataset, generates whole-body motion from text, audio, and spatial-temporal control signals, with reference-motion conditioning.

  2. GENMO: A GENeralist Model for Human MOtion

    cs.GR 2025-05 conditional novelty 6.0 of 10

    GENMO unifies human motion estimation and generation in one diffusion model by training regression and diffusion modes together, and it uses 2D video data to improve generative diversity.

  3. Sketch2Anim: Towards Transferring Sketch Storyboards into 3D Animation

    cs.GR 2025-04 conditional novelty 6.0 of 10

    Sketch2Anim aligns 2D sketch keyposes and joint trajectories with 3D embeddings and uses a trajectory ControlNet plus keypose adapter to generate 3D motion clips from storyboards.

Pith tools