Pith. sign in

REVIEW 3 cited by

Text-driven Human Motion Generation with Motion Masked Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.19686 v1 pith:W2V2F6GK submitted 2024-09-29 cs.CV

classification cs.CV
keywords motiondiffusionhumanmodelmaskmaskedgenerationlearn
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating plausible and realistic human actions with high diversity. Existing diffusion model-based approaches have outstanding performance in the diversity and multimodality of generation. However, compared to autoregressive methods that train motion encoders before inference, diffusion methods lack in fitting the distribution of human motion features which leads to an unsatisfactory FID score. One insight is that the diffusion model lack the ability to learn the motion relations among spatio-temporal semantics through contextual reasoning. To solve this issue, in this paper, we proposed Motion Masked Diffusion Model \textbf{(MMDM)}, a novel human motion masked mechanism for diffusion model to explicitly enhance its ability to learn the spatio-temporal relationships from contextual joints among motion sequences. Besides, considering the complexity of human motion data with dynamic temporal characteristics and spatial structure, we designed two mask modeling strategies: \textbf{time frames mask} and \textbf{body parts mask}. During training, MMDM masks certain tokens in the motion embedding space. Then, the diffusion decoder is designed to learn the whole motion sequence from masked embedding in each sampling step, this allows the model to recover a complete sequence from incomplete representations. Experiments on HumanML3D and KIT-ML dataset demonstrate that our mask strategy is effective by balancing motion quality and text-motion consistency.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interactive Generative Motion Editing via Scheduled Inpainting

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Scheduled inpainting blends a base motion clip into a diffusion model's denoising process via a user-controlled schedule and spatiotemporal mask, enabling interactive editing of existing animations without retraining.

  2. A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting

    cs.RO 2025-09 conditional novelty 6.0 of 10

    A neural retargeting pipeline maps human motion to humanoid robot motion at 5000+ frames per second using a shared latent space and physics-based fine-tuning, filtering noise and producing physically feasible trajectories.

  3. Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head Animation

    cs.GR 2025-09 conditional novelty 6.0 of 10

    Think2Sing uses LLM-generated, time-aligned motion subtitles and a motion-intensity proxy to guide diffusion-based 3D head animation from singing audio and lyrics.

Pith tools