Pith. sign in

REVIEW 2 cited by

MMM: Generative Masked Motion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03596 v2 pith:ZWYHK62X submitted 2023-12-06 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords motiontokensmaskedtexteditinggenerationmodelsaddition
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-to-motion generation using diffusion and autoregressive models have shown promising results. However, these models often suffer from a trade-off between real-time performance, high fidelity, and motion editability. To address this gap, we introduce MMM, a novel yet simple motion generation paradigm based on Masked Motion Model. MMM consists of two key components: (1) a motion tokenizer that transforms 3D human motion into a sequence of discrete tokens in latent space, and (2) a conditional masked motion transformer that learns to predict randomly masked motion tokens, conditioned on the pre-computed text tokens. By attending to motion and text tokens in all directions, MMM explicitly captures inherent dependency among motion tokens and semantic mapping between motion and text tokens. During inference, this allows parallel and iterative decoding of multiple motion tokens that are highly consistent with fine-grained text descriptions, therefore simultaneously achieving high-fidelity and high-speed motion generation. In addition, MMM has innate motion editability. By simply placing mask tokens in the place that needs editing, MMM automatically fills the gaps while guaranteeing smooth transitions between editing and non-editing parts. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that MMM surpasses current leading methods in generating high-quality motion (evidenced by superior FID scores of 0.08 and 0.429), while offering advanced editing features such as body-part modification, motion in-betweening, and the synthesis of long motion sequences. In addition, MMM is two orders of magnitude faster on a single mid-range GPU than editable motion diffusion models. Our project page is available at \url{https://exitudio.github.io/MMM-page}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MOST improves rare-prompt text-to-motion generation by retrieving key motion clips through a new temporal clip Banzhaf interaction and using them as diffusion prompts.

  2. ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ANT makes text embeddings change across denoising steps and schedules classifier-free guidance to decay, improving text-motion alignment in diffusion text-to-motion models.

Pith tools