Pith. sign in

REVIEW 5 cited by

MoMask: Generative Masked Modeling of 3D Human Motions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.00063 v1 pith:ZQYTIFII submitted 2023-11-29 cs.CV

classification cs.CV
keywords tokensmotionmaskedmomaskgenerationhumantransformerlayer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce MoMask, a novel masked modeling framework for text-driven 3D human motion generation. In MoMask, a hierarchical quantization scheme is employed to represent human motion as multi-layer discrete motion tokens with high-fidelity details. Starting at the base layer, with a sequence of motion tokens obtained by vector quantization, the residual tokens of increasing orders are derived and stored at the subsequent layers of the hierarchy. This is consequently followed by two distinct bidirectional transformers. For the base-layer motion tokens, a Masked Transformer is designated to predict randomly masked motion tokens conditioned on text input at training stage. During generation (i.e. inference) stage, starting from an empty sequence, our Masked Transformer iteratively fills up the missing tokens; Subsequently, a Residual Transformer learns to progressively predict the next-layer tokens based on the results from current layer. Extensive experiments demonstrate that MoMask outperforms the state-of-art methods on the text-to-motion generation task, with an FID of 0.045 (vs e.g. 0.141 of T2M-GPT) on the HumanML3D dataset, and 0.228 (vs 0.514) on KIT-ML, respectively. MoMask can also be seamlessly applied in related tasks without further model fine-tuning, such as text-guided temporal inpainting.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A first large multimodal 4D human–dog interaction dataset (6.8M frames) plus an autoregressive model that generates dog motion from human body/hand gestures and audio.

  2. Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    By fine-tuning an LLM with audio and motion tokens, MECo generates co-speech gestures that follow a user-provided motion example, and reports state-of-the-art FGD and diversity on BEAT2 and ZeroEGGS.

  3. MOST: Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MOST improves rare-prompt text-to-motion generation by retrieving key motion clips through a new temporal clip Banzhaf interaction and using them as diffusion prompts.

  4. ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ANT makes text embeddings change across denoising steps and schedules classifier-free guidance to decay, improving text-motion alignment in diffusion text-to-motion models.

  5. Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.

Pith tools