Pith. sign in

REVIEW 6 cited by

Pay Attention and Move Better: Harnessing Attention for Interactive Motion Generation and Training-free Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18977 v2 pith:U65H2CNH submitted 2024-10-24 cs.CV

classification cs.CV
keywords motionattentioneditinggenerationabilityexplainabilitygoodmechanism
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This research delves into the problem of interactive editing of human motion generation. Previous motion diffusion models lack explicit modeling of the word-level text-motion correspondence and good explainability, hence restricting their fine-grained editing ability. To address this issue, we propose an attention-based motion diffusion model, namely MotionCLR, with CLeaR modeling of attention mechanisms. Technically, MotionCLR models the in-modality and cross-modality interactions with self-attention and cross-attention, respectively. More specifically, the self-attention mechanism aims to measure the sequential similarity between frames and impacts the order of motion features. By contrast, the cross-attention mechanism works to find the fine-grained word-sequence correspondence and activate the corresponding timesteps in the motion sequence. Based on these key properties, we develop a versatile set of simple yet effective motion editing methods via manipulating attention maps, such as motion (de-)emphasizing, in-place motion replacement, and example-based motion generation, etc. For further verification of the explainability of the attention mechanism, we additionally explore the potential of action-counting and grounded motion generation ability via attention maps. Our experimental results show that our method enjoys good generation and editing ability with good explainability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Absolute Coordinates Make Motion Generation Easy

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Using absolute 3D joint coordinates with a plain Transformer and velocity-prediction diffusion outperforms the standard local-relative motion representation, improving fidelity and enabling direct control.

  2. Interactive Generative Motion Editing via Scheduled Inpainting

    cs.GR 2026-07 conditional novelty 6.0 of 10

    Scheduled inpainting blends a base motion clip into a diffusion model's denoising process via a user-controlled schedule and spatiotemporal mask, enabling interactive editing of existing animations without retraining.

  3. Motion4Motion: Motion Transfer Across Subjects at Inference

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free motion transfer across species works by extracting source motion flows, matching semantic points, and injecting them into DiT self-attention via TransPE positional padding.

  4. Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A 7B text-to-motion model trained on the new 2M-clip MotionMillion dataset is reported to generalize zero-shot to complex, out-of-domain prompts.

  5. CoDA: Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects

    cs.GR 2025-05 conditional novelty 6.0 of 10

    CoDA generates coordinated whole-body articulated-object manipulation by optimizing the noise of three decoupled diffusion models, guided by BPS-based end-effector and object trajectories.

  6. From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

    cs.RO 2025-05 reject novelty 4.0 of 10

    A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.

Pith tools