REVIEW 10 cited by
Motion-Agent: A Conversational Framework for Human Motion Generation with LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient conversational framework designed for general human motion generation, editing, and understanding. Motion-Agent employs an open-source pre-trained language model to develop a generative agent, MotionLLM, that bridges the gap between motion and text. This is accomplished by encoding and quantizing motions into discrete tokens that align with the language model's vocabulary. With only 1--3\% of the model's parameters fine-tuned using adapters, MotionLLM delivers performance on par with diffusion models and other transformer-based methods trained from scratch. By integrating MotionLLM with GPT-4 without additional training, Motion-Agent is able to generate highly complex motion sequences through multi-turn conversations, a capability that previous models have struggled to achieve. Motion-Agent supports a wide range of motion-language tasks, offering versatile capabilities for generating and customizing human motion through interactive conversational exchanges. Project page: https://knoxzhao.github.io/Motion-Agent
Forward citations
Cited by 10 Pith papers
-
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.
-
Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
Being-M0.5 combines part-aware residual quantization with a 5M-sequence web-video dataset to reach real-time, part-controllable 3D motion generation, though its state-of-the-art claim does not hold on every standard b...
-
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
A 7B text-to-motion model trained on the new 2M-clip MotionMillion dataset is reported to generalize zero-shot to complex, out-of-domain prompts.
-
Speaking Beyond Language: A Large-Scale Multimodal Dataset for Learning Nonverbal Cues from Video-Grounded Dialogues
VENUS is a large podcast-derived dataset aligning text with 3D facial and body cues, and MARS is an LLM fine-tuned on it to generate both words and nonverbal tokens.
-
MotionLab: Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm
MotionLab unifies text-based and trajectory-based motion generation with text-based editing, trajectory-based editing, motion in-betweening, and style transfer in one flow-based transformer.
-
BiomechGPT: Extending Motion-Language Models to Clinical Motion Understanding
A motion-language model fine-tuned on clinical biomechanics data answers structured questions about movement, scoring highly on activity recognition and gait speed estimation but weaker and underpowered on diagnosis a...
-
From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control
A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.
-
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
A survey proposing a four-part taxonomy for human-centric foundation models and reviewing representative methods in each.
-
Motion Generation: A Survey of Generative Approaches and Benchmarks
A structured survey that categorizes recent motion generation methods by underlying generative approach and compiles datasets, metrics, and statistical trends.
-
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.
Discussion (0). Continue with ORCID to comment.