REVIEW 5 cited by
GENMO: A GENeralist Model for Human MOtion
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Human motion modeling traditionally separates motion generation and estimation into distinct tasks with specialized models. Motion generation models focus on creating diverse, realistic motions from inputs like text, audio, or keyframes, while motion estimation models aim to reconstruct accurate motion trajectories from observations like videos. Despite sharing underlying representations of temporal dynamics and kinematics, this separation limits knowledge transfer between tasks and requires maintaining separate models. We present GENMO, a unified Generalist Model for Human Motion that bridges motion estimation and generation in a single framework. Our key insight is to reformulate motion estimation as constrained motion generation, where the output motion must precisely satisfy observed conditioning signals. Leveraging the synergy between regression and diffusion, GENMO achieves accurate global motion estimation while enabling diverse motion generation. We also introduce an estimation-guided training objective that exploits in-the-wild videos with 2D annotations and text descriptions to enhance generative diversity. Furthermore, our novel architecture handles variable-length motions and mixed multimodal conditions (text, audio, video) at different time intervals, offering flexible control. This unified approach creates synergistic benefits: generative priors improve estimated motions under challenging conditions like occlusions, while diverse video data enhances generation capabilities. Extensive experiments demonstrate GENMO's effectiveness as a generalist framework that successfully handles multiple human motion tasks within a single model.
Forward citations
Cited by 5 Pith papers
-
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
STREAM decouples text (via AdaLN) from music (via energy-based BEAM attention) to generate editable, musically aligned dance motions with a new annotated dataset and editability metric.
-
Direct Clinical Joint Angle Extraction from Parametric Body Model Rotation Matrices
A swing-twist decomposition plus a per-body-model calibration table converts body-model rotation matrices into clinical joint angles at 4.50° MAE on OpenCap LabValidation.
-
Video Generation Models are General-Purpose Vision Learners
A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.
-
Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
A learned human-action manifold, combining 3D pose, 2D keypoints, appearance, and motion derivatives, scores generated videos by distance to real-action centroids and embedding smoothness, beating prior metrics on hum...
-
TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
A reinforcement-learned time optimization policy that adaptively slows upper-body motion clips improves stability and precision of humanoid standing manipulation at a modest time cost.
Discussion (0). Sign in to comment.