REVIEW 3 cited by
MoDiPO: text-to-motion alignment via AI-feedback-driven Direct Preference Optimization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion Models have revolutionized the field of human motion generation by offering exceptional generation quality and fine-grained controllability through natural language conditioning. Their inherent stochasticity, that is the ability to generate various outputs from a single input, is key to their success. However, this diversity should not be unrestricted, as it may lead to unlikely generations. Instead, it should be confined within the boundaries of text-aligned and realistic generations. To address this issue, we propose MoDiPO (Motion Diffusion DPO), a novel methodology that leverages Direct Preference Optimization (DPO) to align text-to-motion models. We streamline the laborious and expensive process of gathering human preferences needed in DPO by leveraging AI feedback instead. This enables us to experiment with novel DPO strategies, using both online and offline generated motion-preference pairs. To foster future research we contribute with a motion-preference dataset which we dub Pick-a-Move. We demonstrate, both qualitatively and quantitatively, that our proposed method yields significantly more realistic motions. In particular, MoDiPO substantially improves Frechet Inception Distance (FID) while retaining the same RPrecision and Multi-Modality performances.
Forward citations
Cited by 3 Pith papers
-
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.
-
AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward
AToM uses GPT-4Vision-generated preference scores to fine-tune MotionGPT with IPO and LoRA, improving event-level alignment for integrity, temporal order, and frequency in text-to-motion generation.
-
MAPF-World: Action World Model for Multi-Agent Path Finding
MAPF-World, an autoregressive action world model that predicts future states and actions, is claimed to beat state-of-the-art learnable MAPF solvers while using 96.5% fewer parameters and 92% less data.
Discussion (0). Continue with ORCID to comment.