REVIEW 7 cited by
Duolando: Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce a novel task within the field of 3D dance generation, termed dance accompaniment, which necessitates the generation of responsive movements from a dance partner, the "follower", synchronized with the lead dancer's movements and the underlying musical rhythm. Unlike existing solo or group dance generation tasks, a duet dance scenario entails a heightened degree of interaction between the two participants, requiring delicate coordination in both pose and position. To support this task, we first build a large-scale and diverse duet interactive dance dataset, DD100, by recording about 117 minutes of professional dancers' performances. To address the challenges inherent in this task, we propose a GPT-based model, Duolando, which autoregressively predicts the subsequent tokenized motion conditioned on the coordinated information of the music, the leader's and the follower's movements. To further enhance the GPT's capabilities of generating stable results on unseen conditions (music and leader motions), we devise an off-policy reinforcement learning strategy that allows the model to explore viable trajectories from out-of-distribution samplings, guided by human-defined rewards. Based on the collected dataset and proposed method, we establish a benchmark with several carefully designed metrics.
Forward citations
Cited by 7 Pith papers
-
MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
MDD is the first dataset to pair text, music, and 3D duet dance motion, enabling two new text-conditioned duet generation tasks.
-
InterPet4D: A Multimodal 4D Human-Pet Interaction Dataset for Pet Motion Generation
A first large multimodal 4D human–dog interaction dataset (6.8M frames) plus an autoregressive model that generates dog motion from human body/hand gestures and audio.
-
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
A hierarchical global-keyframe then local-refinement diffusion pipeline produces stable 720p/30fps music-to-dance videos longer than one minute across five genres.
-
Real-time and Controllable Reactive Motion Synthesis via Intention Guidance
A neural system predicts key-joint intentions from motion history and uses adversarially regularized codebook matching to synthesize controllable, real-time reactive motions.
-
FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation
FlowerDance pairs MeanFlow few-step flow matching with a bidirectional Mamba backbone and physical-consistency losses, reporting state-of-the-art dance quality at 2008 FPS on FineDance and AIST++.
-
Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion Synthesis
Human-X jointly predicts actions and reactions in real time to produce physically plausible human-machine interaction motion.
-
Poly-Autoregressive Prediction for Modeling Interactions
A single transformer training recipe, poly-autoregressive prediction, improves multi-agent ego forecasting over autoregressive baselines on three distinct tasks.
Discussion (0). Continue with ORCID to comment.