Pith. sign in

REVIEW 5 cited by

Controllable Human-Object Interaction Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.03913 v2 pith:GJWOHUNS submitted 2023-12-06 cs.CV

classification cs.CV
keywords objectmotionhumanhuman-objectinteractionwaypointsdescriptionsdiffusion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human motion guided by language descriptions in 3D scenes. We propose Controllable Human-Object Interaction Synthesis (CHOIS), an approach that generates object motion and human motion simultaneously using a conditional diffusion model given a language description, initial object and human states, and sparse object waypoints. Here, language descriptions inform style and intent, and waypoints, which can be effectively extracted from high-level planning, ground the motion in the scene. Naively applying a diffusion model fails to predict object motion aligned with the input waypoints; it also cannot ensure the realism of interactions that require precise hand-object and human-floor contact. To overcome these problems, we introduce an object geometry loss as additional supervision to improve the matching between generated object motion and input object waypoints; we also design guidance terms to enforce contact constraints during the sampling process of the trained diffusion model. We demonstrate that our learned interaction module can synthesize realistic human-object interactions, adhering to provided textual descriptions and sparse waypoint conditions. Additionally, our module seamlessly integrates with a path planning module, enabling the generation of long-term interactions in 3D environments.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Diffgrasp: Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model

    cs.CV 2024-12 conditional novelty 7.0 of 10

    DiffGrasp synthesizes full-body grasping motion sequences with realistic hand-object contact from object shape and motion via a single conditional diffusion model.

  2. InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    InterAct is a unified 21.81-hour 3D human-object interaction benchmark with text annotations, quality-corrected data, and a multi-task model that achieves state-of-the-art results across six generation tasks.

  3. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...

  4. Mimicking-Bench: A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking

    cs.RO 2024-12 conditional novelty 6.0 of 10

    Mimicking-Bench provides six humanoid-scene interaction tasks with 23K human motion references and a retarget-track-imitate pipeline that beats data-free RL on average success.

  5. InterDance:Reactive 3D Dance Generation with Realistic Duet Interactions

    cs.CV 2024-12 conditional novelty 6.0 of 10

    InterDance introduces a 3.93-hour duet dance dataset and a diffusion method with contact and penetration guidance for reactive dance generation.

Pith tools