Pith. sign in

REVIEW 2 cited by

COLLAGE: Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.20502 v1 pith:BS5QKG6R submitted 2024-09-30 cs.LG cs.AIcs.CVcs.GR

classification cs.LGcs.AIcs.CVcs.GR
keywords collaborativediffusionhierarchicalinteractionsmodelcollagedatasetsgenerating
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a novel framework COLLAGE for generating collaborative agent-object-agent interactions by leveraging large language models (LLMs) and hierarchical motion-specific vector-quantized variational autoencoders (VQ-VAEs). Our model addresses the lack of rich datasets in this domain by incorporating the knowledge and reasoning abilities of LLMs to guide a generative diffusion model. The hierarchical VQ-VAE architecture captures different motion-specific characteristics at multiple levels of abstraction, avoiding redundant concepts and enabling efficient multi-resolution representation. We introduce a diffusion model that operates in the latent space and incorporates LLM-generated motion planning cues to guide the denoising process, resulting in prompt-specific motion generation with greater control and diversity. Experimental results on the CORE-4D, and InterHuman datasets demonstrate the effectiveness of our approach in generating realistic and diverse collaborative human-object-human interactions, outperforming state-of-the-art methods. Our work opens up new possibilities for modeling complex interactions in various domains, such as robotics, graphics and computer vision.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Affordance-Aware Articulation Synthesis for Rigged Objects

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A training-free optimization framework synthesizes affordance-aware articulation for arbitrary rigged 3D objects by aligning them with 2D diffusion-inpainted references.

  2. SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis

    cs.CV 2024-12 conditional novelty 6.0 of 10

    SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...

Pith tools