Pith. sign in

REVIEW 1 cited by

Modeling and Driving Human Body Soundfields through Acoustic Primitives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.13083 v2 pith:MEOTKTRN submitted 2024-07-18 cs.SD cs.CVeess.AS

classification cs.SDcs.CVeess.AS
keywords bodyrenderingacousticaudiofullhumanprimitivesmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While rendering and animation of photorealistic 3D human body models have matured and reached an impressive quality over the past years, modeling the spatial audio associated with such full body models has been largely ignored so far. In this work, we present a framework that allows for high-quality spatial audio generation, capable of rendering the full 3D soundfield generated by a human body, including speech, footsteps, hand-body interactions, and others. Given a basic audio-visual representation of the body in form of 3D body pose and audio from a head-mounted microphone, we demonstrate that we can render the full acoustic scene at any point in 3D space efficiently and accurately. To enable near-field and realtime rendering of sound, we borrow the idea of volumetric primitives from graphical neural rendering and transfer them into the acoustic domain. Our acoustic primitives result in an order of magnitude smaller soundfield representations and overcome deficiencies in near-field rendering compared to previous approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Hearing Hands: Generating Sounds from Physical Interactions in 3D Scenes

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A rectified flow model conditioned on 3D hand trajectories and rendered scene video generates realistic hand-scene interaction sounds, with a human study finding near-chance discrimination (47% misclassified).

Pith tools