Pith. sign in

REVIEW 6 cited by

ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.19464 v2 pith:QPCYXW4R submitted 2024-06-27 cs.RO cs.AIcs.CVcs.SDeess.AS

classification cs.ROcs.AIcs.CVcs.SDeess.AS
keywords robotmanipulationdemonstrationsin-the-wildlearningaudiodatainformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Audio signals provide rich information for the robot interaction and object properties through contact. This information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments by learning from diverse in-the-wild human demonstrations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.

  2. S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

    cs.RO 2026-07 conditional novelty 6.0 of 10

    S2A2 adds microphone-array spatial audio and spectrograms to imitation-learning policies, substantially improving success on manipulation tasks where vision alone cannot identify the target or destination.

  3. Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    MuSe adapts vision-only visuomotor policies to force-torque sensing with multi-stage fusion, multisensory future prediction, and experience replay, showing strong performance on contact-rich tasks while preserving ori...

  4. Hearing the Slide: Acoustic-Guided Constraint Learning for Fast Non-Prehensile Transport

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A contact microphone on a robot tray learns a velocity-dependent friction constraint that reduces object displacement during fast non-prehensile transport by an average of 86% in physical experiments.

  5. PRISM: Polynomial Representations for Interaction-Structured Motor Control

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Explicit low-degree factorized polynomial proprioceptive features improve robot RL and imitation policies beyond matched-capacity MLPs and induce sensorless compliance-like contact behavior in simulation.

  6. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

Pith tools