REVIEW 6 cited by
ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Audio signals provide rich information for the robot interaction and object properties through contact. This information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments by learning from diverse in-the-wild human demonstrations.
Forward citations
Cited by 6 Pith papers
-
Transformer Transformer: A Unified Model for Motion-Conditioned Robot Co-design
A single diffusion transformer trains on tokenized robot bodies and motions to generate and optimize robot designs for unseen rewards and trajectories, outpacing evolutionary search in speed and often in reward.
-
S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information
S2A2 adds microphone-array spatial audio and spectrograms to imitation-learning policies, substantially improving success on manipulation tasks where vision alone cannot identify the target or destination.
-
Multisensory Continual Learning: Adapting Pretrained Visuomotor Policies to Force
MuSe adapts vision-only visuomotor policies to force-torque sensing with multi-stage fusion, multisensory future prediction, and experience replay, showing strong performance on contact-rich tasks while preserving ori...
-
Hearing the Slide: Acoustic-Guided Constraint Learning for Fast Non-Prehensile Transport
A contact microphone on a robot tray learns a velocity-dependent friction constraint that reduces object displacement during fast non-prehensile transport by an average of 86% in physical experiments.
-
PRISM: Polynomial Representations for Interaction-Structured Motor Control
Explicit low-degree factorized polynomial proprioceptive features improve robot RL and imitation policies beyond matched-capacity MLPs and induce sensorless compliance-like contact behavior in simulation.
-
Data Pyramid for Embodied Manipulation
Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.
Discussion (0). Continue with ORCID to comment.