Pith. sign in

REVIEW 8 cited by

Learning Visuotactile Skills with Two Multifingered Hands

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16823 v2 pith:YTYALRFZ submitted 2024-04-25 cs.RO cs.AIcs.CVcs.LG

classification cs.ROcs.AIcs.CVcs.LG
keywords multifingereddatahandslearningsystemvisuotactilepolicytouch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Aiming to replicate human-like dexterity, perceptual experiences, and motion patterns, we explore learning from human demonstrations using a bimanual system with multifingered hands and visuotactile data. Two significant challenges exist: the lack of an affordable and accessible teleoperation system suitable for a dual-arm setup with multifingered hands, and the scarcity of multifingered hand hardware equipped with touch sensing. To tackle the first challenge, we develop HATO, a low-cost hands-arms teleoperation system that leverages off-the-shelf electronics, complemented with a software suite that enables efficient data collection; the comprehensive software suite also supports multimodal data processing, scalable policy learning, and smooth policy deployment. To tackle the latter challenge, we introduce a novel hardware adaptation by repurposing two prosthetic hands equipped with touch sensors for research. Using visuotactile data collected from our system, we learn skills to complete long-horizon, high-precision tasks which are difficult to achieve without multifingered dexterity and touch feedback. Furthermore, we empirically investigate the effects of dataset size, sensing modality, and visual input preprocessing on policy learning. Our results mark a promising step forward in bimanual multifingered manipulation from visuotactile data. Videos, code, and datasets can be found at https://toruowo.github.io/hato/ .

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Pixel2Catch: Multi-Agent Sim-to-Real Transfer for Agile Manipulation with a Single RGB Camera

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Pixel2Catch shows that pixel-level bounding-box cues from one RGB camera, with separate arm and hand reinforcement-learning policies, are enough to catch thrown objects in the real world.

  2. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  3. Touch begins where vision ends: Generalizable policies for contact-rich manipulation

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A localize-then-execute policy that combines vision-language reaching, semantic background augmentation, and residual reinforcement learning with tactile sensing reaches about 90% success on millimeter-precision manip...

  4. SCIZOR: A Self-Supervised Approach to Data Curation for Large-Scale Imitation Learning

    cs.RO 2025-05 conditional novelty 6.0 of 10

    SCIZOR filters suboptimal and redundant state-action pairs from robot demonstrations without human labels, improving imitation-learning policy success rates by about 15% on average.

  5. Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.

  6. HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation

    cs.RO 2025-08 conditional novelty 5.0 of 10

    HERMES converts a single human motion demonstration into a deployable mobile bimanual dexterous manipulation policy, using RL, depth-image distillation, and closed-loop PnP pose refinement.

  7. Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization

    cs.RO 2025-07 conditional novelty 5.0 of 10

    Tactile-VLA fuses tactile sensing into a vision-language-action model so force-related instructions and corrective reasoning transfer to new contact-rich tasks with few demonstrations.

  8. mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity

    cs.RO 2025-06 conditional novelty 5.0 of 10

    mimic-one reports up to 93.3% out-of-distribution success on three real-world dexterous tasks using a diffusion policy, a custom 16-DoF hand, and a teleoperation data-collection recipe with self-correction trajectories.

Pith tools