Pith. sign in

REVIEW 28 cited by

DexCap: Scalable and Portable Mocap Data Collection System for Dexterous Manipulation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.07788 v2 pith:AQCOXI4G submitted 2024-03-12 cs.RO cs.AIcs.CVcs.LG

classification cs.ROcs.AIcs.CVcs.LG
keywords datamocapdexcaphandhumanmotiondexterousimitation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Imitation learning from human hand motion data presents a promising avenue for imbuing robots with human-like dexterity in real-world manipulation tasks. Despite this potential, substantial challenges persist, particularly with the portability of existing hand motion capture (mocap) systems and the complexity of translating mocap data into effective robotic policies. To tackle these issues, we introduce DexCap, a portable hand motion capture system, alongside DexIL, a novel imitation algorithm for training dexterous robot skills directly from human hand mocap data. DexCap offers precise, occlusion-resistant tracking of wrist and finger motions based on SLAM and electromagnetic field together with 3D observations of the environment. Utilizing this rich dataset, DexIL employs inverse kinematics and point cloud-based imitation learning to seamlessly replicate human actions with robot hands. Beyond direct learning from human motion, DexCap also offers an optional human-in-the-loop correction mechanism during policy rollouts to refine and further improve task performance. Through extensive evaluation across six challenging dexterous manipulation tasks, our approach not only demonstrates superior performance but also showcases the system's capability to effectively learn from in-the-wild mocap data, paving the way for future data collection methods in the pursuit of human-level robot dexterity. More details can be found at https://dex-cap.github.io

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 28 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Handroid: Bridging Dexterous Hand and Humanoid

    cs.RO 2026-07 conditional novelty 7.0 of 10

    A single 27-DoF body doubles as an anthropomorphic dexterous hand and a 0.33 m desktop humanoid, with a unified control stack for manipulation, locomotion, and embodiment switching.

  2. Feel the Force: Contact-Driven Learning from Humans

    cs.RO 2025-06 conditional novelty 7.0 of 10

    FeelTheForce trains a robot policy on human tactile demonstrations, predicting desired contact forces and using a PD controller to track them on the robot gripper, achieving 77% success across five force-sensitive tasks.

  3. DexDirect: Direct Kinesthetic Arm Guidance for Efficient Dexterous Demonstration Collection

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A hybrid kinesthetic-arm-plus-webcam-hand teleoperation interface achieved 17x/3x higher demonstration throughput than vision baselines and trained a 90%-success pick-and-place policy in a ten-person study.

  4. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  5. Scale Up Strategically: Learning Compositional Generalization via Bias-Aware Evaluation and Data Collection for Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot instruction-following policies consistently over-rely on color and under-ground verbs and size, and reallocating training demonstrations to under-grounded factors improves compositional generalization with fewer demos.

  6. Towards Human-level Dexterous Teleoperation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A single-stage RL co-tracking controller trained on consecutive human-derived hand–object subgoals achieves ~75% real-robot success on long-horizon dexterous teleoperation where baselines fail.

  7. TactiDex: A Real-World Tactile-Guided Benchmark for Human-Like Dexterous Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A tactile-rich HOI dataset plus a tri-component force reward improves contact fidelity and success of human-to-robot dexterous transfer over kinematic imitation alone.

  8. Cross-Embodiment Robot Manipulation via a Unified Hand Action Space

    cs.RO 2026-07 conditional novelty 6.0 of 10

    UHAS maps hand actions to deformations of a shared unit sphere and recovers joint commands via cascade IK, enabling multi-hand RL, zero-shot transfer, and modest real-world cube reorientation on LEAP and Allegro.

  9. Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    Task-agnostic RL play pretraining on diverse objects yields a reusable dexterous prior that makes sparse-reward assembly learning ~33× more sample-efficient and enables zero-shot sim-to-real transfer on tight insertio...

  10. Act, Sense, Act: Learning Active Perception from Large-Scale Egocentric Human Data

    cs.RO 2026-02 conditional novelty 6.0 of 10

    CoMe-VLA combines cognitive subtask labels and dual-track memory with human egocentric pretraining, reaching 83% mean success on five active-perception manipulation tasks.

  11. Dexplore: Scalable Neural Control for Dexterous Manipulation from Reference-Scoped Exploration

    cs.RO 2025-09 conditional novelty 6.0 of 10

    Dexplore learns dexterous robotic hand control from human MoCap demonstrations by treating them as soft, adaptively shrinking spatial references, then distills the policy into a vision-based controller.

  12. H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation

    cs.RO 2025-07 conditional novelty 6.0 of 10

    Pre-training a diffusion-transformer robot policy on 338K human hand-manipulation episodes, then fine-tuning with modular adapters, improves bimanual manipulation success across simulation and real robots.

  13. RwoR: Generating Robot Demonstrations from Human Hand Collection for Policy Learning without Robot

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A generative model and wrist camera turn human hand videos into robot gripper demonstrations that train manipulation policies at success rates close to those trained on real gripper data.

  14. DexVLG: Dexterous Vision-Language-Grasp Model at Scale

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DexVLG is a vision-language model trained on 170 million simulated dexterous grasps that generates hand poses aligned with language instructions about which part of an object to grasp.

  15. Knowledge-Driven Imitation Learning: Enabling Generalization Across Diverse Conditions

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A semantic keypoint graph matched to novel objects lets imitation-learned manipulation policies generalize with a quarter of the demonstrations.

  16. Vision in Action: Learning Active Perception from Human Demonstrations

    cs.RO 2025-06 conditional novelty 6.0 of 10

    ViA trains bimanual manipulation policies from human demonstrations that include active head-camera movement, using a 6-DoF robot neck and a VR interface with point-cloud rendering, reporting large gains on three occl...

  17. Object-centric 3D Motion Field for Robot Learning from Human Videos

    cs.RO 2025-06 conditional novelty 6.0 of 10

    A policy trained only on human RGBD videos, with a denoised object-centric 3D motion field as action representation, achieves about 55% average success on five real manipulation tasks where prior flow-based methods st...

  18. DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    DexMachina uses decaying virtual object controllers as a curriculum to train bimanual dexterous policies that track demonstrated object states, and reports large gains over baselines on a new six-hand benchmark.

  19. Learning Generalizable Robot Policy with Human Demonstration Video as a Prompt

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A two-stage pipeline trains a robot policy that accepts a human demonstration video as a prompt and generalizes beyond its robot training tasks, with success rates of up to 79 percent on known task variations and unde...

  20. EgoZero: Robot Learning from Smart Glasses

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Robot policies trained only on egocentric human videos from smart glasses transfer zero-shot to a Franka gripper, with 70% success across 7 manipulation tasks.

  21. Object-Focus Actor for Data-efficient Robot Generalization Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    Object-Focus Actor makes dexterous manipulation policies generalize to new object positions and backgrounds by focusing on the hand-object region and using relative poses and actions, needing only 10 to 30 demonstrations.

  22. Adaptive Visuo-Tactile Fusion with Predictive Force Attention for Dexterous Manipulation

    cs.RO 2025-05 conditional novelty 6.0 of 10

    A force-guided attention module and future-force prediction auxiliary task improve visuo-tactile fusion for dexterous manipulation, reaching 93% average success in real robot trials.

  23. CordViP: Correspondence-based Visuomotor Policy for Dexterous Manipulation in Real-World

    cs.RO 2025-02 conditional novelty 6.0 of 10

    CordViP achieves strong real-world dexterous manipulation by feeding a diffusion policy with pose-tracked 3D object models and hand point clouds, pretrained on contact maps and arm-hand coordination.

  24. TacPrint: A Wearable Fingertip Tactile Sensor for Human-to-Robot Contact Reproduction

    cs.RO 2026-07 conditional novelty 5.0 of 10

    A low-cost wearable fingertip sensor estimates dense contact-depth maps from 24 capacitive channels and uses them to substantially improve robot grasping and wiping in human-to-robot replay.

  25. AXIS: A Growable Community-Driven Data Engine for Scalable Robot Manipulation

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Pretraining π0.5 on the crowdsourced AXIS simulation dataset (207 tasks, 50K+ trajectories) raises downstream LIBERO-Plus success from 83.9% to 88.8% as the pretraining corpus grows from none to the full dataset.

  26. The Latent Color Subspace: Emergent Order in High-Dimensional Chaos

    cs.LG 2026-03 unverdicted novelty 5.0 of 10

    FLUX.1’s VAE latent space contains an interpretable Hue–Saturation–Lightness structure that enables training-free color prediction and control via closed-form latent edits.

  27. ImMimic: Cross-Domain Imitation from Human Videos via Mapping and Interpolation

    cs.RO 2025-09 conditional novelty 5.0 of 10

    A co-training framework that maps retargeted human hand trajectories to robot demonstrations with dynamic time warping and MixUp interpolation improves robot manipulation success rates and smoothness across four embodiments.

  28. A Survey: Learning Embodied Intelligence from Physical Simulators and World Models

    cs.RO 2025-07 conditional novelty 4.0 of 10

    Embodied intelligence learning is reviewed through the complementary lenses of physical simulators and world models, with a proposed IR-L0 to IR-L4 robot capability taxonomy.

Pith tools