Pith. sign in

REVIEW 3 major objections 5 minor 48 references

Cross-Embodiment Robot Manipulation via a Unified Hand Action Space

T0 review · 3 major / 5 minor · reviewed 2026-07-12 · grok-4.5

Pith's one-line read Robotic hand actions can be shared across different hands by treating them as deformations of one canonical sphere, then mapping those deformations to each hand's joints.

desk verdict Practical geometric action space that lets one RL policy drive four hands on cube reorientation, with real (if modest) transfer and honest real-robot numbers. read the letter →

arxiv 2607.03570 v1 pith:GFXBSAO7 submitted 2026-07-03 cs.RO

classification cs.RO
keywords cross-embodimentmanipulationunifiedactionspacedexteroushandsspheredeformationcascadeinversekinematicsin-handreorientationreinforcementlearningsim-to-real
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Most robot learning policies are locked to one specific hand design, so skills learned on one platform cannot be reused on another with different fingers or joints. This paper claims that a single shared action space is enough: represent every hand action as a geometric deformation of a common unit sphere, then recover each embodiment's joint angles with a fast cascade inverse-kinematics procedure. Policies trained by reinforcement learning directly in that sphere space learn stable in-hand cube reorientation and transfer across four very different hands (four-finger and five-finger designs alike). A single multi-hand policy matches single-hand performance; policies trained without seeing a target hand still achieve high zero-shot success; and a few hundred finetuning steps recover near-specialist performance. Real-world LEAP and Allegro experiments confirm that the same controllers work on hardware. If the claim holds, future robot foundation models can share data and policies across heterogeneous dexterous hands instead of retraining from scratch for each morphology.

What carries the argument

Unified Hand Action Space (UHAS): dense, configuration-invariant correspondences between hand surface points and a unit sphere, with actions parameterized as sparse lateral and radial deformations of a few driving planes/vectors; Cascade Inverse Kinematics (CIK) then recovers each hand's joint angles by first solving lateral joints via lookup and then cascading encompassing joints along each finger.

What would settle it

Measure how much the open-hand sphere-to-surface correspondences drift under large closed or contacting finger poses; if residual error is large and zero-shot / multi-hand success collapses once those poses dominate, the shared action basis is insufficient.

Watch

Extended reading notes

Core claim

A sphere-based Unified Hand Action Space (UHAS) lets reinforcement-learning policies control multiple robotic hands with different kinematic structures from one shared continuous action representation. Actions are deformations of a canonical unit sphere; a Cascade Inverse Kinematics (CIK) algorithm maps those deformations to executable joint configurations for each hand. On in-hand cube reorientation the shared representation yields multi-hand policies that match single-hand specialists, meaningful zero-shot transfer to unseen hands, and rapid finetuning, both in simulation and on real hardware.

Load-bearing premise

The spherical surface points projected from an open-hand pose stay a faithful, complete description of hand action even after the fingers close and make multi-contact with a free object.

Editorial extensions

If this is right

  • A single policy network can be trained once and deployed, zero-shot or after short finetuning, on hands that were never seen during training.
  • Dexterous datasets collected on heterogeneous platforms can be mixed inside one geometric action space rather than remaining siloed by morphology.
  • Real-time sphere-to-joint mapping (CIK at ~150 Hz) makes the representation practical for closed-loop control on physical multi-finger hands.
  • Cross-morphology transfer (4-finger ↔ 5-finger) becomes a quantitative experimental regime rather than an ad-hoc engineering exercise.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same sphere deformation space could serve as a common interface for teleoperation or human demonstration retargeting, not only RL policies.
  • If the open-hand correspondence assumption is the main bottleneck, denser contact-aware or configuration-dependent sphere mappings would be a natural next design lever.
  • Foundation models that already output continuous actions could adopt UHAS as a drop-in hand head, allowing one VLA to drive many physical hands without embodiment-specific action tokens.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes the Unified Hand Action Space (UHAS), in which dexterous hand actions are parameterized as compact deformations of a canonical unit sphere (driving-plane lateral angles Δθ and driving-vector radial displacements Δr). A Cascade Inverse Kinematics (CIK) procedure maps those deformations to embodiment-specific joint configurations by classifying joints as lateral or encompassing and solving them sequentially. Policies are trained with PPO directly in UHAS for in-hand cube reorientation and evaluated on Allegro, LEAP, Shadow, and MANO hands. Simulation results (Tables 1–3) show multi-hand policies matching single-hand performance, non-trivial zero-shot transfer to held-out hands, and recovery of high success after only 500 finetuning iterations; limited real-world LEAP (and Allegro) deployments are also reported.

Significance. Cross-embodiment action spaces for multi-finger hands remain an open bottleneck for scalable robot learning. UHAS is a concrete, geometry-motivated alternative to joint-space or latent-action heads, with a practical real-time CIK controller (~150 Hz) and a systematic multi-hand evaluation suite (single-hand, multi-hand, zero-shot, morphology-split, and rapid finetune). The ablations on driving vectors, observation points, and domain randomization, together with promised code and data, make the empirical contribution reproducible and useful even if the geometric story is only partially validated. If the shared sphere interface continues to transfer beyond cube reorientation, it is a meaningful step toward multi-hand foundation policies.

major comments (3)
  1. [Section 3.2, Figure 3] Sec. 3.2 and Fig. 3 assert that spherical coordinates (θ, ϕ) obtained by open-hand surface projection are configuration-invariant and therefore a faithful shared action basis. The manuscript never reports correspondence residual, surface-point tracking error, or CIK reconstruction error under closed or multi-contact configurations—the regime in which cube reorientation actually operates. Without that measurement, Tables 1–3 cannot distinguish a transferable geometric action space from an effective shared interface whose semantics are largely supplied by CIK and multi-hand RL. Please either quantify correspondence/CIK error on closed and contacting poses across the four hands, or reframe the geometric claims accordingly and list the unmeasured correspondence fidelity as a limitation.
  2. [Section 4.2, Table 1; Appendix F.1, Table 12] Appendix F.1 / Table 12 shows that one-to-many zero-shot transfer from a single source hand is often weak (e.g., Allegro→Shadow 8.7% success; MANO→Allegro 33.0%). Multi-hand and leave-one-out zero-shot (Table 1) are much stronger. The main text currently emphasizes the latter without reconciling the former. The central claim of a morphology-agnostic geometric space needs an explicit discussion of when transfer fails (exploitative lateral-joint strategies, workspace asymmetry) and what that implies for the necessity of multi-embodiment training versus pure geometric correspondence.
  3. [Abstract; Section 4.3, Table 5; Section 5] Real-world LEAP results (Table 5) peak at mean 2.0 consecutive reorientations (best single-hand model) versus ~9.5–9.8 in simulation; Allegro real-world means are similarly ~2.1 (Table 13). Section 4.3 acknowledges the gap, but the abstract and conclusion still list “successful real-world deployment” alongside the near-ceiling simulation transfer claims without quantifying the drop. Please state the real-world numbers in the abstract/conclusion and clarify which claims are simulation-only versus hardware-validated.
minor comments (5)
  1. [Section 2] The abstract and introduction use both “Unified Hand Action Space (UHAS)” and, once in Related Work, “Universal Hand Action Space.” Standardize the name.
  2. [Section 3.3; Section 4.1] Fig. 4 caption and Sec. 3.3 give the action dimension as “5 + 2×5 = 15,” but the policy description (Sec. 4.1) uses one plane per finger with a duplicated ring plane for 4-finger hands. State the exact action dimension used in all reported experiments.
  3. [Appendix B.1, Table 7] Reward scales (Table 7) and the two joint-position regularizers are important for multi-hand stability; a one-sentence intuition in the main text for why w_lat is four times larger than w_rad would help readers.
  4. [Section 2] Related Work cites RobotFingerPrint and D(R,O) Grasp as geometric precursors; a short explicit contrast (grasp synthesis vs. continuous closed-loop action space + CIK) would sharpen novelty.
  5. [Throughout] Typos / polish: “Repose Cube” vs. “reposing”; “positivex-axis” spacing; “They-axis”; occasional missing spaces after periods in the arXiv text.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: UHAS/CIK are design choices evaluated empirically; success metrics are measured RL outcomes, not quantities forced by construction or self-citation.

full rationale

The paper proposes a geometric action representation (sphere deformations + Cascade IK) and evaluates it with reinforcement learning on in-hand cube reorientation. Sphere construction (r=2l/π, open-hand projection), driving-plane parameterization, joint classification, and CIK are engineering definitions, not fitted parameters that later reappear as predictions. Reported quantities (success rate, average consecutive reorientations in Tables 1–5) are measured from PPO rollouts against independent baselines (joint-space control) and transfer settings (multi-hand, zero-shot, finetune); none equal a fitted constant by construction. Prior geometric work (e.g., RobotFingerPrint) is cited as inspiration for surface–sphere correspondence, not as a uniqueness theorem that forces the present claims. Reward terms and domain randomization are stated independently of the outcome numbers. The derivation chain is therefore self-contained empirical methods work with no self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

The central empirical claim rests on a geometric abstraction (the unit sphere plus fixed correspondences) and a hand-crafted cascade solver rather than on free physical constants. Design choices such as the number of driving vectors, the sphere-radius formula and reward weights are free parameters that affect reported performance; the domain assumption that open-hand projections remain valid under contact is the main unproven premise.

free parameters (4)
  • driving_vectors_per_plane = 2
    Chosen as 2 after ablation (Table 4); directly sets action dimensionality and final success rates.
  • sphere_radius_formula = 2l/π
    r = 2l/π is a hand-chosen geometric heuristic (Sec. 3.1) that places the sphere in the grasp workspace; no derivation from first principles.
  • reward_scales = see Table 7
    Object-distance, orientation, joint-regularization, success-bonus and fall-penalty weights (Table 7) are tuned by hand and shared across all hands.
  • observation_points_per_finger = 2
    Fixed at midpoint + fingertip after ablation (Table 9); changes the observation dimension of the shared policy.
assumptions (4)
  • domain assumption Open-hand surface projections yield configuration-invariant spherical coordinates that remain a valid action basis under closed and contacting finger poses.
    Stated in Sec. 3.2 and Fig. 3(d); never validated under contact.
  • ad hoc to paper Every hand joint can be cleanly classified as either lateral (affects θ) or encompassing (affects r, ϕ) from a single open-hand sweep.
    Appendix A joint-classification procedure; required for the cascade decomposition.
  • domain assumption Fingers are kinematically independent enough that lateral and encompassing solves can be performed per finger without global optimization.
    Explicit in Appendix A; enables the 150 Hz real-time claim.
  • domain assumption Standard PPO with the listed hyper-parameters and domain randomization converges to transferable policies in the sphere action space.
    Training protocol of Sec. 4.1 and Appendix B.
invented entities (2)
  • Unified Hand Action Space (UHAS) independent evidence
    purpose: Provide a morphology-agnostic continuous action representation for multi-finger hands.
    Defined in Sec. 3; independent evidence is the multi-hand and zero-shot experiments, not external measurements.
  • Cascade Inverse Kinematics (CIK) independent evidence
    purpose: Map sphere deformations to executable joint configurations at interactive rates without numerical optimization.
    Introduced in Sec. 3.4 and Appendix A; evidence is the reported 150 Hz rate and successful closed-loop control.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Cross-Embodiment Robot Manipulation via a Unified Hand Action Space." pith.science (2026). https://pith.science/paper/GFXBSAO7

@misc{pith2026260703570,
  author       = {Pith},
  title        = {Pith review of: Cross-Embodiment Robot Manipulation via a Unified Hand Action Space},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GFXBSAO7}},
  note         = {Machine review of arXiv:2607.03570}
}
read the original abstract

Robot manipulation policies are typically tied to specific robotic hand embodiments, limiting the transfer of learned behaviors across platforms with different kinematic structures. In this work, we propose the Unified Hand Action Space (UHAS), a sphere-based unified action representation for cross-embodiment dexterous manipulation. UHAS represents robotic hand actions as geometric deformations of a canonical sphere and uses a Cascade Inverse Kinematics (CIK) algorithm to map the shared representation to embodiment-specific joint configurations. Using reinforcement learning, we train dexterous manipulation policies directly in the proposed action space for in-hand cube reorientation tasks. We evaluate our method in both simulation and real-world experiments across multiple robotic hands, including the Allegro Hand, LEAP Hand, Shadow Hand, and MANO Human Hand. Experimental results demonstrate effective dexterous manipulation, zero-shot transfer to unseen hands, rapid finetuning across embodiments, and successful real-world deployment. Our experiments show that the proposed UHAS representation enables stable dexterous control and cross-embodiment policy transfer across robotic hands.

Figures

Figures reproduced from arXiv: 2607.03570 by the authors.

Figure 1
Figure 1. In our unified hand action space, an action is represented as the deformation of a canonical [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the process of creating a sphere for a robotic hand given its URDF. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Construction of the unified hand surface correspondence. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Sphere deformation parameterization in the Unified Hand Action Space (UHAS). (a) Initial [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: We classify hand joints into (a) lateral joints and (b) encompassing joints and; (c) Illustra [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: (a) Simulation setup with 4 hands (b) Our real-world setup of the LEAP hand [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: An example of a real-world run of our policy for in-hand cube reorientation. [PITH_FULL_IMAGE:figures/full_fig_p008_7.png]
Figure 8
Figure 8. Figure 8: Illustration of the joint classification procedure in the Cascade Inverse Kinematics (CIK) [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Homogeneous observations across different robotic hands. [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Illustration of our real-world setup for the Allegro hand [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 15 linked inside Pith

  1. [1]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware.arXiv preprint arXiv:2304.13705, 2023

  2. [2]

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025

  3. [3]

    Kalashnikov, A

    D. Kalashnikov, A. Irpan, P. Pastor, J. Ibarz, A. Herzog, E. Jang, D. Quillen, E. Holly, M. Kalakrishnan, V . Vanhoucke, and S. Levine. Scalable deep reinforcement learning for vision-based robotic manipulation. InConference on robot learning, pages 651–673. PMLR, 2018

  4. [4]

    Singh, A

    R. Singh, A. Allshire, A. Handa, N. Ratliff, and K. Van Wyk. Dextrah-rgb: Visuomotor policies to grasp anything with dexterous hands.arXiv preprint arXiv:2412.01791, 2024

  5. [5]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. P. Foster, P. R. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An open-source vision-language-action model. InConference on Robot Learning (CoRL), 2024

  6. [6]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky.π 0: A vision-language-action flow model for general robot control.arXiv preprint arxiv:2410....

  7. [7]

    Black, N

    K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fu- sai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke,...

  8. [8]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manju- nath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...

Show all 48 references
  1. [9]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, P. Florence, C. Fu, M. G. Arenas, K. Gopalakrishnan, K. Han, K. Hausman, A. Herzog, J. Hsu, B. Ichter, A. Irpan, N. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. L...

  2. [10]

    Ghosh, H

    D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, R. Doshi, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An open-source generalist robot policy.arXiv preprint arXiv:2405.12213, 2024

  3. [11]

    O. X.-E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Her- zog, A. Irpan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. ...

  4. [12]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...

  5. [13]

    R. Wen, G. Chen, Z. Cui, M. Du, Y . Gou, Z. Han, L. Huang, M. Lei, Y . Li, Z. Li, W. Liu, Y . Liu, X. Ma, H. Niu, Y . Ouyang, Z. Ren, H. Shi, W. Xu, H. Zhang, J. Zhang, X. Zhang, L. Zheng, W. Zhong, Y . Zhou, Z. Zhu, and H. Li. Gr-dexter technical report.arXiv preprint arXiv:2...

  6. [14]

    Zhang, J

    Z. Zhang, J. Pang, Z. Yang, K. Li, M. Liao, S. Zhang, G. Chi, J. Guo, H. ang Gao, M. Shi, D. Ge, Y . Mu, J. Gu, R. Chen, H. Dong, H. Xu, L. Yi, Y . Zhu, H. Zhao, P. Wang, S. Zhang, G. Yao, J. Chen, H. Li, and H. Zhao. Dexora: Open-source vla for high-dof bimanual dexterity. ar...

  7. [15]

    Zheng, D

    R. Zheng, D. Niu, Y . Xie, J. Wang, M. Xu, Y . Jiang, F. Casta ˜neda, F. Hu, Y . L. Tan, L. Fu, T. Darrell, F. Huang, Y . Zhu, D. Xu, and L. Fan. Egoscale: Scaling dexterous manipulation with diverse egocentric human data.arXiv preprint arXiv:2602.16710, 2026

  8. [16]

    K. Shaw, A. Agarwal, and D. Pathak. Leap hand: Low-cost, efficient, and anthropomorphic hand for robot learning.Robotics: Science and Systems (RSS), 2023

  9. [17]

    Allegro hand.https://www.allegrohand.com/

    Wonik Robotics Co., Ltd. Allegro hand.https://www.allegrohand.com/. Accessed: 2025-08-08

  10. [18]

    Romero, D

    J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Modeling and capturing hands and bodies together.ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6), Nov. 2017

  11. [19]

    Shadow dexterous hand.https://www.shadowrobot.com/ products/dexterous-hand/

    The Shadow Robot Company. Shadow dexterous hand.https://www.shadowrobot.com/ products/dexterous-hand/. Accessed: 2025-08-08

  12. [20]

    Schwarke, M

    C. Schwarke, M. Mittal, N. Rudin, D. Hoeller, and M. Hutter. Rsl-rl: A learning library for robotics research.arXiv preprint arXiv:2509.10771, 2025

  13. [21]

    Andrychowicz, B

    M. Andrychowicz, B. Baker, M. Chociej, R. J ´ozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, J. Schneider, S. Sidor, J. Tobin, P. Welinder, L. Weng, and W. Zaremba. Learning dexterous in-hand manipulation.The International Journal of Robotics Rese...

  14. [22]

    Y . Chen, T. Wu, S. Wang, X. Feng, J. Jiang, Z. Lu, S. McAleer, H. Dong, S.-C. Zhu, and Y . Yang. Towards human-level bimanual dexterous manipulation with reinforcement learning. Advances in Neural Information Processing Systems, 35:5150–5163, 2022. 11

  15. [23]

    Handa, A

    A. Handa, A. Allshire, V . Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. V . Wyk, A. Zhurkevich, B. Sundaralingam, Y . Narang, J.-F. Lafleche, D. Fox, and G. State. Dex- treme: Transfer of agile in-hand manipulation from simulation to reality. In2023 IEEE Inte...

  16. [24]

    Y . J. Ma, W. Liang, G. Wang, D.-A. Huang, O. Bastani, D. Jayaraman, Y . Zhu, J. Fan, et al. Eureka: Human-level reward design via coding large language models. InInternational con- ference on learning Representations, volume 2024, pages 26516–26560, 2024

  17. [25]

    Zakka, B

    K. Zakka, B. Tabanpour, Q. Liao, M. Haiderbhai, S. Holt, J. Y . Luo, A. Allshire, E. Frey, K. Sreenath, L. A. Kahrs, C. Sferrazza, Y . Tassa, and P. Abbeel. Mujoco playground, 2025. URLhttps://arxiv.org/abs/2502.08844

  18. [26]

    K. Shaw, S. Bahl, and D. Pathak. Videodex: Learning dexterity from internet videos. In Conference on Robot Learning, pages 654–665. PMLR, 2023

  19. [27]

    Cheng, J

    X. Cheng, J. Li, S. Yang, G. Yang, and X. Wang. Open-television: Teleoperation with immer- sive active visual feedback.arXiv preprint arXiv:2407.01512, 2024

  20. [28]

    C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu. Dexcap: Scalable and portable mocap data collection system for dexterous manipulation.arXiv preprint arXiv:2403.07788, 2024

  21. [29]

    M. Xu, H. Zhang, Y . Hou, Z. Xu, L. Fan, M. Veloso, and S. Song. Dexumi: Using hu- man hand as the universal manipulation interface for dexterous manipulation.arXiv preprint arXiv:2505.21864, 2025

  22. [30]

    K. Li, P. Li, T. Liu, Y . Li, and S. Huang. Maniptrans: Efficient dexterous bimanual manipula- tion transfer via residual learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6991–7003, 2025

  23. [31]

    T. Tao, M. K. Srirama, J. J. Liu, K. Shaw, and D. Pathak. Dexwild: Dexterous human interac- tions for in-the-wild robot policies.arXiv preprint arXiv:2505.07813, 2025

  24. [32]

    Doshi, H

    R. Doshi, H. Walke, O. Mees, S. Dasari, and S. Levine. Scaling cross-embodied learn- ing: One policy for manipulation, navigation, locomotion and aviation.arXiv preprint arXiv:2408.11812, 2024

  25. [33]

    Zheng, J

    J. Zheng, J. Li, D. Liu, Y . Zheng, Z. Wang, Z. Ou, Y . Liu, J. Liu, Y .-Q. Zhang, and X. Zhan. Universal actions for enhanced embodied foundation models. InProceedings of the Computer Vision and Pattern Recognition Conference, pages 22508–22519, 2025

  26. [34]

    Jiang, Y

    G. Jiang, Y . Liang, J. Ye, J.-Y . Huang, C. Jing, R. Duan, P. Abbeel, X. Wang, and X. Zou. Cross-hand latent representation for vision-language-action models.arXiv preprint arXiv:2603.10158, 2026

  27. [35]

    Z. Wei, Y . Yao, and M. Ding. One hand to rule them all: Canonical representations for unified dexterous manipulation.arXiv preprint arXiv:2602.16712, 2026

  28. [36]

    Z. Wei, Z. Xu, J. Guo, Y . Hou, C. Gao, Z. Cai, J. Luo, and L. Shao. D (r, o) grasp: A unified representation of robot and object interaction for cross-embodiment dexterous grasping.arXiv preprint arXiv:2410.01702, 2024

  29. [37]

    Khargonkar, L

    N. Khargonkar, L. F. Casas, , B. Prabhakaran, and Y . Xiang. Robotfingerprint: Unified gripper coordinate space for multi-gripper grasp synthesis. 2024

  30. [38]

    X. Fei, Z. Xu, H. Fang, T. Zhang, and L. Shao. T (r, o) grasp: Efficient graph diffusion of robot-object spatial transformation for cross-embodiment dexterous grasping.arXiv preprint arXiv:2510.12724, 2025. 12

  31. [39]

    Z. Wu, R. A. Potamias, X. Zhang, Z. Zhang, J. Deng, and S. Luo. Cedex: Cross-embodiment dexterous grasp generation at scale from human-like contact representations.arXiv preprint arXiv:2509.24661, 2025

  32. [40]

    Y . Wu, Y . Lin, W. Lao, Y . Lin, Y .-L. Wei, W.-S. Zheng, and A. Wu. Dexgrasp-zero: A morphology-aligned policy for zero-shot cross-embodiment dexterous grasping.arXiv preprint arXiv:2603.16806, 2026

  33. [41]

    X. He, A. Polavaram, Y . Cao, O. Deshmukh, T. Wang, X. Zhou, and K. Fang. Generate, transfer, adapt: Learning functional dexterous grasping from a single human demonstration. arXiv preprint arXiv:2601.05243, 2026

  34. [42]

    Mittal, P

    M. Mittal, P. Roth, J. Tigue, A. Richard, O. Zhang, P. Du, A. Serrano-Mu ˜noz, X. Yao, R. Zurbr ¨ugg, N. Rudin, L. Wawrzyniak, M. Rakhsha, A. Denzler, E. Heiden, A. Borovicka, O. Ahmed, I. Akinola, A. Anwar, M. T. Carlson, J. Y . Feng, A. Garg, R. Gasoto, L. Gulich, Y . Guo, M...

  35. [43]

    Apriltag.https://github.com/AprilRobotics/apriltag

    AprilRobotics. Apriltag.https://github.com/AprilRobotics/apriltag

  36. [44]

    Park and P

    Y . Park and P. Agrawal. Aprilcube: 3d-printable fiducial targets for reliable 6-dof pose estima- tion, 2026. URLhttps://github.com/younghyopark/aprilcube. 13 A Cascade Inverse Kinematics The Cascade Inverse Kinematics (CIK) algorithm maps a deformed canonical sphere to embodi...

  37. [45]

    Sweep the lateral joint across its full range of motion in uniform steps

  38. [46]

    For each sampled value, fix the lateral joint and solve the encompassing joints on theun- deformedreference sphere. 14

  39. [47]

    Record the resulting azimuthal angleθ fingertip of the corresponding fingertip (surface point) on the canonical sphere frame

  40. [48]

    During inference, the policy outputs a lateral deformation∆θfor the driving plane aligned with the finger

    Store the mapping: lateral joint valueq lateral 7→θ fingertip. During inference, the policy outputs a lateral deformation∆θfor the driving plane aligned with the finger. We compute the fingertip target angle θfingertip =θ initial + ∆θ, whereθ initial corresponds to the neutral...

Pith tools

Reviewed July 12, 2026 · model on record in the stance chip above.