Pith. sign in

REVIEW 4 major objections 7 minor 1 cited by

mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity

T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A 16-DoF tendon-driven hand paired with a diffusion policy and a curated data protocol reaches 93.3% out-of-distribution success on real-world dexterous manipulation tasks.

desk verdict A well-built dexterous hand system with a useful data recipe, but the headline success rates are underdetermined by missing trial counts and error bars, and the self-correction claim is mislabeled. read the letter →

arxiv 2506.11916 v1 pith:WAZOPPKW submitted 2025-06-13 cs.RO

classification cs.RO
keywords dexterousmanipulationdiffusionpolicyimitationlearningtendon-drivenrobotichandteleoperationself-correctiondatascalingrobotgeneralization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that a complete recipe—new 16-DoF tendon-driven hand, a two-interface teleoperation pipeline, a diffusion policy with relative Cartesian end-effector actions, and a data protocol that rotates task configurations and adds self-correction trajectories—can make real-world dexterous manipulation sample-efficient and generalizable. The evidence comes from real-robot evaluation on bread pick-and-place, bottle sorting, and battery insertion, where the full recipe reaches 93.3% out-of-distribution success and self-correction data adds up to 33.3 percentage points. The authors argue the recipe's components are all necessary: ablations removing any one (absolute actions, wrong base frame, single task config, unfiltered data, no self-correction trajectories) drop success to between 5% and 56.7%.

What carries the argument

The load-bearing machinery is the relative end-effector action representation together with the self-correction data protocol. The policy outputs a sequence of Cartesian target poses for the arm, expressed relative to the last observed proprioceptive end-effector pose; using any other base frame (the first commanded target pose, or absolute poses) mismatches the conditioning at inference and sharply hurts success. The hand joint angles are kept absolute so grasp configurations are not lost. The data protocol's second half—changing task configuration roughly every 100 episodes, labeling and filtering, then collecting reset-scene recovery demonstrations for documented failure modes—is what turns a plain imitation policy into one that visibly re-orients a bottle or recovers a failed grasp.

What would settle it

A direct test would be to train a fresh policy on Bottle Sorting with the full protocol but with self-correction trajectories withheld and all other factors identical; the paper reports 75.0% with them and 41.7% without, so a replication that fails to reproduce a substantial gap would refute the claimed boost. Alternatively, varying the task-config change frequency (e.g., every 30 vs. 300 episodes) and measuring success would test whether the transferred scaling-law ratio is actually optimal for this hand.

Watch

Extended reading notes

Core claim

The central claim is that a particular combination of hardware, data curation, and generative control yields general-purpose dexterity in the real world. The paper introduces the mimic-one recipe: a 16-DoF tendon-driven hand with wide-angle wrist cameras, teleoperation through gloves or a VR headset, and a UNet diffusion policy that predicts 48-step action chunks at 15 Hz from raw images and proprioception. Three representation choices carry much of the generalization: Cartesian target end-effector poses expressed relative to the last observed proprioceptive pose (not the first commanded pose), continuous 6D rotations, and absolute hand joint angles. The data protocol rotates task configurations roughly every 100 episodes following a scaling-law observation from earlier imitation-learning work, filters demonstrations by human labeling, and collects targeted self-correction trajectories after identifying failure modes. Across the three tasks, the recipe achieves success rates of 93.3%, 75.0%, and 37.5%, with corresponding self-correction boosts of +26.6, +33.3, and +25.0 percentage points.

Load-bearing premise

The protocol assumes that the ideal task-configuration mix measured in earlier imitation-learning studies transfers to this 16-DoF hand and these tasks, so it fixes a change of setting roughly every 100 episodes; if the optimal ratio is different here, the reported diversity and generalization would change.

Editorial extensions

If this is right

  • With the same protocol, increasing the number of curated demonstrations from 20% to 100% of the dataset raised Bread Pick success from 22.5% to 93.3%, so the recipe scales with data.
  • Adding self-correction trajectories gave success boosts of +26.6, +33.3, and +25.0 percentage points across the three tasks, so targeted failure-recovery data is a reliable lever.
  • The relative end-effector action representation with the correct base pose improved Bread Pick success from 23.0% and 30.0% in the absolute and wrong-base variants to 93.3%, so this representation choice is essential for generalization.
  • Removing data diversity (single task config) or data filtering dropped success to 15.0% and 5.0%, so the curation protocol is load-bearing.
  • Policies run at 15 Hz with a 48-step action chunk covering 3.2 seconds, so the recipe is compatible with real-time control on this hardware.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the same protocol could be evaluated on a simpler two-finger gripper platform, and if the relative-action and config-diversity gains transfer, the recipe generalizes beyond this specific hand.
  • Beyond the paper: a multi-task version sharing representations across task families might lower the total demonstration count, since the paper trains one policy per task and notes this is less data-efficient.
  • Beyond the paper: the self-correction loop could be automated by detecting failure modes from the paper's taxonomy and synthesizing reset scenes, rather than collecting them by hand, which would cut human cost.
  • Beyond the paper: if the scaling trend continues beyond the full dataset, measuring where success rates plateau would show when the current protocol's data budget saturates.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes mimic-one, an integrated recipe for dexterous manipulation with a new 16-DoF tendon-driven hand mounted on a Franka arm. The recipe includes a teleoperation pipeline (glove and Vision Pro), a data collection protocol with task-config randomization and explicit collection of self-correction recovery trajectories, and a diffusion-policy architecture with relative Cartesian end-effector actions, absolute hand joint angles, and 6D rotation representations. The authors report real-world success rates on three tasks (bread pick-and-place, bottle sorting, battery insertion), scaling trends with dataset size, and an ablation study of recipe components.

Significance. The combination of a specifically designed 16-DoF hand, a practical teleoperation interface, and a diffusion-based policy with carefully chosen action representations is a useful engineering contribution. The data-collection protocol, especially the systematic collection of failure-recovery trajectories, is a practical idea that could benefit the community even beyond this hardware. However, the empirical claims rest on a small number of trials without statistical evidence, and the 'emergent self-correction' framing is at odds with the explicit training on recovery data. The manuscript does not release code or data, which limits reproducibility. If the authors address the statistical grounding and reframe the self-correction claim, this could be a valuable systems paper.

major comments (4)
  1. [Section 4, Figs. 5–6] The reported success rates (e.g., 22.5%, 93.3%, 75.0%, 37.5%, +26.6%, +33.3%, +25.0%) are presented without trial counts, confidence intervals, or any indication of run-to-run variability. The fractional values imply very small evaluation sets (93.3% = 14/15, 66.7% = 2/3, 37.5% = 3/8). With these sample sizes, the differences that are claimed as 'clear scaling trends' and 'self-correction boosts' may be within sampling noise; for example, the difference between 66.7% and 93.3% is not statistically significant (Fisher exact p≈0.16). Please report the number of trials per condition and per-task raw results, and include confidence intervals or exact binomial tests.
  2. [Abstract, Section 3.5, Section 4] The paper repeatedly describes self-correcting behaviors as 'emergent' (Abstract: 'emergent self-correcting behaviors'; Section 4; Conclusion: 'emerging self-corrective behaviors'). However, the data collection protocol in Section 3.5 explicitly includes a step that identifies common failure modes and then collects additional 'self-correction trajectories' from reset scenes (steps 4–5 and Fig. 3d). The observed performance boost from adding these trajectories is a direct, expected outcome of training on recovery demonstrations, not an emergent phenomenon. Either remove the term 'emergent' or provide evidence of self-correction in a policy that was not trained on recovery data.
  3. [Section 3.5, step 1.3] The task-config change interval of 'approximately every 100 episodes' is transferred from the scaling-law analysis of [22] without validation for the 16-DoF hand, the three task families, or the specific policy architecture used here. Because this interval determines the diversity of the training data, the reported generalization and scaling results may hinge on an untested assumption. Please provide an ablation on this interval or explicitly discuss the sensitivity of the results to this hyperparameter.
  4. [Section 4, Fig. 5] The caption states that 'dashed bars include self-corrected successes; solid bars count any error as failure,' while the text interprets the dashed-vs-solid comparison as training with versus without self-correction trajectories. This conflates the training-data condition with the evaluation criterion. If the two bars differ in both training data and evaluation protocol, the reported '+33.3% self-correction boost' is not a clean measurement of the value of the added recovery data. Please use a consistent evaluation protocol for both conditions, or clearly separate the two factors.
minor comments (7)
  1. [Section 1] The claim of 'state-of-the-art' in the Introduction is not supported by comparisons with existing imitation-learning methods on the same tasks; the ablations only consider variants of the proposed recipe.
  2. [Figure 6] Figure 6 lists nine numerical success rates but the text and legend describe six experimental conditions; the mapping between bars and conditions is unclear and should be fixed.
  3. [Section 3.2] The notation in Section 3.2 introduces v^h_i and v^r_i without defining i; specify that i indexes the 15 key-vectors.
  4. [Section 3.3 and Appendix A] The observation and action horizons (H_o=2, H_a=48, 15Hz) appear only in Appendix A; consider giving them in Section 3.3 to make the recipe self-contained.
  5. [Section 3.1] The hardware section would benefit from a table with camera specifications (resolution, field of view) and hand dimensions/weight.
  6. [Abstract and Section 1] The phrase 'high-frequency generative control' is used in the abstract and introduction, but the inference rate is 15Hz; clarify what 'high-frequency' means in this context.
  7. [Section 4] There is no mention of whether the evaluation rollouts were performed in one session or across multiple days; environmental drift could affect the results. Report the evaluation protocol in detail.

Circularity Check

1 steps flagged · score 3.0 of 10

Minor circularity: the reported self-correction boost is by construction because recovery trajectories are explicit training inputs, though the core recipe is otherwise independent.

  1. fitted input called prediction [Section 3.5 step (5); Section 4, Fig. 5 caption and text]
    "(5) Repeat — collect self-correction trajectories: based on the common failure modes. The robot and workspace are positioned in a failure state before starting the episode collection (Fig. 3.d). ... Dashed bars include self-corrected successes; solid bars count any error as failure. ... Success rates increased substantially when allowing for recovery behaviors: Bread Pick (+26.6%), Bottle Sort (+33.3%), and Battery Insertion (+25.0%)."

    The paper presents the +26.6/+33.3/+25.0 boosts as evidence of 'emergent self-correcting behaviors' in the Abstract and Conclusion, but those behaviors are trained from explicitly collected recovery demonstrations, not emergent. Section 3.5 step (5) says the robot and workspace are deliberately placed in a failure state and correction episodes are recorded, and Fig. 5's evaluation counts 'self-corrected successes' as successes. Thus the reported improvement is the direct consequence of adding self-correction trajectories to the training set and then scoring the policy in a way that rewards exactly those trajectories. The claim that the boost arises from 'emergent' behavior is a renaming of the input (recovery data) as an output (self-correction), not an independent prediction.

full rationale

The core mimic-one derivation—teleoperation, diffusion policy, relative action representations, data curation, and scaling evaluation—does not reduce to its inputs by construction. The architecture is a standard Diffusion Policy [17] with UMI-style relative actions [20]; the task-config ratio is imported from an external scaling-law study [22]; and the ablations compare real trained variants. The only substantive circularity is the self-correction framing: the paper collects targeted recovery demonstrations and then labels the resulting policy behavior 'emergent self-correcting behavior,' and the reported boost is measured with a success criterion that explicitly counts self-corrected successes. This is a partial, framing-level circularity rather than a fatal one. Trial counts and confidence intervals are absent, which affects evidential strength but is a statistical reporting issue, not a circularity. Self-citations ([6], [7]) appear only in related-work context and are not load-bearing.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The paper's claims rest on a stack of hand-picked hyperparameters and borrowed design assumptions; the most consequential is the transfer of scaling-law ratios from [22]. No code, data, or hardware artifacts are released, so the exact recipe cannot be audited independently.

free parameters (6)
  • Task config change interval = ~100 episodes
    Set following scaling-law optimal ratio from [22]; no sensitivity analysis for the 16-DoF hand.
  • Observation horizon Ho = 2
    Hand-selected hyperparameter; no ablation reported.
  • Action chunk Ha = 48
    Corresponds to 3.2 s at 15 Hz; chosen for smoothness and inference latency.
  • Control/inference rate = 15 Hz
    Chosen as the real-time policy rate.
  • DDIM denoising steps = 16
    Number of inference-time denoising steps; trade-off between speed and action quality.
  • Action re-computation schedule = after 10th action of 48
    Re-computes the action chunk before full execution; heuristic, no ablation.
assumptions (4)
  • domain assumption CLIP ViT-B/16 visual features are suitable for visuomotor policies on this hand and tasks.
    Pretrained CLIP encoders are used without manipulation-specific fine-tuning; no comparison to other encoders is provided.
  • domain assumption Data scaling-law optimal ratio from [22] transfers to the 16-DoF hand and these tasks.
    Used in Section 3.5 (1.3) to set task config frequency; transfer is not validated in this paper.
  • domain assumption Relative Cartesian action representation with the last observed proprioceptive pose as base frame is correct and stable.
    The authors state this is critical; the ablation confirms the wrong base frame hurts, but no theoretical justification is given beyond the argument in Section 3.4.
  • domain assumption Key-vector retargeting from [24,28] preserves sufficient hand pose information for dexterous task execution.
    Teleoperation relies on this retargeting; its fidelity is not quantitatively evaluated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity." pith.science (2026). https://pith.science/paper/WAZOPPKW

@misc{pith2026250611916,
  author       = {Pith},
  title        = {Pith review of: mimic-one: a Scalable Model Recipe for General Purpose Robot Dexterity},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WAZOPPKW}},
  note         = {Machine review of arXiv:2506.11916}
}
read the original abstract

We present a diffusion-based model recipe for real-world control of a highly dexterous humanoid robotic hand, designed for sample-efficient learning and smooth fine-motor action inference. Our system features a newly designed 16-DoF tendon-driven hand, equipped with wide angle wrist cameras and mounted on a Franka Emika Panda arm. We develop a versatile teleoperation pipeline and data collection protocol using both glove-based and VR interfaces, enabling high-quality data collection across diverse tasks such as pick and place, item sorting and assembly insertion. Leveraging high-frequency generative control, we train end-to-end policies from raw sensory inputs, enabling smooth, self-correcting motions in complex manipulation scenarios. Real-world evaluations demonstrate up to 93.3% out of distribution success rates, with up to a +33.3% performance boost due to emergent self-correcting behaviors, while also revealing scaling trends in policy performance. Our results advance the state-of-the-art in dexterous robotic manipulation through a fully integrated, practical approach to hardware, learning, and real-world deployment.

Figures

Figures reproduced from arXiv: 2506.11916 by the authors.

Figure 1
Figure 1. mimic-one is a scalable model recipe and data collection protocol for general purpose dexterous manipulation. Our system features a newly designed 16-DoF tendon-driven hand, a 7-DoF Franka Emika Panda robot arm, and a VR teleoperation system for data collection. We showcase the framework across difficult, dynamic real-world manipulation scenarios, which highlight the system’s high frequency fine-motor skills, visual… view at source ↗
Figure 2
Figure 2. System overview and teleoperation setup. ( [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. The mimic-one data collection protocol and policy recipe. (a) Teleoperation data collec￾tion. Variation in the setting involve randomize object positions, robot starting pose, “task config”, and distractors. (b) Data labeling and filtering (removing failures, non-stable grasps and suboptimal completions). (c) The diffusion policy model architecture, receiving as conditioning input an ob￾servation horizon with encode… view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Representative rollout sequences across three benchmark tasks. The mimic-one policy [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Task success rates vs. dataset scale. Dashed bars include self-corrected successes; solid bars count any error as failure. Absolute Actions Rel: wrong Base Frame Single Task Config Unfiltered Data No Self-Corr Traj mimic-one 0 20 40 60 80 100 Success Rate (%) 23.0 30.0…
Figure 6
Figure 6. Figure 6: Ablation study on Bread Pick-and￾Place, comparing the full mimic-one recipe (vi) against variants removing key components: (i) Absolute actions, (ii) Incorrect relative base frame, (iii) Single task configuration, (iv) Unfil￾tered data, and (v) No self-correction data.…

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Smooth Operator: A Real-Time Sampling-Based Algorithm for Kinematic Hand Retargeting

    cs.RO 2026-07 unverdicted novelty 6.0 of 10

    Sampling-Based Retargeter (SBR) delivers lower-jitter real-time kinematic hand retargeting and higher task success with less operator fatigue than gradient-based baselines in an 18-person study.

Reference graph

Works this paper leans on

36 extracted references · 10 canonical work pages · cited by 1 Pith paper

  1. [22]

    F. Lin, Y . Hu, P. Sheng, C. Wen, J. You, and Y . Gao. Data Scaling Laws in Imitation Learning for Robotic Manipulation, Oct. 2024. URLhttp://arxiv.org/abs/2410.18647. arXiv:2410.18647 [cs]

  2. [1]

    H. Liu, K. Wu, P. Meusel, G. Hirzinger, M. Jin, Y . Liu, S. Fan, T. Lan, and Z. Chen. A dex- terous humanoid five-fingered robotic hand. InRO-MAN 2008 - The 17th IEEE International Symposium on Robot and Human Interactive Communication, pages 371–376, Aug. 2008. doi:10.1109/ROMAN.2008.4600694. URLhttps://ieeexplore.ieee.org/document/ 4600694. ISSN: 1944-9437

  3. [2]

    H. Liu, K. Wu, P. Meusel, N. Seitz, G. Hirzinger, M. Jin, Y . Liu, S. Fan, T. Lan, and Z. Chen. Multisensory five-finger dexterous hand: The DLR/HIT Hand II. In2008 IEEE/RSJ Interna- tional Conference on Intelligent Robots and Systems, pages 3692–3697, Sept. 2008. doi:10. 1109/IROS.2008.4650624. URLhttps://ieeexplore.ieee.org/document/4650624. ISSN: 2153-0866

  4. [3]

    Grebenstein, A

    M. Grebenstein, A. Albu-Sch ¨affer, T. Bahls, M. Chalon, O. Eiberger, W. Friedl, R. Gruber, S. Haddadin, U. Hagn, R. Haslinger, H. H ¨oppner, S. J ¨org, M. Nickl, A. Nothhelfer, F. Petit, J. Reill, N. Seitz, T. Wimb ¨ock, S. Wolf, T. W ¨usthoff, and G. Hirzinger. The DLR hand arm system. In2011 IEEE International Conference on Robotics and Automation, pag...

  5. [4]

    Xu and E

    Z. Xu and E. Todorov. Design of a highly biomimetic anthropomorphic robotic hand to- wards artificial limb regeneration. In2016 IEEE International Conference on Robotics and Automation (ICRA), pages 3485–3492, May 2016. doi:10.1109/ICRA.2016.7487528. URL https://ieeexplore.ieee.org/abstract/document/7487528

  6. [5]

    U. Kim, D. Jung, H. Jeong, J. Park, H.-M. Jung, J. Cheong, H. R. Choi, H. Do, and C. Park. Integrated linkage-driven dexterous anthropomorphic robotic hand.Nature Communications, 12(1):7177, Dec. 2021. ISSN 2041-1723. doi:10.1038/s41467-021-27261-0. URLhttps:// www.nature.com/articles/s41467-021-27261-0. Publisher: Nature Publishing Group

  7. [6]

    Toshimitsu, B

    Y . Toshimitsu, B. Forrai, B. G. Cangan, U. Steger, M. Knecht, S. Weirich, and R. K. Katzschmann. Getting the Ball Rolling: Learning a Dexterous Policy for a Biomimetic Tendon-Driven Hand with Rolling Contact Joints. In2023 IEEE-RAS 22nd Interna- tional Conference on Humanoid Robots (Humanoids), pages 1–7, Dec. 2023. doi: 10.1109/Humanoids57100.2023.10375...

  8. [7]

    T. J. K. Buchner, S. Rogler, S. Weirich, Y . Armati, B. G. Cangan, J. Ramos, S. T. Twiddy, D. M. Marini, A. Weber, D. Chen, G. Ellson, J. Jacob, W. Zengerle, D. Katalichenko, C. Keny, W. Matusik, and R. K. Katzschmann. Vision-controlled jetting for com- posite systems and robots.Nature, 623(7987):522–530, Nov. 2023. ISSN 1476-

Show all 36 references
  1. [8]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D....

  2. [9]

    Zitkovich, T

    B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, Q. Vuong, V . Vanhoucke, H. Tran, R. Soricut, A. Singh, J. Singh, P. Sermanet, P. R. San- keti, G. Salazar, M. S. Ryoo, K. Reymann, K. Rao, K. Pertsch, I. Mordatch, H. Michalewski, Y ....

  3. [10]

    O. M. Team, D. Ghosh, H. Walke, K. Pertsch, K. Black, O. Mees, S. Dasari, J. Hejna, T. Kreiman, C. Xu, J. Luo, Y . L. Tan, L. Y . Chen, P. Sanketi, Q. Vuong, T. Xiao, D. Sadigh, C. Finn, and S. Levine. Octo: An Open-Source Generalist Robot Policy, May 2024. URL http://arxiv.or...

  4. [11]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An Open-Source Vision-Language-Action Model, Sept

  5. [12]

    Pertsch, K

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. FAST: Efficient Action Tokenization for Vision-Language-Action Models, Jan

  6. [13]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning Fine-Grained Bimanual Manipu- lation with Low-Cost Hardware, Apr. 2023. URLhttp://arxiv.org/abs/2304.13705. arXiv:2304.13705 [cs]

  7. [14]

    Song and S

    Y . Song and S. Ermon. Generative Modeling by Estimating Gradients of the Data Dis- tribution. InAdvances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URLhttps://proceedings.neurips.cc/paper/2019/hash/ 3001ef257407d5a371a96dcd947c7d93-Abs...

  8. [15]

    J. Ho, A. Jain, and P. Abbeel. Denoising Diffusion Probabilistic Models. InAd- vances in Neural Information Processing Systems, volume 33, pages 6840–6851. Curran Associates, Inc., 2020. URLhttps://proceedings.neurips.cc/paper/2020/hash/ 4c5bcfec8584af0d967f1ab10179ca4b-Abstract.html

  9. [16]

    Lipman, R

    Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow Matching for Generative Modeling, Feb. 2023. URLhttp://arxiv.org/abs/2210.02747. arXiv:2210.02747 [cs]

  10. [17]

    C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, Mar. 2024. URLhttp://arxiv. org/abs/2303.04137. arXiv:2303.04137 [cs]

  11. [18]

    Chisari, N

    E. Chisari, N. Heppert, M. Argus, T. Welschehold, T. Brox, and A. Valada. Learning Robotic Manipulation Policies from Point Clouds with Conditional Flow Matching, Sept. 2024. URL http://arxiv.org/abs/2409.07343. arXiv:2409.07343 [cs]

  12. [19]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky. pi0: A Visio...

  13. [20]

    C. Chi, Z. Xu, C. Pan, E. Cousineau, B. Burchfiel, S. Feng, R. Tedrake, and S. Song. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots, Mar. 2024. URLhttp://arxiv.org/abs/2402.10329. arXiv:2402.10329 [cs]. 13

  14. [21]

    H. Ha, Y . Gao, Z. Fu, J. Tan, and S. Song. UMI on Legs: Making Manipulation Policies Mobile with Manipulation-Centric Whole-body Controllers, July 2024. URLhttp://arxiv.org/ abs/2407.10353. arXiv:2407.10353 [cs]

  15. [23]

    T. Lin, Y . Zhang, Q. Li, H. Qi, B. Yi, S. Levine, and J. Malik. Learning Visuotactile Skills with Two Multifingered Hands, May 2024. URLhttp://arxiv.org/abs/2404.16823. arXiv:2404.16823 [cs]

  16. [24]

    K. Shaw, S. Bahl, and D. Pathak. VideoDex: Learning Dexterity from Internet Videos, Dec

  17. [25]

    C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. MimicPlay: Long-Horizon Imitation Learning by Watching Human Play, Feb. 2023. URLhttp://arxiv. org/abs/2302.12422. arXiv:2302.12422 [cs]

  18. [26]

    Kareer, D

    S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu. EgoMimic: Scaling Imitation Learning via Egocentric Video, Oct. 2024. URLhttp: //arxiv.org/abs/2410.24221. arXiv:2410.24221 [cs]

  19. [27]

    J. Luo, Z. Hu, C. Xu, Y . L. Tan, J. Berg, A. Sharma, S. Schaal, C. Finn, A. Gupta, and S. Levine. SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning, Mar. 2025. URLhttp://arxiv.org/abs/2401.16013. arXiv:2401.16013 [cs]

  20. [28]

    Handa, K

    A. Handa, K. V . Wyk, W. Yang, J. Liang, Y .-W. Chao, Q. Wan, S. Birchfield, N. Ratliff, and D. Fox. DexPilot: Vision Based Teleoperation of Dexterous Robotic Hand-Arm System, Oct

  21. [29]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning Transferable Visual Models From Natural Language Supervision.arXiv:2103.00020 [cs], Feb. 2021. URLhttp://arxiv. org/abs/2103.000...

  22. [30]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, June 2021. URL http://arxiv.org/abs/201...

  23. [31]

    Y . Zhou, C. Barnes, J. Lu, J. Yang, and H. Li. On the Continuity of Rotation Repre- sentations in Neural Networks, June 2020. URLhttp://arxiv.org/abs/1812.07035. arXiv:1812.07035 [cs]. 14

  24. [2019]

    arXiv:1910.03135 [cs]

    URLhttp://arxiv.org/abs/1910.03135. arXiv:1910.03135 [cs]

  25. [2022]

    arXiv:2212.04498 [cs, eess]

    URLhttp://arxiv.org/abs/2212.04498. arXiv:2212.04498 [cs, eess]

  26. [2024]

    arXiv:2406.09246 [cs]

    URLhttp://arxiv.org/abs/2406.09246. arXiv:2406.09246 [cs]

  27. [2025]

    arXiv:2501.09747 [cs]

    URLhttp://arxiv.org/abs/2501.09747. arXiv:2501.09747 [cs]

  28. [4687]

    URLhttps://www.nature.com/articles/ s41586-023-06684-3

    doi:10.1038/s41586-023-06684-3. URLhttps://www.nature.com/articles/ s41586-023-06684-3. Publisher: Nature Publishing Group

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.