REVIEW 3 major objections 5 minor 1 cited by
LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LodeStar claims that long-horizon dexterous manipulation can be learned from 15 human demonstrations per task by automatically segmenting skills with foundation models, augmenting each skill with residual-RL synthetic data in simulation…
desk verdict A well-specified dexterity system with a clean integration story, but the 25% and 2x headline gains rest on 20-trial evaluations with no error bars or significance tests — a fixable evidence problem, not a broken method. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the Skill Routing Transformer (SRT): a transformer policy that maps a history of point-cloud observations to a low-level action and a discrete stage choice, either transition or one of the learned skills. It is trained on synthetic transition trajectories generated by sampling state pairs from the augmented termination and initiation sets and planning collision-free motions. The other load-bearing pieces are the per-frame skill discriminators $d_i(s_t)=\mathbb{1}[C^{\text{point}}_i(s_t)\wedge C^{\text{contact}}_i(s_t)]$, built by propagating one manual keypoint annotation across demonstrations and asking a vision-language model for Python score functions, and residual reinforcement learning, where a behavior-cloned diffusion base policy is combined with a PPO residual policy in a domain-randomized physics simulator and the successful rollouts are co-trained with real data. Together these mechanisms convert a handful of demos into a large, varied dataset and a controller that can switch between skills at execution time.
What would settle it
Run LodeStar on a fourth long-horizon dexterous task whose objects are transparent or reflective, without manual mesh separation or opaque tape; if the success rate collapses to the real-only baseline, the claim that the pipeline generalizes beyond hand-prepared scenes is falsified.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that decomposing long-horizon dexterous tasks into frame-level skills—rather than predefined motion primitives or terminal-state conditions—makes synthetic data augmentation tractable, and that chaining the resulting skill policies with a learned routing transformer yields robust real-world execution. Each skill is defined by a per-frame discriminator: a conjunction of a keypoint spatial-relation constraint and a fingertip-contact constraint, generated from tracked keypoints and a vision-language model. The simulation environments are built by real-to-sim transfer, with meshes reconstructed from multi-view images, manually separated into links, and trained under randomized dynamics; a behavior-cloned base policy is refined by a residual PPO policy, and the successful simulated rollouts are co-trained with the real demonstrations. The reported results across three tasks and 20 trials per method show LodeStar-PC beating the best replay-based augmentation baseline by 25% average success, outperforming skill-chaining baselines by over 25%, and reaching 10/20 success under larger initial-state distributions where real-only training with 15 demos scores 0/20.
Load-bearing premise
The argument depends on the hand-built simulations—manually separated meshes, hand-picked randomization ranges, and the same few demonstrations as priors—being close enough to the real contact dynamics that residual-RL policies trained there still work on the physical robot.
Editorial extensions
If this is right
- If LodeStar is right, 15 demonstrations per task can replace hundreds or thousands of teleoperated trajectories for multi-stage dexterous tasks, substantially cutting data-collection cost.
- Robustness under out-of-distribution initial conditions should come from the domain-randomized residual-RL augmentation rather than from collecting more real data.
- Learning transitions in simulation removes the need for online motion planning at execution time, which should reduce deployment latency and hand-off failures during skill changes.
- The SRT's explicit stage prediction gives a natural way to inject human oversight or replanning at skill boundaries without retraining the low-level skills.
Reading between the lines
- Editorial extension: the same segmentation-plus-residual-RL pipeline should apply to bimanual or tool-use tasks whose skills can be recognized from keypoint and contact cues, although the paper only demonstrates single-arm dexterous hands.
- Editorial extension: if the reported OOD gains generalize, simulation augmentation targeted at skill boundaries may be a cheaper route to robustness than scaling real demonstrations, a direction the paper's 15-versus-50 demo comparison hints at but does not fully explore.
- Editorial extension: the reliance on manually separated meshes and opaque tape for transparent objects suggests that automating articulation detection and material handling is the next bottleneck for making the recipe fully hands-off.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. LodeStar proposes a three-stage pipeline for long-horizon dexterous manipulation from a small number of human demonstrations: (1) automatic skill segmentation using foundation models and tracked keypoints; (2) per-skill real-to-sim transfer with domain randomization and residual reinforcement learning to generate synthetic demonstration data, co-trained with the real demonstrations; and (3) a Skill Routing Transformer (SRT) that composes the learned skills by also generating transition motions. The paper reports real-world experiments on three tasks (Liquid Handling, Plant Watering, Light Bulb Assembly) using 15 demonstrations per task and 20 trials per condition, claiming an average 25% improvement over the best baseline (SkillMimicGen), improved out-of-distribution robustness, and positive ablations.
Significance. If the quantitative claims hold, LodeStar would be a valuable systems contribution: it combines automatic segmentation, simulation-based augmentation, and skill chaining in one pipeline, and the three real-world dexterous tasks with multi-fingered hands are nontrivial. The paper compares against several baselines (Real-only, T-STAR, Seq-Dex, MimicGen, SkillMimicGen) and includes a thoughtful limitations section. Its main weakness is statistical: all central comparisons rest on 20 trials per condition with no confidence intervals or significance tests, so the headline 25% improvement, the OOD robustness claims, and the ablation rankings are within binomial sampling noise. The limitations acknowledged in Section 7 (unmodeled dynamic parameters, manual taping of transparent objects, rigid objects only) honestly bound the generality of the results, but they do not fix the statistical grounding of the present comparison.
major comments (3)
- [Section 5.1, Fig. 4, Table 1] Every success-rate comparison in the paper uses 20 trials per condition and no confidence intervals or significance tests. For the OOD comparison in Table 1, LodeStar's 10/20 versus SkillMimicGen's 5/20 has a two-sided Fisher exact p of about 0.19, and the 8/20 versus 4/20 comparison has a p of about 0.30; both differences are well within binomial sampling noise. Consequently, the central claim that LodeStar-PC 'boosts the average performance by 25%' is not statistically substantiated. Please report per-condition success counts with binomial confidence intervals, increase the number of trials where feasible, apply an appropriate statistical test, and temper the wording if the trial budget cannot be increased.
- [Table 2] The ablation results in Table 2 use the same 20-trial granularity, so pairwise differences such as 10/20 versus 6/20 or 10/20 versus 4/20 are not statistically significant at the 0.05 level. The text's statements that removing components 'leads to 20% and 30% drops' conflate percentage-point differences with relative improvements; the data support at most directional trends, not component-level quantitative claims. Please report uncertainty on the ablations and use consistent absolute/relative language.
- [Section 5.1 and Section 6] The headline numbers are ambiguous. Section 5.1 says LodeStar-PC 'boosts the average performance by 25%' compared with SkillMimicGen, while Section 6 concludes that LodeStar 'achieves 2 times higher success rate compared to the best baseline'; these are different quantities unless the best baseline average is exactly 25%. The main text also does not provide a table of per-task success counts for Fig. 4, so the aggregate improvement cannot be checked against the raw data. Please define the metric precisely, report per-task counts, and align the relative-versus-absolute wording in the abstract, Sections 5 and 6.
minor comments (5)
- [Section 6] The sentence 'we presents LODE STAR' contains a typo; it should read 'we present LODE STAR'.
- [Fig. 4] Figure 4 shows bars without error bars or confidence intervals; adding binomial confidence intervals would help readers see the sampling uncertainty that the current text omits.
- [Appendix C.2] The statement 'we try our best to ensure consistent initial conditions for the evaluation of different methods' is not a reproducible evaluation protocol; please specify how initial object and robot poses were sampled and whether the same pose sets were used for every method.
- [Section 7] The limitations listed in Section 7, including unmodeled dynamic parameters, manual taping of transparent objects, and restriction to rigid objects, should be reflected in the abstract and conclusion, where the claims of 'robustness' currently appear without these caveats.
- [Table 2] The 'Simulation' column in Table 2 is not defined in the main text; please state what metric is being reported there and how it relates to the real-world success-rate metric.
Circularity Check
No significant circularity: the central 25% performance claim is an empirical comparison against external baselines, and self-citations are not load-bearing.
full rationale
This is an empirical systems paper rather than a derivation, and I found no circular step that reduces a prediction to its inputs. The central claim that LodeStar-PC boosts average performance by 25% over SkillMimicGen (Section 5.1, Fig. 4) is evaluated against external baselines on real-world rollouts, with each method trained from the same 15 demonstrations and assessed over 20 trials (Section C.2). The reported success criteria are task-completion outcomes that are independent of the method's internal skill discriminators: for Liquid Handling the tip must fall into the container, for Plant Watering the bottle must be placed in the cardboard box, and for Light Bulb Assembly the bulb must illuminate (Section C.2). The skill discriminators themselves are generated by an off-the-shelf VLM from language instructions and tracked keypoints (Sections 4.1 and B.2.2), so the evaluation is not definitionally tied to the segmentation. The synthetic data are produced by residual RL with sparse rewards and then co-trained with real demonstrations (Sections 4.2 and B.4), and the SRT policy is tested by whether it completes the full long-horizon task (Section 4.3). The ablations in Table 2 compare component variants against the full system; these are contributions of individual components, not renamed fits of the headline result. The self-citations, such as [93] for unsupervised skill discovery and [114] for progressive exploration, are background technique references and are not used to justify the central performance claim or to forbid alternative approaches. The limitations stated in Section 7, including unmodeled dynamic parameters, manually taped transparent objects, and rigid-object-only experiments, honestly restrict generality but do not indicate circularity. The absence of confidence intervals or significance tests for the 20-trial success rates is a statistical power concern, not a circularity defect. Overall, the central claim is self-contained against external benchmarks and receives a low circularity score.
Assumptions & free parameters
free parameters (3)
- Domain randomization ranges =
e.g., object mass scaling U(0.5,1.5), friction scaling U(0.7,1.3), gravity scaling U(0.9,1.1) (Table 3)
- Number of synthetic demonstrations per skill =
1000 per skill
- Flying point augmentation probability =
0.5%
assumptions (4)
- domain assumption 15 demonstrations per task are sufficient priors for segmentation and for initial and terminal state distributions.
- domain assumption Isaac Gym environments built from real-to-sim transfer faithfully model contact dynamics for residual RL to transfer.
- domain assumption VLM-generated discriminators from OpenAI o3 correctly identify skill boundaries and contact constraints.
- domain assumption Off-the-shelf perception models (DIFT, Co-Tracker, SAM2, FoundationPose) provide sufficient accuracy for keypoint, mask, and pose estimation.
Cite this review
Pith. "Pith review of LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations." pith.science (2026). https://pith.science/paper/P6IV7O7V
@misc{pith2026250817547,
author = {Pith},
title = {Pith review of: LodeStar: Long-horizon Dexterity via Synthetic Data Augmentation from Human Demonstrations},
year = {2026},
howpublished = {\url{https://pith.science/paper/P6IV7O7V}},
note = {Machine review of arXiv:2508.17547}
}
read the original abstract
Developing robotic systems capable of robustly executing long-horizon manipulation tasks with human-level dexterity is challenging, as such tasks require both physical dexterity and seamless sequencing of manipulation skills while robustly handling environment variations. While imitation learning offers a promising approach, acquiring comprehensive datasets is resource-intensive. In this work, we propose a learning framework and system LodeStar that automatically decomposes task demonstrations into semantically meaningful skills using off-the-shelf foundation models, and generates diverse synthetic demonstration datasets from a few human demos through reinforcement learning. These sim-augmented datasets enable robust skill training, with a Skill Routing Transformer (SRT) policy effectively chaining the learned skills together to execute complex long-horizon manipulation tasks. Experimental evaluations on three challenging real-world long-horizon dexterous manipulation tasks demonstrate that our approach significantly improves task performance and robustness compared to previous baselines. Videos are available at lodestar-robot.github.io.
Forward citations
Cited by 1 Pith paper
-
Scaling Cross-Embodiment World Models for Dexterous Manipulation
A single particle-based world model trained on many simulated robot hands and real human hands can plan dexterous manipulation on robot hands it never trained on.
Reference graph
Works this paper leans on
- [1]
-
[2]
E. Jang, A. Irpan, M. Khansari, D. Kappler, F. Ebert, C. Lynch, S. Levine, and C. Finn. Bc-z: Zero-shot task generalization with robotic imitation learning. In Conference on Robot Learning, pages 991–1002. PMLR, 2022
2022
-
[3]
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. arXiv preprint arXiv:2108.03298, 2021
arXiv 2021
-
[4]
Ravichandar, A
H. Ravichandar, A. S. Polydoros, S. Chernova, and A. Billard. Recent advances in robot learning from demonstration. Annual review of control, robotics, and autonomous systems, 3 (1):297–330, 2020
2020
-
[5]
O’Neill, A
A. O’Neill, A. Rehman, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6892–6903. IEEE, 2024
2024
-
[6]
Mandlekar, S
A. Mandlekar, S. Nasiriany, B. Wen, I. Akinola, Y . Narang, L. Fan, Y . Zhu, and D. Fox. Mimicgen: A data generation system for scalable robot learning using human demonstrations. In 7th Annual Conference on Robot Learning, 2023
2023
- [7]
-
[8]
C. Garrett, A. Mandlekar, B. Wen, and D. Fox. Skillmimicgen: Automated demonstration generation for efficient skill learning and deployment. arXiv preprint arXiv:2410.18907 , 2024
arXiv 2024
Show all 146 references
-
[9]
Z. Xue, S. Deng, Z. Chen, Y . Wang, Z. Yuan, and H. Xu. Demogen: Synthetic demonstration generation for data-efficient visuomotor policy learning. arXiv preprint arXiv:2502.16932, 2025
2025 arXiv
-
[10]
Y . Xu, W. Wan, J. Zhang, H. Liu, Z. Shan, H. Shen, R. Wang, H. Geng, Y . Weng, J. Chen, et al. Unidexgrasp: Universal robotic dexterous grasping via learning diverse proposal gener- ation and goal-conditioned policy. InProceedings of the IEEE/CVF Conference on Computer Vision...
2023
-
[11]
T. G. W. Lum, M. Matak, V . Makoviychuk, A. Handa, A. Allshire, T. Hermans, N. D. Ratliff, and K. Van Wyk. Dextrah-g: Pixels-to-action dexterous arm-hand grasping with geometric fabrics. arXiv preprint arXiv:2407.02274, 2024
2024 arXiv
-
[12]
H.-S. Fang, H. Yan, Z. Tang, H. Fang, C. Wang, and C. Lu. Anydexgrasp: General dex- terous grasping for different hands with human-level learning efficiency. arXiv preprint arXiv:2502.16420, 2025
2025 arXiv
-
[13]
O. M. Andrychowicz, B. Baker, M. Chociej, R. Jozefowicz, B. McGrew, J. Pachocki, A. Petron, M. Plappert, G. Powell, A. Ray, et al. Learning dexterous in-hand manipulation. The International Journal of Robotics Research, 39(1):3–20, 2020
2020
-
[14]
H. Qi, A. Kumar, R. Calandra, Y . Ma, and J. Malik. In-hand object rotation via rapid motor adaptation. In Conference on Robot Learning, pages 1722–1732. PMLR, 2023. 10
2023
-
[15]
J. Wang, Y . Yuan, H. Che, H. Qi, Y . Ma, J. Malik, and X. Wang. Lessons from learning to spin” pens”. arXiv preprint arXiv:2407.18902, 2024
2024 arXiv
-
[16]
Konidaris and A
G. Konidaris and A. Barto. Skill discovery in continuous reinforcement learning domains using skill chaining. Advances in neural information processing systems, 22, 2009
2009
-
[17]
Konidaris, S
G. Konidaris, S. Kuindersma, R. Grupen, and A. Barto. Robot learning from demonstration by constructing skill trees. The International Journal of Robotics Research, 31(3):360–375, 2012
2012
-
[18]
Clegg, W
A. Clegg, W. Yu, J. Tan, C. K. Liu, and G. Turk. Learning to dress: Synthesizing human dressing motion via deep reinforcement learning. ACM Transactions on Graphics (TOG), 37 (6):1–10, 2018
2018
-
[19]
Bagaria and G
A. Bagaria and G. Konidaris. Option discovery using deep skill chaining. In International Conference on Learning Representations, 2019
2019
-
[20]
Y . Lee, J. Yang, and J. J. Lim. Learning to coordinate manipulation skills via skill behav- ior diversification. In International Conference on Learning Representations , 2020. URL https://openreview.net/forum?id=ryxB2lBtvH
2020
-
[21]
X. B. Peng, M. Chang, G. Zhang, P. Abbeel, and S. Levine. Mcp: Learning composable hierarchical control with multiplicative compositional policies. In H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alch´e-Buc, E. Fox, and R. Garnett, editors, Advances in Neural Infor- mation ...
2019
-
[22]
J. Gu, D. S. Chaplot, H. Su, and J. Malik. Multi-skill mobile manipulation for object rear- rangement. arXiv preprint arXiv:2209.02778, 2022
2022 arXiv
-
[23]
Y . Lee, J. J. Lim, A. Anandkumar, and Y . Zhu. Adversarial skill chaining for long-horizon robot manipulation via terminal state regularization. arXiv preprint arXiv:2111.07999, 2021
2021 arXiv
-
[24]
Z. Chen, Z. Ji, J. Huo, and Y . Gao. Scar: Refining skill chaining for long-horizon robotic manipulation via dual regularization. Advances in Neural Information Processing Systems , 37:111679–111714, 2024
2024
-
[25]
Y . Chen, C. Wang, L. Fei-Fei, and C. K. Liu. Sequential dexterity: Chaining dexterous policies for long-horizon manipulation. arXiv preprint arXiv:2309.00987, 2023
2023 arXiv
-
[26]
Karaev, I
N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. Cotracker: It is better to track together. In European Conference on Computer Vision, pages 18–35. Springer, 2024
2024
-
[27]
L. Tang, M. Jia, Q. Wang, C. P. Phoo, and B. Hariharan. Emergent correspondence from image diffusion. Advances in Neural Information Processing Systems, 36:1363–1389, 2023
2023
-
[28]
OpenAI o3: Introducing a new generation of reasoning models
OpenAI. OpenAI o3: Introducing a new generation of reasoning models. https://openai. com/index/introducing-o3-and-o4-mini/ , 2025
2025
-
[29]
J. K. Salisbury and J. J. Craig. Articulated hands: Force control and kinematic issues. The International journal of Robotics research, 1(1):4–17, 1982
1982
-
[30]
Mordatch, Z
I. Mordatch, Z. Popovi ´c, and E. Todorov. Contact-invariant optimization for hand manipula- tion. In Proceedings of the ACM SIGGRAPH/Eurographics symposium on computer anima- tion, pages 137–144, 2012
2012
-
[31]
Akkaya, M
I. Akkaya, M. Andrychowicz, M. Chociej, M. Litwin, B. McGrew, A. Petron, A. Paino, M. Plappert, G. Powell, R. Ribas, et al. Solving rubik’s cube with a robot hand. arXiv preprint arXiv:1910.07113, 2019. 11
1910 arXiv
-
[32]
Z. He, B. Ai, Y . Liu, W. Wan, H. I. Christensen, and H. Su. Learning dexterous deformable object manipulation through cross-embodiment dynamics learning. In 3rd RSS Workshop on Dexterous Manipulation: Learning and Control with Diverse Data, 2025
2025
-
[33]
R. Fearing. Implementing a force strategy for object re-orientation. In Proceedings. 1986 IEEE International Conference on Robotics and Automation, volume 3, pages 96–102. IEEE, 1986
1986
-
[34]
Han and J
L. Han and J. C. Trinkle. Dextrous manipulation by rolling and finger gaiting. InProceedings. 1998 IEEE International Conference on Robotics and Automation (Cat. No. 98CH36146) , volume 1, pages 730–735. IEEE, 1998
1998
-
[35]
D. Rus. In-hand dexterous manipulation of piecewise-smooth 3-d objects. The International Journal of Robotics Research, 18(4):355–381, 1999
1999
-
[36]
Bai and C
Y . Bai and C. K. Liu. Dexterous manipulation using both palm and fingers. In2014 IEEE In- ternational Conference on Robotics and Automation (ICRA), pages 1560–1565. IEEE, 2014
2014
-
[37]
W. Wan, H. Geng, Y . Liu, Z. Shan, Y . Yang, L. Yi, and H. Wang. Unidexgrasp++: Improving dexterous grasping policy learning via geometry-aware curriculum and iterative generalist- specialist learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision,...
2023
-
[38]
Handa, A
A. Handa, A. Allshire, V . Makoviychuk, A. Petrenko, R. Singh, J. Liu, D. Makoviichuk, K. Van Wyk, A. Zhurkevich, B. Sundaralingam, et al. Dextreme: Transfer of agile in-hand manipulation from simulation to reality. In 2023 IEEE International Conference on Robotics and Automat...
2023
-
[39]
Y . Liu, Y . Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 21013– 21022, 2022
2022
-
[40]
K. Shaw, Y . Li, J. Yang, M. K. Srirama, R. Liu, H. Xiong, R. Mendonca, and D. Pathak. Bimanual dexterity for complex tasks. arXiv preprint arXiv:2411.13677, 2024
2024 arXiv
-
[41]
Lin, Z.-H
T. Lin, Z.-H. Yin, H. Qi, P. Abbeel, and J. Malik. Twisting lids off with two hands. arXiv preprint arXiv:2403.02338, 2024
2024 arXiv
-
[42]
Y . Chen, C. Wang, Y . Yang, and C. K. Liu. Object-centric dexterous manipulation from human motion data. arXiv preprint arXiv:2411.04005, 2024
2024 arXiv
-
[43]
Z.-H. Yin, C. Wang, L. Pineda, F. Hogan, K. Bodduluri, A. Sharma, P. Lancaster, I. Prasad, M. Kalakrishnan, J. Malik, et al. Dexteritygen: Foundation controller for unprecedented dexterity. arXiv preprint arXiv:2502.04307, 2025
2025 arXiv
-
[44]
Bousmalis, A
K. Bousmalis, A. Irpan, P. Wohlhart, Y . Bai, M. Kelcey, M. Kalakrishnan, L. Downs, J. Ibarz, P. Pastor, K. Konolige, et al. Using simulation and domain adaptation to improve efficiency of deep robotic grasping. In 2018 IEEE international conference on robotics and automation ...
2018
-
[45]
Nasiriany, A
S. Nasiriany, A. Maddukuri, L. Zhang, A. Parikh, A. Lo, A. Joshi, A. Mandlekar, and Y . Zhu. Robocasa: Large-scale simulation of everyday tasks for generalist robots. arXiv preprint arXiv:2406.02523, 2024
2024 arXiv
-
[46]
J. Wang, Y . Qin, K. Kuang, Y . Korkmaz, A. Gurumoorthy, H. Su, and X. Wang. Cyberdemo: Augmenting simulated human demonstration for real-world dexterous manipulation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 17952–17963, 2024. 12
2024
-
[47]
L. Wang, J. Zhao, Y . Du, E. H. Adelson, and R. Tedrake. Poco: Policy composition from and for heterogeneous robot learning. arXiv preprint arXiv:2402.02511, 2024
2024 arXiv
-
[48]
Ankile, A
L. Ankile, A. Simeonov, I. Shenfeld, M. Torne, and P. Agrawal. From imitation to refinement– residual rl for precise assembly. arXiv preprint arXiv:2407.16677, 2024
2024 arXiv
-
[49]
Maddukuri, Z
A. Maddukuri, Z. Jiang, L. Y . Chen, S. Nasiriany, Y . Xie, Y . Fang, W. Huang, Z. Wang, Z. Xu, N. Chernyadev, et al. Sim-and-real co-training: A simple recipe for vision-based robotic manipulation. arXiv preprint arXiv:2503.24361, 2025
2025 arXiv
-
[50]
A. Wei, A. Agarwal, B. Chen, R. Bosworth, N. Pfaff, and R. Tedrake. Empirical analysis of sim-and-real cotraining of diffusion policies for planar pushing from pixels. arXiv preprint arXiv:2503.22634, 2025
2025 arXiv
-
[51]
H. Geng, F. Wang, S. Wei, Y . Li, B. Wang, B. An, C. T. Cheng, H. Lou, P. Li, Y .-J. Wang, et al. Roboverse: Towards a unified platform, dataset and benchmark for scalable and generalizable robot learning. arXiv preprint arXiv:2504.18904, 2025
2025 arXiv
-
[52]
H. Liu, W. Wan, X. Yu, M. Li, J. Zhang, B. Zhao, Z. Chen, Z. Wang, Z. Zhang, and H. Wang. Navid-4d: Unleashing spatial intelligence in egocentric rgb-d videos for vision-and-language navigation. 2025
2025
-
[53]
Henry, M
P. Henry, M. Krainin, E. Herbst, X. Ren, and D. Fox. Rgb-d mapping: Using kinect-style depth cameras for dense 3d modeling of indoor environments. The international journal of Robotics Research, 31(5):647–663, 2012
2012
-
[54]
Tancik, E
M. Tancik, E. Weber, E. Ng, R. Li, B. Yi, T. Wang, A. Kristoffersen, J. Austin, K. Salahi, A. Ahuja, et al. Nerfstudio: A modular framework for neural radiance field development. In ACM SIGGRAPH 2023 conference proceedings, pages 1–12, 2023
2023
-
[55]
Jiang, C.-C
Z. Jiang, C.-C. Hsu, and Y . Zhu. Ditto: Building digital twins of articulated objects from interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5616–5626, 2022
2022
-
[56]
Torne, A
M. Torne, A. Simeonov, Z. Li, A. Chan, T. Chen, A. Gupta, and P. Agrawal. Reconciling reality through simulation: A real-to-sim-to-real approach for robust manipulation. arXiv preprint arXiv:2403.03949, 2024
2024 arXiv
-
[57]
Patel, X
S. Patel, X. Yin, W. Huang, S. Garg, H. Nayyeri, L. Fei-Fei, S. Lazebnik, and Y . Li. A real-to- sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards. arXiv preprint arXiv:2502.08643, 2025
2025 arXiv
-
[58]
W. Ye, F. Liu, Z. Ding, Y . Gao, O. Rybkin, and P. Abbeel. Video2policy: Scaling up manip- ulation tasks in simulation through internet videos. arXiv preprint arXiv:2502.09886, 2025
2025 arXiv
-
[59]
H. Xia, E. Su, M. Memmel, A. Jain, R. Yu, N. Mbiziwo-Tiapo, A. Farhadi, A. Gupta, S. Wang, and W.-C. Ma. Drawer: Digital reconstruction and articulation with environment realism. arXiv preprint arXiv:2504.15278, 2025
2025 arXiv
-
[60]
Z. Chen, A. Walsman, M. Memmel, K. Mo, A. Fang, K. Vemuri, A. Wu, D. Fox, and A. Gupta. Urdformer: A pipeline for constructing articulated simulation environments from real-world images. arXiv preprint arXiv:2405.11656, 2024
2024 arXiv
-
[61]
Y . Wang, Z. Xian, F. Chen, T.-H. Wang, Y . Wang, K. Fragkiadaki, Z. Erickson, D. Held, and C. Gan. Robogen: Towards unleashing infinite data for automated robot learning via generative simulation. arXiv preprint arXiv:2311.01455, 2023
2023 arXiv
-
[62]
T. Dai, J. Wong, Y . Jiang, C. Wang, C. Gokmen, R. Zhang, J. Wu, and L. Fei-Fei. Automated creation of digital cousins for robust policy learning.arXiv preprint arXiv:2410.07408, 2024. 13
2024 arXiv
-
[63]
Mandi, Y
Z. Mandi, Y . Weng, D. Bauer, and S. Song. Real2code: Reconstruct articulated objects via code generation. arXiv preprint arXiv:2406.08474, 2024
2024 arXiv
-
[64]
Tobin, R
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017
2017
-
[65]
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-real transfer of robotic control with dynamics randomization. In 2018 IEEE international conference on robotics and automation (ICRA), pages 3803–3810. IEEE, 2018
2018
-
[66]
Chebotar, A
Y . Chebotar, A. Handa, V . Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox. Clos- ing the sim-to-real loop: Adapting simulation randomization with real world experience. In 2019 International Conference on Robotics and Automation (ICRA), pages 8973–8979. IEEE, 2019
2019
-
[67]
Mehta, M
B. Mehta, M. Diaz, F. Golemo, C. J. Pal, and L. Paull. Active domain randomization. In Conference on Robot Learning, pages 1162–1176. PMLR, 2020
2020
-
[68]
Kumar, Z
A. Kumar, Z. Fu, D. Pathak, and J. Malik. Rma: Rapid motor adaptation for legged robots. arXiv preprint arXiv:2107.04034, 2021
2021 arXiv
-
[69]
Evans, A
B. Evans, A. Thankaraj, and L. Pinto. Context is everything: Implicit identification for dy- namics adaptation. In 2022 International Conference on Robotics and Automation (ICRA) , pages 2642–2648. IEEE, 2022
2022
-
[70]
Muratore, T
F. Muratore, T. Gruner, F. Wiese, B. Belousov, M. Gienger, and J. Peters. Neural posterior domain randomization. In Conference on robot learning, pages 1532–1542. PMLR, 2022
2022
-
[71]
Huang, X
P. Huang, X. Zhang, Z. Cao, S. Liu, M. Xu, W. Ding, J. Francis, B. Chen, and D. Zhao. What went wrong? closing the sim-to-real gap via differentiable causal discovery. In Conference on Robot Learning, pages 734–760. PMLR, 2023
2023
-
[72]
A. Z. Ren, H. Dai, B. Burchfiel, and A. Majumdar. Adaptsim: Task-driven simulation adap- tation for sim-to-real transfer. arXiv preprint arXiv:2302.04903, 2023
2023 arXiv
-
[73]
Memmel, A
M. Memmel, A. Wagenmaker, C. Zhu, P. Yin, D. Fox, and A. Gupta. Asid: Active exploration for system identification in robotic manipulation. arXiv preprint arXiv:2404.12308, 2024
2024 arXiv
-
[74]
K. Rao, C. Harris, A. Irpan, S. Levine, J. Ibarz, and M. Khansari. Rl-cyclegan: Reinforcement learning aware simulation-to-real. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11157–11166, 2020
2020
-
[75]
D. Ho, K. Rao, Z. Xu, E. Jang, M. Khansari, and Y . Bai. Retinagan: An object-aware approach to sim-to-real transfer. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 10920–10926. IEEE, 2021
2021
-
[76]
R. S. Sutton, D. Precup, and S. Singh. Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning. Artificial intelligence , 112(1-2):181–211, 1999
1999
-
[77]
Konidaris
G. Konidaris. On the necessity of abstraction. Current opinion in behavioral sciences , 29: 1–7, 2019
2019
-
[78]
Lee, S.-H
Y . Lee, S.-H. Sun, S. Somasundaram, E. S. Hu, and J. J. Lim. Composing complex skills by learning transition policies. In International conference on learning representations, 2019
2019
-
[79]
T. D. Kulkarni, K. Narasimhan, A. Saeedi, and J. Tenenbaum. Hierarchical deep reinforce- ment learning: Integrating temporal abstraction and intrinsic motivation. Advances in neural information processing systems, 29, 2016. 14
2016
-
[80]
J. Oh, S. Singh, H. Lee, and P. Kohli. Zero-shot task generalization with multi-task deep reinforcement learning. In International Conference on Machine Learning, pages 2661–2670. PMLR, 2017
2017
-
[81]
Y . Zhu, J. Tremblay, S. Birchfield, and Y . Zhu. Hierarchical planning for long-horizon manip- ulation with geometric and symbolic scene graphs. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 6541–6548. Ieee, 2021
2021
-
[82]
Silver, A
T. Silver, A. Athalye, J. B. Tenenbaum, T. Lozano-P ´erez, and L. P. Kaelbling. Learning neuro-symbolic skills for bilevel planning. arXiv preprint arXiv:2206.10680, 2022
2022 arXiv
-
[83]
Nasiriany, H
S. Nasiriany, H. Liu, and Y . Zhu. Augmenting reinforcement learning with behavior prim- itives for diverse manipulation tasks. In 2022 International Conference on Robotics and Automation (ICRA), pages 7477–7484. IEEE, 2022
2022
-
[84]
Cheng and D
S. Cheng and D. Xu. League: Guided skill learning and abstraction for long-horizon manip- ulation. IEEE Robotics and Automation Letters, 8(10):6451–6458, 2023
2023
-
[85]
Y . Guo, B. Tang, I. Akinola, D. Fox, A. Gupta, and Y . Narang. Srsa: Skill retrieval and adaptation for robotic assembly tasks. arXiv preprint arXiv:2503.04538, 2025
2025 arXiv
-
[86]
Pastor, H
P. Pastor, H. Hoffmann, T. Asfour, and S. Schaal. Learning and generalization of motor skills by learning from demonstration. In 2009 IEEE international conference on robotics and automation, pages 763–768. IEEE, 2009
2009
-
[87]
Z. Su, O. Kroemer, G. E. Loeb, G. S. Sukhatme, and S. Schaal. Learning manipulation graphs from demonstrations using multimodal sensory signals. In 2018 IEEE international conference on robotics and automation (ICRA), pages 2758–2765. IEEE, 2018
2018
-
[88]
D. Xu, S. Nair, Y . Zhu, J. Gao, A. Garg, L. Fei-Fei, and S. Savarese. Neural task program- ming: Learning to generalize across hierarchical tasks. In 2018 IEEE international confer- ence on robotics and automation (ICRA), pages 3795–3802. IEEE, 2018
2018
-
[89]
T. Kipf, Y . Li, H. Dai, V . Zambaldi, A. Sanchez-Gonzalez, E. Grefenstette, P. Kohli, and P. Battaglia. Compile: Compositional imitation learning and execution. In International Conference on Machine Learning, pages 3418–3428. PMLR, 2019
2019
-
[90]
Y . Zhu, P. Stone, and Y . Zhu. Bottom-up skill discovery from unsegmented demonstrations for long-horizon robot manipulation. IEEE Robotics and Automation Letters , 7(2):4126–4133, 2022
2022
-
[91]
S. Li, Z. Huang, T. Chen, T. Du, H. Su, J. B. Tenenbaum, and C. Gan. Dexdeform: Dexterous deformable object manipulation with human demonstrations and differentiable physics.arXiv preprint arXiv:2304.03223, 2023
2023 arXiv
-
[92]
M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song. Xskill: Cross embodiment skill discovery. In Conference on robot learning, pages 3536–3555. PMLR, 2023
2023
-
[93]
W. Wan, Y . Zhu, R. Shah, and Y . Zhu. Lotus: Continual imitation learning for robot ma- nipulation through unsupervised skill discovery. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 537–544. IEEE, 2024
2024
-
[94]
Z. Wang, J. Hu, C. Chuck, S. Chen, R. Mart ´ın-Mart´ın, A. Zhang, S. Niekum, and P. Stone. Skild: Unsupervised skill discovery guided by factor interactions. arXiv preprint arXiv:2410.18416, 2024
2024 arXiv
-
[95]
Schmidhuber
J. Schmidhuber. Towards compositional learning with dynamic neural networks . Inst. f ¨ur Informatik, 1990. 15
1990
-
[96]
A. S. Vezhnevets, S. Osindero, T. Schaul, N. Heess, M. Jaderberg, D. Silver, and K. Kavukcuoglu. Feudal networks for hierarchical reinforcement learning. In International conference on machine learning, pages 3540–3549. PMLR, 2017
2017
-
[97]
Bacon, J
P.-L. Bacon, J. Harb, and D. Precup. The option-critic architecture. In Proceedings of the AAAI conference on artificial intelligence, volume 31, 2017
2017
-
[98]
Nachum, S
O. Nachum, S. S. Gu, H. Lee, and S. Levine. Data-efficient hierarchical reinforcement learn- ing. Advances in neural information processing systems, 31, 2018
2018
-
[99]
Eysenbach, A
B. Eysenbach, A. Gupta, J. Ibarz, and S. Levine. Diversity is all you need: Learning skills without a reward function. arXiv preprint arXiv:1802.06070, 2018
2018 arXiv
-
[100]
Hausman, J
K. Hausman, J. T. Springenberg, Z. Wang, N. Heess, and M. Riedmiller. Learning an embed- ding space for transferable robot skills. In International Conference on Learning Represen- tations, 2018
2018
-
[101]
U. A. Mishra, S. Xue, Y . Chen, and D. Xu. Generative skill chaining: Long-horizon skill planning with diffusion models. InConference on Robot Learning, pages 2905–2925. PMLR, 2023
2023
-
[102]
U. A. Mishra, Y . Chen, and D. Xu. Generative factor chaining: Coordinated manipulation with diffusion-based factor graph. In 8th Annual Conference on Robot Learning, 2024
2024
-
[103]
Haldar and L
S. Haldar and L. Pinto. Point policy: Unifying observations and actions with key points for robot manipulation. arXiv preprint arXiv:2502.20391, 2025
2025 arXiv
-
[104]
Sundaresan, H
P. Sundaresan, H. Hu, Q. Vuong, J. Bohg, and D. Sadigh. What’s the move? hybrid imitation learning via salient points. arXiv preprint arXiv:2412.05426, 2024
2024 arXiv
-
[105]
Zhang, Y
Z. Zhang, Y . Li, O. Bastani, A. Gupta, D. Jayaraman, Y . J. Ma, and L. Weihs. Universal visual decomposer: Long-horizon manipulation made easy. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 6973–6980. IEEE, 2024
2024
-
[106]
Huang, C
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation. arXiv preprint arXiv:2409.01652 , 2024
2024 arXiv
-
[107]
Huang, C
W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei. V oxposer: Composable 3d value maps for robotic manipulation with language models.arXiv preprint arXiv:2307.05973, 2023
2023 arXiv
-
[108]
AR Code. AR Code. https://ar-code.com/, 2022
2022
-
[109]
J. Bohg, K. Hausman, B. Sankaran, O. Brock, D. Kragic, S. Schaal, and G. S. Sukhatme. Interactive perception: Leveraging action in perception and perception in action.IEEE Trans- actions on Robotics, 33(6):1273–1291, 2017
2017
-
[110]
Pfaff, E
N. Pfaff, E. Fu, J. Binagia, P. Isola, and R. Tedrake. Scalable real2sim: Physics-aware asset generation via robotic pick-and-place setups. arXiv preprint arXiv:2503.00370, 2025
2025 arXiv
-
[111]
B. Wen, W. Yang, J. Kautz, and S. Birchfield. Foundationpose: Unified 6d pose estimation and tracking of novel objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17868–17879, 2024
2024
-
[112]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimiza- tion algorithms. arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[113]
A. M. Saxe, J. L. McClelland, and S. Ganguli. Exact solutions to the nonlinear dynamics of learning in deep linear neural networks. arXiv preprint arXiv:1312.6120, 2013. 16
2013 arXiv
-
[114]
X. Yuan, T. Mu, S. Tao, Y . Fang, M. Zhang, and H. Su. Policy decorator: Model-agnostic online refinement for large policy model. arXiv preprint arXiv:2412.13630, 2024
2024 arXiv
-
[115]
K. Shaw, A. Agarwal, and D. Pathak. Leap hand: Low-cost, efficient, and anthropomorphic hand for robot learning. Robotics: Science and Systems (RSS), 2023
2023
-
[116]
Makoviychuk, L
V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv preprint arXiv:2108.10470, 2021
2021 arXiv
-
[117]
R. Ding, Y . Qin, J. Zhu, C. Jia, S. Yang, R. Yang, X. Qi, and X. Wang. Bunny-visionpro: Real-time bimanual dexterous teleoperation for imitation learning. 2024. URL https:// arxiv.org/abs/2407.03162
2024 arXiv
-
[118]
H. Zhu, A. Gupta, A. Rajeswaran, S. Levine, and V . Kumar. Dexterous manipulation with deep reinforcement learning: Efficient, general, and low-cost. In 2019 International Confer- ence on Robotics and Automation (ICRA), pages 3651–3657. IEEE, 2019
2019
-
[119]
M. Ahn, H. Zhu, K. Hartikainen, H. Ponte, A. Gupta, S. Levine, and V . Kumar. Robel: Robotics benchmarks for learning with low-cost robots. In Conference on robot learning , pages 1300–1313. PMLR, 2020
2020
-
[120]
X. Li, T. Zhao, X. Zhu, J. Wang, T. Pang, and K. Fang. Planning-guided diffusion policy learn- ing for generalizable contact-rich bimanual manipulation. arXiv preprint arXiv:2412.02676, 2024
2024 arXiv
-
[121]
C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu. Dexcap: Scalable and portable mocap data collection system for dexterous manipulation. arXiv preprint arXiv:2403.07788, 2024
2024 arXiv
-
[122]
Y . Qin, W. Yang, B. Huang, K. Van Wyk, H. Su, X. Wang, Y .-W. Chao, and D. Fox. Anyteleop: A general vision-based dexterous robot arm-hand teleoperation system. In Robotics: Science and Systems, 2023
2023
-
[123]
K. Zakka. Mink: Python inverse kinematics based on MuJoCo, July 2024. URL https: //github.com/kevinzakka/mink
2024
-
[124]
https://github.com/xArm-Developer/xarm_ros2
GitHub - xArm-Developer/xarm ros2: ROS2 developer packages for robotic products from UFACTORY — github.com. https://github.com/xArm-Developer/xarm_ros2. [Ac- cessed 04-05-2025]
2025
-
[125]
N. Ravi, V . Gabeur, Y .-T. Hu, R. Hu, C. Ryali, T. Ma, H. Khedr, R. R ¨adle, C. Rolland, L. Gustafson, E. Mintun, J. Pan, K. V . Alwala, N. Carion, C.-Y . Wu, R. Girshick, P. Doll´ar, and C. Feichtenhofer. Sam 2: Segment anything in images and videos. arXiv preprint arXiv:240...
2024 arXiv
-
[126]
C. Wang, H. Fang, H.-S. Fang, and C. Lu. Rise: 3d perception makes real-world robot imitation simple and effective. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2870–2877. IEEE, 2024
2024
-
[127]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv preprint arXiv:2304.13705, 2023
2023 arXiv
-
[128]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Dif- fusion policy: Visuomotor policy learning via action diffusion. The International Journal of Robotics Research, page 02783649241273668, 2023
2023
-
[129]
Sundaralingam, S
B. Sundaralingam, S. K. S. Hari, A. Fishman, C. Garrett, K. Van Wyk, V . Blukis, A. Millane, H. Oleynikova, A. Handa, F. Ramos, et al. curobo: Parallelized collision-free minimum-jerk robot motion generation. arXiv preprint arXiv:2310.17274, 2023. 17
-
[130]
Haldar, Z
S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learn- ing. arXiv preprint arXiv:2406.07539, 2024
2024 arXiv
-
[131]
Karpathy
A. Karpathy. minGPT: A minimal PyTorch reimplementation of the OpenAI GPT. https: //github.com/karpathy/minGPT, 2021. 18 Supplementary Material A Real Robot Platform In this section, we provide details about our real robot platform, including the hardware setup, teleoperation ...
2021
-
[132]
grasp the pipette from the rack,
-
[133]
aspirate liquid from the reagent bottle,
-
[134]
dispense the liquid into the test tube,
-
[135]
put the pipette back onto the rack, and
-
[136]
dispose of the used tip into the container below. A.3.2 Plant Watering In the Plant Watering task, there are 1) a spray nozzle, 2) a spray bottle body, 3) a nozzle stand to hold the spray nozzle vertically, 4) a container to hold the bottle body, 5) an open cardboard box, and ...
-
[137]
grasp the spray nozzle from the nozzle stand,
-
[138]
insert it onto the spray bottle body,
-
[139]
securely twist it in place,
-
[140]
grasp the assembled spray bottle,
-
[141]
press the trigger in front of the plant while holding the bottle, and
-
[142]
A.3.3 Light Bulb Assembly In the Light Bulb Assembly task, the workspace consists of 1) a light bulb and 2) a bulb base
place the spray bottle into the cardboard box. A.3.3 Light Bulb Assembly In the Light Bulb Assembly task, the workspace consists of 1) a light bulb and 2) a bulb base. To accomplish the task, the robot needs to
-
[143]
grasp the light bulb,
-
[144]
reorient the bulb in-hand for insertion,
-
[145]
insert the light bulb into the bulb base,
-
[146]
precisely screw the bulb into the base until illumination. B Simulation Training Details In this section, we introduce more details about training in simulation, including the used simulator, task designs, skill segmentation, real-to-sim transfer, residual reinforcement learni...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.