REVIEW 4 major objections 5 minor 1 cited by
Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning
T0 review · 4 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read SPI-Active identifies legged-robot physical parameters via massive parallel sampling and uses Fisher-information-optimal command sequences to collect informative real-world data, improving sim-to-real transfer on quadrupeds and a humanoid.
desk verdict New derivative-free SysID with FIM-driven active exploration worth taking seriously, but the headline task gains conflate inertia identification with the actuator-model change, and Appendix A.5 admits the identified parameters are not physical — major revision, not rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
First, the robot walks and jumps using existing control policies while recording what actually happens. The method then uses a computer search (CMA-ES) to find mass, inertia, and motor-saturation numbers that make thousands of short simulated replays of those recorded motions match the real recordings as closely as possible. Because the search only needs to run simulations and compare states, it works even though the physics simulator cannot be differentiated. Second, to make the data more useful, the method designs its own exciting motions: it adjusts the velocity and gait commands of a pre-trained multi-behavioral policy so that the resulting trajectories are as informative as possible, using a Fisher Information criterion. These optimized motions are played on the real robot, and the identification is repeated on the new data.
The paper reports that policies trained with the corrected simulation jump and turn more accurately, and track velocities better, on a Unitree Go2 and a G1 humanoid. The largest gains come on high-torque tasks like jumping, where modeling motor torque saturation matters.
Extended reading notes
Core claim
The paper's central claim is that SPI-Active 'robustly identifies key physical parameters through massive parallel sampling, minimizing state prediction errors between simulated and real-world trajectories,' and that this 'enables precise sim-to-real transfer of learned policies to the real world, outperforming baselines by 42-63% in various locomotion tasks.' If true, the method needs no differentiable simulator or torque sensing.
Load-bearing premise
The identified parameters are treated as valid for downstream tasks even though they are not physically accurate: the identified Go2 mass is 9.36 kg versus an expected 11.62 kg with the 4.7 kg payload, and Appendix A.5 admits Ixx and Iyy are 'unexpectedly large' and suggest 'overfitting or parameter coupling.' The load-bearing assumption is that these effective parameters, fitted on exploration trajectories, generalize across the jump and tracking task distribution, i.e., that the chosen model class (base inertia plus per-joint tanh motor gains) absorbs all significant sim-to-real discrepancies.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes SPI-Active, a two-stage sampling-based system identification method for legged robots. Stage 1 uses CMA-ES and GPU-parallel rollouts to fit base-link mass-inertia parameters and per-joint tanh motor gains by minimizing multi-step prediction error against real-world trajectories. Stage 2 actively designs data-collection commands by optimizing the trace of the inverse Fisher Information matrix of a command-conditioned pre-trained policy. The authors evaluate on Unitree Go2 and G1, reporting improved open-loop prediction and downstream task performance over vanilla, heavy DR, and gradient-based baselines, and claim 42-63% improvement. The identified parameters are then used to train RL policies deployed zero-shot.
Significance. If the identification pipeline, rather than the actuator-model change, is the cause of the reported gains, the paper would provide a practical, sensor-light alternative to differentiable simulators for sim2real in legged robots. Strengths include real-world experiments across multiple tasks and platforms, open-loop evaluation without global feedback, and released code. However, the experimental design does not separate the contribution of the identified inertial parameters from the addition of the tanh actuator model, and the paper's own appendix admits the identified inertial parameters are not physically accurate. These gaps currently prevent the paper from supporting its central causal claim.
major comments (4)
- [5.5, Table 2; Table 1(b)] The main comparison does not isolate the effect of system identification or active exploration from the introduction of the tanh motor model. In Table 1(b), Vanilla uses the nominal URDF and no tanh actuator model (Appendix A.4.3), while SPI and SPI-Active use both identified inertial parameters and per-joint tanh gains. Table 2 shows that the per-joint tanh motor model alone improves Forward Jump error by 45% relative to the ideal-torque baseline, nearly matching the 52% improvement attributed to SPI-Active in Table 1(b). No condition evaluates identified inertia with ideal torque or nominal inertia with the tanh model. Therefore the reported sim-to-real gains cannot be attributed to the proposed identification and active-exploration pipeline; the actuator model is a substantial confound.
- [A.5, Table 9] Appendix A.5 explicitly states that the identified Ixx and Iyy are 'unexpectedly large and not fully supported by the payload geometry,' and suggests 'possible overfitting or parameter coupling.' Additionally, the identified mass of 9.363 kg is about 2.4 kg below the expected 11.62 kg (default 6.921 kg plus the 4.7 kg payload). This admission contradicts the white-box 'physical parameter identification' claim and creates a generalization concern: if the parameters are an overfit to the exploration trajectory distribution, they may not transfer to the jump and tracking tasks used for evaluation. The paper needs either to constrain the identification to physically plausible values or to reframe the claim as fitting an effective surrogate model.
- [Abstract and Table 1(b)] The abstract's claim of outperforming baselines by '42-63%' is not supported by the full table. Attitude Tracking improves by only about 27% (0.73 vs 1.00), and the humanoid column Jhvt has no SPI-Active entry. The humanoid result is also explicitly limited to Stage 1 in the Limitations section. The quantitative claim should be revised to reflect all reported results, and the missing humanoid entry should be either provided or clearly explained.
- [4, Eq. (7)] The FIM-based active exploration uses the current estimate theta_hat1 as a surrogate for theta_star and a Gaussian process-noise assumption to obtain Eq. (7). Since the estimator is not unbiased (as indicated by the systematic mass error in A.5), the Cramer-Rao lower bound motivation is not directly applicable. The paper does not quantify how accurately the surrogate FIM predicts the actual estimation error reduction; without this, the active-exploration advantage over random commands in Fig. 6(c) is an isolated result on one task. A more direct evaluation, such as comparing the resulting parameter covariance or prediction error on a held-out trajectory set, would strengthen the claim.
minor comments (5)
- [Figure 4 caption] The caption contains a stray 'uni00ad' artifact; it should read 'real-world'.
- [Eq. (4)] The notation for x^r and x in the prediction cost is not fully defined; please clarify that x^r denotes the real state and x denotes the simulated state.
- [Table 3] The normalization procedure for the cost coefficients is described only in prose; adding the exact normalization constants or formulas would improve reproducibility.
- [5.3] The sentence reporting SPI improvements (19.6%, 39.9%, 35.9%) omits the Attitude Tracking result and should be tied explicitly to the entries in Table 1(b).
- [Table 1(b) caption] The missing SPI-Active entry for Jhvt should be explained in the caption or the text, since active exploration was not run for the humanoid.
Circularity Check
No significant circularity: the identification, active exploration, and downstream evaluation form an open-loop empirical pipeline; only minor non-load-bearing self-citations are present.
full rationale
This paper's derivation chain is not circular. The system identification objective (Eq. 4) minimizes multi-step simulation prediction error against real trajectories; the active exploration objective (Eqs. 6-7) optimizes command sequences using the current estimate theta1 as a surrogate, then collects new real data and refits with the same SPI procedure. The final downstream policies are trained in the identified simulation and deployed on hardware; the reported task metrics are measured real-world outcomes, not quantities derived from the fitted parameters by construction. The validation data for Table 1(a) is a separately teleoperated 60 s dataset, distinct from the fitting clips. No equation in the paper reduces a claimed prediction to a fitted input. The self-citations ([44], [63], [71], etc.) appear in related-work and implementation contexts (asymmetric actor-critic, sampling-based optimization inspiration, training framework) and are not used to justify the core claim that the identified parameters improve sim-to-real transfer. Appendix A.5's admission that Ixx/Iyy are unexpectedly large and may reflect overfitting or parameter coupling is a physical-plausibility caveat, not a circularity. The experimental confound noted in Section 5.5, namely that the Vanilla baseline omits the tanh actuator model while SPI/SPI-Active include it, so part of the gain may come from the actuator model rather than inertial identification, is a validity and ablation concern but does not make the derivation circular, because the claimed improvements are empirical hardware results rather than consequences of the estimator by construction. Score 1 reflects only the presence of minor, non-load-bearing self-citations.
Assumptions & free parameters
free parameters (8)
- Mass of base link m (Go2) =
9.363 kg
- Center of mass r (Go2) =
(0.004, -0.005, -0.020) m
- Inertia diag(I) (Go2) =
(0.391, 0.515, 0.396) kg m^2
- Tanh motor gain kappa_hip =
22.553
- Tanh motor gain kappa_thigh =
24.969
- Tanh motor gain kappa_calf =
23.523
- SPI cost weights (Table 3) =
position 4.0, velocity 2.0, quaternion 2.0, ang vel 0.5, joint pos 3.0, joint vel 0.1, torque 0.01, mass reg 0.01, CoM…
- DR ranges (Table 8) =
nominal vs heavy ranges (e.g., mass U(0.8,1.2)x, CoM U(-0.1,0.1))
assumptions (6)
- domain assumption The chosen model class (base-link inertia plus tanh motor gain per joint) is sufficient to represent the dominant sim-to-real discrepancies for the tested tasks.
- domain assumption Isaac Gym's contact and actuator simulation, after parameter identification, accurately predicts real-world state trajectories over H-step horizons.
- domain assumption The FIM approximation in Eq. (7) (Gaussian process noise, first-order sensitivity, surrogate theta_1) is a valid objective for experiment design on the real system.
- standard math CMA-ES converges to a good optimizer of J(theta) and of the FIM objective in the sampled parameter ranges.
- standard math The log-Cholesky parameterization of the pseudo-inertia matrix covers all physically valid mass-inertia sets.
- domain assumption The pre-trained multi-behavioral policy can safely and accurately track optimized command sequences on hardware.
Cite this review
Pith. "Pith review of Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning." pith.science (2026). https://pith.science/paper/QYUUUBXR
@misc{pith2026250514266,
author = {Pith},
title = {Pith review of: Sampling-Based System Identification with Active Exploration for Legged Robot Sim2Real Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/QYUUUBXR}},
note = {Machine review of arXiv:2505.14266}
}
read the original abstract
Sim-to-real discrepancies hinder learning-based policies from achieving high-precision tasks in the real world. While Domain Randomization (DR) is commonly used to bridge this gap, it often relies on heuristics and can lead to overly conservative policies with degrading performance when not properly tuned. System Identification (Sys-ID) offers a targeted approach, but standard techniques rely on differentiable dynamics and/or direct torque measurement, assumptions that rarely hold for contact-rich legged systems. To this end, we present SPI-Active (Sampling-based Parameter Identification with Active Exploration), a two-stage framework that estimates physical parameters of legged robots to minimize the sim-to-real gap. SPI-Active robustly identifies key physical parameters through massive parallel sampling, minimizing state prediction errors between simulated and real-world trajectories. To further improve the informativeness of collected data, we introduce an active exploration strategy that maximizes the Fisher Information of the collected real-world trajectories via optimizing the input commands of an exploration policy. This targeted exploration leads to accurate identification and better generalization across diverse tasks. Experiments demonstrate that SPI-Active enables precise sim-to-real transfer of learned policies to the real world, outperforming baselines by 42-63% in various locomotion tasks.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Hold My Beer: Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control
A slow-fast two-agent reinforcement learning architecture with separate upper- and lower-body policies reduces end-effector shaking during humanoid locomotion.
Reference graph
Works this paper leans on
-
[1]
J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning quadrupedal locomo- tion over challenging terrain. Science Robotics, 5(47):eabc5986, Oct. 2020. ISSN 2470-9476. doi:10.1126/scirobotics.abc5986. URL https://www.science.org/doi/10.1126/ scirobotics.abc5986
-
[2]
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning ro- bust perceptive locomotion for quadrupedal robots in the wild. Science Robotics , 7(62): eabk2822, Jan. 2022. ISSN 2470-9476. doi:10.1126/scirobotics.abk2822. URL https://www.science.org/doi/10.1126/scirobotics.abk2822
-
[3]
Y . Yang, G. Shi, X. Meng, W. Yu, T. Zhang, J. Tan, and B. Boots. CAJun: Continuous Adaptive Jumping using a Learned Centroidal Controller, 2023. URL https://arxiv.org/abs/ 2306.09557
arXiv 2023
- [4]
-
[5]
T. He, C. Zhang, W. Xiao, G. He, C. Liu, and G. Shi. Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion, 2024. URL https://arxiv.org/abs/2401.175 83
work page 2024
-
[6]
D. Hoeller, N. Rudin, D. Sako, and M. Hutter. ANYmal parkour: Learning agile navigation for quadrupedal robots. Science Robotics, 9(88):eadi7566, Mar. 2024. ISSN 2470-9476. doi: 10.1126/scirobotics.adi7566. URL https://www.science.org/doi/10.1126/sc irobotics.adi7566
- [7]
-
[8]
F. Muratore, F. Ramos, G. Turk, W. Yu, M. Gienger, and J. Peters. Robot Learning from Randomized Simulations: A Review, 2021. URL https://arxiv.org/abs/2111.0 0956
work page 2021
Show all 80 references
-
[9]
Loquercio, E
A. Loquercio, E. Kaufmann, R. Ranftl, A. Dosovitskiy, V . Koltun, and D. Scaramuzza. Deep Drone Racing: From Simulation to Reality With Domain Randomization. IEEE Transactions on Robotics, 36(1):1–14, Feb. 2020. ISSN 1552-3098, 1941-0468. doi:10.1109/TRO.2019.2 942989. URL htt...
2020
-
[10]
X. B. Peng, M. Andrychowicz, W. Zaremba, and P. Abbeel. Sim-to-Real Transfer of Robotic Control with Dynamics Randomization. In 2018 IEEE International Conference on Robotics and Automation (ICRA), pages 3803–3810, Brisbane, QLD, May 2018. IEEE. ISBN 9781538630815. doi:10.1109...
2018
-
[11]
Tobin, R
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 23–30, Vancouver, BC, Sep...
2017
-
[12]
C. An, C. Atkeson, and J. Hollerbach. Estimation of inertial parameters of rigid body links of manipulators. In 1985 24th IEEE Conference on Decision and Control , pages 990–995, Fort Lauderdale, FL, USA, Dec. 1985. IEEE. doi:10.1109/CDC.1985.268648. URL http: //ieeexplore.iee...
1985
-
[13]
Mayeda, K
H. Mayeda, K. Yoshida, and K. Osuka. Base parameters of manipulator dynamic models. IEEE Transactions on Robotics and Automation, 6(3):312–321, June 1990. ISSN 2374-958X. doi:10.1109/70.56663. URL https://ieeexplore.ieee.org/document/56663/. 10
1990 doi
-
[14]
T. Lee, J. Kwon, P. M. Wensing, and F. C. Park. Robot Model Identification and Learning: A Modern Perspective. Annual Review of Control, Robotics, and Autonomous Systems , 7(1): 311–334, July 2024. ISSN 2573-5144. doi:10.1146/annurev-control-061523-102310. URL https://www.annu...
2024 doi
-
[15]
M. P. Deisenroth and C. E. Rasmussen. Pilco: a model-based and data-efficient approach to policy search. In Proceedings of the 28th International Conference on International Confer- ence on Machine Learning , ICML’11, page 465–472, Madison, WI, USA, 2011. Omnipress. ISBN 9781450306195
2011
-
[16]
G. Shi, X. Shi, M. O’Connell, R. Yu, K. Azizzadenesheli, A. Anandkumar, Y . Yue, and S.- J. Chung. Neural lander: Stable drone landing control using learned dynamics. In 2019 International Conference on Robotics and Automation (ICRA) , pages 9784–9790, 2019. doi: 10.1109/ICRA....
2019
-
[17]
Hwangbo, J
J. Hwangbo, J. Lee, A. Dosovitskiy, D. Bellicoso, V . Tsounis, V . Koltun, and M. Hut- ter. Learning agile and dynamic motor skills for legged robots. Science Robotics , 4(26): eaau5872, Jan. 2019. ISSN 2470-9476. doi:10.1126/scirobotics.aau5872. URL https: //www.science.org/d...
2019 doi
-
[18]
Kumar, Z
A. Kumar, Z. Fu, D. Pathak, and J. Malik. RMA: Rapid Motor Adaptation for Legged Robots,
-
[19]
P. Wu, W. Xie, J. Cao, H. Lai, and W. Zhang. Loopsr: Looping sim-and-real for lifelong policy adaptation of legged robots, 2024. URL https://arxiv.org/abs/2409.17992
2024
-
[20]
T. B. Sch ¨on, A. Wills, and B. Ninness. System identification of nonlinear state-space models. Automatica, 47(1):39–49, Jan. 2011. ISSN 00051098. doi:10.1016/j.automatica.2010.10.013. URL https://linkinghub.elsevier.com/retrieve/pii/S000510981000 4279
2011 doi
-
[21]
Pfaff, E
N. Pfaff, E. Fu, J. Binagia, P. Isola, and R. Tedrake. Scalable Real2Sim: Physics-Aware Asset Generation Via Robotic Pick-and-Place Setups, 2025. URL https://arxiv.org/abs/ 2503.00370
2025 arXiv
-
[22]
Gautier and W
M. Gautier and W. Khalil. A direct determination of minimum inertial parameters of robots. In Proceedings. 1988 IEEE International Conference on Robotics and Automation, pages 1682– 1687, Philadelphia, PA, USA, 1988. IEEE Comput. Soc. Press. ISBN 9780818608520. doi: 10.1109/RO...
1988
-
[23]
J. Tan, T. Zhang, E. Coumans, A. Iscen, Y . Bai, D. Hafner, S. Bohez, and V . Vanhoucke. Sim-to-Real: Learning Agile Locomotion For Quadruped Robots. In Robotics: Science and Systems XIV. Robotics: Science and Systems Foundation, June 2018. ISBN 9780992374747. doi:10.15607/RSS...
2018 doi
-
[24]
Grandia, E
R. Grandia, E. Knoop, M. Hopkins, G. Wiedebach, J. Bishop, S. Pickles, D. M ¨uller, and M. B ¨acher. Design and Control of a Bipedal Robotic Character. In Robotics: Science and Systems XX. Robotics: Science and Systems Foundation, July 2024. ISBN 9798990284807. doi:10.15607/RS...
2024 doi
-
[25]
Memmel, A
M. Memmel, A. Wagenmaker, C. Zhu, D. Fox, and A. Gupta. ASID: Active Exploration for System Identification in Robotic Manipulation. Jan. 2024. URL https://openreview .net/forum?id=pdhMe50hZI. 11
2024
-
[26]
T. Lee, B. D. Lee, and F. C. Park. Optimal excitation trajectories for mechanical systems identification. Automatica, 131:109773, Sept. 2021. ISSN 00051098. doi:10.1016/j.automati ca.2021.109773. URL https://linkinghub.elsevier.com/retrieve/pii/S 0005109821002934
2021
-
[27]
W. Yu, V . C. V . Kumar, G. Turk, and C. K. Liu. Sim-to-real transfer for biped locomotion. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 3503–3510, 2019. URL https://ieeexplore.ieee.org/document/8967890
2019
-
[28]
Mozifian, J
M. Mozifian, J. C. G. Higuera, D. Meger, and G. Dudek. Learning domain randomization distributions for training robust locomotion policies, 2019. URL https://arxiv.org/ abs/1906.00410
2019 arXiv
-
[29]
Siekmann, Y
J. Siekmann, Y . Godse, A. Fern, and J. Hurst. Sim-to-real learning of all common bipedal gaits via periodic reward composition. In 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021. URL https://ieeexplore.ieee.org/document/9 561248
2021
-
[30]
T. He, Z. Luo, W. Xiao, C. Zhang, K. Kitani, C. Liu, and G. Shi. Learning human-to-humanoid real-time whole-body teleoperation. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8944–8951. IEEE, 2024
2024
-
[31]
Sadeghi and S
F. Sadeghi and S. Levine. CAD2RL: Real Single-Image Flight without a Single Real Image,
-
[32]
Mehta, M
B. Mehta, M. Diaz, F. Golemo, C. J. Pal, and L. Paull. Active domain randomization, 2019. URL https://arxiv.org/abs/1904.04762
2019 arXiv
-
[33]
Pinto, J
L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta. Robust adversarial reinforcement learning. In International Conference on Machine Learning, pages 2817–2826. PMLR, 2017
2017
-
[34]
Chebotar, A
Y . Chebotar, A. Handa, V . Makoviychuk, M. Macklin, J. Issac, N. Ratliff, and D. Fox. Closing the Sim-to-Real Loop: Adapting Simulation Randomization with Real World Experience. In 2019 International Conference on Robotics and Automation (ICRA) , pages 8973–8979, May
2019
-
[35]
Kumar, Z
A. Kumar, Z. Li, J. Zeng, D. Pathak, K. Sreenath, and J. Malik. Adapting Rapid Motor Adap- tation for Bipedal Robots, 2022. URL https://arxiv.org/abs/2205.15299
2022 arXiv
-
[36]
H. Qi, A. Kumar, R. Calandra, Y . Ma, and J. Malik. In-Hand Object Rotation via Rapid Motor Adaptation, 2022. URL https://arxiv.org/abs/2210.04887
2022 arXiv
-
[37]
G. B. Margolis, X. Fu, Y . Ji, and P. Agrawal. Learning to See Physical Properties with Active Sensing Motor Policies, 2023. URL https://arxiv.org/abs/2311.01405
2023 arXiv
-
[38]
T. He, Z. Luo, X. He, W. Xiao, C. Zhang, W. Zhang, K. M. Kitani, C. Liu, and G. Shi. Om- nih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning. In 8th Annual Conference on Robot Learning, 2024
2024
-
[39]
A. Bose, S. S. Du, and M. Fazel. Offline multi-task transfer rl with representational penaliza- tion, 2024. URL https://arxiv.org/abs/2402.12570
2024 arXiv
-
[40]
W. Yu, J. Tan, C. K. Liu, and G. Turk. Preparing for the unknown: Learning a universal policy with online system identification, 2017. URLhttps://arxiv.org/abs/1702.02453
2017 arXiv
-
[41]
Huang, R
K. Huang, R. Rana, A. Spitzer, G. Shi, and B. Boots. Datt: Deep adaptive trajectory tracking for quadrotor control. In Conference on Robot Learning, pages 326–340. PMLR, 2023. 12
2023
-
[42]
Smith, J
L. Smith, J. C. Kew, X. B. Peng, S. Ha, J. Tan, and S. Levine. Legged robots that keep on learning: Fine-tuning locomotion policies in the real world, 2021. URL https://arxiv. org/abs/2110.05457
2021 arXiv
-
[43]
N. Fey, G. B. Margolis, M. Peticco, and P. Agrawal. Bridging the sim-to-real gap for athletic loco-manipulation, 2025. URL https://arxiv.org/abs/2502.10894
2025 arXiv
-
[44]
T. He, J. Gao, W. Xiao, Y . Zhang, Z. Wang, J. Wang, Z. Luo, G. He, N. Sobanbab, C. Pan, Z. Yi, G. Qu, K. Kitani, J. Hodgins, L. J. Fan, Y . Zhu, C. Liu, and G. Shi. Asap: Aligning simulation and real-world physics for learning agile humanoid whole-body skills, 2025. URL https...
2025 arXiv
-
[45]
L. Ljung. System Identification. In J. J. Benedetto, A. Proch ´azka, J. Uhl´ıˇr, P. W. J. Rayner, and N. G. Kingsbury, editors, Signal Analysis and Prediction, pages 163–173. Birkh¨auser Boston, Boston, MA, 1998. ISBN 9781461272731 9781461217688. doi:10.1007/978-1-4612-1768-8
1998 doi
-
[46]
H. Natke. System identification. Automatica, 28(5):1069–1071, Sept. 1992. ISSN 00051098. doi:10.1016/0005-1098(92)90167-E. URL https://linkinghub.elsevier.com/ retrieve/pii/000510989290167E
1992
-
[47]
O’Connell, G
M. O’Connell, G. Shi, X. Shi, K. Azizzadenesheli, A. Anandkumar, Y . Yue, and S.-J. Chung. Neural-fly enables rapid learning for agile flight in strong winds. Science Robotics, 7(66): eabm6597, 2022
2022
-
[48]
Gautier, P.-O
M. Gautier, P.-O. Vandanjon, and A. Janot. Dynamic identification of a 6 dof robot without joint position data. 2011 IEEE International Conference on Robotics and Automation , pages 234–239, 2011. URL https://api.semanticscholar.org/CorpusID:167553 17
2011
-
[49]
URL http://link.springer.com/10.1007/978-1-4612-1768-8_11
-
[50]
Y . Han, J. Wu, C. Liu, and Z. Xiong. An iterative approach for accurate dynamic model identification of industrial robots. IEEE Transactions on Robotics , 36(5):1577–1594, 2020. doi:10.1109/TRO.2020.2990368
2020
-
[51]
Masuda and K
S. Masuda and K. Takahashi. Sim-to-Real Transfer of Compliant Bipedal Locomotion on Torque Sensor-Less Gear-Driven Humanoid, 2022. URL https://arxiv.org/abs/22 04.03897
2022
-
[52]
Zhang, D
B. Zhang, D. Haugk, and R. Vasudevan. System Identification For Constrained Robots, Aug
-
[53]
Janot, P.-O
A. Janot, P.-O. Vandanjon, and M. Gautier. A generic instrumental variable approach for industrial robot identification. IEEE Transactions on Control Systems Technology, 22(1):132– 145, 2014. doi:10.1109/TCST.2013.2246163
2014
-
[54]
Bombois, M
X. Bombois, M. Gevers, R. Hildebrand, and G. Solari. Optimal experiment design for open and closed-loop system identification. Communications in Information and Systems , 11(3): 197–224, Mar. 2011. ISSN 2163-4548. doi:10.4310/CIS.2011.v11.n3.a1. URL https: //link.intlpress.com...
2011
-
[55]
Gerencser and H
L. Gerencser and H. Hjalmarsson. Adaptive input design in system identification. In Pro- ceedings of the 44th IEEE Conference on Decision and Control , pages 4988–4993, Seville, Spain, 2005. IEEE. ISBN 9780780395671. doi:10.1109/CDC.2005.1582952. URL http://ieeexplore.ieee.org...
2005
-
[56]
Sathyanarayan and I
H. Sathyanarayan and I. Abraham. Exciting Contact Modes in Differentiable Simulations for Robot Learning, 2024. URL https://arxiv.org/abs/2411.10935
2024 arXiv
-
[57]
Liang, S
J. Liang, S. Saxena, and O. Kroemer. Learning Active Task-Oriented Exploration Policies for Bridging the Sim-to-Real Gap, 2020. URL https://arxiv.org/abs/2006.01952
2020 arXiv
-
[58]
Gevers, A
M. Gevers, A. S. Bazanella, X. Bombois, and L. Miskovic. Identification and the Information Matrix: How to Get Just Sufficiently Rich? IEEE Transactions on Automatic Control, 54(12): 2828–2840, Dec. 2009. ISSN 0018-9286, 1558-2523. doi:10.1109/TAC.2009.2034199. URL http://ieee...
2009
-
[59]
Leboutet, J
Q. Leboutet, J. Roux, A. Janot, J. R. Guadarrama-Olvera, and G. Cheng. Inertial Parameter Identification in Robotics: A Survey. Applied Sciences, 11(9):4303, May 2021. ISSN 2076-
2021
-
[60]
C. D. Sousa and R. Cortesao. Inertia Tensor Properties in Robot Dynamics Identification: A Linear Matrix Inequality Approach. IEEE/ASME Transactions on Mechatronics, 24(1):406– 411, Feb. 2019. ISSN 1083-4435, 1941-014X. doi:10.1109/TMECH.2019.2891177. URL https://ieeexplore.ie...
2019
-
[61]
Rucker and P
C. Rucker and P. M. Wensing. Smooth parameterization of rigid-body inertia. IEEE Robotics and Automation Letters, 7(2):2771–2778, 2022. doi:10.1109/LRA.2022.3144517
2022
-
[62]
Grandia, E
R. Grandia, E. Knoop, M. A. Hopkins, G. Wiedebach, J. Bishop, S. Pickles, D. M ¨uller, and M. B¨acher. Design and control of a bipedal robotic character.arXiv preprint arXiv:2501.05204, 2025
2025 arXiv
-
[63]
Wagenmaker, G
A. Wagenmaker, G. Shi, and K. G. Jamieson. Optimal exploration for model-based rl in non- linear systems. Advances in Neural Information Processing Systems, 36:15406–15455, 2023
2023
-
[64]
N. Hansen. The CMA Evolution Strategy: A Tutorial, Mar. 2023. URL http://arxiv. org/abs/1604.00772. arXiv:1604.00772
2023 arXiv
-
[65]
Makoviychuk, L
V . Makoviychuk, L. Wawrzyniak, Y . Guo, M. Lu, K. Storey, M. Macklin, D. Hoeller, N. Rudin, A. Allshire, A. Handa, and G. State. Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning, Aug. 2021. URL http://arxiv.org/abs/2108.104
2021
-
[66]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal Policy Opti- mization Algorithms, Aug. 2017. URL http://arxiv.org/abs/1707.06347 . arXiv:1707.06347
2017 arXiv
-
[67]
Le Lidec, I
Q. Le Lidec, I. Kalevatykh, I. Laptev, C. Schmid, and J. Carpentier. Differentiable simulation for physical system identification. IEEE Robotics and Automation Letters , 6(2):3413–3420,
-
[68]
Todorov, T
E. Todorov, T. Erez, and Y . Tassa. Mujoco: A physics engine for model-based control. In 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems , pages 5026–
2012
-
[69]
C. Pan, Z. Yi, G. Shi, and G. Qu. Model-based diffusion for trajectory optimization. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL https://openreview.net/forum?id=BJndYScO6o
2024
-
[70]
Akiba, S
T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama. Optuna: A next-generation hyperpa- rameter optimization framework, 2019. URL https://arxiv.org/abs/1907.10902
2019 arXiv
-
[71]
C. L. Lab. Humanoidverse: A multi-simulator framework for humanoid robot sim-to-real learning. https://github.com/LeCAR-Lab/HumanoidVerse, 2025. 14
2025
-
[72]
X. B. Peng, P. Abbeel, S. Levine, and M. van de Panne. Deepmimic: example-guided deep reinforcement learning of physics-based character skills.ACM Trans. Graph., 37(4), July 2018. ISSN 0730-0301. doi:10.1145/3197517.3201311. URL https://doi.org/10.1145/ 3197517.3201311. 15 A A...
2018
-
[74]
doi:10.1109/LRA.2021.3062323
2021
-
[77]
G. B. Margolis and P. Agrawal. Walk these ways: Tuning robot control for generalization with multiplicity of behavior. Conference on Robot Learning, 2022
2022
-
[2016]
URL https://arxiv.org/abs/1611.04201
-
[2019]
URL https://ieeexplore.ieee.org/do cument/8793789/
doi:10.1109/ICRA.2019.8793789. URL https://ieeexplore.ieee.org/do cument/8793789/. ISSN: 2577-087X
2019
-
[2021]
URL https://arxiv.org/abs/2107.04034
- [2024]
-
[3417]
URL https://www.mdpi.com/2076-3417/11/9 /4303
doi:10.3390/app11094303. URL https://www.mdpi.com/2076-3417/11/9 /4303
-
[5033]
doi:10.1109/IROS.2012.6386109
IEEE, 2012. doi:10.1109/IROS.2012.6386109
2012
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.