REVIEW 3 major objections 4 minor 1 cited by
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
T0 review · 3 major / 4 minor · reviewed 2026-08-01 · deepseek-v4-flash
Pith's one-line read Embodied AI trustworthiness is a system property, not a model score
desk verdict A serious, honest position piece that gives the field a useful four-layer vocabulary; the T0–T5 ladder is a scaffold awaiting thresholds, so its comparative-evaluation claim is still a promise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the bounded trustworthiness claim, expressed as the tuple ⟨system version, task, embodiment, operating domain, authority, evidence⟩, together with the four-layer framework that supports it. The four layers—model, system, evidence, deployment—are functional responsibilities rather than software modules, and the paper stresses cross-layer failure propagation through the semantic–physical gap, the action–consequence gap, and cross-layer non-compositionality. The key quantitative instrument is the safe-success rate, SSR = N(task completed ∧ no unacceptable violation)/N(evaluated trials), which separates four outcome classes: safe success, safe failure, unsafe success, and u
What would settle it
A system that meets all T4 assessment dimensions but whose claim lapses after a small unmonitored change—for example, a camera recalibration—without any deployment-layer detection would falsify the claim that a bounded claim's validity is maintained by the four layers. Concretely, audit a deployed robot across such a change and check whether its monitored assumptions trigger revalidation before harm occurs.
Extended reading notes
Core claim
The central claim is that trustworthiness is an end-to-end property of a deployed system, not an attribute of an individual model, component, or benchmark score. The paper defines trustworthy embodied intelligence as sustained safe success: reliable completion of intended tasks while physical, semantic, procedural, and operational risks remain within acceptable bounds. It identifies the model layer, which proposes actions with calibrated uncertainty and safety preferences; the system layer, which realizes authorized actions dependably through sensing, computing, control, hardware safeguards, fault containment, and fallback; the evidence layer, which substantiates bounded claims through evalu
Load-bearing premise
The framework assumes that 'unacceptable violation' and 'acceptable residual risk' can be specified and measured for each application; without that, the safe-success rate and T-level assessments cannot be applied.
Editorial extensions
If this is right
- Evaluation of embodied systems should report safe success, unsafe success, safe failure, and unsafe failure, rather than a single completion rate.
- A highly capable model can still be at T1 or T2 if system safeguards, evidence, or deployment governance are missing; capability alone does not raise a TEI level.
- Deployment should define an explicit operational boundary and use runtime admission, boundary monitoring, intervention, and change control to preserve the validity of the trustworthiness claim.
- The T0–T5 hierarchy offers a common structure for comparative evaluation, bounded deployment, and future standardization, while remaining non-normative and requiring domain-specific profiles.
Reading between the lines
- Editorial: If the safe-success-rate categories were widely adopted, benchmark suites would need to include scenario-conditioned risk reporting and severity weighting, so that a minor safety-filter activation and a high-energy collision are not counted equally.
- Editorial: The T-level framework implies a staged-approval model for regulators: a low T-level would authorize only tightly constrained deployments, while T4–T5 would be needed for open-ended operation.
- Editorial: The paper leaves open the 'assurance-preserving reconfiguration' problem — deciding which tool, payload, or controller changes are minor versus claim-invalidating. An automated impact-analysis tool for revalidation triggers would be a concrete next step.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position/survey paper argues that trustworthiness for embodied intelligence cannot be established by any single model, component, or benchmark score. It defines trustworthy embodied intelligence as “sustained safe success,” organized around four interdependent layers—model, system, evidence, and deployment—and proposes a non-normative T0–T5 hierarchy for grading the strength of bounded deployment claims. The paper reviews a broad literature spanning embodied AI, robotics, control, dependable computing, fault tolerance, and autonomous driving, and it connects these fields through the four-layer framework and the proposed TEI levels.
Significance. If accepted, the framework could provide a useful common vocabulary and architectural reference for safety and assurance in embodied AI, and it appropriately stresses bounded deployment claims, assurance cases, and lifecycle governance. The paper is careful to describe the hierarchy as analytical rather than a certification scheme, and it repeatedly acknowledges that thresholds and acceptable-residual-risk criteria are domain-specific. Its strengths include a broad cross-layer synthesis, explicit treatment of failure propagation, and alignment with existing standards and assurance concepts. However, the central evaluative claim—that the T0–T5 hierarchy supports comparative evaluation and bounded deployment—is not yet operationalized. The only formal evaluative object, the safe-success rate in Eq. (1), depends on an uninterpreted primitive “unacceptable violation,” and the T-level assessment in Section 8.1 depends on undefined thresholds for “adequate support.” The paper is therefore best read as a template or agenda for future standardization rather than a complete comparative instrument.
major comments (3)
- [§8.1–8.5 and Eq. (1)] The T-level assignment function is underdetermined. Section 8.1 says a TEI level is the least adequately supported of five dimensions, but “adequately supported” is not defined, and Section 8.5 defers quantitative thresholds and acceptable residual risk to domain-specific profiles. Concretely, the same system with the same evidence could be T2 under one permissible profile (e.g., one severe violation per 10^3 trials is acceptable) and T4 under another (one per 10^6 trials), because the least-supported dimension changes. Since the paper supplies no constraint on choosing these cutoffs, the claimed functions of comparative evaluation and bounded deployment (Section 1.2, contribution 4) are not well-defined. The authors acknowledge this deferral, but acknowledgement does not remove the tension; the paper should either provide a working example of a profile and its induced T-level, or explic
- [§6.2, Eq. (1)] The safe-success rate is introduced as the primary joint outcome, yet it is an unweighted trial fraction. The same equation treats a high-energy collision and a minor safety-filter activation as equivalent failures, while Section 6.2 subsequently states that these should not be weighted equally. The paper says SSR should be interpreted together with component outcomes, but then SSR is not itself the “primary joint outcome” in any decision-relevant sense. A severity- or exposure-weighted statistic, or a vector of scenario-conditioned rates with the violation predicate made explicit, would be more consistent with the stated safety goals. As written, Eq. (1) is a useful definitional starting point but not a measurable metric without a specification of the unacceptable-violation predicate.
- [§3.2 and §8.1] The four-layer decomposition and the five assessment dimensions are asserted as jointly necessary, but no derivation or empirical evidence is provided for their completeness. This is load-bearing for the paper’s central claim that “no single layer can establish end-to-end trustworthiness” and that a TEI level cannot exceed the least adequately supported dimension. I am not asking for a formal completeness proof in a survey, but the paper should clarify whether these are normative proposals or intended descriptive claims about existing systems. If the latter, at least one illustrative application of the framework to a concrete system would help show that the dimensions can be jointly assessed and that the hierarchy yields stable classifications.
minor comments (4)
- [§1] “Weusetrustworthy embodied intelligenceto” is missing spaces; please correct “We use trustworthy embodied intelligence to”.
- [§6.2, Eq. (1)] The notation N(task completed ∧ no unacceptable violation) / N(evaluated trials) uses the same symbol N for both the numerator and denominator; consider N_safe_success / N_total for clarity.
- [§8.4, Table 4] The table aligns AgiBot G1–G5, SAE L0–L5, and TEI T0–T5 row by row. The text cautions that the alignment is illustrative only, but the visual presentation still invites cross-hierarchy equivalence inferences. Consider adding an explicit sentence that the rows are not intended to imply comparable levels of capability, autonomy, or trustworthiness.
- [Appendix D and [146]] Reference [146] is first-party company material, and Appendix D correctly notes that internal and industry materials should be labeled as first-party sources. In the main text, however, the discussion around Section 8.4 and the roadmap (Section 9.2) cites company material without that label at the point of use. Adding an explicit first-party marker at the citation site would be more transparent.
Circularity Check
No significant circularity: TEI is a proposed definition; T-level grading is admittedly underdetermined, and self-citations are disclosed and non-load-bearing.
full rationale
This is a position/survey paper, not a predictive derivation. Its central claim — 'Trustworthiness is therefore an end-to-end property of a deployed system, not an attribute of an individual model, component, or benchmark score' — is introduced as a definition of 'sustained safe success', and the four-layer framework and T0–T5 hierarchy are organizing proposals rather than results fitted to data. Eq. 1 defines the safe-success rate as N(task completed ∧ no unacceptable violation)/N(evaluated trials); 'no unacceptable violation' is a declared primitive, and the paper explicitly defers operational thresholds to 'domain-specific standards' (Section 8.5: 'Domain-specific standards must define the quantitative thresholds, acceptable residual risk, required evidence, and independent assessment needed to substantiate each level'). That deferral makes the hierarchy underdetermined as a comparative instrument (a correctness/scope limitation, acknowledged in Sections 8.5 and 10), but it is not a circular reduction: the level definitions do not presuppose the levels they are supposed to grade. The relation to AgiBot/SAE taxonomies is declared 'illustrative only' (Section 8.4), so this is not renaming. Several supporting artifacts (SafeDojo [58], RoboDojo [98], RM-Bench [118], RobotWin [94,95], UniVTac [121]) come from the same labs, and [146] is first-party company material; however, these are used as illustrative examples/benchmarks, and the paper explicitly labels [146] as 'Company material; source supplied in Chinese' and instructs in Appendix D that 'Internal and industry materials should remain explicitly identified as first-party sources and should not be treated as independent validation.' No fitted input is renamed as a prediction, and no uniqueness/ansatz is imported via self-citation. The central definitional claim is therefore self-contained, and the self-citation pattern is disclosed and non-load-bearing.
Assumptions & free parameters
assumptions (4)
- domain assumption Task capability and safety are jointly necessary and not substitutable.
- ad hoc to paper Assumptions and failures propagate across layers, so no single layer can establish end-to-end trustworthiness.
- ad hoc to paper The four-layer decomposition (model, system, evidence, deployment) is complete for embodying trustworthiness mechanisms.
- ad hoc to paper An ordinal T0–T5 hierarchy can meaningfully grade bounded trustworthiness claims across five jointly necessary dimensions.
Cite this review
Pith. "Pith review of Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels." pith.science (2026). https://pith.science/paper/X6EBHBAP
@misc{pith2026260726121,
author = {Pith},
title = {Pith review of: Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels},
year = {2026},
howpublished = {\url{https://pith.science/paper/X6EBHBAP}},
note = {Machine review of arXiv:2607.26121}
}
read the original abstract
Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, traceability, and structured assurance arguments. The deployment layer maintains claim validity through runtime monitoring, authority management, intervention, incident response, and controlled updates. Because assumptions and failures propagate across these layers, neither model capability, isolated safeguards, nor benchmark performance alone can establish end-to-end trustworthiness. Drawing on embodied AI, robotics, control, dependable computing, distributed systems, and autonomous driving, we further propose a non-normative hierarchy of trustworthiness levels. This hierarchy grades the strength of bounded deployment claims across task capability, safety, system assurance, operational governance, and supporting evidence, providing a basis for bounded deployment, comparative evaluation, research prioritization, and future standardization.
Forward citations
Cited by 1 Pith paper
-
XPolicyLab: A Unified Standard and Open Ecosystem for Robot Policy Evaluation and Deployment
XPolicyLab is a unified open ecosystem whose adapter contract and dependency-isolated serving reduce robot policy-environment integration from pairwise O(NM) work to O(N+M), cutting a representative integration from o...
Reference graph
Works this paper leans on
-
[146]
Xspark AI: A New Paradigm for Trustworthy Physical Intelligence, July 2026
Xspark AI. Xspark AI: A New Paradigm for Trustworthy Physical Intelligence, July 2026. Company material; source supplied in Chinese
2026
-
[1]
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakr- ishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022
arXiv 2022
-
[2]
Palm-e: An embodied multimodal language model.arXiv preprint arXiv:2303.03378, 2023
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model.arXiv preprint arXiv:2303.03378, 2023
arXiv 2023
-
[3]
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Brianna Zitkovich, Tianhe Yu, Sichun Xu, Peng Xu, Ted Xiao, Fei Xia, Jialin Wu, Paul Wohlhart, Stefan Welker, Ayzaan Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. InConference on Robot Learning, pages 2165–2183. PMLR, 2023
2023
-
[4]
Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
arXiv 2024
-
[5]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.𝜋0: A vision-language-action flow model for general robot control.arXiv preprint arXiv:2410.24164, 2024
arXiv 2024
-
[6]
Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023
Danijar Hafner, Jurgis Pasukonis, Jimmy Ba, and Timothy Lillicrap. Mastering diverse domains through world models.arXiv preprint arXiv:2301.04104, 2023
arXiv 2023
-
[7]
Tactile robotics: An outlook.IEEE Transactions on Robotics, 2025
Shan Luo, Nathan F Lepora, Wenzhen Yuan, Kaspar Althoefer, Gordon Cheng, and Ravinder Dahiya. Tactile robotics: An outlook.IEEE Transactions on Robotics, 2025
2025
Show all 153 references
-
[8]
Expressive whole-body control for humanoid robots.arXiv preprint arXiv:2402.16796, 2024
Xuxin Cheng, Yandong Ji, Junming Chen, Ruihan Yang, Ge Yang, and Xiaolong Wang. Expressive whole-body control for humanoid robots.arXiv preprint arXiv:2402.16796, 2024
2024 arXiv
-
[9]
A comprehensive survey on physical risk control in the era of foundation model-enabled robotics.arXiv preprint arXiv:2505.12583, 2025
Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura, Gang Yan, Shu Morikuni, Ryosuke Takanami, Alfredo Solano, Tatsuya Matsushima, Akiko Murakami, et al. A comprehensive survey on physical risk control in the era of foundation model-enabled robotics.arXiv preprint a...
2025 arXiv
-
[10]
Safety in embodied ai: A survey of risks, attacks, and defenses.arXiv preprint arXiv:2605.02900, 2026
Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia, Yixu Wang, Xin Wang, Ye Sun, Yunhan Zhao, Ming Wen, Jiayu Li, et al. Safety in embodied ai: A survey of risks, attacks, and defenses.arXiv preprint arXiv:2605.02900, 2026
2026 arXiv
-
[11]
What breaks embodied ai security: Llm vulnerabilities, cps flaws, or something else?High-Confidence Computing, page 100403, 2026
Boyang Ma, Hechuan Guo, Peizhuo Lv, Minghui Xu, Xuelong Dai, YeChao Zhang, Yijun Yang, and Yue Zhang. What breaks embodied ai security: Llm vulnerabilities, cps flaws, or something else?High-Confidence Computing, page 100403, 2026
2026
-
[12]
Sim-to-real transfer of robotic control with dynamics randomization.arXiv preprint arXiv:1710.06537, 2017
Xue Bin Peng, Marcin Andrychowicz, Wojciech Zaremba, and Pieter Abbeel. Sim-to-real transfer of robotic control with dynamics randomization.arXiv preprint arXiv:1710.06537, 2017
2017 arXiv
-
[13]
Domain randomization for transferring deep neural networks from simulation to the real world
Josh Tobin, Rachel Fong, Alex Ray, Jonas Schneider, Wojciech Zaremba, and Pieter Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30...
2017
-
[14]
Safe embodied ai for long-horizon tasks: A cross-layer analysis of robotic manipulation.arXiv preprint arXiv:2606.05660, 2026
Dabin Kim, Daemin Park, Sangyub Lee, Jinsik Kim, Yeongtak Oh, Jongho Shin, and Sungroh Yoon. Safe embodied ai for long-horizon tasks: A cross-layer analysis of robotic manipulation.arXiv preprint arXiv:2606.05660, 2026
2026 arXiv
-
[15]
Basic concepts and taxonomy of dependable and secure computing.IEEE transactions on dependable and secure computing, 1(1):11–33, 2004
Algirdas Avizienis, J-C Laprie, Brian Randell, and Carl Landwehr. Basic concepts and taxonomy of dependable and secure computing.IEEE transactions on dependable and secure computing, 1(1):11–33, 2004
2004
-
[16]
Control barrier functions: Theory and applications
Aaron D Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada. Control barrier functions: Theory and applications. In2019 18th European control conference (ECC), pages 3420–
-
[17]
Harnessing embodied agents: Runtime governance for policy-constrained execution.arXiv preprint arXiv:2604.07833, 2026
Xue Qin, Simin Luan, John See, Zeyd Boukhers, Cong Yang, and Zhijun Li. Harnessing embodied agents: Runtime governance for policy-constrained execution.arXiv preprint arXiv:2604.07833, 2026
2026 arXiv
-
[18]
Md Muzakkir Quamar and Ali Nasir. Review on fault diagnosis and fault-tolerant control scheme for robotic manipulators: Recent advances in ai, machine learning, and digital twin.arXiv preprint arXiv:2402.02980, 2024
2024 arXiv
-
[19]
Robot collisions: A survey on detection, isolation, and identification.IEEE Transactions on Robotics, 33(6):1292–1312, 2017
Sami Haddadin, Alessandro De Luca, and Alin Albu-Schäffer. Robot collisions: A survey on detection, isolation, and identification.IEEE Transactions on Robotics, 33(6):1292–1312, 2017
2017
-
[20]
Physical human-robot interaction: A critical review of safety constraints.arXiv preprint arXiv:2601.19462, 2026
Riccardo Zanella, Federico Califano, and Stefano Stramigioli. Physical human-robot interaction: A critical review of safety constraints.arXiv preprint arXiv:2601.19462, 2026
2026
-
[21]
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. On calibration of modern neural networks. In International conference on machine learning, pages 1321–1330. PMLR, 2017
2017
-
[22]
A baseline for detecting misclassified and out-of-distribution examples in neural networks.arXiv preprint arXiv:1610.02136, 2016
Dan Hendrycks and Kevin Gimpel. A baseline for detecting misclassified and out-of-distribution examples in neural networks.arXiv preprint arXiv:1610.02136, 2016
2016 arXiv
-
[23]
Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691, 2022
Michael Ahn, Anthony Brohan, Noah Brown, Yevgen Chebotar, Omar Cortes, Byron David, Chelsea Finn, Chuyuan Fu, Keerthana Gopalakrishnan, Karol Hausman, et al. Do as i can, not as i say: Grounding language in robotic affordances.arXiv preprint arXiv:2204.01691, 2022. 30
2022 arXiv
-
[24]
Concrete problems in ai safety.arXiv preprint arXiv:1606.06565, 2016
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané. Concrete problems in ai safety.arXiv preprint arXiv:1606.06565, 2016
2016 arXiv
-
[25]
Zhiwei Gao, Carlo Cecati, and Steven X Ding. A survey of fault diagnosis and fault-tolerant techniques—part i: Fault diagnosis with model-based and signal-based approaches.IEEE transactions on industrial electronics, 62(6): 3757–3767, 2015
2015
-
[26]
A survey of algorithms for black-box safety validation of cyber-physical systems.Journal of Artificial Intelligence Research, 2021
Anthony Corso, Robert J Moss, Mark Koren, Ritchie Lee, and Mykel J Kochenderfer. A survey of algorithms for black-box safety validation of cyber-physical systems.Journal of Artificial Intelligence Research, 2021
2021
-
[27]
Survey on scenario-based safety assessment of automated vehicles.IEEE access, 8:87456–87477, 2020
Stefan Riedmaier, Thomas Ponn, Dieter Ludwig, Bernhard Schick, and Frank Diermeyer. Survey on scenario-based safety assessment of automated vehicles.IEEE access, 8:87456–87477, 2020
2020
-
[28]
Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?Transportation research part A: policy and practice, 94:182–193, 2016
Nidhi Kalra and Susan M Paddock. Driving to safety: How many miles of driving would it take to demonstrate autonomous vehicle reliability?Transportation research part A: policy and practice, 94:182–193, 2016
2016
-
[29]
Meaningful human control over autonomous systems: A philosophical account.Frontiers in Robotics and AI, 5, 2018
Filippo Santoni de Sio and Jeroen Van den Hoven. Meaningful human control over autonomous systems: A philosophical account.Frontiers in Robotics and AI, 5, 2018
2018
-
[30]
Using simplicity to control complexity.IEEE Software, 18(4):20, 2001
Lui Sha. Using simplicity to control complexity.IEEE Software, 18(4):20, 2001
2001
-
[31]
Leveson.Engineering a safer world: systems thinking applied to safety
Nancy G. Leveson.Engineering a safer world: systems thinking applied to safety. Engineering systems. MIT press, Cambridge (Mass.), 2011. ISBN 9780262016629 9781628703399
2011
-
[32]
ISO 21448:2022: Road vehicles — Safety of the intended function- ality
International Organization for Standardization. ISO 21448:2022: Road vehicles — Safety of the intended function- ality. International Standard, 2022
2022
-
[33]
Assuring safety-critical machine learning enabled systems: Challenges and promise
Alwyn E Goodloe. Assuring safety-critical machine learning enabled systems: Challenges and promise. In2022 IEEE International Symposium on Software Reliability Engineering Workshops (ISSREW), pages 326–332. IEEE, 2022
2022
-
[34]
Physically grounded vision-language models for robotic manipulation
Jensen Gao, Bidipta Sarkar, Fei Xia, Ted Xiao, Jiajun Wu, Brian Ichter, Anirudha Majumdar, and Dorsa Sadigh. Physically grounded vision-language models for robotic manipulation. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 12462–12469. IEEE, 2024
2024
-
[35]
Verifiably following complex robot instructions with foundation models
Benedict Quartey, Eric Rosen, Stefanie Tellex, and George Konidaris. Verifiably following complex robot instructions with foundation models. In2025 IEEE International Conference on Robotics and Automation (ICRA), pages 1–8. IEEE, 2025
2025
-
[36]
3d-vitac: Learning fine-grained manipulation with visuo-tactile sensing.arXiv preprint arXiv:2410.24091, 2024
Binghao Huang, Yixuan Wang, Xinyi Yang, Yiyue Luo, and Yunzhu Li. 3d-vitac: Learning fine-grained manipulation with visuo-tactile sensing.arXiv preprint arXiv:2410.24091, 2024
2024 arXiv
-
[37]
Adversarial patch.arXiv preprint arXiv:1712.09665, 2017
Tom B Brown, Dandelion Mané, Aurko Roy, Martín Abadi, and Justin Gilmer. Adversarial patch.arXiv preprint arXiv:1712.09665, 2017
2017 arXiv
-
[38]
Robust physical-world attacks on deep learning visual classification
Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. Robust physical-world attacks on deep learning visual classification. InProceedings of the IEEE conference on computer vision and pattern recogniti...
2018
-
[39]
Spatiotemporal attacks for embodied agents
Aishan Liu, Tairan Huang, Xianglong Liu, Yitao Xu, Yuqing Ma, Xinyun Chen, Stephen J Maybank, and Dacheng Tao. Spatiotemporal attacks for embodied agents. InEuropean Conference on Computer Vision, pages 122–138. Springer, 2020
2020
-
[40]
Badnets: Identifying vulnerabilities in the machine learning model supply chain.arXiv preprint arXiv:1708.06733, 2017
Tianyu Gu, Brendan Dolan-Gavitt, and Siddharth Garg. Badnets: Identifying vulnerabilities in the machine learning model supply chain.arXiv preprint arXiv:1708.06733, 2017
2017 arXiv
-
[41]
Deep evidential regression.Advances in neural information processing systems, 33:14927–14937, 2020
Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. Deep evidential regression.Advances in neural information processing systems, 33:14927–14937, 2020
2020
-
[42]
Grounding language with visual affordances over unstructured data.arXiv preprint arXiv:2210.01911, 2022
Oier Mees, Jessica Borja-Diaz, and Wolfram Burgard. Grounding language with visual affordances over unstructured data.arXiv preprint arXiv:2210.01911, 2022
2022 arXiv
-
[43]
Code as policies: Language model programs for embodied control
Jacky Liang, Wenlong Huang, Fei Xia, Peng Xu, Karol Hausman, Brian Ichter, Pete Florence, and Andy Zeng. Code as policies: Language model programs for embodied control. In2023 IEEE International conference on robotics and automation (ICRA), pages 9493–9500. IEEE, 2023
2023
-
[44]
Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, and Mario Fritz. Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection. InProceedings of the 16th ACM workshop on artificial intelligenc...
2023
-
[45]
Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043, 2023
2023 arXiv
-
[46]
Badrobot: Jailbreaking embodied llms in the physical world.arXiv preprint arXiv:2407.20242, 2024
Hangtao Zhang, Chenyu Zhu, Xianlong Wang, Ziqi Zhou, Changgan Yin, Minghui Li, Lulu Xue, Yichen Wang, Shengshan Hu, Aishan Liu, et al. Badrobot: Jailbreaking embodied llms in the physical world.arXiv preprint arXiv:2407.20242, 2024
2024 arXiv
-
[47]
Goal-oriented backdoor attack against vision-language-action models via physical objects.arXiv preprint arXiv:2510.09269, 2025
Zirun Zhou, Zhengyang Xiao, Haochuan Xu, Jing Sun, Di Wang, and Jingfeng Zhang. Goal-oriented backdoor attack against vision-language-action models via physical objects.arXiv preprint arXiv:2510.09269, 2025
2025
-
[48]
Integrated task and motion planning.Annual review of control, robotics, and autonomous systems, 4 (1):265–293, 2021
Caelan Reed Garrett, Rohan Chitnis, Rachel Holladay, Beomjoon Kim, Tom Silver, Leslie Pack Kaelbling, and Tomás Lozano-Pérez. Integrated task and motion planning.Annual review of control, robotics, and autonomous systems, 4 (1):265–293, 2021. 31
2021
-
[49]
Task and motion planning with large language models for object rearrangement
Yan Ding, Xiaohan Zhang, Chris Paxton, and Shiqi Zhang. Task and motion planning with large language models for object rearrangement. In2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2086–2092. IEEE, 2023
-
[50]
Constrained policy optimization
Joshua Achiam, David Held, Aviv Tamar, and Pieter Abbeel. Constrained policy optimization. InInternational conference on machine learning, pages 22–31. Pmlr, 2017
2017
-
[51]
Reward constrained policy optimization.arXiv preprint arXiv:1805.11074, 2018
Chen Tessler, Daniel J Mankowitz, and Shie Mannor. Reward constrained policy optimization.arXiv preprint arXiv:1805.11074, 2018
2018 arXiv
-
[52]
A comprehensive survey on safe reinforcement learning.Journal of Machine Learning Research, 16(1):1437–1480, 2015
Javier Garcıa and Fernando Fernández. A comprehensive survey on safe reinforcement learning.Journal of Machine Learning Research, 16(1):1437–1480, 2015
2015
-
[53]
Safety-oriented human-robot collaboration in construction through human preference alignment.Journal of Intelligent Construction, 3(3):1–15, 2025
Mao Tian and Zhengbo Zou. Safety-oriented human-robot collaboration in construction through human preference alignment.Journal of Intelligent Construction, 3(3):1–15, 2025
2025
-
[54]
Constrained diffusers for safe planning and control.Advances in Neural Information Processing Systems, 38:34965–34998, 2026
Jichen Zhang, Liqun Zhao, Antonis Papachristodoulou, and Jack Umenberger. Constrained diffusers for safe planning and control.Advances in Neural Information Processing Systems, 38:34965–34998, 2026
2026
-
[55]
Safe model-based reinforcement learning with an uncertainty-aware reachability certificate.IEEE Transactions on Automation Science and Engineering, 21(3):4129–4142, 2023
Dongjie Yu, Wenjun Zou, Yujie Yang, Haitong Ma, Shengbo Eben Li, Yuming Yin, Jianyu Chen, and Jingliang Duan. Safe model-based reinforcement learning with an uncertainty-aware reachability certificate.IEEE Transactions on Automation Science and Engineering, 21(3):4129–4142, 2023
2023
-
[56]
A predictive safety filter for learning-based control of constrained nonlinear dynamical systems.Automatica, 129:109597, 2021
Kim Peter Wabersich and Melanie N Zeilinger. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems.Automatica, 129:109597, 2021
2021
-
[57]
Unisim: A neural closed-loop sensor simulator
Ze Yang, Yun Chen, Jingkang Wang, Sivabalan Manivasagam, Wei-Chiu Ma, Anqi Joyce Yang, and Raquel Urtasun. Unisim: A neural closed-loop sensor simulator. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1389–1399, 2023
2023
-
[58]
Safedojo: Safe reinforcement learning for vla via interactive world model.arXiv preprint arXiv:2606.20698, 2026
Kai Tang, Peidong Jia, Zhong Chu, Jixian Wu, Rui Ma, Jiajun Cao, Fangyuan Zhao, Sixiang Chen, Yichen Guo, Xiaowei Chi, et al. Safedojo: Safe reinforcement learning for vla via interactive world model.arXiv preprint arXiv:2606.20698, 2026
2026 arXiv
-
[59]
Poisoning attacks against support vector machines.arXiv preprint arXiv:1206.6389, 2012
Battista Biggio, Blaine Nelson, and Pavel Laskov. Poisoning attacks against support vector machines.arXiv preprint arXiv:1206.6389, 2012
2012 arXiv
-
[60]
Robots that ask for help: Uncertainty alignment for large language model planners.arXiv preprint arXiv:2307.01928, 2023
Allen Z Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, et al. Robots that ask for help: Uncertainty alignment for large language model planners.arXiv preprint arXiv:2307.01928, 2023
2023 arXiv
-
[61]
Sensor-enabled safety systems for human–robot collaboration: A review.IEEE Sensors Journal, 25(1):65–88, 2024
Constantin Scholz, Hoang-Long Cao, Emil Imrith, Nima Roshandel, Hamed Firouzipouyaei, Aleksander Burkiewicz, Milan Amighi, Sebastien Menet, Dylan Warawout Sisavath, Antonio Paolillo, et al. Sensor-enabled safety systems for human–robot collaboration: A review.IEEE Sensors Jour...
2024
-
[62]
Robot operating system 2: Design, architecture, and uses in the wild.Science robotics, 7(66):eabm6074, 2022
Steven Macenski, Tully Foote, Brian Gerkey, Chris Lalancette, and William Woodall. Robot operating system 2: Design, architecture, and uses in the wild.Science robotics, 7(66):eabm6074, 2022
2022
-
[63]
A survey of real-time support, analysis, and advancements in ros 2.arXiv preprint arXiv:2601.10722, 2025
Daniel Casini, Jian-Jia Chen, Jing Li, Federico Reghenzani, and Harun Teper. A survey of real-time support, analysis, and advancements in ros 2.arXiv preprint arXiv:2601.10722, 2025
2025 arXiv
-
[64]
Priority inheritance protocols: An approach to real-time synchronization.IEEE Transactions on computers, 39(9):1175–1185, 1990
Lui Sha, Ragunathan Rajkumar, and John P Lehoczky. Priority inheritance protocols: An approach to real-time synchronization.IEEE Transactions on computers, 39(9):1175–1185, 1990
1990
-
[65]
End-to-end timing analysis and optimization of multi-executor ros 2 systems
Harun Teper, Tobias Betz, Mario Günzel, Dominic Ebner, Georg Von Der Brüggen, Johannes Betz, and Jian-Jia Chen. End-to-end timing analysis and optimization of multi-executor ros 2 systems. In2024 IEEE 30th Real-Time and Embedded Technology and Applications Symposium (RTAS), pa...
2024
-
[66]
Timing analysis and priority-driven enhancements of ros 2 multi-threaded executors
Hoora Sobhani, Hyunjong Choi, and Hyoseung Kim. Timing analysis and priority-driven enhancements of ros 2 multi-threaded executors. In2023 IEEE 29th Real-Time and Embedded Technology and Applications Symposium (RTAS), pages 106–118. IEEE, 2023
2023
-
[67]
Series elastic actuators
Gill A Pratt and Matthew M Williamson. Series elastic actuators. InProceedings 1995 IEEE/RSJ international conferenceonintelligentrobotsandsystems.Humanrobotinteractionandcooperativerobots, volume1, pages399–406. IEEE, 1995
1995
-
[68]
Impedance control: An approach to manipulation
Neville Hogan. Impedance control: An approach to manipulation. In1984 American control conference, pages 304–313. IEEE, 1984
1984
-
[69]
Coboskin: Soft robot skin with variable stiffness for safer human–robot collaboration.IEEE Transactions on Industrial Electronics, 68(4):3303–3314, 2020
Gaoyang Pang, Geng Yang, Wenzheng Heng, Zhiqiu Ye, Xiaoyan Huang, Hua-Yong Yang, and Zhibo Pang. Coboskin: Soft robot skin with variable stiffness for safer human–robot collaboration.IEEE Transactions on Industrial Electronics, 68(4):3303–3314, 2020
2020
-
[70]
A review on fault detection and diagnosis of industrial robots and multi-axis machines.Results in Engineering, 23:102397, 2024
Ameer H Sabry and Ungku Anisa Bte Ungku Amirulddin. A review on fault detection and diagnosis of industrial robots and multi-axis machines.Results in Engineering, 23:102397, 2024
2024
-
[71]
Review of fault-tolerant control systems used in robotic manipulators.Applied Sciences, 13(4):2675, 2023
Andrzej Milecki and Patryk Nowak. Review of fault-tolerant control systems used in robotic manipulators.Applied Sciences, 13(4):2675, 2023
2023
-
[72]
Fault-tolerant control of robot manipulators with sensory faults using unbiased active inference
Mohamed Baioumy, Corrado Pezzato, Riccardo Ferrari, Carlos Hernandez Corbato, and Nick Hawes. Fault-tolerant control of robot manipulators with sensory faults using unbiased active inference. In2021 European Control Conference (ECC), pages 1119–1125. IEEE, 2021. 32
2021
-
[73]
Safe, passive control for mechanical systems with application to physical human-robot interactions
Wenceslao Shaw Cortez, Christos K Verginis, and Dimos V Dimarogonas. Safe, passive control for mechanical systems with application to physical human-robot interactions. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 3836–3842. IEEE, 2021
2021
-
[74]
Cyber security of robots: A comprehensive survey
Alessio Botta, Sayna Rotbei, Stefania Zinno, and Giorgio Ventre. Cyber security of robots: A comprehensive survey. Intelligent Systems with Applications, 18:200237, 2023
2023
-
[75]
Zero trust architecture.NIST special publication, 800 (207):1–52, 2020
Scott Rose, Oliver Borchert, Stu Mitchell, and Sean Connelly. Zero trust architecture.NIST special publication, 800 (207):1–52, 2020
2020
-
[76]
Sros2: Usable cyber security tools for ros 2
Victor Mayoral-Vilches, Ruffin White, Gianluca Caiazza, and Mikael Arguedas. Sros2: Usable cyber security tools for ros 2. In2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 11253–11259. IEEE, 2022
2022
-
[77]
On the (in) security of secure ros2
Gelei Deng, Guowen Xu, Yuan Zhou, Tianwei Zhang, and Yang Liu. On the (in) security of secure ros2. InProceedings of the 2022 ACM SIGSAC Conference on Computer and Communications Security, pages 739–753, 2022
2022
-
[78]
The byzantine generals problem.ACM Trans
Leslie Lamport, Robert Shostak, and Marshall Pease. The byzantine generals problem.ACM Trans. Program. Lang. Syst., 4(3):382–401, July 1982. ISSN 0164-0925. doi: 10.1145/357172.357176. URLhttps://doi.org/10. 1145/357172.357176
1982
-
[79]
An investigation of byzantine threats in multi-robot systems
Gelei Deng, Yuan Zhou, Yuan Xu, Tianwei Zhang, and Yang Liu. An investigation of byzantine threats in multi-robot systems. InProceedings of the 24th international symposium on research in attacks, intrusions and defenses, pages 17–32, 2021
2021
-
[80]
An overview on multi-agent consensus under adversarial attacks.Annual Reviews in Control, 53:252–272, 2022
Hideaki Ishii, Yuan Wang, and Shuai Feng. An overview on multi-agent consensus under adversarial attacks.Annual Reviews in Control, 53:252–272, 2022
2022
-
[81]
Assuring the machine learning lifecycle: Desiderata, methods, and challenges, arxiv.arXiv preprint arXiv:1905.04223, 2019
Rob Ashmore, Radu Calinescu, and Colin Paterson. Assuring the machine learning lifecycle: Desiderata, methods, and challenges, arxiv.arXiv preprint arXiv:1905.04223, 2019
1905 arXiv
-
[82]
Toward verified artificial intelligence.Communications of the ACM, 65(7):46–55, 2022
Sanjit A Seshia, Dorsa Sadigh, and S Shankar Sastry. Toward verified artificial intelligence.Communications of the ACM, 65(7):46–55, 2022
2022
-
[83]
Guidance on the assurance of machine learning in autonomous systems (amlas).arXiv preprint arXiv:2102.01564, 2021
Richard Hawkins, Colin Paterson, Chiara Picardi, Yan Jia, Radu Calinescu, and Ibrahim Habli. Guidance on the assurance of machine learning in autonomous systems (amlas).arXiv preprint arXiv:2102.01564, 2021
2021 arXiv
-
[84]
Safety assurance of machine learning for perception functions
Simon Burton, Christian Hellert, Fabian Hüger, Michael Mock, and Andreas Rohatschek. Safety assurance of machine learning for perception functions. InDeep Neural Networks and Data for Automated Driving: Robustness, Uncertainty Quantification, and Insights Towards Safety, pages...
2022
-
[85]
Benchmarking safe exploration in deep reinforcement learning.arXiv preprint arXiv:1910.01708, 7(1):2, 2019
Alex Ray, Joshua Achiam, and Dario Amodei. Benchmarking safe exploration in deep reinforcement learning.arXiv preprint arXiv:1910.01708, 7(1):2, 2019
1910 arXiv
-
[86]
Safety gymnasium: A unified safe reinforcement learning benchmark.Advances in Neural Information Processing Systems, 36:18964–18993, 2023
Jiaming Ji, Borong Zhang, Jiayi Zhou, Xuehai Pan, Weidong Huang, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, and Yaodong Yang. Safety gymnasium: A unified safe reinforcement learning benchmark.Advances in Neural Information Processing Systems, 36:18964–18993, 2023
2023
-
[87]
Hamilton-jacobi reachability: A brief overview and recent advances
Somil Bansal, Mo Chen, Sylvia Herbert, and Claire J Tomlin. Hamilton-jacobi reachability: A brief overview and recent advances. In2017 IEEE 56th annual conference on decision and control (CDC), pages 2242–2253. IEEE, 2017
2017
-
[88]
Simple and scalable predictive uncertainty estimation using deep ensembles.Advances in neural information processing systems, 30, 2017
Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. Simple and scalable predictive uncertainty estimation using deep ensembles.Advances in neural information processing systems, 30, 2017
2017
-
[89]
Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift
Jasper Snoek, Yaniv Ovadia, Emily Fertig, Balaji Lakshminarayanan, Sebastian Nowozin, D Sculley, Joshua V Dillon, Jie Ren, and Zachary Nado. Can you trust your model’s uncertainty? evaluating predictive uncertainty under dataset shift. 2019
2019
-
[90]
Angelopoulos and Stephen Bates
Anastasios N. Angelopoulos and Stephen Bates. A gentle introduction to conformal prediction and distribution-free uncertainty quantification.arXiv preprint arXiv: 2107.07511, 2021
2021 arXiv
-
[91]
Wilds: A benchmark of in-the-wild distribution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. InInternational conference on machine learning, pa...
2021
-
[92]
Robot learning from randomized simulations: A review.Frontiers in Robotics and AI, 9:799893, 2022
Fabio Muratore, Fabio Ramos, Greg Turk, Wenhao Yu, Michael Gienger, and Jan Peters. Robot learning from randomized simulations: A review.Frontiers in Robotics and AI, 9:799893, 2022
2022
-
[93]
Formal scenario-based testing of autonomous vehicles: From simulation to the real world
Daniel J Fremont, Edward Kim, Yash Vardhan Pant, Sanjit A Seshia, Atul Acharya, Xantha Bruso, Paul Wells, Steve Lemke, Qiang Lu, and Shalin Mehta. Formal scenario-based testing of autonomous vehicles: From simulation to the real world. In2020 IEEE 23rd International Conference...
2020
-
[94]
Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint arXiv:2506.18088, 2025
Tianxing Chen, Zanxin Chen, Baijun Chen, Zijian Cai, Yibin Liu, Zixuan Li, Qiwei Liang, Xianliang Lin, Yiheng Ge, Zhenyu Gu, et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation.arXiv preprint ar...
2025 arXiv
-
[95]
Robotwin: Dual-arm robot benchmark with generative digital twins
Yao Mu, Tianxing Chen, Zanxin Chen, Shijia Peng, Zhiqian Lan, Zeyu Gao, Zhixuan Liang, Qiaojun Yu, Yude Zou, Mingkun Xu, Lunkai Lin, Zhiqiang Xie, Mingyu Ding, and Ping Luo. Robotwin: Dual-arm robot benchmark with generative digital twins. InProceedings of the Computer Vision ...
2025
-
[96]
Dexgarmentlab: Dexterous garment manipulation environment with generalizable policy, 2025
Yuran Wang, Ruihai Wu, Yue Chen, Jiarui Wang, Jiaqi Liang, Ziyu Zhu, Haoran Geng, Jitendra Malik, Pieter Abbeel, and Hao Dong. Dexgarmentlab: Dexterous garment manipulation environment with generalizable policy, 2025. URLhttps://arxiv.org/abs/2505.11032
2025
-
[97]
Autobio: A simulation and benchmark for robotic automation in digital biology laboratory.arXiv preprint arXiv:2505.14030, 2025
Zhiqian Lan, Yuxuan Jiang, Ruiqi Wang, Xuanbing Xie, Rongkui Zhang, Yicheng Zhu, Peihang Li, Tianshuo Yang, Tianxing Chen, Haoyu Gao, et al. Autobio: A simulation and benchmark for robotic automation in digital biology laboratory.arXiv preprint arXiv:2505.14030, 2025
2025 arXiv
-
[98]
Robodojo: A unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.arXiv preprint arXiv:2607.04434, 2026
Tianxing Chen, Yue Chen, Zixuan Li, Junyuan Tang, Kailun Su, Weijie Wan, Baijun Chen, Haoran Lu, Haowen Yan, Honghao Su, et al. Robodojo: A unified sim-and-real benchmark for comprehensive evaluation of generalist robot manipulation policies.arXiv preprint arXiv:2607.04434, 2026
2026 arXiv
-
[99]
Robothor: An open simulation-to-real embodied ai platform
Matt Deitke, Winson Han, Alvaro Herrasti, Aniruddha Kembhavi, Eric Kolve, Roozbeh Mottaghi, Jordi Salvador, Dustin Schwenk, Eli VanderBilt, Matthew Wallingford, et al. Robothor: An open simulation-to-real embodied ai platform. InProceedings of the IEEE/CVF conference on comput...
2020
-
[100]
Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024
Xuanlin Li, Kyle Hsu, Jiayuan Gu, Karl Pertsch, Oier Mees, Homer Rich Walke, Chuyuan Fu, Ishikaa Lunawat, Isabel Sieh, Sean Kirmani, et al. Evaluating real-world robot manipulation policies in simulation.arXiv preprint arXiv:2405.05941, 2024
2024 arXiv
-
[101]
Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, 44(10-11):1863–1891, 2025
Minho Heo, Youngwoon Lee, Doohyun Lee, and Joseph J Lim. Furniturebench: Reproducible real-world benchmark for long-horizon complex manipulation.The International Journal of Robotics Research, 44(10-11):1863–1891, 2025
2025
-
[102]
Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025
Pranav Atreya, Karl Pertsch, Tony Lee, Moo Jin Kim, Arhan Jain, Artur Kuramshin, Clemens Eppner, Cyrus Neary, Edward Hu, Fabio Ramos, et al. Roboarena: Distributed real-world evaluation of generalist robot policies.arXiv preprint arXiv:2506.18123, 2025
2025
-
[103]
Robochallenge: Large-scale real-robot evaluation of embodied policies.arXiv preprint arXiv:2510.17950, 2025
Adina Yakefu, Bin Xie, Chongyang Xu, Enwen Zhang, Erjin Zhou, Fan Jia, Haitao Yang, Haoqiang Fan, Haowei Zhang, Hongyang Peng, et al. Robochallenge: Large-scale real-robot evaluation of embodied policies.arXiv preprint arXiv:2510.17950, 2025
2025
-
[104]
Fault injection techniques and tools.Computer, 30(4): 75–82, 1997
Mei-Chen Hsueh, Timothy K Tsai, and Ravishankar K Iyer. Fault injection techniques and tools.Computer, 30(4): 75–82, 1997
1997
-
[105]
Verifai: A toolkit for the formal design and analysis of artificial intelligence-based systems
Tommaso Dreossi, Daniel J Fremont, Shromona Ghosh, Edward Kim, Hadi Ravanbakhsh, Marcell Vazquez-Chanlatte, and Sanjit A Seshia. Verifai: A toolkit for the formal design and analysis of artificial intelligence-based systems. In International Conference on Computer Aided Verifi...
2019
-
[106]
Scenic: a language for scenario specification and scene generation
Daniel J Fremont, Tommaso Dreossi, Shromona Ghosh, Xiangyu Yue, Alberto L Sangiovanni-Vincentelli, and Sanjit A Seshia. Scenic: a language for scenario specification and scene generation. InProceedings of the 40th ACM SIGPLAN conference on programming language design and imple...
2019
-
[107]
Adaptive stress testing for autonomous vehicles
Mark Koren, Saud Alsaif, Ritchie Lee, and Mykel J Kochenderfer. Adaptive stress testing for autonomous vehicles. In2018 IEEE Intelligent Vehicles Symposium (IV), pages 1–7. IEEE, 2018
2018
-
[108]
Advsim: Generating safety-critical scenarios for self-driving vehicles
Jingkang Wang, Ava Pun, James Tu, Sivabalan Manivasagam, Abbas Sadat, Sergio Casas, Mengye Ren, and Raquel Urtasun. Advsim: Generating safety-critical scenarios for self-driving vehicles. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, page...
2021
-
[109]
Deeptest: Automated testing of deep-neural-network-driven autonomous cars
Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. Deeptest: Automated testing of deep-neural-network-driven autonomous cars. InProceedings of the 40th international conference on software engineering, pages 303–314, 2018
2018
-
[110]
Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020
Stephen James, Zicong Ma, David Rovick Arrojo, and Andrew J Davison. Rlbench: The robot learning benchmark & learning environment.IEEE Robotics and Automation Letters, 5(2):3019–3026, 2020
2020
-
[111]
Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning
Tianhe Yu, Deirdre Quillen, Zhanpeng He, Ryan Julian, Karol Hausman, Chelsea Finn, and Sergey Levine. Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning. InConference on robot learning, pages 1094–1100. PMLR, 2020
2020
-
[112]
robosuite: A modular simulation framework and benchmark for robot learning
Yuke Zhu, Josiah Wong, Ajay Mandlekar, Roberto Martín-Martín, Abhishek Joshi, Kevin Lin, Abhiram Maddukuri, Soroush Nasiriany, and Yifeng Zhu. robosuite: A modular simulation framework and benchmark for robot learning. arXiv preprint arXiv:2009.12293, 2020
2009 arXiv
-
[113]
Calvin: Abenchmarkforlanguage-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022
OierMees, LukasHermann, ErickRosete-Beas, andWolframBurgard. Calvin: Abenchmarkforlanguage-conditioned policy learning for long-horizon robot manipulation tasks.IEEE Robotics and Automation Letters, 7(3):7327–7334, 2022
2022
-
[114]
Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023
Bo Liu, Yifeng Zhu, Chongkai Gao, Yihao Feng, Qiang Liu, Yuke Zhu, and Peter Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023
2023
-
[115]
Maniskill2: A unified benchmark for generalizable manipulation skills.arXiv preprint arXiv:2302.04659, 2023
Jiayuan Gu, Fanbo Xiang, Xuanlin Li, Zhan Ling, Xiqiang Liu, Tongzhou Mu, Yihe Tang, Stone Tao, Xinyue Wei, Yunchao Yao, et al. Maniskill2: A unified benchmark for generalizable manipulation skills.arXiv preprint arXiv:2302.04659, 2023. 34
2023 arXiv
-
[116]
Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation.arXiv preprint arXiv:2403.09227, 2024
Chengshu Li, Ruohan Zhang, Josiah Wong, Cem Gokmen, Sanjana Srivastava, Roberto Martín-Martín, Chen Wang, Gabrael Levine, Wensi Ai, Benjamin Martinez, et al. Behavior-1k: A human-centered, embodied ai benchmark with 1,000 everyday activities and realistic simulation.arXiv prep...
2024 arXiv
-
[117]
Imdl-benco: A comprehensive benchmark and codebase for image manipulation detection & localization.Advances in Neural Information Processing Systems, 37:134591–134613, 2024
Xiaochen Ma, Xuekang Zhu, Lei Su, Bo Du, Zhuohang Jiang, Bingkui Tong, Zeyu Lei, Xinyu Yang, Chi-Man Pun, Jiancheng Lv, et al. Imdl-benco: A comprehensive benchmark and codebase for image manipulation detection & localization.Advances in Neural Information Processing Systems, ...
2024
-
[118]
Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026
Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, et al. Rmbench: Memory-dependent robotic manipulation benchmark with insights into policy design.arXiv preprint arXiv:2603.01229, 2026
2026 arXiv
-
[119]
Eventvla: Event-driven visual evidence memory for long-horizon vision-language-action policies.arXiv preprint arXiv:2606.20092, 2026
Ganlin Yang, Zhangzheng Tu, Yuqiang Yang, Sitong Mao, Junyi Dong, Tianxing Chen, Jiaqi Peng, Jing Xiong, Jiafei Cao, Jifeng Dai, et al. Eventvla: Event-driven visual evidence memory for long-horizon vision-language-action policies.arXiv preprint arXiv:2606.20092, 2026
2026 arXiv
-
[120]
Yiting Chen, Kenneth Kimble, Edward H. Adelson, Tamim Asfour, Podshara Chanrungmaneekul, Sachin Chitta, Yash Chitambar, Ziyang Chen, Ken Goldberg, Danica Kragic, Hui Li, Xiang Li, Yunzhu Li, Aaron Prather, Nancy Pollard, Maximo A. Roa-Garzon, Robert Seney, Shuo Sha, Shihefeng ...
2026
-
[121]
Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.10093, 2026
Baijun Chen, Weijie Wan, Tianxing Chen, Xianda Guo, Congsheng Xu, Yuanyang Qi, Haojie Zhang, Longyan Wu, Tianling Xu, Zixuan Li, et al. Univtac: A unified simulation platform for visuo-tactile manipulation data generation, learning, and benchmarking.arXiv preprint arXiv:2602.1...
2026
-
[122]
The goal structuring notation–a safety argument notation
Tim Kelly and Rob Weaver. The goal structuring notation–a safety argument notation. InProceedings of the dependable systems and networks 2004 workshop on assurance cases, volume 6. Citeseer Princeton, NJ, 2004
2004
-
[123]
Safe reinforcement learning via shielding
Mohammed Alshiekh, Roderick Bloem, Rüdiger Ehlers, Bettina Könighofer, Scott Niekum, and Ufuk Topcu. Safe reinforcement learning via shielding. InProceedings of the AAAI conference on artificial intelligence, volume 32, 2018
2018
-
[124]
SAE international, 2021
On-Road Automated Driving (ORAD) Committee.Taxonomy and definitions for terms related to driving automation systems for on-road motor vehicles. SAE international, 2021
2021
-
[125]
How many operational design domains, objects, and events?Safeai@ aaai, 4 (4), 2019
Philip Koopman, Frank Fratrik, et al. How many operational design domains, objects, and events?Safeai@ aaai, 4 (4), 2019
2019
-
[126]
Run-time monitoring of machine learning for robotic perception: A survey of emerging trends.IEEE Access, 9:20067–20075, 2021
Quazi Marufur Rahman, Peter Corke, and Feras Dayoub. Run-time monitoring of machine learning for robotic perception: A survey of emerging trends.IEEE Access, 9:20067–20075, 2021
2021
-
[127]
A system-level view on out-of-distribution data in robotics.arXiv preprint arXiv:2212.14020, 2022
Rohan Sinha, Apoorva Sharma, Somrita Banerjee, Thomas Lew, Rachel Luo, Spencer M Richards, Yixiao Sun, Edward Schmerling, and Marco Pavone. A system-level view on out-of-distribution data in robotics.arXiv preprint arXiv:2212.14020, 2022
2022 arXiv
-
[128]
A brief account of runtime verification.The journal of logic and algebraic programming, 78(5):293–303, 2009
Martin Leucker and Christian Schallhart. A brief account of runtime verification.The journal of logic and algebraic programming, 78(5):293–303, 2009
2009
-
[129]
Formal specification and verification of autonomous robotic systems: A survey.ACM Computing Surveys (CSUR), 52(5):1–41, 2019
Matt Luckcuck, Marie Farrell, Louise A Dennis, Clare Dixon, and Michael Fisher. Formal specification and verification of autonomous robotic systems: A survey.ACM Computing Surveys (CSUR), 52(5):1–41, 2019
2019
-
[130]
Anomaly detection: A survey.ACM computing surveys (CSUR), 41(3):1–58, 2009
Varun Chandola, Arindam Banerjee, and Vipin Kumar. Anomaly detection: A survey.ACM computing surveys (CSUR), 41(3):1–58, 2009
2009
-
[131]
MIT press, 1992
Thomas B Sheridan.Telerobotics, automation, and human supervisory control. MIT press, 1992
1992
-
[132]
A model for types and levels of human interaction with automation.IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans, 30 (3):286–297, 2000
Raja Parasuraman, Thomas B Sheridan, and Christopher D Wickens. A model for types and levels of human interaction with automation.IEEE Transactions on systems, man, and cybernetics-Part A: Systems and Humans, 30 (3):286–297, 2000
2000
-
[133]
The protection of information in computer systems.Proceedings of the IEEE, 63(9):1278–1308, 1975
Jerome H Saltzer and Michael D Schroeder. The protection of information in computer systems.Proceedings of the IEEE, 63(9):1278–1308, 1975
1975
-
[134]
Ironies of automation
Lisanne Bainbridge. Ironies of automation. InAnalysis, design and evaluation of man–machine systems, pages 129–135. Elsevier, 1983
1983
-
[135]
Toward a theory of situation awareness in dynamic systems
Mica R Endsley. Toward a theory of situation awareness in dynamic systems. InSituational awareness, pages 9–42. Routledge, 2017
2017
-
[136]
Takeover time in highly automated vehicles: noncritical transitions to and from manual control.Human factors, 59(4):689–705, 2017
Alexander Eriksson and Neville A Stanton. Takeover time in highly automated vehicles: noncritical transitions to and from manual control.Human factors, 59(4):689–705, 2017
2017
-
[137]
ISO/TS 15066:2016: Robots and robotic devices — Collaborative robots
International Organization for Standardization. ISO/TS 15066:2016: Robots and robotic devices — Collaborative robots. Technical Specification, 2016
2016
-
[138]
Examining accident reports involving autonomous vehicles in california.PLoS one, 2017
Francesca Favaro, Nazanin Nader, Sky Eurich, Michelle Tripp, and Naresh Varadaraju. Examining accident reports involving autonomous vehicles in california.PLoS one, 2017
2017
-
[139]
Hidden technical debt in machine learning systems.Advances in neural information processing systems, 28, 2015
David Sculley, Gary Holt, Daniel Golovin, Eugene Davydov, Todd Phillips, Dietmar Ebner, Vinay Chaudhary, Michael Young, Jean-Francois Crespo, and Dan Dennison. Hidden technical debt in machine learning systems.Advances in neural information processing systems, 28, 2015. 35
2015
-
[140]
Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges.arXiv preprint arXiv:1907.00182, 1(2):8, 2019
Timothée Lesort, Vincenzo Lomonaco, Andrei Stoian, Davide Maltoni, David Filliat, Natalia Díaz-Rodríguez, et al. Continual learning for robotics: Definition, framework, learning strategies, opportunities and challenges.arXiv preprint arXiv:1907.00182, 1(2):8, 2019
1907 arXiv
-
[141]
The vision of autonomic computing.Computer, 36(1):41–50, 2003
Jeffrey O Kephart and David M Chess. The vision of autonomic computing.Computer, 36(1):41–50, 2003
2003
-
[142]
Toward trustworthy ai development: mechanisms for supporting verifiable claims.arxiv, 2020:2004–07213, 2020
Miles Brundage, Shahar Avin, Jasmine Wang, Haydn Belfield, Gretchen Krueger, Gillian Hadfield, Heidy Khlaaf, Jingying Yang, Helen Toner, Ruth Fong, et al. Toward trustworthy ai development: mechanisms for supporting verifiable claims.arxiv, 2020:2004–07213, 2020
2020
-
[143]
Trustworthy ai: From principles to practices.ACM Computing Surveys, 55(9):1–46, 2023
Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, and Bowen Zhou. Trustworthy ai: From principles to practices.ACM Computing Surveys, 55(9):1–46, 2023
2023
-
[144]
A survey on artificial intelligence assurance.Journal of Big Data, 8(1), 2021
Feras A Batarseh, Laura Freeman, and Chih-Hao Huang. A survey on artificial intelligence assurance.Journal of Big Data, 8(1), 2021
2021
-
[145]
Embodied ai in action: Insights from sae world congress 2026 on safety, trust, robotics, and real-world deployment
Jan-Mou Li, Paul Schmitt, Wei Tong, Majed Mohammed, Akshay Chalana, Arpan Kusari, and Edward Griffor. Embodied ai in action: Insights from sae world congress 2026 on safety, trust, robotics, and real-world deployment. arXiv preprint arXiv:2605.10653, 2026
2026 arXiv
-
[147]
Open x-embodiment: Robotic learning datasets and rt-x models
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models. In2024 IEEE International Conference on Robotics and Aut...
2024
-
[148]
Foundation models in robotics: Applications, challenges, and the future
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, et al. Foundation models in robotics: Applications, challenges, and the future. The International Journal of Robotics Research, 44(5):7...
2025
-
[149]
Embodied agent interface: Benchmarking llms for embodied decision making
Manling Li, Shiyu Zhao, Qineng Wang, Kangrui Wang, Yu Zhou, Sanjana Srivastava, Cem Gokmen, Tony Lee, Li Erran Li, Ruohan Zhang, et al. Embodied agent interface: Benchmarking llms for embodied decision making. arXiv preprint arXiv:2410.07166, 2024
2024 arXiv
-
[150]
Safety assurance of machine learning for autonomous systems.Reliability Engineering and System Safety, 2025
Colin Paterson, Richard David Hawkins, Chiara Picardi, Yan Jia, Radu Calinescu, and Ibrahim Habli. Safety assurance of machine learning for autonomous systems.Reliability Engineering and System Safety, 2025
2025
-
[151]
ISO 10218-1:2025: Robotics — Safety requirements — Part 1: Industrial robots
International Organization for Standardization. ISO 10218-1:2025: Robotics — Safety requirements — Part 1: Industrial robots. International Standard, 3rd edition, 2025
2025
-
[152]
IEC 61508: Functional safety of electrical/electronic/programmable electronic safety-related systems
International Electrotechnical Commission. IEC 61508: Functional safety of electrical/electronic/programmable electronic safety-related systems. Technical Report IEC 61508, International Electrotechnical Commission, 2010. Edition 2.0
2010
-
[153]
Artificial intelligence risk management framework (ai rmf 1.0), 2023
Elham Tabassi. Artificial intelligence risk management framework (ai rmf 1.0), 2023. URLhttps://tsapps. nist.gov/publication/get_pdf.cfm?pub_id=936225. 36 Appendix A Suggested Literature Search Protocol Future updates of this review may use a structured search across IEEE Xplo...
2023
Reviewed August 1, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.