{"id":"65c814cb-5218-4fd4-a722-196b08c504b6","arxiv_id":"2607.26121","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A four-layer systems framework and T0–T5 hierarchy for grading and maintaining bounded trustworthiness claims in embodied AI systems.","lead":"This paper defines trustworthy embodied intelligence as sustained safe success and organizes the mechanisms into four layers—model, system, evidence, and deployment—plus a T0–T5 grading hierarchy. It argues that capability, safety, assurance, and governance must be jointly evaluated for any bounded deployment claim.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"T-level grading is underdetermined: Sections 8.1-8.5 provide no operational criterion for 'acceptable residual risk' or 'adequate support,' so a fixed system can receive different T-levels under permissible profiles, undercutting the claimed comparative evaluation.","rationale":"The paper's contribution is a definitional framework; its central claim is that trustworthiness is an end-to-end property and that its strength can be graded. I looked first for whether the four-layer decomposition could be falsified. It is presented as an organizing structure with concrete cross-layer failure examples, and the paper repeatedly hedges that levels are non-normative and domain-specific. The main live threat to the usefulness of the central claim is not internal inconsistency but underdetermination: the grading machinery has no semantics until thresholds and violation criteria are supplied. The reader's weakest assumption identified exactly this. I considered whether the non-systematic literature base (Appendix D) was more load-bearing; it matters for the survey's completeness but not for the logic of the end-to-end claim. I also considered whether absence of formal verification or empirical validation was a defect; this is a position paper, so that absence is not itself a flaw. The concern I raise is concrete and testable: construct two permissible domain profiles and see whether the same system can be assigned different levels. If it cannot, the framework is better defined than it appears; if it can, the paper's comparative-evaluation claim needs a qualifier. Thus the reader's CONDITIONAL verdict stands; no adjustment needed.","tokens_in":31609,"tokens_out":6804,"duration_ms":69171,"concrete_test":"Test the ordinal underdetermination formally. Fix a system with an evidence vector e = (capability, safety, system assurance, evidence, governance) and define 'adequate support' as e_i ≥ c_i. Show that for a fixed e, the paper's rule 'level cannot exceed the least adequately supported dimension' (Section 8.1) yields different T-levels as c_i vary over the ranges implied by existing safety standards (e.g., ISO 10218/TS 15066 thresholds vs stricter automotive SOTIF thresholds). If two c vectors both permissible under Section 8.5 assign the same system different T-levels, the hierarchy cannot support the claimed comparative evaluation without additional constraints. Equivalently, ask two independent evaluators to produce T-level assignments for the same documented robot evaluation using only the paper's definitions; disagreement would demonstrate the missing operationalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The definition of TEI is sustained safe success — 'reliable completion ... while maintaining risk within acceptable bounds' (Section 1). The only formal evaluative object, the safe-success rate (Eq. 1, Section 6.2), conditions on 'no unacceptable violation', so 'unacceptable violation' is a primitive. Section 8.1 then grades T0-T5 by the least adequately supported of five dimensions, and the level definitions (Section 8.2) use phrases like 'explicitly specified and validated conditions', 'defined set of foreseeable failures', and 'accepted residual risk'. Section 8.5 defers the quantitative thresholds, acceptable residual risk, required evidence, and independent assessment to 'domain-specific profiles'. This deferral is reasonable for a non-normative proposal, but it leaves the hierarchy's principal claimed functions — comparative evaluation and bounded deployment (Section 1.2) — underdetermined. Concretely, the assignment function T(s_1,...,s_5) is undefined without cutoffs c_i for adequacy. The same system with the same evidence can be T2 under one reasonable profile (e.g., one severe violation per 10^3 trials acceptable) and T4 under another (one per 10^6 trials), because the least-supported dimension changes. Since the paper supplies no constraint on choosing these cutoffs, the T-level ordering is not a well-defined comparative instrument; it is a template for future standards. The paper acknowledges this (Sections 8.5, 10), but acknowledgement does not remove the tension with the claim that the hierarchy 'supports comparative evaluation'. Appendix D's admission that the review is a 'structured synthesis of the supplied seed literature' also weakens the empirical basis for the claimed completeness of the four-layer decomposition, but the threshold issue is the more direct threat to the central grading claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This position/survey paper argues that trustworthiness for embodied intelligence cannot be established by any single model, component, or benchmark score. It defines trustworthy embodied intelligence as “sustained safe success,” organized around four interdependent layers—model, system, evidence, and deployment—and proposes a non-normative T0–T5 hierarchy for grading the strength of bounded deployment claims. The paper reviews a broad literature spanning embodied AI, robotics, control, dependable computing, fault tolerance, and autonomous driving, and it connects these fields through the four-layer framework and the proposed TEI levels.","tokens_in":31984,"tokens_out":4216,"duration_ms":48551,"significance":"If accepted, the framework could provide a useful common vocabulary and architectural reference for safety and assurance in embodied AI, and it appropriately stresses bounded deployment claims, assurance cases, and lifecycle governance. The paper is careful to describe the hierarchy as analytical rather than a certification scheme, and it repeatedly acknowledges that thresholds and acceptable-residual-risk criteria are domain-specific. Its strengths include a broad cross-layer synthesis, explicit treatment of failure propagation, and alignment with existing standards and assurance concepts. However, the central evaluative claim—that the T0–T5 hierarchy supports comparative evaluation and bounded deployment—is not yet operationalized. The only formal evaluative object, the safe-success rate in Eq. (1), depends on an uninterpreted primitive “unacceptable violation,” and the T-level assessment in Section 8.1 depends on undefined thresholds for “adequate support.” The paper is therefore best read as a template or agenda for future standardization rather than a complete comparative instrument.","major_comments":[{"comment":"The T-level assignment function is underdetermined. Section 8.1 says a TEI level is the least adequately supported of five dimensions, but “adequately supported” is not defined, and Section 8.5 defers quantitative thresholds and acceptable residual risk to domain-specific profiles. Concretely, the same system with the same evidence could be T2 under one permissible profile (e.g., one severe violation per 10^3 trials is acceptable) and T4 under another (one per 10^6 trials), because the least-supported dimension changes. Since the paper supplies no constraint on choosing these cutoffs, the claimed functions of comparative evaluation and bounded deployment (Section 1.2, contribution 4) are not well-defined. The authors acknowledge this deferral, but acknowledgement does not remove the tension; the paper should either provide a working example of a profile and its induced T-level, or explic","section":"§8.1–8.5 and Eq. (1)"},{"comment":"The safe-success rate is introduced as the primary joint outcome, yet it is an unweighted trial fraction. The same equation treats a high-energy collision and a minor safety-filter activation as equivalent failures, while Section 6.2 subsequently states that these should not be weighted equally. The paper says SSR should be interpreted together with component outcomes, but then SSR is not itself the “primary joint outcome” in any decision-relevant sense. A severity- or exposure-weighted statistic, or a vector of scenario-conditioned rates with the violation predicate made explicit, would be more consistent with the stated safety goals. As written, Eq. (1) is a useful definitional starting point but not a measurable metric without a specification of the unacceptable-violation predicate.","section":"§6.2, Eq. (1)"},{"comment":"The four-layer decomposition and the five assessment dimensions are asserted as jointly necessary, but no derivation or empirical evidence is provided for their completeness. This is load-bearing for the paper’s central claim that “no single layer can establish end-to-end trustworthiness” and that a TEI level cannot exceed the least adequately supported dimension. I am not asking for a formal completeness proof in a survey, but the paper should clarify whether these are normative proposals or intended descriptive claims about existing systems. If the latter, at least one illustrative application of the framework to a concrete system would help show that the dimensions can be jointly assessed and that the hierarchy yields stable classifications.","section":"§3.2 and §8.1"}],"minor_comments":[{"comment":"“Weusetrustworthy embodied intelligenceto” is missing spaces; please correct “We use trustworthy embodied intelligence to”.","section":"§1"},{"comment":"The notation N(task completed ∧ no unacceptable violation) / N(evaluated trials) uses the same symbol N for both the numerator and denominator; consider N_safe_success / N_total for clarity.","section":"§6.2, Eq. (1)"},{"comment":"The table aligns AgiBot G1–G5, SAE L0–L5, and TEI T0–T5 row by row. The text cautions that the alignment is illustrative only, but the visual presentation still invites cross-hierarchy equivalence inferences. Consider adding an explicit sentence that the rows are not intended to imply comparable levels of capability, autonomy, or trustworthiness.","section":"§8.4, Table 4"},{"comment":"Reference [146] is first-party company material, and Appendix D correctly notes that internal and industry materials should be labeled as first-party sources. In the main text, however, the discussion around Section 8.4 and the roadmap (Section 9.2) cites company material without that label at the point of use. Adding an explicit first-party marker at the citation site would be more transparent.","section":"Appendix D and [146]"}],"recommendation":"major_revision","confidential_remarks":"The paper is a thoughtful position piece rather than a falsifiable technical contribution. In my view, the underdetermination of the T-level assignment is the main issue: it is not fatal for a survey, but it directly affects the paper’s stated contribution of comparative evaluation. I would be satisfied by a major revision that either (a) adds a concrete worked example of a domain profile and the resulting T-level, or (b) rephrases the fourth contribution to avoid claiming an evaluative instrument and instead presents the hierarchy as a standardization template."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here's my read.\n\nThe paper is a well-organized survey and position piece, not a technical result. What's genuinely new is the package: sustained safe success as the target definition, the four-layer decomposition (model, system, evidence, deployment), and a T0–T5 ladder explicitly distinguished from SAE and AgiBot taxonomies. The central claim that trustworthiness is a property of a bounded, deployed claim rather than a model or benchmark score is right and worth saying. The synthesis draws effectively on robotics, control, dependable computing, and autonomous-driving assurance.\n\nNow the soft spots, in proportion. The stress-test concern is legitimate and the biggest one: the T-level assignment is underdetermined. Eq. (1)'s safe-success rate leaves 'unacceptable violation' as a primitive, and Sections 8.1–8.5 defer quantitative thresholds, acceptable residual risk, and evidence sufficiency to 'domain-specific profiles.' The same system could be T2 under one reasonable profile and T4 under another. The paper is explicit that the hierarchy is non-normative and awaits such thresholds, so this is an honest limitation rather than a hidden one. But it does undercut the stated contribution that the hierarchy 'supports comparative evaluation.' As it stands, it is a template for future standards, not a working comparative instrument. That's a significant gap, not fatal.\n\nSecond, Appendix D admits the review is a 'structured synthesis of the supplied seed literature and project materials rather than a complete systematic review.' Honest, but it limits reproducibility: a reader cannot reconstruct coverage from the stated protocol. Relatedly, the citation base leans heavily on the authors' own benchmarks (SafeDojo, RoboDojo, RM-Bench, UniVTac) and includes one first-party company roadmap (ref [146]). Self-citation is not disqualifying, but in a survey claiming to map the field, the balance should be examined.\n\nThere is also no empirical or formal demonstration that the four layers are complete or jointly necessary. That is acceptable for a proposal, and the paper doesn't overclaim much. The framework is a proposed ontology, not a validated result.\n\nWho benefits: anyone working on embodied-AI safety, assurance, evaluation, or standardization who needs a shared vocabulary and a structured way to debate bounded deployment claims. It deserves a serious referee; the synthesis is useful and the authors are transparent about their own limits.\n\nMy recommendation: send it out. Ask the authors to either sharpen the T-level assignment with at least a concrete worked example of threshold-setting, or soften the 'comparative evaluation' claim to 'template for future comparative evaluation.' Keep the review scope clear: this is a position survey, and the framework's value will be proven by adoption, not by proof.","headline":"A serious, honest position piece that gives the field a useful four-layer vocabulary; the T0–T5 ladder is a scaffold awaiting thresholds, so its comparative-evaluation claim is still a promise.","tokens_in":32669,"tokens_out":3614,"would_cite":true,"duration_ms":39087,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Embodied AI trustworthiness is a system property, not a model score","keywords":["embodied intelligence","trustworthy AI","safety","sustained safe success","assurance","deployment governance","robot safety","evaluation"],"falsifier":"A system that meets all T4 assessment dimensions but whose claim lapses after a small unmonitored change—for example, a camera recalibration—without any deployment-layer detection would falsify the claim that a bounded claim's validity is maintained by the four layers. Concretely, audit a deployed robot across such a change and check whether its monitored assumptions trigger revalidation before harm occurs.","tokens_in":31469,"feed_emoji":"🤖","tokens_out":3903,"duration_ms":38258,"temperature":0.7,"pith_summary":"This paper proposes that an embodied AI system—a robot or autonomous agent that acts in the physical world—should be called trustworthy only when the whole deployed system reliably completes its tasks while keeping risk within acceptable bounds over time. It defines this goal as \"sustained safe success\" and argues that trustworthiness is an end-to-end property of the deployed system, not of any single model, component, or benchmark score. The paper organizes the supporting mechanisms into four interdependent layers—model, system, evidence, and deployment—and proposes a six-level graded hierarchy (T0–T5) for expressing the strength of a system's bounded trustworthiness claim. A sympathetic reader would care because as embodied AI moves from digital inference to physical interaction, task completion alone can conceal unsafe behavior, and this framework supplies a common structure for comparing, evaluating, and governing deployment claims.","feed_headline":"Trustworthy robots need four layers, not just a good model","feed_subtitle":"A new T0–T5 scale grades embodied systems by sustained safe success across model, system, evidence, and deployment.","key_machinery":"The central object is the bounded trustworthiness claim, expressed as the tuple ⟨system version, task, embodiment, operating domain, authority, evidence⟩, together with the four-layer framework that supports it. The four layers—model, system, evidence, deployment—are functional responsibilities rather than software modules, and the paper stresses cross-layer failure propagation through the semantic–physical gap, the action–consequence gap, and cross-layer non-compositionality. The key quantitative instrument is the safe-success rate, SSR = N(task completed ∧ no unacceptable violation)/N(evaluated trials), which separates four outcome classes: safe success, safe failure, unsafe success, and u","core_discovery":"The central claim is that trustworthiness is an end-to-end property of a deployed system, not an attribute of an individual model, component, or benchmark score. The paper defines trustworthy embodied intelligence as sustained safe success: reliable completion of intended tasks while physical, semantic, procedural, and operational risks remain within acceptable bounds. It identifies the model layer, which proposes actions with calibrated uncertainty and safety preferences; the system layer, which realizes authorized actions dependably through sensing, computing, control, hardware safeguards, fault containment, and fallback; the evidence layer, which substantiates bounded claims through evalu","pith_inferences":["Editorial: If the safe-success-rate categories were widely adopted, benchmark suites would need to include scenario-conditioned risk reporting and severity weighting, so that a minor safety-filter activation and a high-energy collision are not counted equally.","Editorial: The T-level framework implies a staged-approval model for regulators: a low T-level would authorize only tightly constrained deployments, while T4–T5 would be needed for open-ended operation.","Editorial: The paper leaves open the 'assurance-preserving reconfiguration' problem — deciding which tool, payload, or controller changes are minor versus claim-invalidating. An automated impact-analysis tool for revalidation triggers would be a concrete next step."],"forward_implications":["Evaluation of embodied systems should report safe success, unsafe success, safe failure, and unsafe failure, rather than a single completion rate.","A highly capable model can still be at T1 or T2 if system safeguards, evidence, or deployment governance are missing; capability alone does not raise a TEI level.","Deployment should define an explicit operational boundary and use runtime admission, boundary monitoring, intervention, and change control to preserve the validity of the trustworthiness claim.","The T0–T5 hierarchy offers a common structure for comparative evaluation, bounded deployment, and future standardization, while remaining non-normative and requiring domain-specific profiles."],"fun_headline_variants":["Robot trust is not one benchmark — it's four layers","Trustworthy AI needs sustained safe success, not just task wins","Grading trust: four layers and a T0–T5 scale for robots","Robot trustworthiness: end-to-end, not a single metric","Four layers stand between a robot and trustworthiness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The framework assumes that 'unacceptable violation' and 'acceptable residual risk' can be specified and measured for each application; without that, the safe-success rate and T-level assessments cannot be applied.","fun_headline_variants_meta":{"raw":{"variants":["Robot trust is not one benchmark — it's four layers","Trustworthy AI needs sustained safe success, not just task wins","Grading trust: four layers and a T0–T5 scale for robots","Robot trustworthiness: end-to-end, not a single metric","Four layers stand between a robot and trustworthiness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000847,"raw_usage":{"total_tokens":3529,"prompt_tokens":756,"completion_tokens":2773,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":500,"completion_tokens_details":{"reasoning_tokens":2701}},"tokens_in":500,"tokens_out":2773,"duration_ms":15056,"temperature":1.0,"reasoning_tokens":2701,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:44:10.044573+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A system that meets all T4 assessment dimensions but whose claim lapses after a small unmonitored change—for example, a camera recalibration—without any deployment-layer detection would falsify the claim that a bounded claim's validity is maintained by the four layers. Concretely, audit a deployed robot across such a change and check whether its monitored assumptions trigger revalidation before harm occurs.","supporting_citations":[],"review_version":1}