{"id":"4a685bf3-ceb5-48a8-bb4c-a0d04a800c4e","arxiv_id":"2511.15620","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Because robustness under the EU AI Act is context-sensitive, the paper argues, horizontal standards must be complemented by domain-specific specifications and a dynamic repository of assessment practices.","lead":"Robustness is a legal requirement for high-risk AI under the EU AI Act, but the law does not say what 'robust' means or how to test it. This paper argues that the answer depends on the use case, the data and the model, and proposes a layered European standards system with a living repository of tests and best practices.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'cannot' in 'horizontal standards alone cannot fulfil M/593' is not established: a horizontal standard could in principle embed context-sensitive conditional specifications, so the necessity of the proposed layered architecture is overstated.","rationale":"The reader's weakest assumption was empirical: ISO/IEC 24029-3 and related JTC 21 deliverables might turn out to be sufficiently detailed. I agree that this is uncertain, but the more fundamental problem is conceptual: the paper moves from context-sensitivity to the impossibility of horizontal standards fulfilling M/593 without considering that 'horizontal' describes scope, not specificity. A horizontal standard can contain conditional, domain-specific provisions. M/593 itself asks for a 'range of technical options' and notes that vertical specifications should be developed 'where appropriate'—it does not require a separate document hierarchy. The paper's own evidence (Table 1, §5.3) shows that robustness evaluations differ across two medical imaging tasks, but this only demonstrates context-sensitivity, not that a single standard could not enumerate both options. Therefore, the central claim as stated ('horizontal standards alone cannot fulfil') is stronger than what the authors prove. The paper would be more accurate claiming that current horizontal standards and the planned JTC 21 deliverables are likely insufficient and that a layered architecture is one way to achieve the needed specificity. This does not undermine the value of the conceptual framework (robustness dimensions, contextual drivers), but it requires moderating the policy conclusion and addressing the horizontal/vertical dichotomy explicitly. Because the paper already acknowledges uncertainty ('its scope and level of specificity remain to be seen', §6), the verdict remains CONDITIONAL: the paper is conditionally acceptable if the authors soften the categorical claim and/or provide an argument for why embedding context-sensitivity in a single horizontal standard is infeasible. My concrete test—designing a horizontal standard with conditional domain-specific clauses—would settle whether the categorical claim is defensible. I partially agree with the reader: we both identify the insufficiency premise as load-bearing, but I locate the core weakness in the logical inference rather than solely in the empirical trajectory of Part 3.","tokens_in":13671,"tokens_out":7621,"duration_ms":84991,"concrete_test":"Draft a minimal horizontal robustness standard section containing, for two high-risk domains (e.g., medical imaging and employment screening), a conditional perturbation taxonomy, candidate tests/metrics, and threshold-setting guidance, organized as clauses triggered by use-case, data, and model descriptors. Check this design against M/593's requirement to provide 'a range of technical options that providers can assess and implement in light of the purpose of their system.' If such a single horizontal standard can be written without internal contradiction and satisfies the M/593 criteria, the paper's claim that horizontal standards alone cannot fulfil the mandate is falsified. If it cannot, the concern is settled in the paper's favor.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference—from robustness being context-sensitive to the impossibility of horizontal standards alone satisfying M/593—relies on conflating horizontal scope with generic content. A standard that applies across sectors is not logically barred from containing detailed conditional specifications: a single horizontal document could include, for each high-risk domain, a perturbation taxonomy, candidate tests/metrics, and threshold-setting guidance, plus the 'range of technical options' M/593 requires. The authors' evidence addresses the current JTC 21 work programme (Part 3 unknown; vision/NLP specs not detailed), which is an empirical claim about present drafts, not a structural proof about the horizontal category. The paper itself hedges: 'its scope and level of specificity remain to be seen' (§6). Without an argument that a single horizontal document cannot carry such content (e.g., due to legal format constraints, consensus process, or maintainability), the categorical 'cannot' is too strong. The practical recommendation (layered architecture) may still be good, but its necessity is not demonstrated. If a well-designed horizontal standard with embedded domain-specific annexes can satisfy M/593, then the diagnosis reduces to 'current drafts are insufficient,' not 'horizontal standards alone cannot fulfil the mandate.' That distinction matters because the latter justifies restructuring the deliverables; the former only justifies improving them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that robustness of AI systems under the EU AI Act is inherently context-sensitive: what should remain stable ('robustness of what'), which perturbations matter ('robustness to what'), and the operational environment all shape how robustness should be assessed. The authors identify use case, data, and model as three contextual drivers and illustrate their interaction by comparing two medical-imaging robustness studies. They then argue that current and planned harmonised standards—mostly horizontal in scope—do not provide the detailed, context-specific technical options that Standardisation Request M/593 requires, and they propose a multi-layered framework: horizontal standards set common principles; domain-specific standards and a lifecycle-oriented perturbation taxonomy operationalise them; and a dynamic, stakeholder-fed repository of practices, benchmarks, and sandboxes addresses both context-dependence and standards obsolescence. The paper is a policy-analytic proposal rather than an empirical study; its evidence base is the AI Act, JRC reports, M/593, and the existing JTC 21 work programme.","tokens_in":13988,"tokens_out":3875,"duration_ms":44408,"significance":"If the paper's central claim is accepted, it would have concrete consequences for the EU standardisation process: CEN/CENELEC JTC 21 and ISO/IEC would need to restructure the robustness deliverables (the 24029 series and successors) into a layered, context-sensitive architecture rather than generic horizontal documents. The paper makes a useful conceptual contribution by translating philosophical critiques of robustness (e.g., Freiesleben & Grote) into a practical standardisation vocabulary, and it grounds its diagnosis in specific legal and documentary sources, including Article 15, M/593, and the JRC report. The two-case comparison is a helpful illustration. The authors are also candid about the main uncertainty, explicitly noting that the scope of ISO/IEC 24029-3 'remains to be seen.' The framework's policy recommendations are plausible, but their necessity and feasibility are not fully established in the current text.","major_comments":[{"comment":"The paper's load-bearing claim is that 'horizontal standards alone cannot fulfil' M/593, but the evidence presented supports a weaker claim: the current and planned horizontal deliverables are, at present, insufficiently detailed. The authors themselves write that Part 3's 'scope and level of specificity remain to be seen.' The argument conflates horizontal scope with generic content: a single horizontal standard could in principle contain conditional, domain-specific annexes, sector-specific perturbation taxonomies, and a 'range of technical options' as required by M/593. To make the categorical claim, the authors need either (i) an institutional or legal argument that a single horizontal standard cannot carry such content (e.g., drafting constraints, consensus process, maintainability, scope limitations), or (ii) a reformulation of the conclusion as 'the currently planned horizontal de","section":"§6, 'Current standardisation directions'"},{"comment":"The proposed dynamic repository is central to the framework, but its institutional and legal feasibility is asserted rather than demonstrated. The authors state that ESOs would maintain a provider-fed repository of 'informative methods' and that this would 'reduce the interpretative burden, mitigate arbitrariness and address obsolescence,' but no evidence or pilot is provided. It is also unclear what legal status 'informative' contributions would have within the harmonised standards regime: if they are not part of the harmonised standard, it is not shown that they can confer presumption of conformity or meaningfully constrain providers' choices. The paper should either add a governance sketch—who verifies proposals, what conflict-of-interest rules apply, how updates are timed, how the repository interacts with the HAS assessment—or explicitly mark this as an open design question requirin","section":"§7.3, 'Dynamic repository of best practices'"}],"minor_comments":[{"comment":"Typo: 'berobust' should be 'be robust' in the abstract.","section":"Abstract"},{"comment":"The two-case comparison is illustrative but not controlled: the two studies differ simultaneously in task, data sources, model architectures, and perturbation types. This makes it hard to isolate the effect of any single 'contextual driver.' Please state explicitly that the table is intended as an illustration, not as evidence for the independent influence of each driver.","section":"§5.3, Table 1"},{"comment":"The manuscript uses 'horizontal' in at least two senses: the AI Act's horizontal approach (obligations across all AI systems) and horizontal standards (cross-sector technical documents). This is a natural distinction, but the paper would benefit from an explicit clarification early on, since the central argument depends on separating the legal scope of the Act from the content of standards.","section":"§6, terminology"},{"comment":"Reference [44] is 'Submitted to ICLR 2025,' which is not a stable citation. If the paper is under review, cite the arXiv version or omit the venue.","section":"References"},{"comment":"The discussion of benchmarks and sandboxes is interesting but somewhat diffuse. Consider condensing and clearly distinguishing the two mechanisms: benchmarks as ex-ante evaluation tools and sandboxes as ex-ante regulatory experimentation environments.","section":"§7.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a thoughtful policy analysis with a sound conceptual core, but the central 'cannot' claim is overstated relative to the evidence. The reviewer's main concern is not a stylistic quibble: it determines whether the paper's recommendation should be 'restructure the standards' or 'improve the current standards.' A major revision is appropriate. The authors can address it by softening the necessity claim and/or adding an institutional argument, and by acknowledging the repository proposal is a design hypothesis requiring piloting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: this is a solid, honest policy paper. It argues that robustness under the EU AI Act is context-sensitive and that current horizontal standardisation work under M/593 is unlikely to give providers workable assessment methods. That argument is well-grounded in the legislation and the standardisation request, and the paper is careful to note where it cannot see the future (its own line on ISO/IEC 24029-3: 'its scope and level of specificity remain to be seen'). The three-driver framework (use case, data, model) is a useful synthesis, and the comparison of two medical imaging robustness evaluations in Table 1 genuinely shows how the same broad task can require very different tests and metrics.\n\nThe genuinely new element is the proposal for a multi-layered architecture: horizontal principles, domain-specific standards, a lifecycle-mapped perturbation taxonomy, and a dynamic, provider-fed repository of practices and benchmarks. That package is not assembled this way elsewhere.\n\nWhere it wobbles: the necessity claim. The paper's wording is usually 'seems unlikely' or 'difficult to argue' rather than a hard 'cannot,' so the stress-test note that calls it categorical overshoots. But the underlying worry is fair. The paper never shows why a single horizontal standard could not carry domain-specific annexes or conditional specifications that satisfy M/593's demand for a range of technical options. What it actually establishes is that the current and planned JTC 21 drafts are not detailed enough yet. That supports 'improve the standards,' not strictly 'restructure the deliverables.' The distinction matters because the latter is a stronger institutional ask. A modest revision reframing the framework as one viable way to meet M/593, rather than the only way, would make the argument robust.\n\nThe other soft spot is the dynamic repository and the claim that it will reduce interpretative burden and arbitrariness. That is asserted as a design rationale, not demonstrated. The paper mentions similar efforts (OECD, AIME) but gives no pilot evidence or discussion of governance problems in a harmonised-standards regime. That's a feasibility gap, not a fatal one.\n\nOverall, the paper is worth a serious referee. It is accurate on the legal materials, honest about uncertainty, and the proposal is concrete enough to be actionable. For CEN/CENELEC JTC 21 participants, regulators, and AI governance researchers, this is a useful contribution. I'd suggest revision, not rejection.","headline":"Useful, honest standards paper; main soft spot is the step from 'current drafts are insufficient' to 'a layered architecture is required.'","tokens_in":14454,"tokens_out":3074,"would_cite":true,"duration_ms":32007,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Robustness of an AI system is context-dependent, so the EU AI Act's horizontal standards need to be supplemented with layered, domain-specific specifications.","keywords":["robustness","EU AI Act","standardisation","context-sensitivity","harmonised standards","perturbation taxonomy","performance metrics","AI lifecycle"],"falsifier":"Inspect the final published text of the next robustness assessment standard in the series (the part expected to cover adversarial robustness and distribution shift): if it contains detailed, domain-specific test procedures and metrics for multiple sectors, the claim that horizontal standards alone cannot fulfill the standardisation mandate loses its factual basis.","tokens_in":13531,"feed_emoji":"🧩","tokens_out":4641,"duration_ms":46470,"temperature":0.7,"pith_summary":"This paper claims that whether an AI system counts as robust is not a fixed property but depends on which performance aspects must remain stable, which perturbations it must withstand, and the operational environment. Because of this context-sensitivity, the generic“horizontal” standards being developed to implement the EU AI Act are unlikely to give providers enough concrete direction, leaving room for arbitrary choices. The paper proposes a multi-layered standardisation framework: horizontal standards set common principles, domain-specific standards identify risks across the AI lifecycle, and a dynamic repository lets providers share best practices, benchmarks, and new methods. If the argument holds, standardisation bodies should restructure their robustness work from horizontal coverage into a layered, context-sensitive architecture.","feed_headline":"Robustness is context-specific; AI standards need layers","feed_subtitle":"Horizontal AI rules are too vague; add domain-specific layers plus a living best-practice repository.","key_machinery":"The key machinery is the decomposition of robustness into “robustness of what” versus “robustness to what” and the three contextual drivers (use case, data, model). This decomposition converts an abstract property into an assessment pipeline: the drivers narrow down the relevant perturbation classes, which in turn determine the tests, metrics, and benchmarks to use. The proposed standardisation framework—horizontal layer, domain-specific layer, and a dynamic repository of practices—is the institutional translation of that pipeline.","core_discovery":"The paper's central claim is that robustness is a relational, context-dependent property, not an intrinsic one. It unpacks this into two dimensions—“robustness of what” (which performance metrics and baselines must remain stable) and “robustness to what” (which classes of perturbations, from adversarial attacks to natural distribution shifts, are relevant)—shaped by three contextual drivers: the use case (domain, task, deployment environment), the data (quantity, quality, type), and the model (learning paradigm, architecture, training configuration). From this it follows that a single horizontal standard cannot specify appropriate tests, metrics, or thresholds across all systems. The paper t","pith_inferences":["The same context-sensitivity argument likely applies to other horizontal AI Act requirements such as accuracy or cybersecurity, so the proposed layered architecture could serve as a template for those standards too.","The repository component raises governance questions the paper leaves open: who decides which “informative” methods are admitted, how they gain de facto authority without being mandatory, and how to avoid capture by well-resourced providers.","The paper's two medical-imaging examples show that even within one domain, perturbation types can differ drastically; a testable extension would be to build a small prototype taxonomy for one sector and check whether providers select different tests and thresholds than they do under the current generic standard."],"forward_implications":["If robustness is context-dependent, a single generic robustness standard either stays too vague to guide compliance or becomes prescriptive in ways that fail in many deployment contexts.","Standardisation bodies should extend the robustness work programme with vertical, domain-specific specifications rather than only adding another horizontal part.","Providers would need to document and justify their choices of perturbations, tests, and thresholds in light of the system's intended use, giving auditors a clear trail.","A shared, updated perturbation taxonomy mapped to lifecycle stages could reduce the interpretative burden and make conformity assessments more comparable across providers.","A dynamic repository of informative methods would let new techniques enter official guidance quickly, reducing the risk of standards becoming obsolete soon after publication."],"fun_headline_variants":["AI robustness is context-specific; layer the standards","EU AI Act: robust standards must be domain-specific","One robustness standard can't cover all AI contexts","Context-sensitive layering key to AI Act robustness"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument rests on the premise that the planned horizontal robustness standards will actually turn out to be too vague to provide workable guidance; the paper itself notes that the specificity of the upcoming robustness standard part remains to be seen.","fun_headline_variants_meta":{"raw":{"variants":["AI robustness is context-specific; layer the standards","EU AI Act: robust standards must be domain-specific","One robustness standard can't cover all AI contexts","Context-sensitive layering key to AI Act robustness"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00049,"raw_usage":{"total_tokens":2290,"prompt_tokens":827,"completion_tokens":1463,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":571,"completion_tokens_details":{"reasoning_tokens":1403}},"tokens_in":571,"tokens_out":1463,"duration_ms":11566,"temperature":1.0,"reasoning_tokens":1403,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:19:40.265461+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the final published text of the next robustness assessment standard in the series (the part expected to cover adversarial robustness and distribution shift): if it contains detailed, domain-specific test procedures and metrics for multiple sectors, the claim that horizontal standards alone cannot fulfill the standardisation mandate loses its factual basis.","supporting_citations":[],"review_version":1}