{"id":"d1d700c3-01d5-4a9f-8c2d-3d88c32946ba","arxiv_id":"2508.14119","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fabric is a public repository of 20 deployed AI systems with co-designed workflow diagrams, plus a four-level human oversight and five-level institutional oversight taxonomy.","lead":"Researchers interviewed 20 practitioners about live AI systems and drew workflow diagrams with them to show where human and institutional oversight actually sits. They released Fabric, a public repository of these 20 use cases with a simple classification of oversight patterns.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A pattern claim in §5.1 contradicts the paper's own Table 2: not all conditionally autonomous use cases follow only ad-hoc oversight; No. 17 and No. 19 have stricter institutional oversight. The central 'common patterns' analysis needs revalidation.","rationale":"The reader's weakest assumption was that practitioner self-reports may be incomplete or inaccurate, which the authors disclose. That concern is real but not the most decisive one for the paper's central contribution: the repository's 'common patterns of governance.' A more checkable and more damaging issue is internal: the paper's own pattern summary in Section 5.1 does not match the classifications it publishes in Table 2 and Appendix B. The conditionally autonomous use cases No. 17 and No. 19 are explicitly assigned stricter institutional oversight levels than 'ad-hoc practice' (organization policy, industry standard, regulation). Therefore the claimed pattern—'all conditionally autonomous AI use cases, except the financial one, follow ad-hoc practice'—is either false or so ambiguous as to be unverifiable. The same coding fragility appears in No. 13, where an 'Autonomous AI' label coexists with a workflow in which a healthcare provider reviews the AI-generated report; under the paper's own Table 1 definition, 'Autonomous AI' requires that no human can change the final output. This is not an attack on the qualitative method or the honesty of the authors; it is a request for the central pattern claims to be re-derived from the repository with explicit coding rules. The concrete test of independent re-coding would settle whether the mismatch is a real inconsistency or a wording artifact. If the mismatch stands, the paper's conclusions about governance patterns are not supported without revision. The repository itself can remain a useful contribution, so conditional acceptance is appropriate: the authors should correct the pattern claims, refine the definitions, and ideally publish an inter-coder reliability check or a transparent coding matrix.","tokens_in":23502,"tokens_out":6795,"duration_ms":71931,"concrete_test":"Take the 20 repository entries in Appendix B as the only data, and have two independent coders (blind to Section 4/5 labels) re-classify each use case using Table 1's definitions of 'final output' and oversight levels. Then (1) compute inter-coder agreement (e.g., Cohen's kappa) for the four human-oversight levels and the five institutional-oversight levels; and (2) check the exact Section 5.1 bullets against the reproduced Table 2, especially No. 17 and No. 19. If the reproduced classification places any stricter institutional level on No. 17 or No. 19, or if No. 13 is coded as anything other than Autonomous AI under a literal reading, the pattern claims must be revised or the definitions refined before the repository can support the stated conclusions.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing weakness is not the disclosed reliance on practitioner self-reports but an internal inconsistency in the pattern analysis that is checkable from the paper's own appendix. Section 5.1 states: 'All conditionally autonomous AI use cases, except for the more high-risk financial use case, follow the least strict level of institutional oversight: ad-hoc practice.' Appendix B/Table 2 contradict this. OriginTrail DKG (No. 17) is listed as Conditionally Autonomous AI with ad-hoc practice *and* organization policy; Call Center Virtual Assistant (No. 19) is Conditionally Autonomous AI with ad-hoc practice, organization best practice, organization policy, industry standard, and regulation. The only way to rescue the sentence is to read 'follow' as 'include at least ad-hoc,' but that reading is not what it says and would trivialize the claimed contrast with autonomous AI systems, all of which are said to follow regulation. Since Section 5.1's bullet points are the paper's headline 'patterns of oversight,' this mismatch means the central empirical claim is not currently supported by the repository as published. The same classification fragility appears in Mental Health Triage Tool No. 13, labeled Autonomous AI even though its output section says the report 'is reviewed by the healthcare provider with the patient'—a human may affect the final output under Table 1's definition. That suggests the coding of human-oversight levels is not robustly operationalized.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Fabric, a public repository of 20 deployed AI use cases gathered through semi-structured interviews and co-designed workflow diagrams. It defines four levels of human oversight (Autonomous AI, Conditionally Autonomous AI, Human-Approved AI, Human-Led with AI-Assistance) and five levels of institutional oversight (ad-hoc practice through regulation), assigns each use case to these categories, and reports cross-cutting 'patterns of oversight,' such as autonomous systems following regulation and conditionally autonomous systems following ad-hoc practice. The paper is explicitly qualitative and descriptive; it does not claim to measure governance effectiveness, and it discloses limits around selection bias, evolving interview questions, and reliance on practitioner self-reports.","tokens_in":23869,"tokens_out":7308,"duration_ms":77533,"significance":"If the classification and pattern claims hold, Fabric is a useful empirical resource: it provides workflow-level, cross-sectoral documentation of governance in real deployments, complements risk- and incident-focused repositories, and offers a taxonomy plus a practitioner questionnaire. The paper's strengths include the public repository, the co-design methodology with practitioner validation, and a candid limitations statement. The main weakness is internal consistency: a headline pattern in §5.1 is contradicted by the paper's own Table 2, and at least one autonomous-use-case classification is questionable. Because the pattern claims are the paper's central empirical contribution, these issues must be fixed before publication.","major_comments":[{"comment":"The bullet 'All conditionally autonomous AI use cases, except for the more high-risk financial use case, follow the least strict level of institutional oversight: ad-hoc practice' is contradicted by Table 2. OriginTrail DKG (No. 17) is listed with ad-hoc practice and organization policy; Call Center Virtual Assistant (No. 19) is listed with ad-hoc practice, organization best practice, organization policy, industry standard, and regulation. Only Personalized Feedback Assessor (No. 4) has ad-hoc practice alone. The sentence is only salvageable by reading 'follow' as 'include at least,' which is not what it says and would trivialize the claim. Since this bullet is a headline 'pattern of oversight,' the pattern analysis is not currently supported by the published repository.","section":"§5.1, Table 2"},{"comment":"The Mental Health Triage Tool No. 13 is classified as Autonomous AI, but its output description states 'The report is reviewed by the healthcare provider with the patient.' Under Table 1's definition, Autonomous AI means 'AI output is the final output' with no human able to change it. If a provider reviews the report with the patient, either the provider can influence the final output (making the classification at least Conditionally Autonomous), or the term 'reviewed' is being used in a non-standard way. This coding needs to be justified or corrected, because the 'all autonomous AI use cases follow regulation' pattern in §5.1 and §5.3 depends on the membership of this set.","section":"Appendix B, No. 13 and §4.3"},{"comment":"The paper's central contribution is the set of 'common patterns of oversight.' Because the taxonomy was induced from and then applied to the same 20 cases, the pattern statements are descriptive, not predictive, and therefore they must be exactly consistent with Table 2. The two mismatches above mean that the repository, as currently coded, does not support the headline patterns without re-analysis or re-coding. This is fixable, but it is load-bearing rather than cosmetic.","section":"§5.1–§5.3, overall pattern claims"}],"minor_comments":[{"comment":"The checkmark placement is visually ambiguous; rows for No. 8, No. 16, and No. 20 have variable numbers of checkmarks without clear column alignment. Please use explicit column entries (e.g., '✓' or '—') for every cell.","section":"Table 2"},{"comment":"'Credit Lending Classifier No. 6' appears to be a numbering error: No. 6 is Insurance Claims Classifier, while the credit lending use case is No. 7 (see §4.3 and Table 2).","section":"§4.2, Regulation subsection"},{"comment":"The text refers to 'Figure 1 in Appendix B' for the diagram legend, but the legend in Appendix B is Figure 5; Figure 1 is the main-text example workflow. Update the cross-reference.","section":"§3.2/Appendix B cross-reference"},{"comment":"In the first row, 'Is the AI output the final output?' appears to have a checkmark under Human-Approved AI, which conflicts with Table 1's definition that human approval is required before the output becomes final. Please clarify the questionnaire table or the checkmark alignment.","section":"Appendix C, Table 4"},{"comment":"There are typos in the conclusion: 'part ofy Fabric' and 'thefabric of our everyday life.' Please proofread the final section.","section":"§6, miscellaneous"}],"recommendation":"major_revision","confidential_remarks":"This is a useful qualitative resource, but the internal inconsistency in §5.1 versus Table 2 is exactly the kind of checkable error that must be resolved before the central empirical claim is published. The repository itself has value, and the fix is local, so I do not recommend rejection. However, if the authors are unwilling to re-code or substantially reword the pattern claims, the paper's main contribution would not be supported by its own appendix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Fabric gives the field something it did not have: a public repository of 20 deployed AI systems with co-designed workflow diagrams showing where human and institutional oversight actually sit, not just incident reports or policy frameworks. The interview methodology is clearly described, the authors disclose selection bias, evolving interview questions, and their inability to verify practitioner accounts, and they are explicit that this is an initial repository rather than an effectiveness study. That is honest, reproducible qualitative work and deserves credit.\n\nThe soft spot is the pattern analysis. Section 5.1 says all conditionally autonomous AI use cases except the high-risk financial one follow only the least strict institutional oversight, ad-hoc practice. Table 2 contradicts that. OriginTrail DKG (No. 17) is conditionally autonomous with both ad-hoc practice and organization policy; Call Center Virtual Assistant (No. 19) is conditionally autonomous with ad-hoc practice, organization best practice, organization policy, industry standard, and regulation. Only Personalized Feedback Assessor (No. 4) is ad-hoc-only. The only way to rescue the sentence is to read \"follow\" as \"include at least ad-hoc,\" which trivializes the claimed contrast with autonomous systems. Since that bullet is a headline empirical result, it needs to be revalidated and probably reworded to describe the distribution rather than a uniform pattern.\n\nThere is a smaller coding concern: Mental Health Triage Tool (No. 13) is labeled Autonomous AI, but its own workflow says the report is reviewed by the healthcare provider with the patient. That may be compatible if the review cannot change the output, but the paper's definition asks whether a human can affect the final output, so the example needs clarification or recoding. This is a minor issue compared to the §5.1 contradiction, but it suggests the human-oversight labels are not always robustly operationalized.\n\nThe self-report reliance and small sample are real limitations, but the authors flag them and they do not undermine the repository's value as a descriptive artifact. The repository itself is the contribution; the pattern claims are the part that is currently overextended.\n\nI would send this to peer review. A serious referee should push on Table 2 versus §5.1 and ask for either corrected pattern claims or a more careful statement of what the patterns mean. The article is worth engaging with, and with revision it could become a solid empirical reference for the AI governance community.","headline":"Fabric is a genuinely useful empirical artifact—the first workflow-level repository of deployed AI governance cases—but the headline pattern claims contain a checkable contradiction with the paper's own table and need rework before they can be cited.","tokens_in":24311,"tokens_out":2073,"would_cite":true,"duration_ms":26678,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces Fabric, a public repository of 20 deployed AI use cases, and claims that co-designed workflow diagrams reveal common patterns of human and institutional oversight.","keywords":["AI governance","human oversight","institutional oversight","deployed AI systems","workflow diagrams","semi-structured interviews","co-design","public repository"],"falsifier":"An independent audit of a sample of the 20 organizations that compared actual workflows, logs, or escalation behaviour against the co-designed diagrams and oversight labels would settle the central claim. If, for example, systems labelled autonomous turn out to route borderline cases to humans, or systems labelled human-approved turn out to be rubber-stamped without review, the governance patterns would not describe how oversight actually operates.","tokens_in":23444,"feed_emoji":"🗂️","tokens_out":8993,"duration_ms":85090,"temperature":0.7,"pith_summary":"This paper is trying to show what AI governance actually looks like when systems are in operation. It releases Fabric, a public repository built from semi-structured interviews with practitioners about 20 real deployed AI systems, each accompanied by a co-designed diagram of the AI workflow and its oversight checkpoints. From these cases, the paper derives two independent axes of governance: four levels of human oversight (autonomous AI, conditionally autonomous AI, human-approved AI, and human-led with AI assistance) and five levels of institutional oversight (ad-hoc practice up to regulation). The authors argue that this workflow-level, empirical view fills a gap left by incident databases and documentation standards, which capture risks and model details but not how governance is enacted in practice. If the repository is right, researchers and practitioners gain a concrete, cross-sectoral map of governance patterns and gaps they can build on.","feed_headline":"Mapped: 20 real AI deployments and their governance","feed_subtitle":"Interviews with practitioners yield a public repository of 20 use cases with workflow diagrams and oversight labels.","key_machinery":"The central object is the co-designed AI workflow diagram: a decision flowchart built together with the practitioner during the interview, showing inputs, AI components, decision points, and feedback loops. Working alongside it is a two-axis oversight taxonomy. The human oversight axis runs from autonomous AI (AI output is final) through conditionally autonomous AI, human-approved AI, to human-led with AI assistance; the institutional oversight axis runs from ad-hoc practice, organization best practice, organization policy, and industry standard to regulation. The diagrams anchor the taxonomy in concrete workflows, and the taxonomy gives the diagrams a comparative structure across sectors.","core_discovery":"The central claim is that a repository of 20 co-designed AI workflow diagrams, each annotated with human and institutional oversight labels, can surface common patterns in how deployed AI systems are governed. The paper reports that all four autonomous AI use cases in the sample fall under the strictest institutional oversight, regulation; that the most common human oversight pattern is human-led with AI assistance, where a person decides whether and how to use the AI output; and that conditionally autonomous systems tend to rely on the least formal institutional oversight, ad-hoc practice. These patterns are meant to be the beginning of an extendable resource, not a final verdict on governa","pith_inferences":["Because the 20 cases were recruited through personal and professional networks with snowball sampling, the observed patterns may shift once the repository includes a more representative set of organizations; the autonomous-AI-under-regulation correlation in particular should be treated as sample-bound until then.","The authors state they cannot audit practitioners' claims, so the repository is best read as a record of what organizations say they do; a verification study comparing diagrams against logs, audits, or direct observation would test whether the governance patterns hold up.","The two-axis taxonomy could be turned into a coding scheme for existing incident databases: applying the same oversight labels to documented failures would allow a direct test of whether certain governance patterns are associated with fewer or less severe incidents.","Re-interviewing the same practitioners over time and diffing the workflow diagrams would reveal how governance evolves as systems are updated, procured, or scaled."],"forward_implications":["Researchers can compare governance structure against system risk and sector using the repository's consistent labels and diagrams.","The oversight levels can be added to model cards and other system documentation, making governance visible before deployment rather than after the fact.","The paper's questionnaire gives practitioners a practical way to self-assess their human oversight level and start conversations about institutional oversight.","The observed patterns become testable hypotheses: for instance, that fully autonomous deployments only appear under regulatory pressure, and that lightly governed conditional autonomy is common in lower-risk settings.","As more use cases are contributed, the repository could support empirical study of which governance mechanisms are associated with better deployment outcomes."],"supporting_citations":[{"why":"Supplies the semi-structured interview method used to collect practitioner accounts of deployed AI workflows.","marker":"(Gubrium and Holstein 2002)"},{"why":"Supplies the co-design/prototyping methodology behind the joint construction of workflow diagrams with practitioners.","marker":"(Buchenau and Suri 2000)"},{"why":"Supplies the snowball sampling technique used to expand the practitioner interview pool beyond personal networks.","marker":"(Biernacki and Waldorf 1981)"},{"why":"Establishes the incident-cataloguing approach that Fabric contrasts with by documenting governance rather than failures.","marker":"(McGregor 2021)"},{"why":"Provides the AI risk repository and taxonomy that motivate the paper's focus on under-documented deployment-time governance.","marker":"(Slattery et al. 2024)"},{"why":"Model cards, cited as transparency documentation that stops short of capturing workflow-level oversight.","marker":"(Mitchell et al. 2019)"},{"why":"Datasheets for datasets, cited as another documentation practice that does not record how oversight is embedded in deployment.","marker":"(Gebru et al. 2021)"},{"why":"The EU AI Act, cited as the risk-level framework the paper uses to characterize many of its use cases as high-risk.","marker":"(European Parliament and Council of the European Union 2024)"},{"why":"Supports the claim that human-led oversight can act as a safeguard when humans decide whether to incorporate AI output.","marker":"(Steyvers et al. 2022)"}],"fun_headline_variants":["20 AI deployments mapped with oversight labels","Fabric repository: 20 use cases, governance gaps","AI governance patterns from 20 real deployments","Diagrams reveal AI oversight in practice","New repo maps human oversight of deployed AI"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The repository's patterns rest on the assumption that practitioners' descriptions of their deployed workflows and governance are truthful and complete, since the paper states it has no way to verify or audit the correctness of those claims.","fun_headline_variants_meta":{"raw":{"variants":["20 AI deployments mapped with oversight labels","Fabric repository: 20 use cases, governance gaps","AI governance patterns from 20 real deployments","Diagrams reveal AI oversight in practice","New repo maps human oversight of deployed AI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000141,"raw_usage":{"total_tokens":957,"prompt_tokens":657,"completion_tokens":300,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":401,"completion_tokens_details":{"reasoning_tokens":233}},"tokens_in":401,"tokens_out":300,"duration_ms":4328,"temperature":1.0,"reasoning_tokens":233,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:09:27.968683+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent audit of a sample of the 20 organizations that compared actual workflows, logs, or escalation behaviour against the co-designed diagrams and oversight labels would settle the central claim. If, for example, systems labelled autonomous turn out to route borderline cases to humans, or systems labelled human-approved turn out to be rubber-stamped without review, the governance patterns would not describe how oversight actually operates.","supporting_citations":[],"review_version":1}