{"id":"d459241e-63e3-43f6-94b8-17bd20490f9a","arxiv_id":"2502.07794","paper_version":1,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A perspective calling for adaptive, globally coordinated regulation of LLM-based medical devices, arguing that the existing total product life cycle approach is a poor fit.","lead":"This paper argues that current medical device regulations, designed around a total product life cycle, fit large language models poorly because their outputs are unpredictable and their uses are broad. It calls for global regulatory science collaboration, adaptive policies, and sandboxes to govern LLM-based health tools.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central premise that LLM outputs are intrinsically non-deterministic at temperature zero is asserted, not demonstrated; it is implementation-dependent and testable.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the Section 4 assertion that LLM outputs are non-deterministic in nature even at temperature zero. This premise matters because the paper argues that TPLC-style pre-market verification is fundamentally ill-suited to LLM-based devices; if deterministic inference is achievable by configuration, then at least part of that argument collapses into a weaker, implementation-dependent claim. However, the paper is a perspective and global call to action rather than a formal empirical or formal contribution. Even if the non-determinism sentence is overstated, much of the TPLC critique survives on other grounds: broad functionality, off-label use, data provenance, evaluation gaps, and post-market monitoring challenges. The recommended verdict therefore remains UNCHANGED, consistent with the reader's UNVERDICTED assessment, but the factual claim should be revised and supported before publication. The proposed experiment would settle whether the premise is true in controlled settings and would give the authors an evidence base for either the strong or the weak version of the claim.","tokens_in":12805,"tokens_out":2930,"duration_ms":29003,"concrete_test":"Run a controlled reproducibility experiment: for a representative open-weight LLM (e.g., Llama-3-70B) with temperature=0, top_p=1, fixed seed, and identical input/output schemas on a pinned software/hardware stack, generate 100 responses per prompt across 10 prompts and compute the exact-match rate. Then repeat across different batch sizes and GPU types. If exact-match is 100% on the fixed stack, the 'non-deterministic in nature' premise is false in the controlled regime; if it is <100%, the paper's claim about deployed behavior is supported. Report both controlled and cross-environment rates.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's core justification for adapting or moving beyond TPLC verification rests partly on the claim in the 'Ambiguity in Medical Device Definitions' section: 'Unlike predictive models, LLM outputs are non-deterministic in nature, even when the model temperature is set to zero.' This is stated without citation or measurement. In fact, greedy decoding at temperature zero with fixed infrastructure, fixed weights, fixed seed (where supported) can yield bit-identical outputs; observed API nondeterminism typically comes from batched execution, floating-point non-associativity, load balancing, or software updates, not from an intrinsic property of LLMs. The paper conflates 'non-deterministic in current deployments' with 'non-deterministic in nature.' The stronger claim is load-bearing because it is invoked to argue that pre-market verification of safety cannot be relied upon. The weaker, defensible claim—that outputs vary with prompts, contexts, model updates, and implementation choices—already supports much of the TPLC critique, but the sentence as written overstates the technological constraint. This should be corrected and supported, or qualified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This perspective argues that existing medical device regulatory frameworks, in particular the U.S. FDA's Total Product Life Cycle (TPLC) approach, are poorly suited to generative AI and large language models in healthcare. It identifies several classes of challenge: ambiguity in medical device definitions, lack of robust evaluation methods, difficulties in monitoring and enforcement, and ethical and equity concerns. It then proposes a global regulatory science agenda centered on adaptive regulatory approaches, regulatory sandboxes, supply-chain oversight, and international harmonization, with attention to health equity. The paper is a call to action rather than an empirical study, and its claims are supported primarily by cited literature and illustrative examples.","tokens_in":12922,"tokens_out":5374,"duration_ms":51643,"significance":"The essay is timely and well-referenced; it brings together recent evaluation studies (e.g., Hager et al.), reporting guidelines, and regulatory initiatives across the US, EU, UK, and Singapore. Its main value is as a synthesis and agenda-setting piece for regulatory science. It does not provide new measurements, formal models, or outcome data, and the central recommendation is programmatic. Nevertheless, the paper makes a useful contribution by framing the regulatory problem and identifying concrete gaps, such as predicate creep, data provenance, and supply-chain vulnerabilities, that are often discussed separately. If revised to qualify its stronger empirical claims, it would be a serviceable perspective for stimulating discussion at the intersection of AI, medicine, and regulation.","major_comments":[{"comment":"The sentence 'Unlike predictive models, LLM outputs are non-deterministic in nature, even when the model temperature is set to zero' is an overstatement that is not supported by the cited literature. With fixed weights, fixed infrastructure, and greedy decoding, LLM inference can be bit-identical across runs; observed API-level nondeterminism typically arises from batching, floating-point non-associativity, load balancing, or software updates. The paper should replace this with the weaker, defensible claim that LLM outputs are highly sensitive to prompts, contexts, model versions, and implementation choices, and support that claim with citations. This correction matters because the sentence is invoked to justify revising risk classification and control for LLM-based devices; however, the weaker claim is sufficient for the paper's TPLC critique, so the issue is locally fixable rather than fatal to the overall argument.","section":"Ambiguity in Medical Device Definitions"},{"comment":"The proposal for regulatory sandboxes and adaptive policies would be strengthened by specifying measurable success criteria and data-collection mechanisms. As written, the sandbox discussion relies on the OECD characterization and examples such as MHRA's AI Airlock and Singapore's IMDA sandbox, but it does not say how a sandbox's outcome would be evaluated, how results would generalize across jurisdictions, or what would count as failure triggering withdrawal of a policy. For a 'global call for action,' the absence of an evaluation design is a nontrivial gap that leaves the central recommendation untestable.","section":"Applying Adaptive Regulatory Approaches"}],"minor_comments":[{"comment":"In the paragraph on TPLC, 'are adaption' should be 'are adopting.'","section":"Introduction"},{"comment":"In the Europe row, 'provsions' should be 'provisions.'","section":"Table 1"},{"comment":"The heading misspells 'Health' as 'Heath.'","section":"Advancing the Collective Goals of Heath Equity"},{"comment":"In the LMIC paragraph, 'the authors note raised the need' should be 'the authors noted the need.'","section":"Future Directions for Regulatory Science and Regulators"},{"comment":"The phrase 'lends in to' should be 'lends itself to.'","section":"Beyond Medical Device Regulation: Responsible AI in Health Product Development"},{"comment":"The phrase 'challenges that that fall' should be 'challenges that fall.'","section":"Conclusion"},{"comment":"The phrase 'close to two-third' should be 'close to two-thirds.'","section":"Challenges to Monitoring and Regulatory Enforcement"},{"comment":"Reference formatting is inconsistent, with some entries using abbreviated author names (e.g., 'H-G, E., et al.' and 'Alan, B.'); a consistent author-year style would improve readability.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper contains a notable number of self-citations by members of the author group, but these are used for background claims about ethics, evaluation, and reporting rather than as the sole basis for the central argument. The manuscript is a perspective and should be evaluated as such; if the journal expects primary empirical or technical contributions, the fit may be a concern. The main technical issue, the overbroad non-determinism claim, is correctable without undermining the paper's overall message."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a policy perspective, not a research result. The genuinely new bit is the 'Software as a Medical Service' label and the push to bring health-services regulation and supply-chain thinking into the LLM-device conversation. Everything else—TPLC limits, EU AI Act, sandboxes, model cards, equity—is already in the cited literature, assembled here into a readable overview. If you need a bird's-eye map of why TPLC doesn't fit LLM-based devices, this is a decent place to start.\n\nWhat it does well: it's well-referenced, fairly balanced, and gives concrete examples (clinical scribes, 510(k) predicate creep, CrowdStrike). The healthcare AI supply chain figure and the global sandbox idea are useful framing devices. The group clearly knows the regulatory landscape, and the paper doesn't oversell any particular fix; it repeatedly calls for evidence and collaboration.\n\nThe soft spot is one load-bearing sentence in the 'Ambiguity in Medical Device Definitions' section: 'Unlike predictive models, LLM outputs are non-deterministic in nature, even when the model temperature is set to zero.' That overstates the technical reality. Greedy decoding at zero temperature can be bit-reproducible under fixed infrastructure; the observed nondeterminism comes from batching, floating-point non-associativity, load balancing, or updates. That is implementation variability, not an intrinsic property. The paper's core TPLC critique doesn't need the strong claim—the weaker point that outputs vary with prompts, contexts, and model versions already undermines static verification. The sentence should be qualified, or supported with a citation. A conscientious referee would ask for that.\n\nOther issues are minor: the 'global call to action' lists aspirations without pilots or outcome metrics, but that's normal for a perspective. There are a few typos ('adaption,' 'that that,' 'Heath Equity') and some repetition. Self-citations are present but mostly for background claims; I don't see a circle.\n\nBottom line: this deserves peer review as a commentary. It is not a research paper, and shouldn't be judged as one. With a clarifying revision on nondeterminism and some tightening, it would be a useful contribution to the regulatory science literature. I'd send it to a serious referee rather than desk reject.","headline":"A solid regulatory perspective whose technical premise overreaches; the TPLC argument survives once 'non-deterministic' is qualified.","tokens_in":13541,"tokens_out":2683,"would_cite":false,"duration_ms":24926,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the total product life cycle framework cannot adequately regulate LLM-based medical devices.","keywords":["generative AI","large language models","medical device regulation","total product life cycle","regulatory sandboxes","adaptive regulation","health equity","regulatory science"],"falsifier":"Run the same medically relevant prompt through the same model hundreds of times with temperature set to zero and identical sampling configuration, and measure the distribution of outputs; if outputs are identical across runs and stable across controlled implementations, the premise of intrinsic non-determinism would be falsified, and the case for abandoning TPLC-style verification would weaken.","tokens_in":12566,"feed_emoji":"⚖️","tokens_out":8953,"duration_ms":78272,"temperature":0.7,"pith_summary":"This perspective paper argues that the total product life cycle (TPLC) framework—the standard regulatory structure that follows a medical device from design through post-market monitoring—cannot adequately govern LLM-based medical devices. The authors identify three structural reasons: LLM outputs are non-deterministic even when temperature is set to zero; LLM functions are broad and generalist, making intended-use classification ambiguous; and LLM systems are complex composites whose behavior shifts with prompts, model versions, and third-party modifications. Because of this, premarket verification cannot certify safety at a single time point, and automated post-market monitoring cannot capture long-form outputs. The paper concludes that regulators should move toward adaptive policies, regulatory sandboxes, international harmonization, and supply-chain oversight, developed through global multidisciplinary regulatory science. A sympathetic reader would care because the current regulatory toolkit was built for deterministic, single-purpose software, and the safe deployment of generative AI in medicine depends on whether that toolkit is replaced in time.","feed_headline":"Total product life cycle regulation cannot tame LLM medical devices","feed_subtitle":"Non-deterministic outputs and broad functions demand adaptive, globally harmonized rules, the paper argues.","key_machinery":"The central object is the total product life cycle (TPLC) framework, a regulatory structure that tracks a software medical device from planning and design through verification, deployment, and post-market monitoring. The paper uses TPLC as the baseline that LLM-based devices fail, and the load-bearing property is the assertion that LLM outputs are non-deterministic even when the temperature is set to zero. That property is what makes one-time verification insufficient and continuous monitoring necessary but hard. The proposed replacement mechanism is the regulatory sandbox—a constrained real-world environment in which developers and regulators test policies iteratively before full approval—paired with adaptive policies that tighten or loosen restrictions as real-world evidence accumulates. The machinery also includes international harmonization of evaluation metrics and data provenance standards as the coordinating layer.","core_discovery":"The central claim is that TPLC, as currently practiced, is the wrong instrument for LLM-based medical devices. The paper's argument turns on three properties of LLMs: outputs are non-deterministic by nature even at zero temperature; functionality is not tied to a single intended use but spans summarization, diagnosis suggestions, and documentation; and integration is layered, so fine-tuning, retrieval-augmented generation, prompt variation, and base-model updates all change behavior after approval. Each of these properties breaks a different phase of the TPLC: verification and validation cannot be completed once, operation and monitoring cannot be automated for free-form outputs, and classification and enforcement cannot rely on intended-use definitions or predicate-based clearance. The paper's positive thesis is that regulatory science must innovate through adaptive regulation and regulatory sandboxes, with global harmonization of standards and deliberate attention to health equity, so that governance can be tested and revised in real-world settings rather than fixed at approval.","pith_inferences":["If non-determinism is the true crux, then a direct measurement campaign—running identical prompts on identical models with temperature zero, fixed seeds, and controlled decoding—would either confirm or undercut the paper's central premise, which is asserted rather than demonstrated.","The same TPLC mismatch likely extends to high-stakes LLM deployment outside medicine, such as legal advice or financial decisions, where a single approval-time evaluation cannot guarantee behavior across future prompts and versions.","Regulatory sandboxes could be designed as comparative experiments across jurisdictions, publishing outcomes with common metrics so that different governance policies can be evaluated against each other.","The paper's observation that product and service regulation are converging in health suggests a broader shift: as LLM-based agents deliver services rather than fixed functions, product-centric regulation will increasingly need service-centric oversight."],"forward_implications":["Regulators would need to treat LLM-based medical devices as a distinct regulatory category rather than fitting them into existing single-purpose software pathways.","Accelerated approval routes based on substantial equivalence to predicate devices would have to be revisited, because fine-tuning and retrieval-augmented generation can change a model's risk profile without a new submission.","Post-market surveillance of LLM tools would require human-in-the-loop review of long-form outputs and new incident-reporting mechanisms, since automated drift detection cannot capture hallucinations or redaction errors.","Global harmonization of evaluation metrics, dataset-provenance standards, and risk classification would become a prerequisite for managing cross-border deployment of LLM medical tools.","Regulatory sandboxes and adaptive policies would become standard tools for generating real-world evidence before and after approval."],"supporting_citations":[{"why":"Supplies the total product life cycle framing that the paper argues is insufficient for LLM-based devices.","marker":"[1,2]"},{"why":"The advisory committee executive summary that makes TPLC the bedrock recommendation for regulating generative AI-enabled devices, which the paper challenges.","marker":"[3]"},{"why":"Prior analysis of ethical and regulatory challenges of large language models in medicine, establishing the structural differences from approved AI technologies.","marker":"[5]"},{"why":"Real-world evaluation study showing LLMs underperform clinicians on diagnosis and guideline adherence, supporting the lack of robust performance evaluation.","marker":"[7]"},{"why":"Large-scale audit of dataset licensing and attribution documenting frequent license omissions, supporting the data-provenance challenge.","marker":"[11]"},{"why":"Comparative analysis showing most AI-based devices are cleared via accelerated predicate-based pathways, supporting the enforcement concern.","marker":"[18]"},{"why":"Evidence of predicate creep and recall risks in accelerated clearance pathways, supporting the claim that LLM devices could be cleared against non-LLM predicates.","marker":"[20,21]"},{"why":"Case study of adaptive regulation for advanced biotherapeutics, providing the template for adaptive approaches the paper endorses.","marker":"[39]"},{"why":"Example of an international regulators' forum harmonizing medical device regulation, cited as the model for global coordination.","marker":"[57]"}],"fun_headline_variants":["TPLC can't regulate non-deterministic LLM devices","LLM medical devices demand adaptive global rules","Old device rules fail new LLM medicine","Regulatory rethink needed for LLM health tools"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument leans on the claim that LLM outputs are inherently non-deterministic even when the model temperature is set to zero, so that no fixed verification can pin down their behavior; this is asserted rather than measured.","fun_headline_variants_meta":{"raw":{"variants":["TPLC can't regulate non-deterministic LLM devices","LLM medical devices demand adaptive global rules","Old device rules fail new LLM medicine","Regulatory rethink needed for LLM health tools"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000344,"raw_usage":{"total_tokens":1884,"prompt_tokens":934,"completion_tokens":950,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":550,"completion_tokens_details":{"reasoning_tokens":890}},"tokens_in":550,"tokens_out":950,"duration_ms":6859,"temperature":1.0,"reasoning_tokens":890,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:55:19.940167+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same medically relevant prompt through the same model hundreds of times with temperature set to zero and identical sampling configuration, and measure the distribution of outputs; if outputs are identical across runs and stable across controlled implementations, the premise of intrinsic non-determinism would be falsified, and the case for abandoning TPLC-style verification would weaken.","supporting_citations":[{"cited_title":"Total Product Lifecycle Considerations for Generative AI- Enabled Devices","cited_arxiv_id":null,"evidence_quote":"The advisory committee executive summary that makes TPLC the bedrock recommendation for regulating generative AI-enabled devices, which the paper challenges."},{"cited_title":"Ethical and regulatory challenges of large language models in medicine","cited_arxiv_id":null,"evidence_quote":"Prior analysis of ethical and regulatory challenges of large language models in medicine, establishing the structural differences from approved AI technologies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Real-world evaluation study showing LLMs underperform clinicians on diagnosis and guideline adherence, supporting the lack of robust performance evaluation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Large-scale audit of dataset licensing and attribution documenting frequent license omissions, supporting the data-provenance challenge."},{"cited_title":"& Kerstin N, V","cited_arxiv_id":null,"evidence_quote":"Comparative analysis showing most AI-based devices are cleared via accelerated predicate-based pathways, supporting the enforcement concern."},{"cited_title":"& Farid, S.S","cited_arxiv_id":null,"evidence_quote":"Case study of adaptive regulation for advanced biotherapeutics, providing the template for adaptive approaches the paper endorses."},{"cited_title":"Working Groups: Artificial Intelligence/Machine Learning-enabled","cited_arxiv_id":null,"evidence_quote":"Example of an international regulators' forum harmonizing medical device regulation, cited as the model for global coordination."}],"review_version":1}