{"id":"63c49818-14af-4e91-8031-e858b5356c2f","arxiv_id":"2508.14118","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"AI hallucinations in medical devices are defined as plausible errors, either impactful or benign, to guide device evaluation.","lead":"This paper proposes a new definition of AI hallucinations in medical devices: plausible errors that can be impactful or benign. It reviews how such errors appear in imaging, synthetic data, and language models, and how to quantify and reduce them.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The definition of hallucination depends on a plausibility threshold tau that the paper declares unmeasured and out of scope, leaving the central classification axis non-operational.","rationale":"The reader's weakest-assumption analysis correctly identifies the operationalization of plausibility, and specifically the unmeasured threshold tau, as the main barrier to accepting the paper's central claim. I agree with that assessment. The paper is coherent in its conceptual structure: it distinguishes hallucinations from non-hallucination errors by plausibility, and then subdivides hallucinations by impact. The examples in imaging, synthetic data, and language models are useful and show that the authors have thought carefully about the phenomena. The paper also deserves credit for explicitly acknowledging the limitation rather than hiding it. However, the definition's practical value depends on the ability to label errors consistently in real evaluations. Without a procedure for estimating tau or specifying the reference observer, the same output could be classified differently by different users or tasks, which undermines the stated goal of facilitating device evaluation across product areas. The proposed reader study is a direct check of whether plausibility can be treated as a measurable, reproducible axis. Because the paper's own text declares threshold-determination studies out of scope, the verdict should remain CONDITIONAL rather than moving to ACCEPT or REJECT. The current conditional verdict matches the strength of the evidence: the conceptual definition is sound, but its central operational requirement is unfulfilled.","tokens_in":18802,"tokens_out":3013,"duration_ms":33515,"concrete_test":"Run a multi-reader study on a fixed set of AI reconstruction errors, such as the false bowel loops and plaque-like features in Figure 2(b), with radiologists, residents, and non-experts. Each reader rates plausibility on a continuous scale and classifies each error as a hallucination or a non-hallucination error; include conventional artifacts from Figure 2(a) as controls. Compute inter-rater agreement (e.g., Cohen's kappa) and examine whether plausibility ratings for known hallucinations separate from ratings for conventional artifacts across expertise levels. If kappa is below 0.4 or plausibility distributions overlap substantially, tau cannot serve as a stable practical classification threshold and the definition's practical claim fails; if agreement is high and separation is clear, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that hallucinations can be defined practically as 'plausible errors' (Section 1) and that device evaluation should classify errors along plausibility and impact axes. The load-bearing condition is that tau in Figure 1 can separate hallucinations from non-hallucination errors in real evaluations. The paper itself states tau is 'likely observer and task-specific' and that 'studies to determine this threshold... are outside of the scope of this work' (Section 1). Section 3 repeats that 'the determination of an artifact as a hallucination is difficult to formally define, as it relies on the artifact being plausible,' and the stability-based methods discussed 'do not provide the necessary measure of plausibility.' The hallucination index and related metrics measure distributional divergence, not the observer-relative plausibility that the definition requires. Because plausibility is never operationalized, the definition does not yet support the claimed practical and universal use: a given error can be a hallucination for a naive user and a non-hallucination error for an expert, and nothing in the framework fixes the reference observer. This is an openly acknowledged limitation rather than a hidden contradiction, but it is the precise point at which the central construction would need independent support to become usable in regulatory or clinical evaluation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript argues that the term 'hallucination' lacks a universally recognized definition in AI/ML-enabled medical devices and proposes one: hallucinations are a subset of errors, specifically errors that are plausible, and can be subdivided into impactful hallucinations and benign hallucinations. Non-hallucination errors are characterized by their obviousness and traceability to device artifacts or pre-specified failure modes. The paper applies this definition to three areas: imaging devices, synthetic image generation, and language and multimodal devices, and it reviews existing approaches for quantifying and mitigating hallucinations. It explicitly concedes that the plausibility threshold tau is observer- and task-specific and that studies to determine this threshold are outside the scope of the work.","tokens_in":18981,"tokens_out":4056,"duration_ms":42410,"significance":"If the proposed definition were operationalized, it would give device developers and regulators a common vocabulary across imaging, generative, and language domains, and it would direct evaluation toward plausibility and impact axes, which is a potentially valuable contribution. The paper's strengths include its candid acknowledgment of the unmeasured threshold, its use of concrete examples grounded in external empirical work, and its synthesis of stability-based quantification methods. The definition builds on Xu et al. without circularity, and the paper does not derive its central claim from itself. However, as presented, the central classification axis is non-operational, and the abstract's claim that the definition is 'practical and universal' is not yet supported, so substantial revision is needed before the paper can stand as a foundation for device evaluation.","major_comments":[{"comment":"The definition rests entirely on the plausibility threshold tau, yet the manuscript states that tau is 'likely observer and task-specific' and that studies to determine it are 'outside of the scope of this work.' Because no method, protocol, or reference observer is given to fix tau, the same erroneous output can be classified as a hallucination for one user and a non-hallucination error for another, and nothing in the framework decides between these classifications. This directly undercuts the abstract's claim that the definition is 'practical and universal' and is a load-bearing gap rather than a cosmetic one. The authors should either provide an operational protocol (for example, reader studies with a specified expertise level and a calibration procedure for tau) or explicitly revise the claim to present the definition as a conceptual framework pending empirical calibration.","section":"§1, Fig. 1"},{"comment":"The quantification section acknowledges that worst-case perturbation methods 'do not provide the necessary measure of plausibility' and that the hallucination index, which computes the Hellinger distance between ground-truth and reconstructed distributions, still leaves 'the cut-off that dichotomizes faithful and hallucinated reconstruction' nontrivial. Since the proposed taxonomy is defined by plausibility, none of the surveyed metrics operationalizes the definition's central construct. The manuscript should state clearly which of these metrics, if any, could serve as a proxy for plausibility and under what assumptions, or explain why the taxonomy does not require a metric to be useful in evaluation.","section":"§3"},{"comment":"The classification of the artifacts in Fig. 2(a) as non-hallucination errors rests on their 'obviousness' and 'traceability' to imaging-system limitations, but the manuscript offers no decision procedure for either property. This is the second axis of the dichotomy and is as underspecified as tau. The authors should specify observable criteria or a study design that would determine when an error is obvious or traceable, or state that the taxonomy is intended only as a retrospective classification.","section":"§2.1, Fig. 2"}],"minor_comments":[{"comment":"The sentence 'The driving force in the technological advancement of medical imaging has been less radiation‡ and saving scan time' places the footnote marker awkwardly; consider rewriting as 'lower radiation dose and shorter scan time' with the footnote attached to the relevant term.","section":"§2.1"},{"comment":"The caption states that 'Unmitigated impactful errors are colored yellow,' but the figure itself has no legend and the text does not define 'unmitigated' in this context; please add a legend or clarify the color scheme.","section":"Fig. 1 caption"},{"comment":"The phrase 'as previously mentioned in the section 2.1' should be 'as previously mentioned in Section 2.1,' and the manuscript should use consistent capitalization for figure references such as 'fig. 2' versus 'Fig. 2.'","section":"§3"},{"comment":"The court case references are not formatted consistently (for example, 'Ko v. li, Inc.' uses an uppercase 'I' in 'li'); please use a consistent legal citation style.","section":"References"},{"comment":"The phrase 'hallucinations must be expected' is italicized without explanation; if emphasis is intended, please state the reasoning in words, and otherwise remove the emphasis.","section":"§2.2.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a position paper from FDA authors and fits a journal such as Artificial Intelligence in the Life Sciences if framed as a perspective. The central operationalization gap is acknowledged by the authors, but it is load-bearing for the 'practical and universal' claim, so the revision should either add a concrete strategy for measuring plausibility or temper the claim. No circularity concerns arise from the citation pattern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper worth knowing about: an FDA-affiliated group proposes a definition of hallucination for AI-based medical devices. The novel part is modest but real: hallucinations are defined as a subset of errors—specifically, plausible errors—and split into impactful and benign subtypes. Non-hallucination errors are the obvious, traceable-to-hardware ones. That framing is useful for device evaluation, and the paper applies it coherently across imaging, synthetic-data generation, and LLM-based devices with well-chosen examples. The review is solid, the attribution is fair, and the prose is candid; the authors explicitly flag the main weakness.\n\nThe soft spot is exactly where the stress-test points. The load-bearing axis is plausibility, gated by a threshold tau that the paper admits is observer- and task-specific and provides no way to measure or estimate. Section 1 says studies to determine tau are out of scope; Section 3 repeats that plausibility is hard to formally define and that stability-based metrics don't capture it. Without an operationalized tau, the same error can be a hallucination for a naive user and a plain artifact for an expert, and the framework doesn't fix a reference observer. This keeps the definition from being the \"practical and universal\" tool the title claims. The authors acknowledge this, so it's an honest limitation rather than a hidden one, but it's a genuine gap: the central classification axis is not yet usable in regulatory or clinical evaluations.\n\nI don't think this is fatal. The conceptual contribution can be valuable even without tau nailed down—it organizes the discussion, gives regulators a shared language, and separates impact from plausibility in a way that existing frameworks don't. The citation pattern is healthy; self-citations are to prior work on hallucination in imaging and don't create a circular argument. The paper would be stronger if it framed itself as a position piece proposing a definition that needs empirical calibration, rather than a universal definition. If the authors do that, it's a reasonable candidate for a serious referee; conditional accept seems right. It's not a methods paper and doesn't deliver new measurements, so I wouldn't expect it to drive new results on its own, but it's a useful reference for anyone working on medical AI evaluation.\n\nI'd bring it to a reading group on AI safety/medical devices, though I wouldn't cite it in my own next paper. Send it out for review.","headline":"A useful position piece that adds plausibility and impact axes to the hallucination definition, but leaves the key threshold tau unmeasured and therefore stops short of its practical/universal claim.","tokens_in":19500,"tokens_out":2527,"would_cite":false,"duration_ms":25674,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes that a hallucination in a medical device is any output error that is plausible to the user, split into impactful and benign hallucinations.","keywords":["hallucinations","deep learning","artificial intelligence","generative models","medical imaging","plausibility","error taxonomy","device evaluation"],"falsifier":"A multi-reader study in which clinicians at different training levels independently classify the same set of medical-device errors as plausible versus obvious, and then repeat the task for different clinical tasks, would settle the point: if the placement of $\\tau$ varies so widely across readers or tasks that no stable boundary emerges, the proposed definition cannot ground a practical evaluation method.","tokens_in":18602,"feed_emoji":"🩺","tokens_out":6251,"duration_ms":59486,"temperature":0.7,"pith_summary":"The paper argues that the word 'hallucination', when applied to medical devices, should not mean 'any error made by an AI'. It proposes a definition: a hallucination is a plausible error, one that a relevant observer could mistake for a correct result, and it can be either impactful or benign. Errors that are obvious or traceable to a conventional device artifact are not hallucinations. The motivation is that plausible errors form a new risk vector: they can fool experts and bypass the guardrails built for conventional devices, so evaluation should classify errors along plausibility and impact rather than only counting errors. If accepted, the definition gives device developers and regulators a common language for measuring and mitigating this failure mode.","feed_headline":"AI device hallucinations: definition by plausibility and impact","feed_subtitle":"A practical taxonomy separates plausible AI errors from obvious artifacts to guide medical device evaluation.","key_machinery":"The load-bearing object is the plausibility-impact error diagram of Figure 1, organized around a plausibility axis with a threshold $\\tau$. An error above $\\tau$ is labeled a hallucination; an error below it is a non-hallucination error, characterized by obviousness and traceability to device artifacts or pre-specified failure modes. The paper stresses that $\\tau$ is a continuum and is observer- and task-specific, and it treats the user of the device, who may be an expert, a patient, or an algorithmic interpreter, as part of the definition. This axis does the work of separating hallucinations from conventional artifacts and explains why a model can produce fewer impactful errors yet still cause worse patient outcomes: plausible errors escape clinician intuition and existing risk mitigation.","core_discovery":"The central claim is that hallucination in a medical device is best defined as a subset of error: an error that is plausible to the intended user, with two subtypes, impactful and benign, and with a separate category of non-hallucination errors that are obvious and traceable to device artifacts or pre-specified failure modes. The paper grounds this in a ground-truth-function account in which hallucinations are an unavoidable property of data-driven models, and it shows the definition at work across imaging, synthetic image generation, and language and multimodal devices. Examples include AI super-resolution outputs that add bowel loops or plaque-like features absent from the reference, conditional generators that insert tumors or histological features that do not exist in the input, and language models that insert a diagnosis into a summary. The consequence is that device evaluation should measure where errors fall on plausibility and impact axes, and that a device producing fewer errors overall can still be more dangerous if its errors are plausible.","pith_inferences":["An implication the paper leaves implicit: plausibility thresholds could be measured empirically with reader studies using forced-choice judgments at different expertise levels, but the definition requires such studies to become operational.","A testable extension: task-based evaluation metrics for detection, quantification, and classification could convert the impactful-versus-benign distinction from a qualitative label into a measurable downstream performance change.","If the definition is right, one would expect regulatory and industry test reports to separate 'hallucination' counts from 'artifact' counts, and to specify the observer population used for plausibility judgments.","The claim that hallucinations cannot be fully removed suggests that certification criteria should focus on bounding the rate of impactful plausible errors within a task-specific tolerance rather than requiring error-free outputs."],"forward_implications":["Device evaluation would report errors classified along plausibility and impact, not only frequency; a model with fewer total errors could still be higher-risk if its errors are plausible.","A hallucination that is benign in the original task can become impactful if the output is later reused in patient-care decisions, so evaluation should track downstream use.","Stability measurements under small input perturbations can serve as a proxy for hallucination propensity, but they do not by themselves provide the plausibility information needed for classification.","Because hallucinations are argued to be intrinsic to neural-network methods, mitigation strategies such as null-space constraints, noise injection, retrieval augmentation, and conformal redaction can reduce but not eliminate them."],"supporting_citations":[{"why":"Supplies the ground-truth-function account that hallucinations are unavoidable, which the paper extends to a plausibility-based definition.","marker":"[76]"},{"why":"Provides an evidence-based definition of hallucination that the proposed definition aligns with.","marker":"[1]"},{"why":"Establishes the stability-versus-performance trade-off in neural network reconstruction that motivates measuring hallucinations through instability.","marker":"[26]"},{"why":"Gives the worst-case small-perturbation method used to demonstrate that reconstruction networks are unstable.","marker":"[27]"},{"why":"Introduces hallucination maps for isolating artifacts associated with imperfect priors in tomographic reconstruction.","marker":"[45]"},{"why":"Provides the hallucination index, a Hellinger-distance metric used to evaluate generative reconstruction models.","marker":"[123]"},{"why":"Documents distribution-matching losses hallucinating features in medical image translation, used as a conditional-generation example.","marker":"[63]"},{"why":"Offers survey evidence that plausibility of health information differs between patients and professionals, supporting the observer-dependence of the definition.","marker":"[2]"}],"fun_headline_variants":["Hallucination in medical AI: plausible errors, impactful or benign","Define medical AI hallucination: plausible errors, not artifacts","Plausibility determines if an AI error is a true hallucination","Medical AI hallucinations: plausible, impactful, or benign errors","Medical AI hallucination: plausible errors, not obvious glitches"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole taxonomy rests on the idea that plausibility can be treated as a measurable axis with a threshold $\\tau$, but the paper gives no procedure for estimating $\\tau$ and notes it is observer- and task-specific; if that boundary cannot be pinned down, the hallucination label cannot be applied consistently.","fun_headline_variants_meta":{"raw":{"variants":["Hallucination in medical AI: plausible errors, impactful or benign","Define medical AI hallucination: plausible errors, not artifacts","Plausibility determines if an AI error is a true hallucination","Medical AI hallucinations: plausible, impactful, or benign errors","Medical AI hallucination: plausible errors, not obvious glitches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000747,"raw_usage":{"total_tokens":3267,"prompt_tokens":825,"completion_tokens":2442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":441,"completion_tokens_details":{"reasoning_tokens":2356}},"tokens_in":441,"tokens_out":2442,"duration_ms":18134,"temperature":1.0,"reasoning_tokens":2356,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:16:45.916879+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A multi-reader study in which clinicians at different training levels independently classify the same set of medical-device errors as plausible versus obvious, and then repeat the task for different clinical tasks, would settle the point: if the placement of $\\tau$ varies so widely across readers or tasks that no stable boundary emerges, the proposed definition cannot ground a practical evaluation method.","supporting_citations":[{"cited_title":"Hallucination index: An image quality metric for generative reconstruction models,","cited_arxiv_id":null,"evidence_quote":"Provides the hallucination index, a Hellinger-distance metric used to evaluate generative reconstruction models."}],"review_version":2}