Pith. sign in

REVIEW 3 major objections 5 minor 140 references

Hallucinations in medical devices

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This paper proposes that a hallucination in a medical device is any output error that is plausible to the user, split into impactful and benign hallucinations.

desk verdict A useful position piece that adds plausibility and impact axes to the hallucination definition, but leaves the key threshold tau unmeasured and therefore stops short of its practical/universal claim. read the letter →

arxiv 2508.14118 v1 pith:MKNQSWD5 submitted 2025-08-18 eess.IV cs.CV

classification eess.IVcs.CV
keywords hallucinationsdeeplearningartificialintelligencegenerativemodelsmedicalimagingplausibilityerrortaxonomydeviceevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the word 'hallucination', when applied to medical devices, should not mean 'any error made by an AI'. It proposes a definition: a hallucination is a plausible error, one that a relevant observer could mistake for a correct result, and it can be either impactful or benign. Errors that are obvious or traceable to a conventional device artifact are not hallucinations. The motivation is that plausible errors form a new risk vector: they can fool experts and bypass the guardrails built for conventional devices, so evaluation should classify errors along plausibility and impact rather than only counting errors. If accepted, the definition gives device developers and regulators a common language for measuring and mitigating this failure mode.

What carries the argument

The load-bearing object is the plausibility-impact error diagram of Figure 1, organized around a plausibility axis with a threshold $\tau$. An error above $\tau$ is labeled a hallucination; an error below it is a non-hallucination error, characterized by obviousness and traceability to device artifacts or pre-specified failure modes. The paper stresses that $\tau$ is a continuum and is observer- and task-specific, and it treats the user of the device, who may be an expert, a patient, or an algorithmic interpreter, as part of the definition. This axis does the work of separating hallucinations from conventional artifacts and explains why a model can produce fewer impactful errors yet still cause worse patient outcomes: plausible errors escape clinician intuition and existing risk mitigation.

What would settle it

A multi-reader study in which clinicians at different training levels independently classify the same set of medical-device errors as plausible versus obvious, and then repeat the task for different clinical tasks, would settle the point: if the placement of $\tau$ varies so widely across readers or tasks that no stable boundary emerges, the proposed definition cannot ground a practical evaluation method.

Watch

Extended reading notes

Core claim

The central claim is that hallucination in a medical device is best defined as a subset of error: an error that is plausible to the intended user, with two subtypes, impactful and benign, and with a separate category of non-hallucination errors that are obvious and traceable to device artifacts or pre-specified failure modes. The paper grounds this in a ground-truth-function account in which hallucinations are an unavoidable property of data-driven models, and it shows the definition at work across imaging, synthetic image generation, and language and multimodal devices. Examples include AI super-resolution outputs that add bowel loops or plaque-like features absent from the reference, conditional generators that insert tumors or histological features that do not exist in the input, and language models that insert a diagnosis into a summary. The consequence is that device evaluation should measure where errors fall on plausibility and impact axes, and that a device producing fewer errors overall can still be more dangerous if its errors are plausible.

Load-bearing premise

The whole taxonomy rests on the idea that plausibility can be treated as a measurable axis with a threshold $\tau$, but the paper gives no procedure for estimating $\tau$ and notes it is observer- and task-specific; if that boundary cannot be pinned down, the hallucination label cannot be applied consistently.

Editorial extensions

If this is right

  • Device evaluation would report errors classified along plausibility and impact, not only frequency; a model with fewer total errors could still be higher-risk if its errors are plausible.
  • A hallucination that is benign in the original task can become impactful if the output is later reused in patient-care decisions, so evaluation should track downstream use.
  • Stability measurements under small input perturbations can serve as a proxy for hallucination propensity, but they do not by themselves provide the plausibility information needed for classification.
  • Because hallucinations are argued to be intrinsic to neural-network methods, mitigation strategies such as null-space constraints, noise injection, retrieval augmentation, and conformal redaction can reduce but not eliminate them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit: plausibility thresholds could be measured empirically with reader studies using forced-choice judgments at different expertise levels, but the definition requires such studies to become operational.
  • A testable extension: task-based evaluation metrics for detection, quantification, and classification could convert the impactful-versus-benign distinction from a qualitative label into a measurable downstream performance change.
  • If the definition is right, one would expect regulatory and industry test reports to separate 'hallucination' counts from 'artifact' counts, and to specify the observer population used for plausibility judgments.
  • The claim that hallucinations cannot be fully removed suggests that certification criteria should focus on bounding the rate of impactful plausible errors within a task-specific tolerance rather than requiring error-free outputs.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript argues that the term 'hallucination' lacks a universally recognized definition in AI/ML-enabled medical devices and proposes one: hallucinations are a subset of errors, specifically errors that are plausible, and can be subdivided into impactful hallucinations and benign hallucinations. Non-hallucination errors are characterized by their obviousness and traceability to device artifacts or pre-specified failure modes. The paper applies this definition to three areas: imaging devices, synthetic image generation, and language and multimodal devices, and it reviews existing approaches for quantifying and mitigating hallucinations. It explicitly concedes that the plausibility threshold tau is observer- and task-specific and that studies to determine this threshold are outside the scope of the work.

Significance. If the proposed definition were operationalized, it would give device developers and regulators a common vocabulary across imaging, generative, and language domains, and it would direct evaluation toward plausibility and impact axes, which is a potentially valuable contribution. The paper's strengths include its candid acknowledgment of the unmeasured threshold, its use of concrete examples grounded in external empirical work, and its synthesis of stability-based quantification methods. The definition builds on Xu et al. without circularity, and the paper does not derive its central claim from itself. However, as presented, the central classification axis is non-operational, and the abstract's claim that the definition is 'practical and universal' is not yet supported, so substantial revision is needed before the paper can stand as a foundation for device evaluation.

major comments (3)
  1. [§1, Fig. 1] The definition rests entirely on the plausibility threshold tau, yet the manuscript states that tau is 'likely observer and task-specific' and that studies to determine it are 'outside of the scope of this work.' Because no method, protocol, or reference observer is given to fix tau, the same erroneous output can be classified as a hallucination for one user and a non-hallucination error for another, and nothing in the framework decides between these classifications. This directly undercuts the abstract's claim that the definition is 'practical and universal' and is a load-bearing gap rather than a cosmetic one. The authors should either provide an operational protocol (for example, reader studies with a specified expertise level and a calibration procedure for tau) or explicitly revise the claim to present the definition as a conceptual framework pending empirical calibration.
  2. [§3] The quantification section acknowledges that worst-case perturbation methods 'do not provide the necessary measure of plausibility' and that the hallucination index, which computes the Hellinger distance between ground-truth and reconstructed distributions, still leaves 'the cut-off that dichotomizes faithful and hallucinated reconstruction' nontrivial. Since the proposed taxonomy is defined by plausibility, none of the surveyed metrics operationalizes the definition's central construct. The manuscript should state clearly which of these metrics, if any, could serve as a proxy for plausibility and under what assumptions, or explain why the taxonomy does not require a metric to be useful in evaluation.
  3. [§2.1, Fig. 2] The classification of the artifacts in Fig. 2(a) as non-hallucination errors rests on their 'obviousness' and 'traceability' to imaging-system limitations, but the manuscript offers no decision procedure for either property. This is the second axis of the dichotomy and is as underspecified as tau. The authors should specify observable criteria or a study design that would determine when an error is obvious or traceable, or state that the taxonomy is intended only as a retrospective classification.
minor comments (5)
  1. [§2.1] The sentence 'The driving force in the technological advancement of medical imaging has been less radiation‡ and saving scan time' places the footnote marker awkwardly; consider rewriting as 'lower radiation dose and shorter scan time' with the footnote attached to the relevant term.
  2. [Fig. 1 caption] The caption states that 'Unmitigated impactful errors are colored yellow,' but the figure itself has no legend and the text does not define 'unmitigated' in this context; please add a legend or clarify the color scheme.
  3. [§3] The phrase 'as previously mentioned in the section 2.1' should be 'as previously mentioned in Section 2.1,' and the manuscript should use consistent capitalization for figure references such as 'fig. 2' versus 'Fig. 2.'
  4. [References] The court case references are not formatted consistently (for example, 'Ko v. li, Inc.' uses an uppercase 'I' in 'li'); please use a consistent legal citation style.
  5. [§2.2.2] The phrase 'hallucinations must be expected' is italicized without explanation; if emphasis is intended, please state the reasoning in words, and otherwise remove the emphasis.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the proposed definition is a stipulative taxonomy grounded in external frameworks, and no prediction is derived from its own inputs.

full rationale

The paper's central contribution is a proposed definition of hallucinations in medical devices as plausible errors that are either impactful or benign, building on the external theoretical framework of Xu et al. and on empirical studies of imaging, synthetic data, and language models. Because the claim is a definitional taxonomy rather than a derived quantitative result, there is no derivation chain that could collapse into its inputs. The plausibility threshold tau in Figure 1 is explicitly acknowledged as observer- and task-specific and as not determined in this work; that is an acknowledged operational limitation, not a circular reduction, since the definition does not pretend to measure tau from itself. Self-citations to the authors' prior empirical studies (e.g., Deshpande et al. on spatial context and generative model evaluation, and Badano et al. on shape artifacts) are used as supporting examples or as sources for detection methods, not as justifications of the definition's validity. No fitted parameter is renamed as a prediction, no uniqueness theorem from the authors' prior work is invoked, and no ansatz is smuggled in via citation. The examples in Section 2 are applications of the stated definition to concrete cases, not proofs that the definition is true, so the argument is self-contained in the sense required for circularity analysis. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The paper introduces no physical entities. Its load-bearing assumptions are the existence of a ground truth function, the information-loss principle in imaging, the inevitability of hallucinations, and the plausibility-based dichotomy between error types. The only free parameter is the plausibility threshold tau, which is left unmeasured.

free parameters (1)
  • Plausibility threshold tau
    Introduced in Figure 1 and Section 1 as the cut-off for labeling an error a hallucination. The paper states it is observer and task-specific and does not provide a method to estimate it (Section 3).
assumptions (4)
  • domain assumption A ground truth function exists and hallucinations are discrepancies between generated outputs and that function (Xu et al. framework).
    Adopted from Xu et al. [76] as the theoretical basis for the proposed definition, but not proven in this paper (Section 1).
  • domain assumption Information is always lost during the imaging process, and no post-processing can recover diagnostic details that were not measured.
    Used in Section 2.1 to argue that AI reconstructions compensate with data priors, leading to hallucinations.
  • domain assumption Hallucinations are intrinsic to neural network-based methods and cannot be fully removed.
    Presented in the Summary and Section 3, relying on the theoretical result of Xu et al. rather than a new derivation.
  • domain assumption Conventional device errors are obvious and traceable to device artifacts or pre-specified failure modes, while hallucinations are plausible and not traceable in the same way.
    Underpins the dichotomy in Section 1 between hallucination and non-hallucination errors; asserted without systematic evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Hallucinations in medical devices." pith.science (2026). https://pith.science/paper/MKNQSWD5

@misc{pith2026250814118,
  author       = {Pith},
  title        = {Pith review of: Hallucinations in medical devices},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MKNQSWD5}},
  note         = {Machine review of arXiv:2508.14118}
}
read the original abstract

Computer methods in medical devices are frequently imperfect and are known to produce errors in clinical or diagnostic tasks. However, when deep learning and data-based approaches yield output that exhibit errors, the devices are frequently said to hallucinate. Drawing from theoretical developments and empirical studies in multiple medical device areas, we introduce a practical and universal definition that denotes hallucinations as a type of error that is plausible and can be either impactful or benign to the task at hand. The definition aims at facilitating the evaluation of medical devices that suffer from hallucinations across product areas. Using examples from imaging and non-imaging applications, we explore how the proposed definition relates to evaluation methodologies and discuss existing approaches for minimizing the prevalence of hallucinations.

Figures

Figures reproduced from arXiv: 2508.14118 by the authors.

Figure 1
Figure 1. Mock diagram of errors from a conventional and AI-enabled device, plotted [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. An illustration of artifacts that are readily discernible in (a) and non [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

140 extracted references · 44 canonical work pages

  1. [1]

    Measuring short-form factuality in large language models,

    J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus, “Measuring short-form factuality in large language models,”arXiv preprint arXiv:2411.04368, 2024

  2. [2]

    Health online 2013,

    S. Fox and M. Duggan, “Health online 2013,” tech. rep., Pew Research Center, Washington, D.C., Jan. 2013. CONTENTS13

  3. [3]

    No. 54 Civ. 1461,

    Mata v. Avianca, Inc., “No. 54 Civ. 1461,” June 2023. https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts. nysd.575368.54.0_3.pdf

  4. [4]

    ONSC 2766,

    Ko v. li, Inc., “ONSC 2766,” May 2025. https://www.canlii.org/en/on/onsc/doc/2025/2025onsc2766/2025onsc2766.html

  5. [5]

    No. 2:24-cv-05205-FMO-MAA,

    Lacey v. State Farm, “No. 2:24-cv-05205-FMO-MAA,” May 2025. https://www.lawnext.com/wp-content/uploads/2025/05/C.D.-Cal. -24-cv-05205-dckt-000119_000-filed-2025-05-06.pdf

  6. [6]

    The impact of AI errors in a human- in-the-loop process,

    U. Agudo, K. G. Liberal, M. Arrese, and H. Matute, “The impact of AI errors in a human- in-the-loop process,”Cognitive Research: Principles and Implications, vol. 9, no. 1, p. 1, 2024

  7. [7]

    Quantifying the impact of AI recommendations with explanations on prescription decision making,

    M. Nagendran, P. Festor, M. Komorowski, A. C. Gordon, and A. A. Faisal, “Quantifying the impact of AI recommendations with explanations on prescription decision making,”NPJ Digital Medicine, vol. 6, no. 1, p. 206, 2023

  8. [8]

    How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection,

    M. Jacobs, M. F. Pradier, T. H. McCoy Jr, R. H. Perlis, F. Doshi-Velez, and K. Z. Gajos, “How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection,”Translational psychiatry, vol. 11, no. 1, p. 108, 2021

Show all 140 references
  1. [9]

    Humans inherit artificial intelligence biases,

    L. Vicente and H. Matute, “Humans inherit artificial intelligence biases,”Scientific Reports, vol. 13, no. 1, p. 15737, 2023

  2. [10]

    Medical hallucination in foundation models and their impact on healthcare,

    Y. Kim, H. Jeong, S. Chen, S. S. Li, M. Lu, K. Alhamoud, J. Mun, C. Grau, M. Jung, R. R. Gameiro,et al., “Medical hallucination in foundation models and their impact on healthcare,”medRxiv, pp. 2025–02, 2025

  3. [11]

    Solving inverse problems using data- driven models,

    S. Arridge, P. Maass, O. ¨Oktem, and C.-B. Sch¨ onlieb, “Solving inverse problems using data- driven models,”Acta Numerica, vol. 28, pp. 1–174, 2019

  4. [12]

    Deep magnetic resonance image reconstruction: Inverse problems meet neural networks,

    D. Liang, J. Cheng, Z. Ke, and L. Ying, “Deep magnetic resonance image reconstruction: Inverse problems meet neural networks,”IEEE Signal Processing Magazine, vol. 37, no. 1, pp. 141–151, 2020

  5. [13]

    Convolutional neural networks for inverse problems in imaging: A review,

    M. T. McCann, K. H. Jin, and M. Unser, “Convolutional neural networks for inverse problems in imaging: A review,”IEEE Signal Processing Magazine, vol. 34, no. 6, pp. 85–95, 2017

  6. [14]

    Deep learning techniques for inverse problems in imaging,

    G. Ongie, A. Jalal, C. A. Metzler, R. G. Baraniuk, A. G. Dimakis, and R. Willett, “Deep learning techniques for inverse problems in imaging,”IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 1, pp. 39–56, 2020

  7. [15]

    Deep learning for tomographic image reconstruction,

    G. Wang, J. C. Ye, and B. De Man, “Deep learning for tomographic image reconstruction,” Nature machine intelligence, vol. 2, no. 12, pp. 737–748, 2020

  8. [16]

    Deep learning for pet image reconstruction,

    A. J. Reader, G. Corda, A. Mehranian, C. da Costa-Luis, S. Ellis, and J. A. Schnabel, “Deep learning for pet image reconstruction,”IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 5, no. 1, pp. 1–25, 2020

  9. [17]

    Image reconstruction is a new frontier of machine learning,

    G. Wang, J. C. Ye, K. Mueller, and J. A. Fessler, “Image reconstruction is a new frontier of machine learning,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1289–1296, 2018

  10. [18]

    Null-space smoothing of tomographic images using tv norm minimization,

    B. Smith, “Null-space smoothing of tomographic images using tv norm minimization,” in2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and Room-Temperature Semiconductor Detector Workshop (NSS/MIC/RTSD), pp. 1–4, IEEE, 2016

  11. [19]

    Null space and resolution in dynamic computerized tomography,

    B. N. Hahn, “Null space and resolution in dynamic computerized tomography,”Inverse Problems, vol. 32, no. 2, p. 025006, 2016

  12. [20]

    Deep learning-guided image reconstruction from incomplete data,

    B. Kelly, T. P. Matthews, and M. A. Anastasio, “Deep learning-guided image reconstruction from incomplete data,”arXiv preprint arXiv:1709.00584, 2017

  13. [21]

    Deep null space learning for inverse problems: convergence analysis and rates,

    J. Schwab, S. Antholzer, and M. Haltmeier, “Deep null space learning for inverse problems: convergence analysis and rates,”Inverse Problems, vol. 35, p. 025008, jan 2019

  14. [22]

    Improved inversion through use of the null space,

    P. S. Rowbotham and R. G. Pratt, “Improved inversion through use of the null space,” Geophysics, vol. 62, no. 3, pp. 869–883, 1997

  15. [23]

    Nullspace shuttles,

    M. M. Deal and G. Nolet, “Nullspace shuttles,”Geophysical Journal International, vol. 124, pp. 372–380, 02 1996

  16. [24]

    A perspective on deep imaging,

    G. Wang, “A perspective on deep imaging,”IEEE Access, vol. 4, pp. 8914–8924, 2016

  17. [25]

    Image reconstruction: from sparsity to data- adaptive methods and machine learning,

    S. Ravishankar, J. C. Ye, and J. A. Fessler, “Image reconstruction: from sparsity to data- adaptive methods and machine learning,”Proc. IEEE, vol. 108, pp. 86–109, Jan. 2020

  18. [26]

    The troublesome kernel: why deep learning for inverse problems is typically unstable,

    N. Gottschling, V. Antun, B. Adcock, and A. C. Hansen, “The troublesome kernel: why deep learning for inverse problems is typically unstable,”ArXiv, vol. abs/2001.01258, 2020

  19. [27]

    On instabilities of deep learning in image reconstruction and the potential costs of AI,

    V. Antun, F. Renna, C. Poon, B. Adcock, and A. C. Hansen, “On instabilities of deep learning in image reconstruction and the potential costs of AI,”Proceedings of the National Academy of Sciences, vol. 117, no. 48, pp. 30088–30095, 2020. CONTENTS14

  20. [28]

    Applications, promises, and pitfalls of deep learning for fluorescence image reconstruction,

    C. Belthangady and L. A. Royer, “Applications, promises, and pitfalls of deep learning for fluorescence image reconstruction,”Nature methods, vol. 16, no. 12, pp. 1215–1225, 2019

  21. [29]

    The promise and peril of deep learning in microscopy,

    D. P. Hoffman, I. Slavitt, and C. A. Fitzpatrick, “The promise and peril of deep learning in microscopy,”Nature methods, vol. 18, no. 2, pp. 131–132, 2021

  22. [30]

    Machine learning for medical imaging: methodological failures and recommendations for the future,

    G. Varoquaux and V. Cheplygina, “Machine learning for medical imaging: methodological failures and recommendations for the future,”NPJ digital medicine, vol. 5, no. 1, p. 48, 2022

  23. [31]

    Advancing machine learning for mr image reconstruction with an open competition: Overview of the 2019 fastmri challenge,

    F. Knoll, T. Murrell, A. Sriram, N. Yakubova, J. Zbontar, M. Rabbat, A. Defazio, M. J. Muckley, D. K. Sodickson, C. L. Zitnick,et al., “Advancing machine learning for mr image reconstruction with an open competition: Overview of the 2019 fastmri challenge,”Magnetic resonance i...

  24. [32]

    Results of the 2020 fastmri challenge for machine learning mr image reconstruction,

    M. J. Muckley, B. Riemenschneider, A. Radmanesh, S. Kim, G. Jeong, J. Ko, Y. Jun, H. Shin, D. Hwang, M. Mostapha,et al., “Results of the 2020 fastmri challenge for machine learning mr image reconstruction,”IEEE transactions on medical imaging, vol. 40, no. 9, pp. 2306– 2317, 2021

  25. [33]

    Deep learning reconstruction of accelerated mri: False-positive cartilage delamination inserted in mri arthrography under traction,

    W. A. Bosbach, K. C. Merdes, B. Jung, E. Montazeri, S. Anderson, M. Mitrakovic, and K. Daneshvar, “Deep learning reconstruction of accelerated mri: False-positive cartilage delamination inserted in mri arthrography under traction,”Topics in Magnetic Resonance Imaging, vol. 33,...

  26. [34]

    Robust physical-world attacks on deep learning visual classification,

    K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, and D. Song, “Robust physical-world attacks on deep learning visual classification,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1625– 1634, 2018

  27. [35]

    Audio adversarial examples: Targeted attacks on speech-to-text,

    N. Carlini and D. Wagner, “Audio adversarial examples: Targeted attacks on speech-to-text,” in2018 IEEE security and privacy workshops (SPW), pp. 1–7, IEEE, 2018

  28. [36]

    Adversarial attacks on medical machine learning,

    S. G. Finlayson, J. D. Bowers, J. Ito, J. L. Zittrain, A. L. Beam, and I. S. Kohane, “Adversarial attacks on medical machine learning,”Science, vol. 363, no. 6433, pp. 1287–1289, 2019

  29. [37]

    Why deep-learning ais are so easy to fool,

    D. Heavenet al., “Why deep-learning ais are so easy to fool,”Nature, vol. 574, no. 7777, pp. 163–166, 2019

  30. [38]

    The mathematics of adversarial attacks in ai–why deep learning is unstable despite the existence of stable neural networks,

    A. Bastounis, A. C. Hansen, and V. Vlaˇ ci´ c, “The mathematics of adversarial attacks in ai–why deep learning is unstable despite the existence of stable neural networks,”arXiv preprint arXiv:2109.06098, 2021

  31. [39]

    Some investigations on robustness of deep learning in limited angle tomography,

    Y. Huang, T. W¨ urfl, K. Breininger, L. Liu, G. Lauritsch, and A. Maier, “Some investigations on robustness of deep learning in limited angle tomography,” inMedical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, S...

  32. [40]

    Measuring robustness in deep learning based compressive sensing,

    M. Z. Darestani, A. S. Chaudhari, and R. Heckel, “Measuring robustness in deep learning based compressive sensing,” inInternational Conference on Machine Learning, pp. 2433– 2444, PMLR, 2021

  33. [41]

    Solving inverse problems with deep neural networks– robustness included?,

    M. Genzel, J. Macdonald, and M. M¨ arz, “Solving inverse problems with deep neural networks– robustness included?,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 1, pp. 1119–1134, 2022

  34. [42]

    Improving robustness of deep-learning-based image reconstruction,

    A. Raj, Y. Bresler, and B. Li, “Improving robustness of deep-learning-based image reconstruction,” inInternational Conference on Machine Learning, pp. 7932–7942, PMLR, 2020

  35. [43]

    Adversarial robustness of mr image reconstruction under realistic perturbations,

    J. N. Morshuis, S. Gatidis, M. Hein, and C. F. Baumgartner, “Adversarial robustness of mr image reconstruction under realistic perturbations,” inInternational Workshop on Machine Learning for Medical Image Reconstruction, pp. 24–33, Springer, 2022

  36. [44]

    Localized adversarial artifacts for compressed sensing mri,

    R. Alaifari, G. S. Alberti, and T. Gauksson, “Localized adversarial artifacts for compressed sensing mri,”SIAM Journal on Imaging Sciences, vol. 16, no. 4, pp. SC14–SC26, 2023

  37. [45]

    On hallucinations in tomographic image reconstruction,

    S. Bhadra, V. A. Kelkar, F. J. Brooks, and M. A. Anastasio, “On hallucinations in tomographic image reconstruction,”IEEE Transactions on Medical Imaging, vol. 40, pp. 3249–3260, Nov. 2021

  38. [46]

    The difficulty of computing stable and accurate neural networks: On the barriers of deep learning and Smale’s 18th problem,

    M. J. Colbrook, V. Antun, and A. C. Hansen, “The difficulty of computing stable and accurate neural networks: On the barriers of deep learning and Smale’s 18th problem,”Proceedings of the National Academy of Sciences, vol. 119, no. 12, p. e2107151119, 2022

  39. [47]

    Impact of deep learning- based image super-resolution on binary signal detection,

    X. Zhang, V. A. Kelkar, J. Granstedt, H. Li, and M. A. Anastasio, “Impact of deep learning- based image super-resolution on binary signal detection,”Journal of Medical Imaging, vol. 8, no. 6, pp. 065501–065501, 2021

  40. [48]

    Unified SNR analysis of medical imaging systems,

    R. F. Wagner and D. G. Brown, “Unified SNR analysis of medical imaging systems,”Physics in Medicine & Biology, vol. 30, no. 6, p. 489, 1985. CONTENTS15

  41. [49]

    Icru report 54: Medical imaging-the assessment of image quality-isbn 0-913394- 53-x. april 1996, maryland, usa,

    W. Vennart, “Icru report 54: Medical imaging-the assessment of image quality-isbn 0-913394- 53-x. april 1996, maryland, usa,”Radiography, vol. 3, no. 3, pp. 243–244, 1997

  42. [50]

    Model observers for assessment of image quality,

    H. H. Barrett, J. Yao, J. P. Rolland, and K. J. Myers, “Model observers for assessment of image quality,”Proceedings of the National Academy of Sciences, vol. 90, no. 21, pp. 9758–9765, 1993

  43. [52]

    Strategies for reducing radiation dose in ct,

    C. H. McCollough, A. N. Primak, N. Braun, J. Kofler, L. Yu, and J. Christner, “Strategies for reducing radiation dose in ct,”Radiologic Clinics, vol. 47, no. 1, pp. 27–40, 2009

  44. [53]

    Algorithms for reconstruction with nondiffracting sources,

    A. C. Kak and M. Slaney, “Algorithms for reconstruction with nondiffracting sources,” in Principles of computerized tomographic imaging, ch. 3, pp. 49–112, Philadelphia: SIAM, 2001

  45. [54]

    Acquisition and reconstruction of magnetic resonance imaging,

    A.-F. Santiago and V.-S.-F. Gonzalo, “Acquisition and reconstruction of magnetic resonance imaging,” inStatistical Analysis of Noise in MRI Modeling, Filtering and Estimation, ch. 2, pp. 9–29, Switzerland: Springer International Publishing, 2016

  46. [55]

    Low-dose ct with a residual encoder-decoder convolutional neural network,

    H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, J. Zhou, and G. Wang, “Low-dose ct with a residual encoder-decoder convolutional neural network,”IEEE transactions on medical imaging, vol. 36, no. 12, pp. 2524–2535, 2017

  47. [56]

    Low-dose ct image denoising using a generative adversarial network with wasserstein distance and perceptual loss,

    Q. Yang, P. Yan, Y. Zhang, H. Yu, Y. Shi, X. Mou, M. K. Kalra, Y. Zhang, L. Sun, and G. Wang, “Low-dose ct image denoising using a generative adversarial network with wasserstein distance and perceptual loss,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1348–1357, 2018

  48. [57]

    Deep admm-net for compressive sensing mri,

    J. Sun, H. Li, Z. Xu,et al., “Deep admm-net for compressive sensing mri,”Advances in neural information processing systems, vol. 29, 2016

  49. [58]

    Deep networks and mutual information maximization for cross-modal medical image synthesis,

    R. Vemulapalli, H. V. Nguyen, and S. K. Zhou, “Deep networks and mutual information maximization for cross-modal medical image synthesis,” inDeep Learning for Medical Image Analysis(S. K. Zhou, H. Greenspan, and D. Shen, eds.), ch. 16, pp. 381–403, Oxford: Academic Press, 2023

  50. [59]

    Cross-modality image synthesis from unpaired data using cyclegan: Effects of gradient consistency loss and training data size,

    Y. Hiasa, Y. Otake, M. Takao, T. Matsuoka, K. Takashima, A. Carass, J. L. Prince, N. Sugano, and Y. Sato, “Cross-modality image synthesis from unpaired data using cyclegan: Effects of gradient consistency loss and training data size,” inSimulation and Synthesis in Medical Imag...

  51. [60]

    The data processing inequality and stochastic resonance,

    M. D. McDonnell, N. G. Stocks, C. E. Pearce, and D. Abbott, “The data processing inequality and stochastic resonance,” inNoise in Complex Systems and Stochastic Dynamics, vol. 5114, pp. 249–260, SPIE, 2003

  52. [61]

    On hallucinations in tomographic image reconstruction,

    S. Bhadra, V. A. Kelkar, F. J. Brooks, and M. A. Anastasio, “On hallucinations in tomographic image reconstruction,”IEEE transactions on medical imaging, vol. 40, no. 11, pp. 3249–3260, 2021

  53. [62]

    Null space imaging: nonlinear magnetic encoding fields designed complementary to receiver coil sensitivities for improved acceleration in parallel imaging,

    L. K. Tam, J. P. Stockmann, G. Galiana, and R. T. Constable, “Null space imaging: nonlinear magnetic encoding fields designed complementary to receiver coil sensitivities for improved acceleration in parallel imaging,”Magnetic resonance in medicine, vol. 68, no. 4, pp. 1166– 1...

  54. [64]

    Image artifacts: Appearances, causes, and corrections,

    J. Hsieh, “Image artifacts: Appearances, causes, and corrections,” inComputed tomography: principles, design, artifacts, and recent advances, ch. 7, pp. 207–300, Bellingham, Washington: SPIE press, 2003

  55. [65]

    Artifacts in magnetic resonance imaging,

    K. Krupa and M. Bekiesi´ nska-Figatowska, “Artifacts in magnetic resonance imaging,”Polish journal of radiology, vol. 80, p. 93, 2015

  56. [66]

    P. C. Hansen, J. Jørgensen, and W. R. Lionheart,Computed tomography: algorithms, insight, and just enough theory, ch. 10, pp. 183–209. SIAM, 2021

  57. [67]

    fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data,

    R. Zhao, B. Yaman, Y. Zhang, R. Stewart, A. Dixon, F. Knoll, Z. Huang, Y. W. Lui, M. S. Hansen, and M. P. Lungren, “fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data,”Scientific Data, vol. 9, no. 1, p. 152, 2022

  58. [68]

    Lungx challenge for computerized lung CONTENTS16 nodule classification,

    S. G. Armato III, K. Drukker, F. Li, L. Hadjiiski, G. D. Tourassi, R. M. Engelmann, M. L. Giger, G. Redmond, K. Farahani, J. S. Kirby,et al., “Lungx challenge for computerized lung CONTENTS16 nodule classification,”Journal of Medical Imaging, vol. 3, no. 4, pp. 044506–044506, 2016

  59. [69]

    Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning,

    K. Yan, X. Wang, L. Lu, and R. M. Summers, “Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning,”Journal of medical imaging, vol. 5, no. 3, pp. 036501–036501, 2018

  60. [70]

    No” zero-shot

    V. Udandarao, A. Prabhu, A. Ghosh, Y. Sharma, P. Torr, A. Bibi, S. Albanie, and M. Bethge, “No” zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance,” inThe Thirty-eighth Annual Conference on Neural Information Processing S...

  61. [71]

    Generative adversarial networks in medical image augmentation: a review,

    Y. Chen, X.-H. Yang, Z. Wei, A. A. Heidari, N. Zheng, Z. Li, H. Chen, H. Hu, Q. Zhou, and Q. Guan, “Generative adversarial networks in medical image augmentation: a review,” Computers in Biology and Medicine, vol. 144, p. 105382, 2022

  62. [72]

    Data augmentation for medical imaging: A systematic literature review,

    F. Garcea, A. Serra, F. Lamberti, and L. Morra, “Data augmentation for medical imaging: A systematic literature review,”Computers in Biology and Medicine, vol. 152, p. 106391, 2023

  63. [73]

    Synthetic breast ultrasound images: A study to overcome medical data sharing barriers,

    J. Xu, Q. Hua, X. Jia, Y. Zheng, Q. Hu, B. Bai, J. Miao, L. Zhu, M. Zhang, R. Tao,et al., “Synthetic breast ultrasound images: A study to overcome medical data sharing barriers,” Research, vol. 7, p. 0532, 2024

  64. [74]

    Analyzing gan artifacts for simulating mammograms: application towards finding mammographically-occult cancer,

    J. Lee and R. M. Nishikawa, “Analyzing gan artifacts for simulating mammograms: application towards finding mammographically-occult cancer,” inMedical Imaging 2022: Computer- Aided Diagnosis, vol. 12033, pp. 78–84, SPIE, 2022

  65. [75]

    Selective synthetic augmentation with histogan for improved histopathology image classification,

    Y. Xue, J. Ye, Q. Zhou, L. R. Long, S. Antani, Z. Xue, C. Cornwell, R. Zaino, K. C. Cheng, and X. Huang, “Selective synthetic augmentation with histogan for improved histopathology image classification,”Medical image analysis, vol. 67, p. 101816, 2021

  66. [76]

    Hallucination is inevitable: An innate limitation of large language models,

    Z. Xu, S. Jain, and M. Kankanhalli, “Hallucination is inevitable: An innate limitation of large language models,”arXiv preprint arXiv:2401.11817, 2024

  67. [77]

    A method for evaluating deep generative models of images for hallucinations in high-order spatial context,

    R. Deshpande, M. A. Anastasio, and F. J. Brooks, “A method for evaluating deep generative models of images for hallucinations in high-order spatial context,”Pattern Recognition Letters, vol. 186, pp. 23–29, 2024

  68. [78]

    Assessing the capacity of a denoising diffusion probabilistic model to reproduce spatial context,

    R. Deshpande, M. ¨Ozbey, H. Li, M. A. Anastasio, and F. J. Brooks, “Assessing the capacity of a denoising diffusion probabilistic model to reproduce spatial context,”IEEE Transactions on Medical Imaging, 2024

  69. [79]

    Assessing the ability of generative adversarial networks to learn canonical medical image statistics,

    V. A. Kelkar, D. S. Gotsis, F. J. Brooks, K. Prabhat, K. J. Myers, R. Zeng, and M. A. Anastasio, “Assessing the ability of generative adversarial networks to learn canonical medical image statistics,”IEEE transactions on medical imaging, vol. 42, no. 6, pp. 1799–1808, 2023

  70. [80]

    A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis,

    G. M¨ uller-Franzes, J. M. Niehues, F. Khader, S. T. Arasteh, C. Haarburger, C. Kuhl, T. Wang, T. Han, T. Nolte, S. Nebelung,et al., “A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis,” Sc...

  71. [81]

    Report on the aapm grand challenge on deep generative modeling for learning medical image statistics,

    R. Deshpande, V. A. Kelkar, D. Gotsis, P. Kc, R. Zeng, K. J. Myers, F. J. Brooks, and M. A. Anastasio, “Report on the aapm grand challenge on deep generative modeling for learning medical image statistics,”Medical Physics, vol. 52, no. 1, pp. 4–20, 2025

  72. [82]

    A knowledge-based method for detecting network-induced shape artifacts in synthetic images,

    R. Deshpande, M. Lago, A. Subbaswamy, S. Kahaki, J. G. Delfino, A. Badano, and G. Zamzmi, “A knowledge-based method for detecting network-induced shape artifacts in synthetic images,” inMedical Imaging with Deep Learning, 2025

  73. [83]

    Distribution matching losses can hallucinate features in medical image translation,

    J. P. Cohen, M. Luck, and S. Honari, “Distribution matching losses can hallucinate features in medical image translation,” inMedical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceeding...

  74. [84]

    Cyclegan for virtual stain transfer: Is seeing really believing?,

    J. Vasiljevi´ c, Z. Nisar, F. Feuerhake, C. Wemmert, and T. Lampert, “Cyclegan for virtual stain transfer: Is seeing really believing?,”Artificial Intelligence in Medicine, vol. 133, p. 102420, 2022

  75. [85]

    Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive models,

    S. Bond-Taylor, A. Leach, Y. Long, and C. G. Willcocks, “Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive models,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 11, pp. 7327–7347, 2021

  76. [86]

    Benchmarking large language models for news summarization,

    T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto, “Benchmarking large language models for news summarization,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 39–57, 2024

  77. [87]

    Evaluation of chatgpt as a question answering system for answering complex questions,

    Y. Tan, D. Min, Y. Li, W. Li, N. Hu, Y. Chen, and G. Qi, “Evaluation of chatgpt as a question answering system for answering complex questions,”arXiv preprint arXiv:2303.07992, 2023

  78. [88]

    Multilingual machine CONTENTS17 translation with large language models: Empirical results and analysis,

    W. Zhu, H. Liu, Q. Dong, J. Xu, S. Huang, L. Kong, J. Chen, and L. Li, “Multilingual machine CONTENTS17 translation with large language models: Empirical results and analysis,”arXiv preprint arXiv:2304.04675, 2023

  79. [89]

    Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics,

    A. Pagnoni, V. Balachandran, and Y. Tsvetkov, “Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics,”arXiv preprint arXiv:2104.13346, 2021

  80. [90]

    Challenges in building intelligent open-domain dialog systems,

    M. Huang, X. Zhu, and J. Gao, “Challenges in building intelligent open-domain dialog systems,” ACM Transactions on Information Systems (TOIS), vol. 38, no. 3, pp. 1–32, 2020

  81. [91]

    Language models are unsupervised multitask learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever,et al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019

  82. [92]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray,et al., “Training language models to follow instructions with human feedback,”Advances in neural information processing systems, vol. 35, pp. 27730–27744, 2022

  83. [93]

    Bert: Pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” inProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologie...

  84. [94]

    How much knowledge can you pack into the parameters of a language model?,

    A. Roberts, C. Raffel, and N. Shazeer, “How much knowledge can you pack into the parameters of a language model?,”arXiv preprint arXiv:2002.08910, 2020

  85. [95]

    Black swans and the domains of statistics,

    N. N. Taleb, “Black swans and the domains of statistics,”The american statistician, vol. 61, no. 3, pp. 198–200, 2007

  86. [96]

    Survey of hallucination in natural language generation,

    Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”ACM computing surveys, vol. 55, no. 12, pp. 1–38, 2023

  87. [97]

    Diversifying dialogue generation with non-conversational text,

    H. Su, X. Shen, S. Zhao, X. Zhou, P. Hu, R. Zhong, C. Niu, and J. Zhou, “Diversifying dialogue generation with non-conversational text,”arXiv preprint arXiv:2005.04346, 2020

  88. [98]

    Union: An unreferenced metric for evaluating open-ended story generation,

    J. Guan and M. Huang, “Union: An unreferenced metric for evaluating open-ended story generation,”arXiv preprint arXiv:2009.07602, 2020

  89. [99]

    Retrieval augmentation reduces hallucination in conversation,

    K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, “Retrieval augmentation reduces hallucination in conversation,”arXiv preprint arXiv:2104.07567, 2021

  90. [100]

    Towards conversational diagnostic ai,

    T. Tu, A. Palepu, M. Schaekermann, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, N. Tomasev,et al., “Towards conversational diagnostic ai,”arXiv preprint arXiv:2401.05654, 2024

  91. [101]

    Towards generalist biomedical ai,

    T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P.-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena,et al., “Towards generalist biomedical ai,”Nejm Ai, vol. 1, no. 3, p. AIoa2300138, 2024

  92. [103]

    Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation,

    C.-Y. Li, K.-J. Chang, C.-F. Yang, H.-Y. Wu, W. Chen, H. Bansal, L. Chen, Y.-P. Yang, Y.-C. Chen, S.-P. Chen,et al., “Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation,”Nature Communications, vol. 16, no. 1, p. 2258, 2025

  93. [104]

    Collaboration between clinicians and vision–language models in radiology report generation,

    R. Tanno, D. G. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, C. Lau, T. Tu, S. Azizi,et al., “Collaboration between clinicians and vision–language models in radiology report generation,”Nature Medicine, vol. 31, no. 2, pp. 599–608, 2025

  94. [105]

    Glue: A multi-task benchmark and analysis platform for natural language understanding,

    A. Wang, “Glue: A multi-task benchmark and analysis platform for natural language understanding,”arXiv preprint arXiv:1804.07461, 2018

  95. [106]

    Holistic evaluation of language models,

    P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar,et al., “Holistic evaluation of language models,”arXiv preprint arXiv:2211.09110, 2022

  96. [107]

    Measuring massive multitask language understanding,

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,”arXiv preprint arXiv:2009.03300, 2020

  97. [108]

    Judging llm-as-a-judge with mt-bench and chatbot arena,

    L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing,et al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”Advances in Neural Information Processing Systems, vol. 36, pp. 46595–46623, 2023

  98. [109]

    A review of current trends, techniques, and challenges in large language models (llms),

    R. Patil and V. Gudivada, “A review of current trends, techniques, and challenges in large language models (llms),”Applied Sciences, vol. 14, no. 5, p. 2074, 2024

  99. [110]

    Llama 2: Open foundation and fine-tuned chat models,

    H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale,et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023

  100. [111]

    Mitigating the alignment tax of rlhf,

    Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang,et al., CONTENTS18 “Mitigating the alignment tax of rlhf,”arXiv preprint arXiv:2309.06256, 2023

  101. [112]

    ” do anything now

    X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” inProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 1671– 1685, 2024

  102. [113]

    Fine-tuning aligned language models compromises safety, even when users do not intend to!,

    X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!,”arXiv preprint arXiv:2310.03693, 2023

  103. [114]

    Siren’s song in the ai ocean: a survey on hallucination in large language models,

    Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen,et al., “Siren’s song in the ai ocean: a survey on hallucination in large language models,”arXiv preprint arXiv:2309.01219, 2023

  104. [115]

    The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,

    G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay, “The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,”arXiv preprint arXiv:2306.01116, 2023

  105. [116]

    Toward expert-level medical question answering with large language models,

    K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis,et al., “Toward expert-level medical question answering with large language models,”Nature Medicine, pp. 1–8, 2025

  106. [117]

    Scaling instruction-finetuned language models,

    H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma,et al., “Scaling instruction-finetuned language models,”Journal of Machine Learning Research, vol. 25, no. 70, pp. 1–53, 2024

  107. [118]

    Pmc-llama: toward building open-source language models for medicine,

    C. Wu, W. Lin, X. Zhang, Y. Zhang, W. Xie, and Y. Wang, “Pmc-llama: toward building open-source language models for medicine,”Journal of the American Medical Informatics Association, vol. 31, no. 9, pp. 1833–1843, 2024

  108. [119]

    Vision-language models for medical report generation and visual question answering: A review,

    I. Hartsock and G. Rasool, “Vision-language models for medical report generation and visual question answering: A review,”Frontiers in Artificial Intelligence, vol. 7, p. 1430984, 2024

  109. [120]

    Evaluating object hallucination in large vision-language models,

    Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,”arXiv preprint arXiv:2305.10355, 2023

  110. [121]

    Eyes wide shut? exploring the visual shortcomings of multimodal llms,

    S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9568–9578, 2024

  111. [122]

    Investigating the catastrophic forgetting in multimodal large language models,

    Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, “Investigating the catastrophic forgetting in multimodal large language models,”arXiv preprint arXiv:2309.10313, 2023

  112. [123]

    Hallucination index: An image quality metric for generative reconstruction models,

    M. Tivnan, S. Yoon, Z. Chen, X. Li, D. Wu, and Q. Li, “Hallucination index: An image quality metric for generative reconstruction models,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 449–458, Springer, 2024

  113. [124]

    Language models with conformal factuality guarantees,

    C. Mohri and T. Hashimoto, “Language models with conformal factuality guarantees,”arXiv preprint arXiv:2402.10978, 2024

  114. [125]

    Large language model validity via enhanced conformal prediction methods,

    J. Cherian, I. Gibbs, and E. Candes, “Large language model validity via enhanced conformal prediction methods,”Advances in Neural Information Processing Systems, vol. 37, pp. 114812–114842, 2024

  115. [126]

    Med-halt: Medical domain hallucination test for large language models,

    A. Pal, L. K. Umapathi, and M. Sankarasubbu, “Med-halt: Medical domain hallucination test for large language models,”arXiv preprint arXiv:2307.15343, 2023

  116. [127]

    Cares: A comprehensive benchmark of trustworthiness in medical vision language models,

    P. Xia, Z. Chen, J. Tian, Y. Gong, R. Hou, Y. Xu, Z. Wu, Z. Fan, Y. Zhou, K. Zhu,et al., “Cares: A comprehensive benchmark of trustworthiness in medical vision language models,” Advances in Neural Information Processing Systems, vol. 37, pp. 140334–140365, 2024

  117. [128]

    Detecting and evaluating medical hallucinations in large vision language models,

    J. Chen, D. Yang, T. Wu, Y. Jiang, X. Hou, M. Li, S. Wang, D. Xiao, K. Li, and L. Zhang, “Detecting and evaluating medical hallucinations in large vision language models,”arXiv preprint arXiv:2406.10185, 2024

  118. [129]

    State of what art? a call for multi-prompt llm evaluation,

    M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, and G. Stanovsky, “State of what art? a call for multi-prompt llm evaluation,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 933–949, 2024

  119. [130]

    Adversarial glue: A multi-task benchmark for robustness evaluation of language models,

    B. Wang, C. Xu, S. Wang, Z. Gan, Y. Cheng, J. Gao, A. H. Awadallah, and B. Li, “Adversarial glue: A multi-task benchmark for robustness evaluation of language models,”arXiv preprint arXiv:2111.02840, 2021

  120. [131]

    On the robustness of chatgpt: An adversarial and out-of-distribution perspective,

    J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y. Wang, L. Yang, H. Huang, W. Ye, X. Geng, et al., “On the robustness of chatgpt: An adversarial and out-of-distribution perspective,” arXiv preprint arXiv:2302.12095, 2023

  121. [132]

    Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.,

    B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al., “Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.,” in NeurIPS, 2023

  122. [133]

    Learning a variational network for reconstruction of accelerated mri data,

    K. Hammernik, T. Klatzer, E. Kobler, M. P. Recht, D. K. Sodickson, T. Pock, and F. Knoll, CONTENTS19 “Learning a variational network for reconstruction of accelerated mri data,”Magnetic resonance in medicine, vol. 79, no. 6, pp. 3055–3071, 2018

  123. [134]

    Learn: Learned experts’ assessment-based reconstruction network for sparse- data ct,

    H. Chen, Y. Zhang, Y. Chen, J. Zhang, W. Zhang, H. Sun, Y. Lv, P. Liao, J. Zhou, and G. Wang, “Learn: Learned experts’ assessment-based reconstruction network for sparse- data ct,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1333–1347, 2018

  124. [135]

    Ambientgan: Generative models from lossy measurements,

    A. Bora, E. Price, and A. G. Dimakis, “Ambientgan: Generative models from lossy measurements,” inInternational conference on learning representations, 2018

  125. [136]

    Latent retrieval for weakly supervised open domain question answering,

    K. Lee, M.-W. Chang, and K. Toutanova, “Latent retrieval for weakly supervised open domain question answering,”arXiv preprint arXiv:1906.00300, 2019

  126. [137]

    Retrieval augmented language model pre-training,

    K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang, “Retrieval augmented language model pre-training,” inInternational conference on machine learning, pp. 3929–3938, PMLR, 2020

  127. [138]

    Retrieval-augmented generation for knowledge-intensive nlp tasks,

    P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K¨ uttler, M. Lewis, W.-t. Yih, T. Rockt¨ aschel,et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,”Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020

  128. [139]

    Knowledge conflicts for llms: A survey,

    R. Xu, Z. Qi, Z. Guo, C. Wang, H. Wang, Y. Zhang, and W. Xu, “Knowledge conflicts for llms: A survey,”arXiv preprint arXiv:2403.08319, 2024

  129. [140]

    Knowledge graphs, large language models, and hallucinations: An nlp perspective,

    E. Lavrinovics, R. Biswas, J. Bjerva, and K. Hose, “Knowledge graphs, large language models, and hallucinations: An nlp perspective,”Journal of Web Semantics, vol. 85, p. 100844, 2025

  130. [141]

    Towards mitigating hallucination in large language models via self-reflection,

    Z. Ji, T. Yu, Y. Xu, N. Lee, E. Ishii, and P. Fung, “Towards mitigating hallucination in large language models via self-reflection,”arXiv preprint arXiv:2310.06271, 2023

  131. [142]

    Improving factuality and reasoning in language models through multiagent debate,

    Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving factuality and reasoning in language models through multiagent debate,” inForty-first International Conference on Machine Learning, 2023

  132. [143]

    Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration,

    Z. Wang, S. Mao, W. Wu, T. Ge, F. Wei, and H. Ji, “Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration,” arXiv preprint arXiv:2307.05300, 2023

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.