REVIEW 3 major objections 5 minor 140 references
Hallucinations in medical devices
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper proposes that a hallucination in a medical device is any output error that is plausible to the user, split into impactful and benign hallucinations.
desk verdict A useful position piece that adds plausibility and impact axes to the hallucination definition, but leaves the key threshold tau unmeasured and therefore stops short of its practical/universal claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the plausibility-impact error diagram of Figure 1, organized around a plausibility axis with a threshold $\tau$. An error above $\tau$ is labeled a hallucination; an error below it is a non-hallucination error, characterized by obviousness and traceability to device artifacts or pre-specified failure modes. The paper stresses that $\tau$ is a continuum and is observer- and task-specific, and it treats the user of the device, who may be an expert, a patient, or an algorithmic interpreter, as part of the definition. This axis does the work of separating hallucinations from conventional artifacts and explains why a model can produce fewer impactful errors yet still cause worse patient outcomes: plausible errors escape clinician intuition and existing risk mitigation.
What would settle it
A multi-reader study in which clinicians at different training levels independently classify the same set of medical-device errors as plausible versus obvious, and then repeat the task for different clinical tasks, would settle the point: if the placement of $\tau$ varies so widely across readers or tasks that no stable boundary emerges, the proposed definition cannot ground a practical evaluation method.
Extended reading notes
Core claim
The central claim is that hallucination in a medical device is best defined as a subset of error: an error that is plausible to the intended user, with two subtypes, impactful and benign, and with a separate category of non-hallucination errors that are obvious and traceable to device artifacts or pre-specified failure modes. The paper grounds this in a ground-truth-function account in which hallucinations are an unavoidable property of data-driven models, and it shows the definition at work across imaging, synthetic image generation, and language and multimodal devices. Examples include AI super-resolution outputs that add bowel loops or plaque-like features absent from the reference, conditional generators that insert tumors or histological features that do not exist in the input, and language models that insert a diagnosis into a summary. The consequence is that device evaluation should measure where errors fall on plausibility and impact axes, and that a device producing fewer errors overall can still be more dangerous if its errors are plausible.
Load-bearing premise
The whole taxonomy rests on the idea that plausibility can be treated as a measurable axis with a threshold $\tau$, but the paper gives no procedure for estimating $\tau$ and notes it is observer- and task-specific; if that boundary cannot be pinned down, the hallucination label cannot be applied consistently.
Editorial extensions
If this is right
- Device evaluation would report errors classified along plausibility and impact, not only frequency; a model with fewer total errors could still be higher-risk if its errors are plausible.
- A hallucination that is benign in the original task can become impactful if the output is later reused in patient-care decisions, so evaluation should track downstream use.
- Stability measurements under small input perturbations can serve as a proxy for hallucination propensity, but they do not by themselves provide the plausibility information needed for classification.
- Because hallucinations are argued to be intrinsic to neural-network methods, mitigation strategies such as null-space constraints, noise injection, retrieval augmentation, and conformal redaction can reduce but not eliminate them.
Reading between the lines
- An implication the paper leaves implicit: plausibility thresholds could be measured empirically with reader studies using forced-choice judgments at different expertise levels, but the definition requires such studies to become operational.
- A testable extension: task-based evaluation metrics for detection, quantification, and classification could convert the impactful-versus-benign distinction from a qualitative label into a measurable downstream performance change.
- If the definition is right, one would expect regulatory and industry test reports to separate 'hallucination' counts from 'artifact' counts, and to specify the observer population used for plausibility judgments.
- The claim that hallucinations cannot be fully removed suggests that certification criteria should focus on bounding the rate of impactful plausible errors within a task-specific tolerance rather than requiring error-free outputs.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript argues that the term 'hallucination' lacks a universally recognized definition in AI/ML-enabled medical devices and proposes one: hallucinations are a subset of errors, specifically errors that are plausible, and can be subdivided into impactful hallucinations and benign hallucinations. Non-hallucination errors are characterized by their obviousness and traceability to device artifacts or pre-specified failure modes. The paper applies this definition to three areas: imaging devices, synthetic image generation, and language and multimodal devices, and it reviews existing approaches for quantifying and mitigating hallucinations. It explicitly concedes that the plausibility threshold tau is observer- and task-specific and that studies to determine this threshold are outside the scope of the work.
Significance. If the proposed definition were operationalized, it would give device developers and regulators a common vocabulary across imaging, generative, and language domains, and it would direct evaluation toward plausibility and impact axes, which is a potentially valuable contribution. The paper's strengths include its candid acknowledgment of the unmeasured threshold, its use of concrete examples grounded in external empirical work, and its synthesis of stability-based quantification methods. The definition builds on Xu et al. without circularity, and the paper does not derive its central claim from itself. However, as presented, the central classification axis is non-operational, and the abstract's claim that the definition is 'practical and universal' is not yet supported, so substantial revision is needed before the paper can stand as a foundation for device evaluation.
major comments (3)
- [§1, Fig. 1] The definition rests entirely on the plausibility threshold tau, yet the manuscript states that tau is 'likely observer and task-specific' and that studies to determine it are 'outside of the scope of this work.' Because no method, protocol, or reference observer is given to fix tau, the same erroneous output can be classified as a hallucination for one user and a non-hallucination error for another, and nothing in the framework decides between these classifications. This directly undercuts the abstract's claim that the definition is 'practical and universal' and is a load-bearing gap rather than a cosmetic one. The authors should either provide an operational protocol (for example, reader studies with a specified expertise level and a calibration procedure for tau) or explicitly revise the claim to present the definition as a conceptual framework pending empirical calibration.
- [§3] The quantification section acknowledges that worst-case perturbation methods 'do not provide the necessary measure of plausibility' and that the hallucination index, which computes the Hellinger distance between ground-truth and reconstructed distributions, still leaves 'the cut-off that dichotomizes faithful and hallucinated reconstruction' nontrivial. Since the proposed taxonomy is defined by plausibility, none of the surveyed metrics operationalizes the definition's central construct. The manuscript should state clearly which of these metrics, if any, could serve as a proxy for plausibility and under what assumptions, or explain why the taxonomy does not require a metric to be useful in evaluation.
- [§2.1, Fig. 2] The classification of the artifacts in Fig. 2(a) as non-hallucination errors rests on their 'obviousness' and 'traceability' to imaging-system limitations, but the manuscript offers no decision procedure for either property. This is the second axis of the dichotomy and is as underspecified as tau. The authors should specify observable criteria or a study design that would determine when an error is obvious or traceable, or state that the taxonomy is intended only as a retrospective classification.
minor comments (5)
- [§2.1] The sentence 'The driving force in the technological advancement of medical imaging has been less radiation‡ and saving scan time' places the footnote marker awkwardly; consider rewriting as 'lower radiation dose and shorter scan time' with the footnote attached to the relevant term.
- [Fig. 1 caption] The caption states that 'Unmitigated impactful errors are colored yellow,' but the figure itself has no legend and the text does not define 'unmitigated' in this context; please add a legend or clarify the color scheme.
- [§3] The phrase 'as previously mentioned in the section 2.1' should be 'as previously mentioned in Section 2.1,' and the manuscript should use consistent capitalization for figure references such as 'fig. 2' versus 'Fig. 2.'
- [References] The court case references are not formatted consistently (for example, 'Ko v. li, Inc.' uses an uppercase 'I' in 'li'); please use a consistent legal citation style.
- [§2.2.2] The phrase 'hallucinations must be expected' is italicized without explanation; if emphasis is intended, please state the reasoning in words, and otherwise remove the emphasis.
Circularity Check
No significant circularity: the proposed definition is a stipulative taxonomy grounded in external frameworks, and no prediction is derived from its own inputs.
full rationale
The paper's central contribution is a proposed definition of hallucinations in medical devices as plausible errors that are either impactful or benign, building on the external theoretical framework of Xu et al. and on empirical studies of imaging, synthetic data, and language models. Because the claim is a definitional taxonomy rather than a derived quantitative result, there is no derivation chain that could collapse into its inputs. The plausibility threshold tau in Figure 1 is explicitly acknowledged as observer- and task-specific and as not determined in this work; that is an acknowledged operational limitation, not a circular reduction, since the definition does not pretend to measure tau from itself. Self-citations to the authors' prior empirical studies (e.g., Deshpande et al. on spatial context and generative model evaluation, and Badano et al. on shape artifacts) are used as supporting examples or as sources for detection methods, not as justifications of the definition's validity. No fitted parameter is renamed as a prediction, no uniqueness theorem from the authors' prior work is invoked, and no ansatz is smuggled in via citation. The examples in Section 2 are applications of the stated definition to concrete cases, not proofs that the definition is true, so the argument is self-contained in the sense required for circularity analysis. The appropriate finding is therefore no significant circularity.
Assumptions & free parameters
free parameters (1)
- Plausibility threshold tau
assumptions (4)
- domain assumption A ground truth function exists and hallucinations are discrepancies between generated outputs and that function (Xu et al. framework).
- domain assumption Information is always lost during the imaging process, and no post-processing can recover diagnostic details that were not measured.
- domain assumption Hallucinations are intrinsic to neural network-based methods and cannot be fully removed.
- domain assumption Conventional device errors are obvious and traceable to device artifacts or pre-specified failure modes, while hallucinations are plausible and not traceable in the same way.
Cite this review
Pith. "Pith review of Hallucinations in medical devices." pith.science (2026). https://pith.science/paper/MKNQSWD5
@misc{pith2026250814118,
author = {Pith},
title = {Pith review of: Hallucinations in medical devices},
year = {2026},
howpublished = {\url{https://pith.science/paper/MKNQSWD5}},
note = {Machine review of arXiv:2508.14118}
}
read the original abstract
Computer methods in medical devices are frequently imperfect and are known to produce errors in clinical or diagnostic tasks. However, when deep learning and data-based approaches yield output that exhibit errors, the devices are frequently said to hallucinate. Drawing from theoretical developments and empirical studies in multiple medical device areas, we introduce a practical and universal definition that denotes hallucinations as a type of error that is plausible and can be either impactful or benign to the task at hand. The definition aims at facilitating the evaluation of medical devices that suffer from hallucinations across product areas. Using examples from imaging and non-imaging applications, we explore how the proposed definition relates to evaluation methodologies and discuss existing approaches for minimizing the prevalence of hallucinations.
Figures
Reference graph
Works this paper leans on
-
[1]
Measuring short-form factuality in large language models,
J. Wei, N. Karina, H. W. Chung, Y. J. Jiao, S. Papay, A. Glaese, J. Schulman, and W. Fedus, “Measuring short-form factuality in large language models,”arXiv preprint arXiv:2411.04368, 2024
arXiv 2024
-
[2]
Health online 2013,
S. Fox and M. Duggan, “Health online 2013,” tech. rep., Pew Research Center, Washington, D.C., Jan. 2013. CONTENTS13
2013
-
[3]
No. 54 Civ. 1461,
Mata v. Avianca, Inc., “No. 54 Civ. 1461,” June 2023. https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts. nysd.575368.54.0_3.pdf
2023
-
[4]
ONSC 2766,
Ko v. li, Inc., “ONSC 2766,” May 2025. https://www.canlii.org/en/on/onsc/doc/2025/2025onsc2766/2025onsc2766.html
2025
-
[5]
No. 2:24-cv-05205-FMO-MAA,
Lacey v. State Farm, “No. 2:24-cv-05205-FMO-MAA,” May 2025. https://www.lawnext.com/wp-content/uploads/2025/05/C.D.-Cal. -24-cv-05205-dckt-000119_000-filed-2025-05-06.pdf
2025
-
[6]
The impact of AI errors in a human- in-the-loop process,
U. Agudo, K. G. Liberal, M. Arrese, and H. Matute, “The impact of AI errors in a human- in-the-loop process,”Cognitive Research: Principles and Implications, vol. 9, no. 1, p. 1, 2024
2024
-
[7]
Quantifying the impact of AI recommendations with explanations on prescription decision making,
M. Nagendran, P. Festor, M. Komorowski, A. C. Gordon, and A. A. Faisal, “Quantifying the impact of AI recommendations with explanations on prescription decision making,”NPJ Digital Medicine, vol. 6, no. 1, p. 206, 2023
2023
-
[8]
How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection,
M. Jacobs, M. F. Pradier, T. H. McCoy Jr, R. H. Perlis, F. Doshi-Velez, and K. Z. Gajos, “How machine-learning recommendations influence clinician treatment selections: the example of antidepressant selection,”Translational psychiatry, vol. 11, no. 1, p. 108, 2021
2021
Show all 140 references
-
[9]
Humans inherit artificial intelligence biases,
L. Vicente and H. Matute, “Humans inherit artificial intelligence biases,”Scientific Reports, vol. 13, no. 1, p. 15737, 2023
2023
-
[10]
Medical hallucination in foundation models and their impact on healthcare,
Y. Kim, H. Jeong, S. Chen, S. S. Li, M. Lu, K. Alhamoud, J. Mun, C. Grau, M. Jung, R. R. Gameiro,et al., “Medical hallucination in foundation models and their impact on healthcare,”medRxiv, pp. 2025–02, 2025
2025
-
[11]
Solving inverse problems using data- driven models,
S. Arridge, P. Maass, O. ¨Oktem, and C.-B. Sch¨ onlieb, “Solving inverse problems using data- driven models,”Acta Numerica, vol. 28, pp. 1–174, 2019
2019
-
[12]
Deep magnetic resonance image reconstruction: Inverse problems meet neural networks,
D. Liang, J. Cheng, Z. Ke, and L. Ying, “Deep magnetic resonance image reconstruction: Inverse problems meet neural networks,”IEEE Signal Processing Magazine, vol. 37, no. 1, pp. 141–151, 2020
2020
-
[13]
Convolutional neural networks for inverse problems in imaging: A review,
M. T. McCann, K. H. Jin, and M. Unser, “Convolutional neural networks for inverse problems in imaging: A review,”IEEE Signal Processing Magazine, vol. 34, no. 6, pp. 85–95, 2017
2017
-
[14]
Deep learning techniques for inverse problems in imaging,
G. Ongie, A. Jalal, C. A. Metzler, R. G. Baraniuk, A. G. Dimakis, and R. Willett, “Deep learning techniques for inverse problems in imaging,”IEEE Journal on Selected Areas in Information Theory, vol. 1, no. 1, pp. 39–56, 2020
2020
-
[15]
Deep learning for tomographic image reconstruction,
G. Wang, J. C. Ye, and B. De Man, “Deep learning for tomographic image reconstruction,” Nature machine intelligence, vol. 2, no. 12, pp. 737–748, 2020
2020
-
[16]
Deep learning for pet image reconstruction,
A. J. Reader, G. Corda, A. Mehranian, C. da Costa-Luis, S. Ellis, and J. A. Schnabel, “Deep learning for pet image reconstruction,”IEEE Transactions on Radiation and Plasma Medical Sciences, vol. 5, no. 1, pp. 1–25, 2020
2020
-
[17]
Image reconstruction is a new frontier of machine learning,
G. Wang, J. C. Ye, K. Mueller, and J. A. Fessler, “Image reconstruction is a new frontier of machine learning,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1289–1296, 2018
2018
-
[18]
Null-space smoothing of tomographic images using tv norm minimization,
B. Smith, “Null-space smoothing of tomographic images using tv norm minimization,” in2016 IEEE Nuclear Science Symposium, Medical Imaging Conference and Room-Temperature Semiconductor Detector Workshop (NSS/MIC/RTSD), pp. 1–4, IEEE, 2016
2016
-
[19]
Null space and resolution in dynamic computerized tomography,
B. N. Hahn, “Null space and resolution in dynamic computerized tomography,”Inverse Problems, vol. 32, no. 2, p. 025006, 2016
2016
-
[20]
Deep learning-guided image reconstruction from incomplete data,
B. Kelly, T. P. Matthews, and M. A. Anastasio, “Deep learning-guided image reconstruction from incomplete data,”arXiv preprint arXiv:1709.00584, 2017
2017 arXiv
-
[21]
Deep null space learning for inverse problems: convergence analysis and rates,
J. Schwab, S. Antholzer, and M. Haltmeier, “Deep null space learning for inverse problems: convergence analysis and rates,”Inverse Problems, vol. 35, p. 025008, jan 2019
2019
-
[22]
Improved inversion through use of the null space,
P. S. Rowbotham and R. G. Pratt, “Improved inversion through use of the null space,” Geophysics, vol. 62, no. 3, pp. 869–883, 1997
1997
-
[23]
Nullspace shuttles,
M. M. Deal and G. Nolet, “Nullspace shuttles,”Geophysical Journal International, vol. 124, pp. 372–380, 02 1996
1996
-
[24]
A perspective on deep imaging,
G. Wang, “A perspective on deep imaging,”IEEE Access, vol. 4, pp. 8914–8924, 2016
2016
-
[25]
Image reconstruction: from sparsity to data- adaptive methods and machine learning,
S. Ravishankar, J. C. Ye, and J. A. Fessler, “Image reconstruction: from sparsity to data- adaptive methods and machine learning,”Proc. IEEE, vol. 108, pp. 86–109, Jan. 2020
2020
-
[26]
The troublesome kernel: why deep learning for inverse problems is typically unstable,
N. Gottschling, V. Antun, B. Adcock, and A. C. Hansen, “The troublesome kernel: why deep learning for inverse problems is typically unstable,”ArXiv, vol. abs/2001.01258, 2020
2001 arXiv
-
[27]
On instabilities of deep learning in image reconstruction and the potential costs of AI,
V. Antun, F. Renna, C. Poon, B. Adcock, and A. C. Hansen, “On instabilities of deep learning in image reconstruction and the potential costs of AI,”Proceedings of the National Academy of Sciences, vol. 117, no. 48, pp. 30088–30095, 2020. CONTENTS14
2020
-
[28]
Applications, promises, and pitfalls of deep learning for fluorescence image reconstruction,
C. Belthangady and L. A. Royer, “Applications, promises, and pitfalls of deep learning for fluorescence image reconstruction,”Nature methods, vol. 16, no. 12, pp. 1215–1225, 2019
2019
-
[29]
The promise and peril of deep learning in microscopy,
D. P. Hoffman, I. Slavitt, and C. A. Fitzpatrick, “The promise and peril of deep learning in microscopy,”Nature methods, vol. 18, no. 2, pp. 131–132, 2021
2021
-
[30]
Machine learning for medical imaging: methodological failures and recommendations for the future,
G. Varoquaux and V. Cheplygina, “Machine learning for medical imaging: methodological failures and recommendations for the future,”NPJ digital medicine, vol. 5, no. 1, p. 48, 2022
2022
-
[31]
Advancing machine learning for mr image reconstruction with an open competition: Overview of the 2019 fastmri challenge,
F. Knoll, T. Murrell, A. Sriram, N. Yakubova, J. Zbontar, M. Rabbat, A. Defazio, M. J. Muckley, D. K. Sodickson, C. L. Zitnick,et al., “Advancing machine learning for mr image reconstruction with an open competition: Overview of the 2019 fastmri challenge,”Magnetic resonance i...
2019
-
[32]
Results of the 2020 fastmri challenge for machine learning mr image reconstruction,
M. J. Muckley, B. Riemenschneider, A. Radmanesh, S. Kim, G. Jeong, J. Ko, Y. Jun, H. Shin, D. Hwang, M. Mostapha,et al., “Results of the 2020 fastmri challenge for machine learning mr image reconstruction,”IEEE transactions on medical imaging, vol. 40, no. 9, pp. 2306– 2317, 2021
2020
-
[33]
Deep learning reconstruction of accelerated mri: False-positive cartilage delamination inserted in mri arthrography under traction,
W. A. Bosbach, K. C. Merdes, B. Jung, E. Montazeri, S. Anderson, M. Mitrakovic, and K. Daneshvar, “Deep learning reconstruction of accelerated mri: False-positive cartilage delamination inserted in mri arthrography under traction,”Topics in Magnetic Resonance Imaging, vol. 33,...
2024
-
[34]
Robust physical-world attacks on deep learning visual classification,
K. Eykholt, I. Evtimov, E. Fernandes, B. Li, A. Rahmati, C. Xiao, A. Prakash, T. Kohno, and D. Song, “Robust physical-world attacks on deep learning visual classification,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1625– 1634, 2018
2018
-
[35]
Audio adversarial examples: Targeted attacks on speech-to-text,
N. Carlini and D. Wagner, “Audio adversarial examples: Targeted attacks on speech-to-text,” in2018 IEEE security and privacy workshops (SPW), pp. 1–7, IEEE, 2018
2018
-
[36]
Adversarial attacks on medical machine learning,
S. G. Finlayson, J. D. Bowers, J. Ito, J. L. Zittrain, A. L. Beam, and I. S. Kohane, “Adversarial attacks on medical machine learning,”Science, vol. 363, no. 6433, pp. 1287–1289, 2019
2019
-
[37]
Why deep-learning ais are so easy to fool,
D. Heavenet al., “Why deep-learning ais are so easy to fool,”Nature, vol. 574, no. 7777, pp. 163–166, 2019
2019
-
[38]
The mathematics of adversarial attacks in ai–why deep learning is unstable despite the existence of stable neural networks,
A. Bastounis, A. C. Hansen, and V. Vlaˇ ci´ c, “The mathematics of adversarial attacks in ai–why deep learning is unstable despite the existence of stable neural networks,”arXiv preprint arXiv:2109.06098, 2021
2021 arXiv
-
[39]
Some investigations on robustness of deep learning in limited angle tomography,
Y. Huang, T. W¨ urfl, K. Breininger, L. Liu, G. Lauritsch, and A. Maier, “Some investigations on robustness of deep learning in limited angle tomography,” inMedical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, S...
2018
-
[40]
Measuring robustness in deep learning based compressive sensing,
M. Z. Darestani, A. S. Chaudhari, and R. Heckel, “Measuring robustness in deep learning based compressive sensing,” inInternational Conference on Machine Learning, pp. 2433– 2444, PMLR, 2021
2021
-
[41]
Solving inverse problems with deep neural networks– robustness included?,
M. Genzel, J. Macdonald, and M. M¨ arz, “Solving inverse problems with deep neural networks– robustness included?,”IEEE transactions on pattern analysis and machine intelligence, vol. 45, no. 1, pp. 1119–1134, 2022
2022
-
[42]
Improving robustness of deep-learning-based image reconstruction,
A. Raj, Y. Bresler, and B. Li, “Improving robustness of deep-learning-based image reconstruction,” inInternational Conference on Machine Learning, pp. 7932–7942, PMLR, 2020
2020
-
[43]
Adversarial robustness of mr image reconstruction under realistic perturbations,
J. N. Morshuis, S. Gatidis, M. Hein, and C. F. Baumgartner, “Adversarial robustness of mr image reconstruction under realistic perturbations,” inInternational Workshop on Machine Learning for Medical Image Reconstruction, pp. 24–33, Springer, 2022
2022
-
[44]
Localized adversarial artifacts for compressed sensing mri,
R. Alaifari, G. S. Alberti, and T. Gauksson, “Localized adversarial artifacts for compressed sensing mri,”SIAM Journal on Imaging Sciences, vol. 16, no. 4, pp. SC14–SC26, 2023
2023
-
[45]
On hallucinations in tomographic image reconstruction,
S. Bhadra, V. A. Kelkar, F. J. Brooks, and M. A. Anastasio, “On hallucinations in tomographic image reconstruction,”IEEE Transactions on Medical Imaging, vol. 40, pp. 3249–3260, Nov. 2021
2021
-
[46]
The difficulty of computing stable and accurate neural networks: On the barriers of deep learning and Smale’s 18th problem,
M. J. Colbrook, V. Antun, and A. C. Hansen, “The difficulty of computing stable and accurate neural networks: On the barriers of deep learning and Smale’s 18th problem,”Proceedings of the National Academy of Sciences, vol. 119, no. 12, p. e2107151119, 2022
2022
-
[47]
Impact of deep learning- based image super-resolution on binary signal detection,
X. Zhang, V. A. Kelkar, J. Granstedt, H. Li, and M. A. Anastasio, “Impact of deep learning- based image super-resolution on binary signal detection,”Journal of Medical Imaging, vol. 8, no. 6, pp. 065501–065501, 2021
2021
-
[48]
Unified SNR analysis of medical imaging systems,
R. F. Wagner and D. G. Brown, “Unified SNR analysis of medical imaging systems,”Physics in Medicine & Biology, vol. 30, no. 6, p. 489, 1985. CONTENTS15
1985
-
[49]
Icru report 54: Medical imaging-the assessment of image quality-isbn 0-913394- 53-x. april 1996, maryland, usa,
W. Vennart, “Icru report 54: Medical imaging-the assessment of image quality-isbn 0-913394- 53-x. april 1996, maryland, usa,”Radiography, vol. 3, no. 3, pp. 243–244, 1997
1996
-
[50]
Model observers for assessment of image quality,
H. H. Barrett, J. Yao, J. P. Rolland, and K. J. Myers, “Model observers for assessment of image quality,”Proceedings of the National Academy of Sciences, vol. 90, no. 21, pp. 9758–9765, 1993
1993
-
[52]
Strategies for reducing radiation dose in ct,
C. H. McCollough, A. N. Primak, N. Braun, J. Kofler, L. Yu, and J. Christner, “Strategies for reducing radiation dose in ct,”Radiologic Clinics, vol. 47, no. 1, pp. 27–40, 2009
2009
-
[53]
Algorithms for reconstruction with nondiffracting sources,
A. C. Kak and M. Slaney, “Algorithms for reconstruction with nondiffracting sources,” in Principles of computerized tomographic imaging, ch. 3, pp. 49–112, Philadelphia: SIAM, 2001
2001
-
[54]
Acquisition and reconstruction of magnetic resonance imaging,
A.-F. Santiago and V.-S.-F. Gonzalo, “Acquisition and reconstruction of magnetic resonance imaging,” inStatistical Analysis of Noise in MRI Modeling, Filtering and Estimation, ch. 2, pp. 9–29, Switzerland: Springer International Publishing, 2016
2016
-
[55]
Low-dose ct with a residual encoder-decoder convolutional neural network,
H. Chen, Y. Zhang, M. K. Kalra, F. Lin, Y. Chen, P. Liao, J. Zhou, and G. Wang, “Low-dose ct with a residual encoder-decoder convolutional neural network,”IEEE transactions on medical imaging, vol. 36, no. 12, pp. 2524–2535, 2017
2017
-
[56]
Low-dose ct image denoising using a generative adversarial network with wasserstein distance and perceptual loss,
Q. Yang, P. Yan, Y. Zhang, H. Yu, Y. Shi, X. Mou, M. K. Kalra, Y. Zhang, L. Sun, and G. Wang, “Low-dose ct image denoising using a generative adversarial network with wasserstein distance and perceptual loss,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1348–1357, 2018
2018
-
[57]
Deep admm-net for compressive sensing mri,
J. Sun, H. Li, Z. Xu,et al., “Deep admm-net for compressive sensing mri,”Advances in neural information processing systems, vol. 29, 2016
2016
-
[58]
Deep networks and mutual information maximization for cross-modal medical image synthesis,
R. Vemulapalli, H. V. Nguyen, and S. K. Zhou, “Deep networks and mutual information maximization for cross-modal medical image synthesis,” inDeep Learning for Medical Image Analysis(S. K. Zhou, H. Greenspan, and D. Shen, eds.), ch. 16, pp. 381–403, Oxford: Academic Press, 2023
2023
-
[59]
Cross-modality image synthesis from unpaired data using cyclegan: Effects of gradient consistency loss and training data size,
Y. Hiasa, Y. Otake, M. Takao, T. Matsuoka, K. Takashima, A. Carass, J. L. Prince, N. Sugano, and Y. Sato, “Cross-modality image synthesis from unpaired data using cyclegan: Effects of gradient consistency loss and training data size,” inSimulation and Synthesis in Medical Imag...
2018
-
[60]
The data processing inequality and stochastic resonance,
M. D. McDonnell, N. G. Stocks, C. E. Pearce, and D. Abbott, “The data processing inequality and stochastic resonance,” inNoise in Complex Systems and Stochastic Dynamics, vol. 5114, pp. 249–260, SPIE, 2003
2003
-
[61]
On hallucinations in tomographic image reconstruction,
S. Bhadra, V. A. Kelkar, F. J. Brooks, and M. A. Anastasio, “On hallucinations in tomographic image reconstruction,”IEEE transactions on medical imaging, vol. 40, no. 11, pp. 3249–3260, 2021
2021
-
[62]
Null space imaging: nonlinear magnetic encoding fields designed complementary to receiver coil sensitivities for improved acceleration in parallel imaging,
L. K. Tam, J. P. Stockmann, G. Galiana, and R. T. Constable, “Null space imaging: nonlinear magnetic encoding fields designed complementary to receiver coil sensitivities for improved acceleration in parallel imaging,”Magnetic resonance in medicine, vol. 68, no. 4, pp. 1166– 1...
2012
-
[64]
Image artifacts: Appearances, causes, and corrections,
J. Hsieh, “Image artifacts: Appearances, causes, and corrections,” inComputed tomography: principles, design, artifacts, and recent advances, ch. 7, pp. 207–300, Bellingham, Washington: SPIE press, 2003
2003
-
[65]
Artifacts in magnetic resonance imaging,
K. Krupa and M. Bekiesi´ nska-Figatowska, “Artifacts in magnetic resonance imaging,”Polish journal of radiology, vol. 80, p. 93, 2015
2015
-
[66]
P. C. Hansen, J. Jørgensen, and W. R. Lionheart,Computed tomography: algorithms, insight, and just enough theory, ch. 10, pp. 183–209. SIAM, 2021
2021
-
[67]
fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data,
R. Zhao, B. Yaman, Y. Zhang, R. Stewart, A. Dixon, F. Knoll, Z. Huang, Y. W. Lui, M. S. Hansen, and M. P. Lungren, “fastmri+, clinical pathology annotations for knee and brain fully sampled magnetic resonance imaging data,”Scientific Data, vol. 9, no. 1, p. 152, 2022
2022
-
[68]
Lungx challenge for computerized lung CONTENTS16 nodule classification,
S. G. Armato III, K. Drukker, F. Li, L. Hadjiiski, G. D. Tourassi, R. M. Engelmann, M. L. Giger, G. Redmond, K. Farahani, J. S. Kirby,et al., “Lungx challenge for computerized lung CONTENTS16 nodule classification,”Journal of Medical Imaging, vol. 3, no. 4, pp. 044506–044506, 2016
2016
-
[69]
Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning,
K. Yan, X. Wang, L. Lu, and R. M. Summers, “Deeplesion: automated mining of large-scale lesion annotations and universal lesion detection with deep learning,”Journal of medical imaging, vol. 5, no. 3, pp. 036501–036501, 2018
2018
-
[70]
No” zero-shot
V. Udandarao, A. Prabhu, A. Ghosh, Y. Sharma, P. Torr, A. Bibi, S. Albanie, and M. Bethge, “No” zero-shot” without exponential data: Pretraining concept frequency determines multimodal model performance,” inThe Thirty-eighth Annual Conference on Neural Information Processing S...
2024
-
[71]
Generative adversarial networks in medical image augmentation: a review,
Y. Chen, X.-H. Yang, Z. Wei, A. A. Heidari, N. Zheng, Z. Li, H. Chen, H. Hu, Q. Zhou, and Q. Guan, “Generative adversarial networks in medical image augmentation: a review,” Computers in Biology and Medicine, vol. 144, p. 105382, 2022
2022
-
[72]
Data augmentation for medical imaging: A systematic literature review,
F. Garcea, A. Serra, F. Lamberti, and L. Morra, “Data augmentation for medical imaging: A systematic literature review,”Computers in Biology and Medicine, vol. 152, p. 106391, 2023
2023
-
[73]
Synthetic breast ultrasound images: A study to overcome medical data sharing barriers,
J. Xu, Q. Hua, X. Jia, Y. Zheng, Q. Hu, B. Bai, J. Miao, L. Zhu, M. Zhang, R. Tao,et al., “Synthetic breast ultrasound images: A study to overcome medical data sharing barriers,” Research, vol. 7, p. 0532, 2024
2024
-
[74]
Analyzing gan artifacts for simulating mammograms: application towards finding mammographically-occult cancer,
J. Lee and R. M. Nishikawa, “Analyzing gan artifacts for simulating mammograms: application towards finding mammographically-occult cancer,” inMedical Imaging 2022: Computer- Aided Diagnosis, vol. 12033, pp. 78–84, SPIE, 2022
2022
-
[75]
Selective synthetic augmentation with histogan for improved histopathology image classification,
Y. Xue, J. Ye, Q. Zhou, L. R. Long, S. Antani, Z. Xue, C. Cornwell, R. Zaino, K. C. Cheng, and X. Huang, “Selective synthetic augmentation with histogan for improved histopathology image classification,”Medical image analysis, vol. 67, p. 101816, 2021
2021
-
[76]
Hallucination is inevitable: An innate limitation of large language models,
Z. Xu, S. Jain, and M. Kankanhalli, “Hallucination is inevitable: An innate limitation of large language models,”arXiv preprint arXiv:2401.11817, 2024
2024 arXiv
-
[77]
A method for evaluating deep generative models of images for hallucinations in high-order spatial context,
R. Deshpande, M. A. Anastasio, and F. J. Brooks, “A method for evaluating deep generative models of images for hallucinations in high-order spatial context,”Pattern Recognition Letters, vol. 186, pp. 23–29, 2024
2024
-
[78]
Assessing the capacity of a denoising diffusion probabilistic model to reproduce spatial context,
R. Deshpande, M. ¨Ozbey, H. Li, M. A. Anastasio, and F. J. Brooks, “Assessing the capacity of a denoising diffusion probabilistic model to reproduce spatial context,”IEEE Transactions on Medical Imaging, 2024
2024
-
[79]
Assessing the ability of generative adversarial networks to learn canonical medical image statistics,
V. A. Kelkar, D. S. Gotsis, F. J. Brooks, K. Prabhat, K. J. Myers, R. Zeng, and M. A. Anastasio, “Assessing the ability of generative adversarial networks to learn canonical medical image statistics,”IEEE transactions on medical imaging, vol. 42, no. 6, pp. 1799–1808, 2023
2023
-
[80]
A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis,
G. M¨ uller-Franzes, J. M. Niehues, F. Khader, S. T. Arasteh, C. Haarburger, C. Kuhl, T. Wang, T. Han, T. Nolte, S. Nebelung,et al., “A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis,” Sc...
2023
-
[81]
Report on the aapm grand challenge on deep generative modeling for learning medical image statistics,
R. Deshpande, V. A. Kelkar, D. Gotsis, P. Kc, R. Zeng, K. J. Myers, F. J. Brooks, and M. A. Anastasio, “Report on the aapm grand challenge on deep generative modeling for learning medical image statistics,”Medical Physics, vol. 52, no. 1, pp. 4–20, 2025
2025
-
[82]
A knowledge-based method for detecting network-induced shape artifacts in synthetic images,
R. Deshpande, M. Lago, A. Subbaswamy, S. Kahaki, J. G. Delfino, A. Badano, and G. Zamzmi, “A knowledge-based method for detecting network-induced shape artifacts in synthetic images,” inMedical Imaging with Deep Learning, 2025
2025
-
[83]
Distribution matching losses can hallucinate features in medical image translation,
J. P. Cohen, M. Luck, and S. Honari, “Distribution matching losses can hallucinate features in medical image translation,” inMedical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceeding...
2018
-
[84]
Cyclegan for virtual stain transfer: Is seeing really believing?,
J. Vasiljevi´ c, Z. Nisar, F. Feuerhake, C. Wemmert, and T. Lampert, “Cyclegan for virtual stain transfer: Is seeing really believing?,”Artificial Intelligence in Medicine, vol. 133, p. 102420, 2022
2022
-
[85]
Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive models,
S. Bond-Taylor, A. Leach, Y. Long, and C. G. Willcocks, “Deep generative modelling: A comparative review of vaes, gans, normalizing flows, energy-based and autoregressive models,”IEEE transactions on pattern analysis and machine intelligence, vol. 44, no. 11, pp. 7327–7347, 2021
2021
-
[86]
Benchmarking large language models for news summarization,
T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto, “Benchmarking large language models for news summarization,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 39–57, 2024
2024
-
[87]
Evaluation of chatgpt as a question answering system for answering complex questions,
Y. Tan, D. Min, Y. Li, W. Li, N. Hu, Y. Chen, and G. Qi, “Evaluation of chatgpt as a question answering system for answering complex questions,”arXiv preprint arXiv:2303.07992, 2023
2023 arXiv
-
[88]
Multilingual machine CONTENTS17 translation with large language models: Empirical results and analysis,
W. Zhu, H. Liu, Q. Dong, J. Xu, S. Huang, L. Kong, J. Chen, and L. Li, “Multilingual machine CONTENTS17 translation with large language models: Empirical results and analysis,”arXiv preprint arXiv:2304.04675, 2023
2023 arXiv
-
[89]
Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics,
A. Pagnoni, V. Balachandran, and Y. Tsvetkov, “Understanding factuality in abstractive summarization with frank: A benchmark for factuality metrics,”arXiv preprint arXiv:2104.13346, 2021
2021 arXiv
-
[90]
Challenges in building intelligent open-domain dialog systems,
M. Huang, X. Zhu, and J. Gao, “Challenges in building intelligent open-domain dialog systems,” ACM Transactions on Information Systems (TOIS), vol. 38, no. 3, pp. 1–32, 2020
2020
-
[91]
Language models are unsupervised multitask learners,
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever,et al., “Language models are unsupervised multitask learners,”OpenAI blog, vol. 1, no. 8, p. 9, 2019
2019
-
[92]
Training language models to follow instructions with human feedback,
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray,et al., “Training language models to follow instructions with human feedback,”Advances in neural information processing systems, vol. 35, pp. 27730–27744, 2022
2022
-
[93]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” inProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologie...
2019
-
[94]
How much knowledge can you pack into the parameters of a language model?,
A. Roberts, C. Raffel, and N. Shazeer, “How much knowledge can you pack into the parameters of a language model?,”arXiv preprint arXiv:2002.08910, 2020
2002 arXiv
-
[95]
Black swans and the domains of statistics,
N. N. Taleb, “Black swans and the domains of statistics,”The american statistician, vol. 61, no. 3, pp. 198–200, 2007
2007
-
[96]
Survey of hallucination in natural language generation,
Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii, Y. J. Bang, A. Madotto, and P. Fung, “Survey of hallucination in natural language generation,”ACM computing surveys, vol. 55, no. 12, pp. 1–38, 2023
2023
-
[97]
Diversifying dialogue generation with non-conversational text,
H. Su, X. Shen, S. Zhao, X. Zhou, P. Hu, R. Zhong, C. Niu, and J. Zhou, “Diversifying dialogue generation with non-conversational text,”arXiv preprint arXiv:2005.04346, 2020
2005 arXiv
-
[98]
Union: An unreferenced metric for evaluating open-ended story generation,
J. Guan and M. Huang, “Union: An unreferenced metric for evaluating open-ended story generation,”arXiv preprint arXiv:2009.07602, 2020
2009 arXiv
-
[99]
Retrieval augmentation reduces hallucination in conversation,
K. Shuster, S. Poff, M. Chen, D. Kiela, and J. Weston, “Retrieval augmentation reduces hallucination in conversation,”arXiv preprint arXiv:2104.07567, 2021
2021 arXiv
-
[100]
Towards conversational diagnostic ai,
T. Tu, A. Palepu, M. Schaekermann, K. Saab, J. Freyberg, R. Tanno, A. Wang, B. Li, M. Amin, N. Tomasev,et al., “Towards conversational diagnostic ai,”arXiv preprint arXiv:2401.05654, 2024
2024 arXiv
-
[101]
Towards generalist biomedical ai,
T. Tu, S. Azizi, D. Driess, M. Schaekermann, M. Amin, P.-C. Chang, A. Carroll, C. Lau, R. Tanno, I. Ktena,et al., “Towards generalist biomedical ai,”Nejm Ai, vol. 1, no. 3, p. AIoa2300138, 2024
2024
-
[103]
Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation,
C.-Y. Li, K.-J. Chang, C.-F. Yang, H.-Y. Wu, W. Chen, H. Bansal, L. Chen, Y.-P. Yang, Y.-C. Chen, S.-P. Chen,et al., “Towards a holistic framework for multimodal llm in 3d brain ct radiology report generation,”Nature Communications, vol. 16, no. 1, p. 2258, 2025
2025
-
[104]
Collaboration between clinicians and vision–language models in radiology report generation,
R. Tanno, D. G. Barrett, A. Sellergren, S. Ghaisas, S. Dathathri, A. See, J. Welbl, C. Lau, T. Tu, S. Azizi,et al., “Collaboration between clinicians and vision–language models in radiology report generation,”Nature Medicine, vol. 31, no. 2, pp. 599–608, 2025
2025
-
[105]
Glue: A multi-task benchmark and analysis platform for natural language understanding,
A. Wang, “Glue: A multi-task benchmark and analysis platform for natural language understanding,”arXiv preprint arXiv:1804.07461, 2018
2018 arXiv
-
[106]
Holistic evaluation of language models,
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar,et al., “Holistic evaluation of language models,”arXiv preprint arXiv:2211.09110, 2022
2022 arXiv
-
[107]
Measuring massive multitask language understanding,
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,”arXiv preprint arXiv:2009.03300, 2020
2009 arXiv
-
[108]
Judging llm-as-a-judge with mt-bench and chatbot arena,
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing,et al., “Judging llm-as-a-judge with mt-bench and chatbot arena,”Advances in Neural Information Processing Systems, vol. 36, pp. 46595–46623, 2023
2023
-
[109]
A review of current trends, techniques, and challenges in large language models (llms),
R. Patil and V. Gudivada, “A review of current trends, techniques, and challenges in large language models (llms),”Applied Sciences, vol. 14, no. 5, p. 2074, 2024
2024
-
[110]
Llama 2: Open foundation and fine-tuned chat models,
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale,et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023
2023 arXiv
-
[111]
Mitigating the alignment tax of rlhf,
Y. Lin, H. Lin, W. Xiong, S. Diao, J. Liu, J. Zhang, R. Pan, H. Wang, W. Hu, H. Zhang,et al., CONTENTS18 “Mitigating the alignment tax of rlhf,”arXiv preprint arXiv:2309.06256, 2023
2023 arXiv
-
[112]
” do anything now
X. Shen, Z. Chen, M. Backes, Y. Shen, and Y. Zhang, “” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models,” inProceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 1671– 1685, 2024
2024
-
[113]
Fine-tuning aligned language models compromises safety, even when users do not intend to!,
X. Qi, Y. Zeng, T. Xie, P.-Y. Chen, R. Jia, P. Mittal, and P. Henderson, “Fine-tuning aligned language models compromises safety, even when users do not intend to!,”arXiv preprint arXiv:2310.03693, 2023
2023 arXiv
-
[114]
Siren’s song in the ai ocean: a survey on hallucination in large language models,
Y. Zhang, Y. Li, L. Cui, D. Cai, L. Liu, T. Fu, X. Huang, E. Zhao, Y. Zhang, Y. Chen,et al., “Siren’s song in the ai ocean: a survey on hallucination in large language models,”arXiv preprint arXiv:2309.01219, 2023
2023 arXiv
-
[115]
The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,
G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay, “The refinedweb dataset for falcon llm: outperforming curated corpora with web data, and web data only,”arXiv preprint arXiv:2306.01116, 2023
2023 arXiv
-
[116]
Toward expert-level medical question answering with large language models,
K. Singhal, T. Tu, J. Gottweis, R. Sayres, E. Wulczyn, M. Amin, L. Hou, K. Clark, S. R. Pfohl, H. Cole-Lewis,et al., “Toward expert-level medical question answering with large language models,”Nature Medicine, pp. 1–8, 2025
2025
-
[117]
Scaling instruction-finetuned language models,
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y. Tay, W. Fedus, Y. Li, X. Wang, M. Dehghani, S. Brahma,et al., “Scaling instruction-finetuned language models,”Journal of Machine Learning Research, vol. 25, no. 70, pp. 1–53, 2024
2024
-
[118]
Pmc-llama: toward building open-source language models for medicine,
C. Wu, W. Lin, X. Zhang, Y. Zhang, W. Xie, and Y. Wang, “Pmc-llama: toward building open-source language models for medicine,”Journal of the American Medical Informatics Association, vol. 31, no. 9, pp. 1833–1843, 2024
2024
-
[119]
Vision-language models for medical report generation and visual question answering: A review,
I. Hartsock and G. Rasool, “Vision-language models for medical report generation and visual question answering: A review,”Frontiers in Artificial Intelligence, vol. 7, p. 1430984, 2024
2024
-
[120]
Evaluating object hallucination in large vision-language models,
Y. Li, Y. Du, K. Zhou, J. Wang, W. X. Zhao, and J.-R. Wen, “Evaluating object hallucination in large vision-language models,”arXiv preprint arXiv:2305.10355, 2023
2023 arXiv
-
[121]
Eyes wide shut? exploring the visual shortcomings of multimodal llms,
S. Tong, Z. Liu, Y. Zhai, Y. Ma, Y. LeCun, and S. Xie, “Eyes wide shut? exploring the visual shortcomings of multimodal llms,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 9568–9578, 2024
2024
-
[122]
Investigating the catastrophic forgetting in multimodal large language models,
Y. Zhai, S. Tong, X. Li, M. Cai, Q. Qu, Y. J. Lee, and Y. Ma, “Investigating the catastrophic forgetting in multimodal large language models,”arXiv preprint arXiv:2309.10313, 2023
2023 arXiv
-
[123]
Hallucination index: An image quality metric for generative reconstruction models,
M. Tivnan, S. Yoon, Z. Chen, X. Li, D. Wu, and Q. Li, “Hallucination index: An image quality metric for generative reconstruction models,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 449–458, Springer, 2024
2024
-
[124]
Language models with conformal factuality guarantees,
C. Mohri and T. Hashimoto, “Language models with conformal factuality guarantees,”arXiv preprint arXiv:2402.10978, 2024
2024 arXiv
-
[125]
Large language model validity via enhanced conformal prediction methods,
J. Cherian, I. Gibbs, and E. Candes, “Large language model validity via enhanced conformal prediction methods,”Advances in Neural Information Processing Systems, vol. 37, pp. 114812–114842, 2024
2024
-
[126]
Med-halt: Medical domain hallucination test for large language models,
A. Pal, L. K. Umapathi, and M. Sankarasubbu, “Med-halt: Medical domain hallucination test for large language models,”arXiv preprint arXiv:2307.15343, 2023
2023 arXiv
-
[127]
Cares: A comprehensive benchmark of trustworthiness in medical vision language models,
P. Xia, Z. Chen, J. Tian, Y. Gong, R. Hou, Y. Xu, Z. Wu, Z. Fan, Y. Zhou, K. Zhu,et al., “Cares: A comprehensive benchmark of trustworthiness in medical vision language models,” Advances in Neural Information Processing Systems, vol. 37, pp. 140334–140365, 2024
2024
-
[128]
Detecting and evaluating medical hallucinations in large vision language models,
J. Chen, D. Yang, T. Wu, Y. Jiang, X. Hou, M. Li, S. Wang, D. Xiao, K. Li, and L. Zhang, “Detecting and evaluating medical hallucinations in large vision language models,”arXiv preprint arXiv:2406.10185, 2024
2024 arXiv
-
[129]
State of what art? a call for multi-prompt llm evaluation,
M. Mizrahi, G. Kaplan, D. Malkin, R. Dror, D. Shahaf, and G. Stanovsky, “State of what art? a call for multi-prompt llm evaluation,”Transactions of the Association for Computational Linguistics, vol. 12, pp. 933–949, 2024
2024
-
[130]
Adversarial glue: A multi-task benchmark for robustness evaluation of language models,
B. Wang, C. Xu, S. Wang, Z. Gan, Y. Cheng, J. Gao, A. H. Awadallah, and B. Li, “Adversarial glue: A multi-task benchmark for robustness evaluation of language models,”arXiv preprint arXiv:2111.02840, 2021
2021 arXiv
-
[131]
On the robustness of chatgpt: An adversarial and out-of-distribution perspective,
J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y. Wang, L. Yang, H. Huang, W. Ye, X. Geng, et al., “On the robustness of chatgpt: An adversarial and out-of-distribution perspective,” arXiv preprint arXiv:2302.12095, 2023
2023 arXiv
-
[132]
Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.,
B. Wang, W. Chen, H. Pei, C. Xie, M. Kang, C. Zhang, C. Xu, Z. Xiong, R. Dutta, R. Schaeffer, et al., “Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.,” in NeurIPS, 2023
2023
-
[133]
Learning a variational network for reconstruction of accelerated mri data,
K. Hammernik, T. Klatzer, E. Kobler, M. P. Recht, D. K. Sodickson, T. Pock, and F. Knoll, CONTENTS19 “Learning a variational network for reconstruction of accelerated mri data,”Magnetic resonance in medicine, vol. 79, no. 6, pp. 3055–3071, 2018
2018
-
[134]
Learn: Learned experts’ assessment-based reconstruction network for sparse- data ct,
H. Chen, Y. Zhang, Y. Chen, J. Zhang, W. Zhang, H. Sun, Y. Lv, P. Liao, J. Zhou, and G. Wang, “Learn: Learned experts’ assessment-based reconstruction network for sparse- data ct,”IEEE transactions on medical imaging, vol. 37, no. 6, pp. 1333–1347, 2018
2018
-
[135]
Ambientgan: Generative models from lossy measurements,
A. Bora, E. Price, and A. G. Dimakis, “Ambientgan: Generative models from lossy measurements,” inInternational conference on learning representations, 2018
2018
-
[136]
Latent retrieval for weakly supervised open domain question answering,
K. Lee, M.-W. Chang, and K. Toutanova, “Latent retrieval for weakly supervised open domain question answering,”arXiv preprint arXiv:1906.00300, 2019
1906 arXiv
-
[137]
Retrieval augmented language model pre-training,
K. Guu, K. Lee, Z. Tung, P. Pasupat, and M. Chang, “Retrieval augmented language model pre-training,” inInternational conference on machine learning, pp. 3929–3938, PMLR, 2020
2020
-
[138]
Retrieval-augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. K¨ uttler, M. Lewis, W.-t. Yih, T. Rockt¨ aschel,et al., “Retrieval-augmented generation for knowledge-intensive nlp tasks,”Advances in neural information processing systems, vol. 33, pp. 9459–9474, 2020
2020
-
[139]
Knowledge conflicts for llms: A survey,
R. Xu, Z. Qi, Z. Guo, C. Wang, H. Wang, Y. Zhang, and W. Xu, “Knowledge conflicts for llms: A survey,”arXiv preprint arXiv:2403.08319, 2024
2024 arXiv
-
[140]
Knowledge graphs, large language models, and hallucinations: An nlp perspective,
E. Lavrinovics, R. Biswas, J. Bjerva, and K. Hose, “Knowledge graphs, large language models, and hallucinations: An nlp perspective,”Journal of Web Semantics, vol. 85, p. 100844, 2025
2025
-
[141]
Towards mitigating hallucination in large language models via self-reflection,
Z. Ji, T. Yu, Y. Xu, N. Lee, E. Ishii, and P. Fung, “Towards mitigating hallucination in large language models via self-reflection,”arXiv preprint arXiv:2310.06271, 2023
2023 arXiv
-
[142]
Improving factuality and reasoning in language models through multiagent debate,
Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch, “Improving factuality and reasoning in language models through multiagent debate,” inForty-first International Conference on Machine Learning, 2023
2023
-
[143]
Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration,
Z. Wang, S. Mao, W. Wu, T. Ge, F. Wei, and H. Ji, “Unleashing the emergent cognitive synergy in large language models: A task-solving agent through multi-persona self-collaboration,” arXiv preprint arXiv:2307.05300, 2023
2023 arXiv
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.