Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Best Practices for Large Language Models in Radiology

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This review argues that radiology departments should not pretrain their own large language models, but instead follow a staged escalation from prompting to retrieval-augmented generation to fine-tuning only when needed.

desk verdict A solid, well-organized review of LLM best practices in radiology whose practical escalation ladder is useful and appropriately hedged, but the manuscript is unfinished (citation placeholders, a broken reference) and the ladder's ordering lacks direct radiology evidence. read the letter →

arxiv 2412.01233 v1 pith:4NHTLOTR submitted 2024-12-02 cs.AI

classification cs.AI
keywords largelanguagemodelsradiologypromptengineeringretrieval-augmentedgenerationfine-tuningLoRAmodelevaluationprivacy
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review argues that radiology departments do not need to pretrain their own large language models. The authors recommend an iterative escalation: start with prompt optimization on general-purpose models, add in-context examples and retrieval-augmented generation to supply relevant knowledge, enable tools, and only fine-tune if the simpler steps are insufficient. The payoff is that radiology teams can deploy capable language tools for reporting, summarization, and communication while controlling compute cost, privacy risk, and confabulation. The recommendation rests on a synthesis of recent LLM research and radiology-specific studies rather than a new experiment.

What carries the argument

The organizing mechanism is the iterative optimization ladder (Figure 5). It orders adaptation techniques by increasing complexity and cost: prompting, in-context learning, retrieval-augmented generation, tool access, and fine-tuning. Each higher rung permanently modifies the model or adds persistent infrastructure, so the recommendation is to stop at the lowest rung that meets the task requirement, with prompt engineering remaining relevant even after fine-tuning.

What would settle it

A controlled comparison on a representative radiology task in which a small fine-tuned model clearly beats a well-prompted large model with retrieval augmentation, at equal or lower cost, would undercut the recommended order of escalation.

Watch

Extended reading notes

Core claim

The central claim is that the most effective path to using large language models in radiology follows a cost-ordered adaptation ladder. Its rungs are: iteratively refine the prompt; include a few task examples (in-context learning); ground the model with retrieved external knowledge (retrieval-augmented generation); give the model access to tools; and finally, if needed, fine-tune with curated data, preferably via parameter-efficient methods like LoRA. The paper further claims that pretraining from scratch is impractical for most radiology labs, that open locally hosted models should be favored over closed hosted models for privacy and reproducibility, and that evaluation must combine automated metrics with radiologist review and continue after deployment.

Load-bearing premise

The entire ladder assumes that techniques validated in general-domain LLM research and a small set of early radiology studies will transfer reliably to real radiology workflows across different models, hospitals, and tasks.

Editorial extensions

If this is right

  • Radiology AI projects can skip the multi-million-dollar pretraining step and begin with prompted general-purpose models already available today.
  • Factual reliability can be improved without retraining by grounding outputs in retrieved guidelines and reports, which also makes outputs easier to verify.
  • Privacy-sensitive deployment is feasible with locally hosted open models, since imaging data and notes do not need to be sent to an external API.
  • Fine-tuning, when needed, can be carried out cheaply with LoRA-style methods, making customization attainable for smaller labs.
  • Evaluation that pairs automated scores with radiologist review becomes the gatekeeper for deciding when to escalate to the next rung.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not say how to choose the threshold that triggers a move up the ladder; an operational rule would be to escalate only when the current rung fails an agreed clinical benchmark.
  • The open-vs-closed model recommendation is driven by governance concerns rather than head-to-head accuracy evidence; head-to-head radiology benchmarks could change the balance.
  • As multimodal LLMs improve, the ladder may effectively start at a higher rung, with image-aware prompting already covering tasks like draft report generation without any fine-tuning.
  • The ladder can be tested prospectively: fix an evaluation set, measure each rung's gain, and check whether the gains plateau before fine-tuning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript is a narrative review of best practices for using and adapting large language models (LLMs) in radiology. It covers technical foundations (transformers, training pipeline, multimodal extensions), possible applications (clinical, operational, research, education), desiderata, limitations (confabulation, privacy, bias, reproducibility), adaptation methods (prompting, in-context learning, retrieval-augmented generation, instruction tuning, alignment, parameter-efficient fine-tuning, quantization, distillation), and evaluation strategies (automatic metrics, radiology-specific benchmarks, human expert assessment). The central recommendation, stated in Section 5, is an iterative escalation ladder: start with optimized prompting of general-purpose models, then enrich context through in-context learning and retrieval-augmented generation, enable access to tools, and finally fine-tune if necessary, while preferring local open models and maintaining ongoing human evaluation.

Significance. If the recommendations are accepted, this review could serve as a practical entry point for radiology teams without dedicated machine-learning resources, helping them decide how to deploy LLMs for language-oriented tasks in radiology. The paper is timely and broad, and it has several concrete strengths: Table 1 and Table 2 are actionable and clearly organized; the authors explicitly hedge that effectiveness depends on the model and may change over time; practical cost estimates for pretraining and fine-tuning are useful; and the review covers a wide range of recent radiology-specific and general-domain sources. As a narrative review rather than a primary study, its value lies in synthesis and guidance, and the authors generally distinguish between established practice, emerging evidence, and suggestions based on practical experience.

major comments (3)
  1. [Section 4.1 (Multimodality paragraph) and Section 4.1.1] The manuscript contains unresolved citation placeholders: 'radiological images.?, 90' in Section 4.1 and 'Model Cards and Versioning.?, 64' in Section 4.1.1. These incomplete citations undermine the evidence base for statements about multimodal LLM performance and model cards, and they are not acceptable in a submitted manuscript. Please replace the '?' placeholders with the intended reference keys or remove the citations.
  2. [Section 4.3 and Section 5 (Central recommendation)] The central recommendation is an ordered escalation ladder: prompting, then in-context learning/RAG, then tools, then fine-tuning. The manuscript states in Section 4.3 that 'Prompting will only moderately influence concrete radiology-specific abilities,' but the only comparative citation for the prompting-versus-fine-tuning choice is ref. 120, a general-domain knowledge-injection study. No radiology-specific head-to-head evidence is provided for the ordering. Because the ladder is the paper's main takeaway, the authors should either cite radiology-specific comparisons (e.g., prompting versus fine-tuning for report summarization or information extraction) or explicitly reframe the ordering as a resource-based heuristic rather than an evidence-based effectiveness ranking, and discuss the evidence gap. The existing hedges about model dependence are helpful but do not directly address the ordinal claim.
  3. [Section 4.3 and Section 5] The proposed ladder is text-centric and does not clearly account for the core radiology task of image interpretation. The paper acknowledges that large multimodal models are needed for image-based tasks and advises readers to 'keep an eye' on them, but the central recommendation does not integrate multimodal model selection or vision bridges into the escalation path. For many real-world radiology tasks, prompt optimization and text-based RAG cannot supply visual competence; the appropriate starting point may be selecting a multimodal model or pairing a text LLM with a vision module. Please clarify how the escalation ladder applies to language-only tasks versus image-interpretation tasks, and incorporate the multimodality discussion into the central recommendation.
minor comments (5)
  1. [Section 2.2] There is a typo: 'particulary' should be 'particularly.'
  2. [Table 1, Few-shot prompting row] The example JSON in the few-shot prompting row has mismatched braces, making the example harder to parse; consider formatting it as a complete, valid JSON object.
  3. [Figure 5 caption] The phrase 'The presented method groups increase in complexity' is awkward; consider rewording to 'The presented groups of methods increase in complexity.'
  4. [Section 4.1 and elsewhere] Several author names contain stray spacing or OCR-like artifacts (e.g., 'LLaV A' for LLaVA, 'V o' in reference author names); a careful proofreading pass is needed.
  5. [Footnote 2 in Section 2.4] The cost calculation is written as '$33/8 GPUs/h', which is ambiguous; clarify whether the $33 is per GPU-hour or for a group of GPUs, and make the arithmetic explicit.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: central recommendations are practical heuristics grounded in external evidence, with self-citations used only as illustrative examples.

full rationale

This paper is a narrative review of best practices rather than a derivation, so the usual circularity patterns (self-definitional fits, fitted inputs called predictions, imported uniqueness theorems, ansatz smuggled via citation) do not apply. The central escalation recommendation in Section 5 and Figure 5 — start with prompting, then in-context learning/RAG, then tools, then fine-tuning if necessary — is explicitly hedged: Table 1 notes that 'the effectiveness of some recommendations depends on the model, and best practices may change over time,' and the Conclusions state that 'the effectiveness of these techniques may vary across different LLM types and versions.' The comparative claim about prompting versus fine-tuning is supported by external general-domain work (ref. 120), while the RAG recommendation cites the original external RAG paper (ref. 118) and the few-shot/in-context claim cites the external GPT-3 paper (ref. 10). Self-citations (RadAdapt ref. 12, CheXagent ref. 26, Clinical Text Summarization ref. 52, RaLEs ref. 106, Almanac ref. 107) appear as examples of existing systems, benchmarks, or studies rather than as the load-bearing warrant for the paper's recommendations. No equation is fitted and then renamed as a prediction, and no result is forced by a self-citation chain. The main scientific limitation is evidentiary rather than circular: the ordinal structure of the escalation ladder is not backed by a head-to-head radiology benchmark across all rungs, but the authors acknowledge uncertainty and present the ladder as iterative guidance. Under the hard rule that circularity requires a quotable reduction of a claim to its own inputs, no such step is present; the score reflects only the presence of non-load-bearing self-citations in a review context.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

Free parameters: none, the paper fits no data and introduces no fitted quantities. Axioms: two domain assumptions about transferability of general LLM findings to radiology and about the accuracy of the cited literature. Invented entities: none.

assumptions (2)
  • domain assumption Recommendations validated on general-domain LLMs and a few radiology studies transfer to real radiology deployments.
    The paper builds its advice on general LLM references and early radiology pilots (refs 10, 52, 57, 78, 118); the authors themselves qualify this in Table 1's note and the Conclusions, but the transfer is still assumed for the central ladder.
  • domain assumption The cited empirical results (capabilities, costs, failure modes, evaluation metrics) are accurate.
    Claims such as the cost estimates in the footnotes, LoRA and QLoRA effectiveness, and confabulation statistics rest on sources the paper does not reproduce; two in-text citations are unresolved placeholders.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Best Practices for Large Language Models in Radiology." pith.science (2026). https://pith.science/paper/4NHTLOTR

@misc{pith2026241201233,
  author       = {Pith},
  title        = {Pith review of: Best Practices for Large Language Models in Radiology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4NHTLOTR}},
  note         = {Machine review of arXiv:2412.01233}
}
read the original abstract

At the heart of radiological practice is the challenge of integrating complex imaging data with clinical information to produce actionable insights. Nuanced application of language is key for various activities, including managing requests, describing and interpreting imaging findings in the context of clinical data, and concisely documenting and communicating the outcomes. The emergence of large language models (LLMs) offers an opportunity to improve the management and interpretation of the vast data in radiology. Despite being primarily general-purpose, these advanced computational models demonstrate impressive capabilities in specialized language-related tasks, even without specific training. Unlocking the potential of LLMs for radiology requires basic understanding of their foundations and a strategic approach to navigate their idiosyncrasies. This review, drawing from practical radiology and machine learning expertise and recent literature, provides readers insight into the potential of LLMs in radiology. It examines best practices that have so far stood the test of time in the rapidly evolving landscape of LLMs. This includes practical advice for optimizing LLM characteristics for radiology practices along with limitations, effective prompting, and fine-tuning strategies.

Figures

Figures reproduced from arXiv: 2412.01233 by the authors.

Figure 1
Figure 1. Process of training a large language model (LLM). (1) First, an LLM is pretrained in a self-supervised way to predict [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Simplified example of enhancing model responses by incorporating additional information from external sources with [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Large language models (LLMs) in agentic workflows. In this example, the user asks for lung nodules in a CT scan [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Simplified overview of training and domain adaptation methods for large language models (LLMs). Initially, any LLM [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: Iterative LLM Optimization for Radiology Tasks. When applying a large language model (LLM) to a new problem, [PITH_FULL_IMAGE:figures/full_fig_p014_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine

    cs.AI 2025-02 conditional novelty 3.0 of 10

    A PRISMA-ScR scoping review of 144 studies finds the field shifting from text-only LLMs to multimodal AI in medicine, with evaluation and data diversity still the main bottlenecks.

Reference graph

Works this paper leans on

120 extracted references · 20 canonical work pages · cited by 1 Pith paper

  1. [1]

    Langlotz

    Curtis P. Langlotz. The Radiology Report : A Guide to Thoughtful Communication for Radiologists and Other Medical Professionals . CreateSpace Independent Publishing Platform, November 2015. Google-Books-ID: K3b3jgEACAAJ

  2. [2]

    GPTs are GPTs : An Early Look at the Labor Market Impact Potential of Large Language Models , August 2023

    Tyna Eloundou, Sam Manning, Pamela Mishkin, and Daniel Rock. GPTs are GPTs : An Early Look at the Labor Market Impact Potential of Large Language Models , August 2023. arXiv:2303.10130 [cs, econ, q-fin]

  3. [3]

    Hentel, Beatriu Reig, George Shih, and Linda Moy

    Yiqiu Shen, Laura Heacock, Jonathan Elias, Keith D. Hentel, Beatriu Reig, George Shih, and Linda Moy. ChatGPT and other large language models are double-edged swords. Radiology , 307(2):e230163, 2023. https://doi.org/10.1148/radiol.230163

  4. [4]

    The Hitchhiker 's Guide to the Galaxy

    Douglas Adams. The Hitchhiker 's Guide to the Galaxy . Hitchhiker series. Harmony Books, 1979

  5. [5]

    Chatbots and Large Language Models in Radiology : A Practical Primer for Clinical and Research Applications

    Rajesh Bhayana. Chatbots and Large Language Models in Radiology : A Practical Primer for Clinical and Research Applications . Radiology , 310(1):e232756, January 2024. Publisher: Radiological Society of North America

  6. [6]

    A neural probabilistic language model

    Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. A neural probabilistic language model. Advances in neural information processing systems , 13, 2000

  7. [7]

    A review of recurrent neural networks: LSTM cells and network architectures

    Yong Yu, Xiaosheng Si, Changhua Hu, and Jianxun Zhang. A review of recurrent neural networks: LSTM cells and network architectures. Neural C omputation , 31(7):1235--1270, 2019

  8. [8]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention Is All You Need , December 2017. arXiv:1706.03762 [cs]

Show all 120 references
  1. [9]

    Bert: Pre -training of deep bidirectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre -training of deep bidirectional transformers for language understanding. 2018. arXiv:1810.04805 [cs]

  2. [10]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  3. [11]

    Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei

    Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling Laws for Neural Language Models , January 2020. arXiv:2001.08361 [cs, stat]

  4. [12]

    RadAdapt : Radiology report summarization via lightweight domain adaptation of large language models

    Dave Van Veen, Cara Van Uden, Maayane Attias, Anuj Pareek, Christian Bluethgen, Malgorzata Polacin, Wah Chiu, Jean-Benoit Delbrouck, Juan Zambrano Chaves, Curtis Langlotz, Akshay Chaudhari, and John Pauly. RadAdapt : Radiology report summarization via lightweight domain adapta...

  5. [13]

    Xlnet: Generalized autoregressive pretraining for language understanding

    Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. Xlnet: Generalized autoregressive pretraining for language understanding. Advances in neural information processing systems , 32, 2019

  6. [14]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  7. [15]

    VisualGPT : Data -efficient adaptation of pretrained language models for image captioning

    Jun Chen, Han Guo, Kai Yi, Boyang Li, and Mohamed Elhoseiny. VisualGPT : Data -efficient adaptation of pretrained language models for image captioning. In 2022 IEEE / CVF conference on computer vision and pattern recognition ( CVPR ) , pages 18009--18019, 2022

  8. [16]

    Visual instruction tuning

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. In Thirty-seventh conference on neural information processing systems , 2023

  9. [17]

    Toolformer: Language Models Can Teach Themselves to Use Tools , February 2023

    Timo Schick, Jane Dwivedi-Yu , Roberto Dess \`i , Roberta Raileanu, Maria Lomeli, Luke Zettlemoyer, Nicola Cancedda, and Thomas Scialom. Toolformer: Language Models Can Teach Themselves to Use Tools , February 2023. arXiv:2302.04761 [cs]

  10. [18]

    WebGPT : Browser -assisted question-answering with human feedback, June 2022

    Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. WebG...

  11. [19]

    Brady, Bibb Allen, Jaron Chong, Elmar Kotter, Nina Kottler, John Mongan, Lauren Oakden-Rayner, Daniel Pinto dos Santos, An Tang, Christoph Wald, and John Slavotinek

    Adrian P. Brady, Bibb Allen, Jaron Chong, Elmar Kotter, Nina Kottler, John Mongan, Lauren Oakden-Rayner, Daniel Pinto dos Santos, An Tang, Christoph Wald, and John Slavotinek. Developing, Purchasing , Implementing and Monitoring AI Tools in Radiology : Practical Considerations...

  12. [20]

    Large Language Models : A Guide for Radiologists

    Sunkyu Kim, Choong-kun Lee, and Seung-seob Kim. Large Language Models : A Guide for Radiologists . Korean Journal of Radiology , 25(2):126--133, February 2024

  13. [21]

    Kuo, Subathra Adithan, Eduardo Pontes Reis, Stephen Kwak, Vasantha Kumar Venugopal, Chloe P

    Benjamin Yan, Ruochen Liu, David E. Kuo, Subathra Adithan, Eduardo Pontes Reis, Stephen Kwak, Vasantha Kumar Venugopal, Chloe P. O'Connell, Agustina Saenz, Pranav Rajpurkar, and Michael Moor. Style- Aware Radiology Report Generation with RadGraph and Few - Shot Prompting , Oct...

  14. [22]

    Schmidt, Jarrel C

    Reuben A. Schmidt, Jarrel C. Y. Seah, Ke Cao, Lincoln Lim, Wei Lim, and Justin Yeung. Generating Large Language Models for Detection of Speech Recognition Errors in Radiology Reports . Radiology: Artificial Intelligence , page e230205, January 2024

  15. [23]

    Nestor, Ali Soroush, Pierre A

    Liyan Tang, Zhaoyi Sun, Betina Idnay, Jordan G. Nestor, Ali Soroush, Pierre A. Elias, Ziyang Xu, Ying Ding, Greg Durrett, Justin F. Rousseau, Chunhua Weng, and Yifan Peng. Evaluating large language models on medical evidence summarization. npj Digital Medicine , 6(1):1--8, Aug...

  16. [24]

    Wong, David Blumenthal, Isaac Kohane, and the editors and editorial board of NEJM AI

    Daphne Koller, Andrew Beam, Arjun Manrai, Euan Ashley, Xiaoxuan Liu, Judy Gichoya, Chris Holmes, James Zou, Noa Dagan, Tien Y. Wong, David Blumenthal, Isaac Kohane, and the editors and editorial board of NEJM AI . Why We Support and Encourage the Use of Large Language Models i...

  17. [25]

    Raymond Geis, Adrian P

    J. Raymond Geis, Adrian P. Brady, Carol C. Wu, Jack Spencer, Erik Ranschaert, Jacob L. Jaremko, Steve G. Langer, Andrea Borondy Kitts, Judy Birch, William F. Shields, Robert van den Hoven van Genderen, Elmar Kotter, Judy Wawira Gichoya, Tessa S. Cook, Matthew B. Morgan, An Tan...

  18. [26]

    Tsai, Andrew Johnston, Cameron Olsen, Tanishq Mathew Abraham, Sergios Gatidis, Akshay S

    Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, Emily B. Tsai, Andrew Johnston, Cameron Olsen, Tanishq Mathew Abraham, Sergios Gatidis, Aksh...

  19. [27]

    A generalist agent

    Scott Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-maron, Mai Giménez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vin...

  20. [28]

    Large language model AI chatbots require approval as medical devices

    Stephen Gilbert, Hugh Harvey, Tom Melvin, Erik Vollebregt, and Paul Wicks. Large language model AI chatbots require approval as medical devices. Nature Medicine , pages 1--3, June 2023

  21. [29]

    Klontzas, Renato Cuocolo, Roberto Cannella, and Burak Koçak

    Tugba Akinci D'Antonoli, Arnaldo Stanzione, Christian Bluethgen, Federica Vernuccio, Lorenzo Ugga, Michail E. Klontzas, Renato Cuocolo, Roberto Cannella, and Burak Koçak. Large language models in radiology: fundamentals, applications, ethical considerations, risks, and future ...

  22. [30]

    Omiye, Haiwen Gui, Shawheen J

    Jesutofunmi A. Omiye, Haiwen Gui, Shawheen J. Rezaei, James Zou, and Roxana Daneshjou. Large Language Models in Medicine : The Potentials and Pitfalls . Annals of Internal Medicine , January 2024. Publisher: American College of Physicians

  23. [31]

    A Survey of Large Language Models for Healthcare : From Data , Technology , and Applications to Accountability and Ethics , October 2023

    Kai He, Rui Mao, Qika Lin, Yucheng Ruan, Xiang Lan, Mengling Feng, and Erik Cambria. A Survey of Large Language Models for Healthcare : From Data , Technology , and Applications to Accountability and Ethics , October 2023

  24. [32]

    Yu, Qiang Yang, and Xing Xie

    Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. A Survey on Evaluation of Large Language Models , December 2023. arXiv:2307.03109 [cs]

  25. [33]

    A Survey of Large Language Models , November 2023

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...

  26. [34]

    Smith, Felix Greaves, and Trishan Panch

    Andrew L. Smith, Felix Greaves, and Trishan Panch. Hallucination or Confabulation ? Neuroanatomy as metaphor in Large Language Models . PLOS Digital Health , 2(11):e0000388, November 2023

  27. [35]

    Thornton

    Rami Hatem, Brianna Simmons, and Joseph E. Thornton. Chatbot Confabulations Are Not Hallucinations . JAMA Internal Medicine , 183(10):1177, October 2023

  28. [36]

    Evaluating Object Hallucination in Large Vision - Language Models

    Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen. Evaluating Object Hallucination in Large Vision - Language Models . In Houda Bouamor, Juan Pino, and Kalika Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process...

  29. [37]

    Explainability for Large Language Models : A Survey

    Haiyan Zhao, Hanjie Chen, Fan Yang, Ninghao Liu, Huiqi Deng, Hengyi Cai, Shuaiqiang Wang, Dawei Yin, and Mengnan Du. Explainability for Large Language Models : A Survey . ACM Transactions on Intelligent Systems and Technology , January 2024. Just Accepted

  30. [38]

    S. M. Towhidul Islam Tonmoy, S. M. Mehedi Zaman, Vinija Jain, Anku Rani, Vipula Rawte, Aman Chadha, and Amitava Das. A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models , January 2024. arXiv:2401.01313 [cs]

  31. [39]

    Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. Language models as knowledge bases? In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint confe...

  32. [40]

    Reis-Filho, and Jakob Nikolas Kather

    Daniel Truhn, Jorge S. Reis-Filho, and Jakob Nikolas Kather. Large language models should be used as scientific reasoning engines, not knowledge databases. Nature Medicine , 29(12):2983--2984, December 2023. Publisher: Nature Publishing Group

  33. [41]

    Large language models in medicine

    Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. Large language models in medicine. Nature Medicine , 29(8):1930--1940, August 2023. Number: 8 Publisher: Nature Publishing Group

  34. [42]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Art...

  35. [43]

    Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves, Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G. T. Barrett, Cathy Cheung, Basil Mustafa, Anil Pal...

  36. [44]

    Thomas Savage, Ashwin Nayak, Robert Gallo, Ekanath Rangan, and Jonathan H. Chen. Diagnostic reasoning prompts reveal the potential for large language model interpretability in medicine. npj Digital Medicine , 7(1):1--7, January 2024. Number: 1 Publisher: Nature Publishing Group

  37. [45]

    Gallegos, Ryan A

    Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, Md Mehrab Tanjim, Sungchul Kim, Franck Dernoncourt, Tong Yu, Ruiyi Zhang, and Nesreen K. Ahmed. Bias and Fairness in Large Language Models : A Survey , September 2023. arXiv:2309.00770 [cs]

  38. [46]

    Unveiling Gender Bias in Terms of Profession Across LLMs : Analyzing and Addressing Sociological Implications , August 2023

    Vishesh Thakur. Unveiling Gender Bias in Terms of Profession Across LLMs : Analyzing and Addressing Sociological Implications , August 2023. arXiv:2307.09162 [cs]

  39. [47]

    Bias in artificial intelligence for medical imaging: fundamentals, detection, avoidance, mitigation, challenges, ethics, and prospects

    Burak Ko c ak, Andrea Ponsiglione, Arnaldo Stanzione, Christian Bluethgen, Jo \ a o Santinha, Lorenzo Ugga, Merel Huisman, Michail E Klontzas, Roberto Cannella, and Renato Cuocolo. Bias in artificial intelligence for medical imaging: fundamentals, detection, avoidance, mitigat...

  40. [48]

    Extracting Training Data from Large Language Models

    Nicholas Carlini, Florian Tram \`e r, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss , Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, \'U lfar Erlingsson, Alina Oprea, and Colin Raffel. Extracting Training Data from Large Language Models . In 30th USENIX Security Symp...

  41. [49]

    Are Chatbots Ready for Privacy - Sensitive Applications ? An Investigation into Input Regurgitation and Prompt - Induced Sanitization , May 2023

    Aman Priyanshu, Supriti Vijay, Ayush Kumar, Rakshit Naidu, and Fatemehsadat Mireshghallah. Are Chatbots Ready for Privacy - Sensitive Applications ? An Investigation into Input Regurgitation and Prompt - Induced Sanitization , May 2023. arXiv:2305.15008 [cs]

  42. [50]

    https://openai.com/blog/bug-bounty-program [acc

    Announcing OpenAI ’s Bug Bounty Program . https://openai.com/blog/bug-bounty-program [acc. 2024-02-08]

  43. [51]

    Abusing Images and Sounds for Indirect Instruction Injection in Multi - Modal LLMs , October 2023

    Eugene Bagdasaryan, Tsung-Yin Hsieh, Ben Nassi, and Vitaly Shmatikov. Abusing Images and Sounds for Indirect Instruction Injection in Multi - Modal LLMs , October 2023. arXiv:2307.10490 [cs]

  44. [52]

    Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S

    Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerova, Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Sergios...

  45. [53]

    The Case for Co - Designing Model Architectures with Hardware , January 2024

    Quentin Anthony, Jacob Hatef, Deepak Narayanan, Stella Biderman, Stas Bekman, Junqi Yin, Aamir Shafi, Hari Subramoni, and Dhabaleswar Panda. The Case for Co - Designing Model Architectures with Hardware , January 2024. arXiv:2401.14489 [cs]

  46. [54]

    ZeRO : Memory Optimizations Toward Training Trillion Parameter Models , May 2020

    Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. ZeRO : Memory Optimizations Toward Training Trillion Parameter Models , May 2020. arXiv:1910.02054 [cs, stat]

  47. [55]

    Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Mu...

  48. [56]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging LLM -as-a- Judge with MT - Bench and Chatbot Arena , December 2023. arXiv:2306.05685 [cs]

  49. [57]

    Hashimoto

    Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto . Stanford Alpaca : An instruction-following LLaMA model, 2023. https://github.com/tatsu-lab/stanford\_alpaca [acc. 2024-01-25]

  50. [58]

    Sara Mahdavi, Joelle Barral, Dale Webster, Greg S

    Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y ...

  51. [59]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas ...

  52. [60]

    Tiny Titans : Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization ?, February 2024

    Xue-Yong Fu, Md Tahmid Rahman Laskar, Elena Khasanova, Cheng Chen, and Shashi Bhushan TN. Tiny Titans : Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization ?, February 2024. arXiv:2402.00841 [cs]

  53. [61]

    Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne...

  54. [62]

    Mamba: Linear - Time Sequence Modeling with Selective State Spaces , December 2023

    Albert Gu and Tri Dao. Mamba: Linear - Time Sequence Modeling with Selective State Spaces , December 2023. arXiv:2312.00752 [cs]

  55. [63]

    Shah, David Entwistle, and Michael A

    Nigam H. Shah, David Entwistle, and Michael A. Pfeffer. Creation and Adoption of Large Language Models in Medicine . JAMA , 330(9):866--869, September 2023

  56. [64]

    Llama 2: Open Foundation and Fine - Tuned Chat Models , July 2023

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  57. [65]

    RLHF : Reinforcement Learning from Human Feedback , May 2023

    Chip Huyen. RLHF : Reinforcement Learning from Human Feedback , May 2023

  58. [66]

    DoctorGLM : Fine -tuning your Chinese Doctor is not a Herculean Task , April 2023

    Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen. DoctorGLM : Fine -tuning your Chinese Doctor is not a Herculean Task , April 2023. arXiv:2304.01097 [cs]

  59. [67]

    LLaVA - Med : Training a Large Language -and- Vision Assistant for Biomedicine in One Day , June 2023

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. LLaVA - Med : Training a Large Language -and- Vision Assistant for Biomedicine in One Day , June 2023. arXiv:2306.00890 [cs]

  60. [68]

    Fine-tuned language models for text classification

    Jeremy Howard and Sebastian Ruder. Fine-tuned language models for text classification. CoRR , abs/1801.06146, 2018

  61. [69]

    Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K

    Tianyu Han, Lisa C. Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K. Bressem. MedAlpaca -- An Open - Source Collection of Medical Conversational AI Models and Training Data , October 2023. arXiv:2304.08247 [cs]

  62. [70]

    Charles E. Kahn. The Long Tail . https://pubs.rsna.org/page/ai/blog/2019/4/the\_long\_tail [acc. 2024-02-08]

  63. [71]

    Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M. Saiful Bari, Sheng Shen, Zheng-Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward...

  64. [72]

    Manning, Christopher Ré, Diana Acosta-Navas, Drew A

    Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher Ré, Diana Acosta-Navas...

  65. [73]

    Manning, and Chelsea Finn

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn. Direct Preference Optimization : Your Language Model is Secretly a Reward Model , December 2023. arXiv:2305.18290 [cs]

  66. [74]

    KTO : Model Alignment as Prospect Theoretic Optimization , February 2024

    Kawin Ethayarajh, Winnie Xu, Niklas Muennighoff, Dan Jurafsky, and Douwe Kiela. KTO : Model Alignment as Prospect Theoretic Optimization , February 2024. arXiv:2402.01306 [cs]

  67. [75]

    Self- Alignment with Instruction Backtranslation , August 2023

    Xian Li, Ping Yu, Chunting Zhou, Timo Schick, Luke Zettlemoyer, Omer Levy, Jason Weston, and Mike Lewis. Self- Alignment with Instruction Backtranslation , August 2023. arXiv: 2308.06259v2 [cs]

  68. [76]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. Self-instruct: Aligning language models with self-generated instructions. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Proceedings of the 61st a...

  69. [77]

    Self- Verification Improves Few - Shot Clinical Information Extraction , May 2023

    Zelalem Gero, Chandan Singh, Hao Cheng, Tristan Naumann, Michel Galley, Jianfeng Gao, and Hoifung Poon. Self- Verification Improves Few - Shot Clinical Information Extraction , May 2023. arXiv: 2306.00024v1 [cs]

  70. [78]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA : Low - Rank Adaptation of Large Language Models . June 2021. arXiv:2106.09685v2 [cs]

  71. [79]

    Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Seneviratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Mansf...

  72. [80]

    The Power of Scale for Parameter - Efficient Prompt Tuning , September 2021

    Brian Lester, Rami Al-Rfou, and Noah Constant. The Power of Scale for Parameter - Efficient Prompt Tuning , September 2021. arXiv:2104.08691 [cs]

  73. [81]

    Alisa Liu, Xiaochuang Han, Yizhong Wang, Yulia Tsvetkov, Yejin Choi, and Noah A. Smith. Tuning Language Models by Proxy , January 2024. arXiv:2401.08565 [cs]

  74. [82]

    PEFT : State -of-the-art parameter-efficient fine-tuning methods, 2022

    Sourab Mangrulkar, Sylvain Gugger, Lysandre Debut, Younes Belkada, Sayak Paul, and Benjamin Bossan. PEFT : State -of-the-art parameter-efficient fine-tuning methods, 2022. https://github.com/huggingface/peft [acc. 2024-02-01]

  75. [83]

    QLoRA : Efficient Finetuning of Quantized LLMs , May 2023

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. QLoRA : Efficient Finetuning of Quantized LLMs , May 2023. arXiv:2305.14314 [cs]

  76. [84]

    Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A. Efros. Dataset Distillation , February 2020. arXiv:1811.10959 [cs, stat]

  77. [85]

    Distilling Step -by- Step ! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes , July 2023

    Cheng-Yu Hsieh, Chun-Liang Li, Chih-Kuan Yeh, Hootan Nakhost, Yasuhisa Fujii, Alexander Ratner, Ranjay Krishna, Chen-Yu Lee, and Tomas Pfister. Distilling Step -by- Step ! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes , July 2023. arXiv:2...

  78. [86]

    https://ollama.com [acc

    Ollama. https://ollama.com [acc. 2024-02-01]

  79. [87]

    Lungren, Hoifung Poon, and Akshay S

    Shih-Cheng Huang, Malte Jensen, Serena Yeung-Levy , Matthew P. Lungren, Hoifung Poon, and Akshay S. Chaudhari. Multimodal Foundation Models for Medical Imaging - A Systematic Review and Implementation Guidelines , October 2024. medRxiv:2024.10.23.24316003

  80. [88]

    MM - LLMs : Recent Advances in MultiModal Large Language Models , January 2024

    Duzhen Zhang, Yahan Yu, Chenxing Li, Jiahua Dong, Dan Su, Chenhui Chu, and Dong Yu. MM - LLMs : Recent Advances in MultiModal Large Language Models , January 2024. arXiv:2401.13601 [cs]

  81. [89]

    Can Generalist Foundation Models Outcompete Special - Purpose Tuning ? Case Study in Medicine , November 2023

    Harsha Nori, Yin Tat Lee, Sheng Zhang, Dean Carignan, Richard Edgar, Nicolo Fusi, Nicholas King, Jonathan Larson, Yuanzhi Li, Weishung Liu, Renqian Luo, Scott Mayer McKinney, Robert Osazuwa Ness, Hoifung Poon, Tao Qin, Naoto Usuyama, Chris White, and Eric Horvitz. Can Generali...

  82. [90]

    The Dawn of LMMs : Preliminary Explorations with GPT - 4V (ision), October 2023

    Zhengyuan Yang, Linjie Li, Kevin Lin, Jianfeng Wang, Chung-Ching Lin, Zicheng Liu, and Lijuan Wang. The Dawn of LMMs : Preliminary Explorations with GPT - 4V (ision), October 2023. arXiv:2309.17421 [cs]

  83. [91]

    Sara Mahdavi, Bradley Green, Ewa Dominowska, Blaise Aguera y Arcas, Joelle Barral, Dale Webster, Greg S

    Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Chuck Lau, Ryutaro Tanno, Ira Ktena, Basil Mustafa, Aakanksha Chowdhery, Yun Liu, Simon Kornblith, David Fleet, Philip Mansfield, Sushant Prakash, Renee Wong, Sunny Virmani,...

  84. [92]

    Towards Generalist Foundation Model for Radiology by Leveraging Web -scale 2D & 3D Medical Data , November 2023

    Chaoyi Wu, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie. Towards Generalist Foundation Model for Radiology by Leveraging Web -scale 2D & 3D Medical Data , November 2023. arXiv:2308.02463 [cs]

  85. [93]

    RaDialog : A Large Vision - Language Model for Radiology Report Generation and Conversational Assistance , November 2023

    Chantal Pellegrini, Ege Özsoy, Benjamin Busam, Nassir Navab, and Matthias Keicher. RaDialog : A Large Vision - Language Model for Radiology Report Generation and Conversational Assistance , November 2023. arXiv:2311.18681 [cs]

  86. [94]

    CXR - LLAVA : a multimodal large language model for interpreting chest X -ray images, January 2024

    Seowoo Lee, Jiwon Youn, Hyungjin Kim, Mansu Kim, and Soon Ho Yoon. CXR - LLAVA : a multimodal large language model for interpreting chest X -ray images, January 2024. arXiv:2310.18341 [cs]

  87. [95]

    Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling , May 2023

    Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O'Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal. Pythia: A Suite for Analyzing Large Language Mode...

  88. [96]

    Goodman, Paul H

    Katherine E. Goodman, Paul H. Yi, and Daniel J. Morgan. AI - Generated Clinical Summaries Require More Than Accuracy . JAMA , January 2024

  89. [97]

    BLEU : A method for automatic evaluation of machine translation

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. BLEU : A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics , pages 311--318, 2002

  90. [98]

    Rouge: A package for automatic evaluation of summaries

    Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In Text Summarization Branches Out , pages 74--81, 2004

  91. [99]

    ACI - BENCH : a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation , June 2023

    Wen-wai Yim, Yujuan Fu, Asma Ben Abacha, Neal Snider, Thomas Lin, and Meliha Yetisgen. ACI - BENCH : a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation , June 2023

  92. [100]

    BERTScore : Evaluating text generation with BERT

    Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. BERTScore : Evaluating text generation with BERT . In International Conference on Learning Representations , 2019

  93. [101]

    Manning, and Curtis Langlotz

    Yuhao Zhang, Derek Merck, Emily Tsai, Christopher D. Manning, and Curtis Langlotz. Optimizing the Factual Correctness of a Summary : A Study of Summarizing Radiology Reports . In Dan Jurafsky, Joyce Chai, Natalie Schluter, and Joel Tetreault, editors, Proceedings of the 58th A...

  94. [102]

    Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards

    Jean-Benoit Delbrouck, Pierre Chambon, Christian Bluethgen, Emily Tsai, Omar Almusa, and Curtis Langlotz. Improving the Factual Correctness of Radiology Report Generation with Semantic Rewards . In Yoav Goldberg, Zornitsa Kozareva, and Yue Zhang, editors, Findings of the Assoc...

  95. [103]

    PubMedQA : A dataset for biomedical research question answering

    Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William Cohen, and Xinghua Lu. PubMedQA : A dataset for biomedical research question answering. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natu...

  96. [104]

    Measuring massive multitask language understanding

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR) , 2021

  97. [105]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. GLUE : A multi-task benchmark and analysis platform for natural language understanding. In International conference on learning representations , 2019

  98. [106]

    Chaudhari

    Juan Manuel Zambrano Chaves, Nandita Bhaskhar, Maayane Attias, Jean-Benoit Delbrouck, Daniel Rubin, Andreas Markus Loening, Curtis Langlotz, and Akshay S. Chaudhari. RaLEs : a Benchmark for Radiology Language Evaluations . In Thirty-seventh Conference on Neural Information Pro...

  99. [107]

    Dalal, Jennifer L

    Cyril Zakka, Rohan Shad, Akash Chaurasia, Alex R. Dalal, Jennifer L. Kim, Michael Moor, Robyn Fong, Curran Phillips, Kevin Alexander, Euan Ashley, Jack Boyd, Kathleen Boyd, Karen Hirsch, Curt Langlotz, Rita Lee, Joanna Melia, Joanna Nelson, Karim Sallam, Stacey Tullis, Melissa...

  100. [108]

    The ecological footprint of medical AI

    Daniel Truhn, Gustav Müller-Franzes, and Jakob Nikolas Kather. The ecological footprint of medical AI . European Radiology , August 2023

  101. [109]

    Carbon emissions and large neural network training

    David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean. Carbon emissions and large neural network training. arXiv , 2021. arXiv:2104.10350 [cs]

  102. [110]

    Doo, Vishwa S

    Florence X. Doo, Vishwa S. Parekh, Adway Kanhere, Dharmam Savani, Ali S. Tejani, Amir Sapkota, and Paul H. Yi. Evaluation of Climate - Aware Metrics Tools for Radiology Informatics and Artificial Intelligence : Toward a Potential Radiology Ecolabel . Journal of the American Co...

  103. [111]

    Zamfirescu-Pereira , Richmond Y

    J.D. Zamfirescu-Pereira , Richmond Y. Wong, Bjoern Hartmann, and Qian Yang. Why Johnny Can 't Prompt : How Non-AI Experts Try (and Fail ) to Design LLM Prompts . In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems , CHI '23, pages 1--21, New York, N...

  104. [112]

    Prompt engineering

    Lilian Weng. Prompt engineering. lilianweng.github.io , March 2023. https://lilianweng.github.io/posts/2023-03-15-prompt-engineering/ [acc. 2024-02-01]

  105. [113]

    https://platform.openai.com [acc

    OpenAI Platform . https://platform.openai.com [acc. 2024-02-07]

  106. [114]

    https://www.promptingguide.ai/ [acc

    Prompt Engineering Guide . https://www.promptingguide.ai/ [acc. 2024-02-07]

  107. [115]

    Adams, Daniel Truhn, Felix Busch, Avan Kader, Stefan M

    Lisa C. Adams, Daniel Truhn, Felix Busch, Avan Kader, Stefan M. Niehues, Marcus R. Makowski, and Keno K. Bressem. Leveraging GPT -4 for Post Hoc Transformation of Free -text Radiology Reports into Structured Reporting : A Multilingual Feasibility Study . Radiology , 307(4):e23...

  108. [116]

    Pydantic, February 2024

    Samuel Colvin, David Montague, Adrian Garcia Badaracco, Hasan Ramezani, Eric Jolibois, Marcelo Trylesinski, Sydney Runkle, Terrence Dorsey, Serge Matveenko, pyup.io bot, David Hewitt, Arseny Boykov, Sebastián Ramírez, Viicos , Nikita Grishko, Koudai Aono, Alex Hall, Yurii Kara...

  109. [117]

    The Curious Case of Neural Text Degeneration , February 2020

    Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. The Curious Case of Neural Text Degeneration , February 2020. arXiv:1904.09751 [cs]

  110. [118]

    Retrieval- Augmented Generation for Knowledge - Intensive NLP Tasks , April 2021

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. Retrieval- Augmented Generation for Knowledge - Intensive NLP Tasks , April 2021. arXiv...

  111. [119]

    MedCPT : Contrastive Pre -trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval

    Qiao Jin, Won Kim, Qingyu Chen, Donald C Comeau, Lana Yeganova, W John Wilbur, and Zhiyong Lu. MedCPT : Contrastive Pre -trained Transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval. Bioinformatics , 39(11):btad651, November 2023

  112. [120]

    Fine- Tuning or Retrieval ? Comparing Knowledge Injection in LLMs , January 2024

    Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. Fine- Tuning or Retrieval ? Comparing Knowledge Injection in LLMs , January 2024. arXiv:2312.05934 [cs]

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.