REVIEW 3 major objections 5 minor 55 references
JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read JRadiEvo merges four off-the-shelf models with an evolutionary search over 50 translated samples and claims Japanese chest X-ray reports that beat CheXagent and GPT-4o on ROUGE-L and METEOR with no fine-tuning.
desk verdict A plausible data-efficient adaptation trick, but the evaluation is too thin to support the performance claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is evolutionary model merging: a task vector $\tau_t = \theta_{\mathrm{ft}}^t - \theta_{\mathrm{init}}$ is formed for each source fine-tune; DARE randomly drops and rescales entries; TIES-Merging trims small entries, elects a sign per parameter, and averages only matching-sign task vectors; and the final weights are $\theta_{\mathrm{final}} = \theta_{\mathrm{init}} + \lambda \tau_{\mathrm{merged}}$. The evolutionary search, CMA-ES, optimizes the DARE drop rates, TIES retained percentages, per-task weights, and scaling parameter $\lambda$, treating the whole LLM component as a single merged layer, with ROUGE-L as the fitness function. This lets the pipeline recombine skills—vision-to-text, medical text, and Japanese text—without backpropagation.
What would settle it
Have board-certified radiologists blindly rate JRadiEvo's reports against the GPT-3.5-translated references for factual completeness and correctness, or recompute the ROUGE-L and METEOR comparison using human-translated rather than machine-translated references. If expert preference does not favor JRadiEvo, or the score gap over CheXagent and GPT-4o collapses with human references, the central claim that 50 samples suffice for accurate report generation would fail.
Extended reading notes
Core claim
On the paper's own terms, JRadiEvo establishes that evolutionary optimization of model merging can transplant medical knowledge and Japanese language ability into a vision-language model without any training-time gradient updates. Starting from Bunny-v1_1-Llama-3-8B-V2, the authors form task vectors from the fine-tuned weights of MMed-Llama-3-8B-EnIns, OpenBioLLM-Llama3-8B, and Llama-3-Swallow-8B-Instruct, apply DARE drop-and-rescale and TIES-Merging trim/elect/merge steps, and use CMA-ES to search the recipe parameters against ROUGE-L on the 50 translated reference reports. The optimized model reaches ROUGE-L 0.212 and METEOR 0.191 on the test set, the best among the compared systems, and the learned weights show OpenBioLLM carrying most of the medical signal while MMed-Llama contributes little. The qualitative examples, however, show short template-like sentences that omit several ground-truth findings and, in some cases, introduce abnormalities not present in the reference.
Load-bearing premise
The quantitative claim depends on ROUGE-L and METEOR against GPT-3.5-translated Japanese references being a valid proxy for clinical report quality; the paper's own examples show generated reports that are short, templated, miss reference findings, and occasionally hallucinate abnormalities, so if the metrics reward that style, the 'accurate reports' claim overstates what the scores show.
Editorial extensions
If this is right
- If correct, JRadiEvo shows that a non-English medical vision-language model can be built from 50 translated cases, substantially lowering the data barrier for languages without public radiology report corpora.
- Because the merge involves no gradient updates and the resulting model has 8 billion parameters, the same recipe could run on a single hospital GPU, keeping patient data in-house for privacy-sensitive environments.
- The comparison suggests that for a model of this size, evolutionary merging is a more stable adaptation route than LoRA instruction-tuning, which the authors report suffered catastrophic forgetting when the training set was enlarged to 10,000 samples.
- Direct Japanese generation eliminates the generate-in-English-then-translate workflow that CheXagent requires, making the model immediately usable in Japanese clinical settings.
- The learned merging weights identify OpenBioLLM as the dominant source of medical knowledge and the Japanese model as secondary but necessary, indicating which source models future medical merges should prioritize.
Reading between the lines
- A testable extension the authors do not run is a sample-size curve: repeating the evolutionary merge with 10, 100, and 500 translated examples would show whether the 50-sample result is a threshold or a plateau.
- The same merging recipe should transfer to other low-resource languages and other imaging domains, but such transfer is untested and likely depends on the availability of a language-capable base LLM and domain-specific medical LLMs.
- The paper's own qualitative table suggests the metric lead may reflect fluency and template overlap rather than clinical completeness, since generated reports are generic while references list specific findings; human expert evaluation would be the real test, and the authors themselves note this gap.
- Because the search optimizes ROUGE-L, the model is explicitly tuned to overlap with reference phrasing; a different fitness function, such as clinical entity correctness, might yield different merging weights and more specific reports.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes JRadiEvo, a Japanese chest X-ray report generation model built by evolutionary optimization of model merging. Starting from Bunny-v1_1-Llama-3-8B-V2 as a non-medical vision-language model, the authors merge it with two medical text LLMs (MMed-Llama-3-8B-EnIns, OpenBioLLM-Llama3-8B) and a Japanese LLM (Llama-3-Swallow-8B-Instruct-v0.1) using DARE, TIES-Merging, and CMA-ES, with only 50 human-reviewed, GPT-3.5-translated MIMIC-CXR training reports. The resulting 8B model is evaluated with BLEU, ROUGE-L, and METEOR on a held-out MIMIC-CXR test set whose Japanese references were produced by GPT-3.5, and compared against LoRA instruction-tuned versions of the same VLM, CheXagent, and GPT-4o. The authors report the highest ROUGE-L (0.212) and METEOR (0.191) among the compared models and argue that this demonstrates efficient low-resource adaptation, direct Japanese generation, local deployability, and superiority over instruction-tuning.
Significance. If the empirical claims are validated, the contribution is meaningful: it would be the first demonstration that evolutionary model merging can adapt a multimodal foundation model to a non-English medical domain using only 50 translated examples, avoiding backpropagation and catastrophic forgetting while producing a locally deployable 8B model. The manuscript is also commendable for describing the merge recipe in detail, for including instruction-tuning baselines and recent strong baselines such as CheXagent and GPT-4o, and for analyzing the relative contributions of the merged LLMs. However, the current evaluation does not support the headline claims: the optimization/evaluation separation is not explicitly stated, no test-set size or statistical significance is reported, the reference translations are machine-generated, and the qualitative examples show template-like outputs that miss or hallucinate findings. These are load-bearing issues for the central claim that JRadiEvo generates accurate reports and outperforms leading models.
major comments (3)
- [Section 3.2 and Section 4.1] The evolutionary optimization objective is not fully specified. Section 3.2 states that CMA-ES maximizes the ROUGE-L score between the generated text and "the reference text," but it does not state whether this reference is the 50 translated training reports or the held-out test references introduced in Section 4.1. Since the algorithm runs for 600 iterations, optimizing on the test references would make the Table 2 comparison circular: the model would be selected for the exact metric and references used in the evaluation. The authors must state explicitly that only the 50 training reports were used in the CMA-ES objective, and report the size of the test set used in Table 2.
- [Section 4.2.2, Table 2] The comparison lacks any measure of uncertainty. No test-set size, standard deviation, confidence interval, or significance test is reported, and the largest claimed margins are small: JRadiEvo's ROUGE-L exceeds CheXagent's by 0.013 and its METEOR exceeds GPT-4o's by 0.003. Without repeated evaluations, bootstrap intervals, or paired significance tests, the claim that JRadiEvo "outperformed" these models is not statistically established. The authors should provide error bars or significance tests, and ideally report results over multiple CMA-ES runs or random seeds.
- [Section 4.1 and Table 3] The evaluation metric may not reflect clinical report quality. The test references were translated by GPT-3.5 without the human review applied to the training translations, so the metric scores compare against machine-generated Japanese. More importantly, the qualitative examples in Table 3 contradict the abstract's claim of "accurate" reports: Example 1 misses the ground-truth findings of reduced lung volume, scarring, and fibrosis; Example 2 misses the sternotomy wires and mediastinal clips and instead reports left ventricular enlargement and possible heart failure; Example 3 misses pleural effusion and metastatic nodules. The generated texts are short stereotyped templates, so the ROUGE-L/METEOR advantage may reflect template overlap rather than factual correctness. The authors should add human expert evaluation or clinically grounded metrics (e.g., finding-level factuality or CheXpert-label consistency) before claiming clinical accuracy or superiority.
minor comments (5)
- [Section 4.1 and Section 4.2.1] The base VLM name is inconsistent: Section 4.1 and footnote 2 give "Bunny-v1_1-Llama-3-8B-V2" while Section 4.2.1 and footnote 6 give "Bunny-v1_1-Llama-3-8B-V6"; both footnotes point to the same URL, so the intended version should be clarified.
- [Section 4.1] The description "randomly selected 50 samples from both views" should specify how many AP and PA images were selected, and whether the 50 samples are images, reports, or image-report pairs; this matters because the paper emphasizes that only 50 cases were used.
- [Section 4.2.1] The LoRA baselines are said to be trained on 2,000 translated samples, but the paper does not state whether the same test set and the same GPT-3.5-translated references were used for their evaluation; this should be stated explicitly.
- [Section 4.3, Figure 1] Figure 1 is described as showing "density and weight parameters after optimization," but the caption does not define "density" (presumably the retained percentage k_t) or the units of the weight c_t; axis labels and a legend should be added.
- [Throughout] There are several typographical errors: "from scrach" in Section 1, "was was set" in Section 4.2.1, "and and revised" in Section 4.1, "a extremely limited dataset" in Section 4.2.2, and inconsistent spacing in model names such as "LLaV A" and "Med-PaLM M". A careful proofread is needed.
Circularity Check
No significant circularity: the merge recipe is optimized on 50 training reports and evaluated on a held-out official test split, with no load-bearing self-citation.
full rationale
The central derivation chain is not circular. JRadiEvo is built by merging public models (Bunny, MMed-Llama, OpenBioLLM, Llama-3-Swallow) using TIES-Merging and DARE, with CMA-ES optimizing the merge hyperparameters. Section 3.2 states that the evolutionary algorithm maximizes ROUGE-L between generated text and reference text, and Section 4.1 establishes that the 50 translated samples are drawn from the official MIMIC-CXR training set while the test samples are drawn from the official test set, explicitly stating that this ensures no data leakage. Thus the ROUGE-L-based selection on the training reports and the ROUGE-L-based evaluation on the test reports are not the same data, so Table 2 does not reduce to the optimization objective by construction. The evolutionary merging recipe itself is adopted from prior work by other authors (ref. [13]), so no self-citation carries the load. The qualitative examples in Table 3 may raise concerns about clinical validity of the metric-based claims, and the paper does not report error bars or test-set size, but those are robustness/validity issues rather than circularity. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors, and no known result is relabeled as a new organization. Therefore no circular step is established from the paper's own equations or citations.
Assumptions & free parameters
free parameters (5)
- DARE drop rate alpha =
not reported (optimized by CMA-ES)
- TIES saved rates k_t for each of the four models =
not reported numerically; shown in Figure 1
- Task weights c_t for each model =
not reported numerically; shown in Figure 1
- Scaling parameter lambda =
not reported
- CMA-ES hyperparameters (init, sigma, population size) =
0.5, 1/6, 4+floor(3 ln n)
assumptions (4)
- domain assumption TIES-Merging combined with DARE preserves and combines capabilities of same-base fine-tuned models without destructive interference.
- domain assumption ROUGE-L computed on Japanese tokenized text is a valid optimization target and evaluation metric for radiology report generation.
- domain assumption Machine translation of the test set with GPT-3.5 provides adequate reference translations for evaluating Japanese report generation.
- domain assumption The base VLM Bunny and the Japanese Swallow model have sufficient inherent Japanese capability for the merged model to generate intelligible Japanese.
Cite this review
Pith. "Pith review of JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging." pith.science (2026). https://pith.science/paper/LPVNB3XX
@misc{pith2026241109933,
author = {Pith},
title = {Pith review of: JRadiEvo: A Japanese Radiology Report Generation Model Enhanced by Evolutionary Optimization of Model Merging},
year = {2026},
howpublished = {\url{https://pith.science/paper/LPVNB3XX}},
note = {Machine review of arXiv:2411.09933}
}
read the original abstract
With the rapid advancement of large language models (LLMs), foundational models (FMs) have seen significant advancements. Healthcare is one of the most crucial application areas for these FMs, given the significant time and effort required for physicians to analyze large volumes of patient data. Recent efforts have focused on adapting multimodal FMs to the medical domain through techniques like instruction-tuning, leading to the development of medical foundation models (MFMs). However, these approaches typically require large amounts of training data to effectively adapt models to the medical field. Moreover, most existing models are trained on English datasets, limiting their practicality in non-English-speaking regions where healthcare professionals and patients are not always fluent in English. The need for translation introduces additional costs and inefficiencies. To address these challenges, we propose a \textbf{J}apanese \textbf{Radi}ology report generation model enhanced by \textbf{Evo}lutionary optimization of model merging (JRadiEvo). This is the first attempt to extend a non-medical vision-language foundation model to the medical domain through evolutionary optimization of model merging. We successfully created a model that generates accurate Japanese reports from X-ray images using only 50 translated samples from publicly available data. This model, developed with highly efficient use of limited data, outperformed leading models from recent research trained on much larger datasets. Additionally, with only 8 billion parameters, this relatively compact foundation model can be deployed locally within hospitals, making it a practical solution for environments where APIs and other external services cannot be used due to strict privacy and security requirements.
Figures
Reference graph
Works this paper leans on
-
[1]
On the opportunities and risks of foundation models,
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al., “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021
arXiv 2021
-
[2]
A survey of large language models,
Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong Wen, “A survey of large language models,” arXiv preprint arXiv:2303.18223, 2023
arXiv 2023
- [3]
-
[4]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee, “Visual instruction tuning,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds. 2023, vol. 36, pp. 34892–34916, Curran Associates, Inc
work page 2023
-
[5]
Flamingo: a visual language model for few-shot learning,
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob L Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikoł aj Bi´nk...
work page 2022
-
[6]
Instruction tuning with GPT-4,
Baolin Peng, Chunyuan Li, Pengcheng He, Michel Galley, and Jianfeng Gao, “Instruction tuning with GPT-4,” arXiv preprint arXiv:2304.03277, 2023
arXiv 2023
-
[7]
LLaV A-Med: Training a large language-and- vision assistant for biomedicine in one day,
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao, “LLaV A-Med: Training a large language-and- vision assistant for biomedicine in one day,” in Advances in Neural Information Processing Systems, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, Eds. 2023, vol....
work page 2023
-
[8]
Towards expert-level medical question answering with large language models,
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Le Hou, Kevin Clark, Stephen Pfohl, Heather Cole-Lewis, Darlene Neal, Mike Schaekermann, Amy Wang, Mohamed Amin, Sami Lachgar, Philip Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Aguera y Arcas, Nenad Tomasev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi,...
arXiv 2023
Show all 55 references
-
[9]
Med-Flamingo: a multimodal medical few-shot learner,
Michael Moor, Qian Huang, Shirley Wu, Michihiro Yasunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Eduardo Pontes Reis, and Pranav Rajpurkar, “Med-Flamingo: a multimodal medical few-shot learner,” in Proceedings of the 3rd Machine Learning for Health Symposium, Stefan Hegselma...
2023
-
[10]
MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports,
Alistair E. W. Johnson, Tom J. Pollard, Seth J. Berkowitz, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih-ying Deng, Roger G. Mark, and Steven Horng, “MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports,” Scientific Data, vol. 6...
2019
-
[11]
Preparing a collection of radiology examinations for distribution and retrieval,
Dina Demner-Fushman, Marc Kohli, Marc Rosenman, Sonya Shooshan, Laritza Rodriguez, Sameer Antani, George Thoma, and Clement Mcdonald, “Preparing a collection of radiology examinations for distribution and retrieval,” Journal of the American Medical Informatics Association : JA...
2015
-
[12]
GPT-4 Technical Report,
OpenAI, “GPT-4 Technical Report,” arXiv preprint arXiv:2303.08774, 2024. 10
2024 arXiv
-
[13]
Evolutionary optimization of model merging recipes,
Takuya Akiba, Makoto Shing, Yujin Tang, Qi Sun, and David Ha, “Evolutionary optimization of model merging recipes,” arXiv preprint arXiv:2403.13187, 2024
2024 arXiv
-
[14]
Catastrophic interference in connectionist networks: The sequential learning problem,
Michael McCloskey and Neal J. Cohen, “Catastrophic interference in connectionist networks: The sequential learning problem,” vol. 24 of Psychology of Learning and Motivation , pp. 109–165. Academic Press, 1989
1989
-
[15]
Understanding catastrophic forgetting in language models via implicit inference,
Suhas Kotha, Jacob Mitchell Springer, and Aditi Raghunathan, “Understanding catastrophic forgetting in language models via implicit inference,” in The Twelfth International Conference on Learning Representations, 2024
2024
-
[16]
An empirical study of catastrophic forgetting in large language models during continual fine-tuning,
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang, “An empirical study of catastrophic forgetting in large language models during continual fine-tuning,”arXiv preprint arXiv:2308.08747, 2024
2024 arXiv
-
[17]
Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge,
Yunxiang Li, Zihan Li, Kai Zhang, Ruilong Dan, Steve Jiang, and You Zhang, “Chatdoctor: A medical chat model fine-tuned on a large language model meta-ai (llama) using medical domain knowledge,” Cureus, vol. 15, 06 2023
2023
-
[18]
DoctorGLM: Fine-tuning your chinese doctor is not a herculean task,
Honglin Xiong, Sheng Wang, Yitao Zhu, Zihao Zhao, Yuxiao Liu, Linlin Huang, Qian Wang, and Dinggang Shen, “DoctorGLM: Fine-tuning your chinese doctor is not a herculean task,” ArXiv, vol. abs/2304.01097, 2023
2023 arXiv
-
[19]
BioGPT: generative pre-trained transformer for biomedical text generation and mining,
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu, “BioGPT: generative pre-trained transformer for biomedical text generation and mining,” Briefings in Bioinformatics, vol. 23, no. 6, pp. bbac409, 09 2022
2022
-
[20]
MedAlpaca – an open-source collection of medical conversational ai models and training data,
Tianyu Han, Lisa C. Adams, Jens-Michalis Papaioannou, Paul Grundmann, Tom Oberhauser, Alexander Löser, Daniel Truhn, and Keno K. Bressem, “MedAlpaca – an open-source collection of medical conversational ai models and training data,” arXiv preprint arXiv:2304.08247, 2023
2023 arXiv
-
[21]
Pmc-llama: Towards building open-source language models for medicine,
Chaoyi Wu, Weixiong Lin, Xiaoman Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie, “Pmc-llama: Towards building open-source language models for medicine,” arXiv preprint arXiv:2304.14454, 2023
2023 arXiv
-
[22]
Capabilities of gemini models in medicine,
Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves, Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G. T. Barrett, Cathy Cheung, Basil Mustafa, Anil Pal...
2024 arXiv
-
[23]
Large language models encode clinical knowledge,
Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, Perry Payne, Martin Senevi- ratne, Paul Gamble, Chris Kelly, Abubakr Babiker, Nathanael Schärli, Aakanksha Chowdhery, Philip Man...
2023
-
[24]
Coca: Contrastive captioners are image-text foundation models,
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu, “Coca: Contrastive captioners are image-text foundation models,” Transactions on Machine Learning Research, 2022. 11
2022
-
[25]
BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi, “BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Proceedings of the 39th International Conference on Machine Learning, Kamalika Chaudhuri, Stefanie Jegelka, Le So...
2022
-
[26]
PaLI-X: On scaling up a multilingual vision and language model,
Xi Chen, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Jialin Wu, Carlos Riquelme Ruiz, Sebastian Goodman, Xiao Wang, Yi Tay, Siamak Shakeri, Mostafa Dehghani, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang, ...
2023 arXiv
-
[27]
Cogvlm: Visual expert for pretrained language models,
Weihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong, Ji Qi, Yan Wang, Junhui Ji, Zhuoyi Yang, Lei Zhao, Xixuan Song, Jiazheng Xu, Bin Xu, Juanzi Li, Yuxiao Dong, Ming Ding, and Jie Tang, “Cogvlm: Visual expert for pretrained language models,” 2023
2023
-
[28]
XrayGPT: Chest radiographs summarization using large medical vision-language models,
Omkar Chakradhar Thawakar, Abdelrahman M. Shaker, Sahal Shaji Mullappilly, Hisham Cholakkal, Rao Muhammad Anwer, Salman Khan, Jorma Laaksonen, and Fahad Khan, “XrayGPT: Chest radiographs summarization using large medical vision-language models,” in Proceedings of the 23rd Work...
2024
-
[29]
Towards generalist biomedical ai,
Tao Tu, Shekoofeh Azizi, Danny Driess, Mike Schaekermann, Mohamed Amin, Pi-Chuan Chang, Andrew Carroll, Chuck Lau, Ryutaro Tanno, Ira Ktena, Basil Mustafa, Aakanksha Chowdhery, Yun Liu, Simon Kornblith, David Fleet, Philip Mansfield, Sushant Prakash, Renee Wong, Sunny Virmani,...
2023 arXiv
-
[30]
Chexagent: Towards a foundation model for chest x-ray interpretation,
Zhihong Chen, Maya Varma, Jean-Benoit Delbrouck, Magdalini Paschali, Louis Blankemeier, Dave Van Veen, Jeya Maria Jose Valanarasu, Alaa Youssef, Joseph Paul Cohen, Eduardo Pontes Reis, Emily Tsai, Andrew Johnston, Cameron Olsen, Tanishq Mathew Abraham, Sergios Gatidis, Akshay ...
2024
-
[31]
Sampling generative networks,
Tom White, “Sampling generative networks,” arXiv preprint arXiv:1609.04468, 2016
2016 arXiv
-
[32]
Editing models with task arithmetic,
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Ha- jishirzi, and Ali Farhadi, “Editing models with task arithmetic,” in The Eleventh International Conference on Learning Representations, 2023
2023
-
[33]
TIES-merging: Resolving interference when merging models,
Prateek Yadav, Derek Tam, Leshem Choshen, Colin Raffel, and Mohit Bansal, “TIES-merging: Resolving interference when merging models,” in Thirty-seventh Conference on Neural Infor- mation Processing Systems, 2023
2023
-
[34]
Language models are super mario: Absorbing abilities from homologous models as a free lunch,
Le Yu, Bowen Yu, Haiyang Yu, Fei Huang, and Yongbin Li, “Language models are super mario: Absorbing abilities from homologous models as a free lunch,” in International Conference on Machine Learning. PMLR, 2024
2024
-
[35]
Llama 3 model card,
AI@Meta, “Llama 3 model card,” 2024
2024
-
[36]
ROUGE: A package for automatic evaluation of summaries,
Chin-Yew Lin, “ROUGE: A package for automatic evaluation of summaries,” in Text Summa- rization Branches Out, Barcelona, Spain, July 2004, pp. 74–81, Association for Computational Linguistics. 12
2004
-
[37]
MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs,
Alistair E. W. Johnson, Tom J. Pollard, Nathaniel R. Greenbaum, Matthew P. Lungren, Chih ying Deng, Yifan Peng, Zhiyong Lu, Roger G. Mark, Seth J. Berkowitz, and Steven Horng, “MIMIC-CXR-JPG, a large publicly available database of labeled chest radiographs,” arXiv preprint arX...
1901 arXiv
-
[38]
Hybrid retrieval-generation reinforced agent for medical image report generation,
Yuan Li, Xiaodan Liang, Zhiting Hu, and Eric P Xing, “Hybrid retrieval-generation reinforced agent for medical image report generation,” in Advances in Neural Information Processing Systems, S. Bengio, H. Wallach, H. Larochelle, K. Grauman, N. Cesa-Bianchi, and R. Garnett, Eds...
2018
-
[39]
Generating radiology reports via memory-driven transformer,
Zhihong Chen, Yan Song, Tsung-Hui Chang, and Xiang Wan, “Generating radiology reports via memory-driven transformer,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu, Eds., On...
2020
-
[40]
Improving chest X-ray report generation by leveraging warm starting,
Aaron Nicolson, Jason Dowling, and Bevan Koopman, “Improving chest X-ray report generation by leveraging warm starting,” Artificial Intelligence in Medicine, vol. 144, pp. 102633, 2023
2023
-
[41]
Language models are few-shot learners,
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey ...
2020
-
[42]
Interactive and explainable region-guided radiology report generation,
Tim Tanida, Philip Müller, Georgios Kaissis, and Daniel Rueckert, “Interactive and explainable region-guided radiology report generation,” 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 7433–7442, 2023
2023
-
[43]
BLEU: a method for automatic evaluation of machine translation,
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, “BLEU: a method for automatic evaluation of machine translation,” in Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, USA, 2002, ACL ’02, p. 311–318, Association for Computational ...
2002
-
[44]
Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems,
Michael Denkowski and Alon Lavie, “Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems,” in Proceedings of the Sixth Workshop on Statis- tical Machine Translation, Chris Callison-Burch, Philipp Koehn, Christof Monz, and Omar F. ...
2011
-
[45]
MeCab : Yet another part-of-speech and morphological analyzer,
Taku Kudo, “MeCab : Yet another part-of-speech and morphological analyzer,” 2006
2006
-
[46]
Efficient multimodal learning from data-centric perspective,
Muyang He, Yexin Liu, Boya Wu, Jianhao Yuan, Yueze Wang, Tiejun Huang, and Bo Zhao, “Efficient multimodal learning from data-centric perspective,” arXiv preprint arXiv:2402.11530, 2024
2024 arXiv
-
[47]
Towards building multilingual language model for medicine,
Pengcheng Qiu, Chaoyi Wu, Xiaoman Zhang, Weixiong Lin, Haicheng Wang, Ya Zhang, Yanfeng Wang, and Weidi Xie, “Towards building multilingual language model for medicine,” 2024
2024
-
[48]
OpenBioLLMs: Advancing open-source large language models for healthcare and life sciences,
Malaikannan Sankarasubbu Ankit Pal, “OpenBioLLMs: Advancing open-source large language models for healthcare and life sciences,” https://huggingface.co/aaditya/ OpenBioLLM-Llama3-70B, 2024
2024
-
[49]
Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities,
Kazuki Fujii, Taishi Nakamura, Mengsay Loem, Hiroki Iida, Masanari Ohi, Kakeru Hattori, Hirai Shota, Sakae Mizuki, Rio Yokota, and Naoaki Okazaki, “Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities,” in First Conference on Langua...
2024
-
[50]
75–102, Springer Berlin Heidelberg, Berlin, Heidelberg, 2006
Nikolaus Hansen, The CMA Evolution Strategy: A Comparing Review, pp. 75–102, Springer Berlin Heidelberg, Berlin, Heidelberg, 2006. 13
2006
-
[51]
Optuna: A next-generation hyperparameter optimization framework,
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama, “Optuna: A next-generation hyperparameter optimization framework,” in The 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019, pp. 2623–2631
2019
-
[52]
LoRA: Low-rank adaptation of large language models,
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, “LoRA: Low-rank adaptation of large language models,” in International Conference on Learning Representations, 2022
2022
-
[53]
SGDR: Stochastic gradient descent with warm restarts,
Ilya Loshchilov and Frank Hutter, “SGDR: Stochastic gradient descent with warm restarts,” in International Conference on Learning Representations, 2017
2017
-
[54]
GPT-4o system card,
Open AI, “GPT-4o system card,” 2024
2024
-
[55]
Lora tuning for large-scale japanese foundational models,
Hao Wang, Akifumi Nakamachi, and Toshinori Sato, “Lora tuning for large-scale japanese foundational models,” in Proceedings of the 29th Annual Meeting of the Association for Natural Language Processing, Tokyo, Japan, March 2023, The Association for Natural Language Processing. 14
2023
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.