Pith. sign in

REVIEW 1 cited by

Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.08838 v2 pith:RN4ZWNL3 submitted 2025-05-13 eess.IV cs.AIcs.CV

classification eess.IVcs.AIcs.CV
keywords generationreportstandardizedtextachievesconsistentframeworkimaging
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ultrasound (US) report generation is a challenging task due to the variability of US images, operator dependence, and the need for standardized text. Unlike X-ray and CT, US imaging lacks consistent datasets, making automation difficult. In this study, we propose a unified framework for multi-organ and multilingual US report generation, integrating fragment-based multilingual training and leveraging the standardized nature of US reports. By aligning modular text fragments with diverse imaging data and curating a bilingual English-Chinese dataset, the method achieves consistent and clinically accurate text generation across organ sites and languages. Fine-tuning with selective unfreezing of the vision transformer (ViT) further improves text-image alignment. Compared to the previous state-of-the-art KMVE method, our approach achieves relative gains of about 2\% in BLEU scores, approximately 3\% in ROUGE-L, and about 15\% in CIDEr, while significantly reducing errors such as missing or incorrect content. By unifying multi-organ and multi-language report generation into a single, scalable framework, this work demonstrates strong potential for real-world clinical workflows.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedVQA-TREE: A Multimodal Reasoning and Retrieval Framework for Sarcopenia Prediction

    eess.IV 2025-08 reject novelty 4.0 of 10

    MedVQA-TREE fuses three levels of ultrasound image features with UMLS-guided PubMed retrieval to predict sarcopenia, reporting 99% accuracy on a 24-patient proprietary dataset.

Pith tools