Pith. sign in

REVIEW 3 cited by

MG-3D: Multi-Grained Knowledge-Enhanced 3D Medical Vision-Language Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05876 v1 pith:TO4QVWZV submitted 2024-12-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords medicalmg-3dacrossanalysiscorrelationscross-modaldataimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D medical image analysis is pivotal in numerous clinical applications. However, the scarcity of labeled data and limited generalization capabilities hinder the advancement of AI-empowered models. Radiology reports are easily accessible and can serve as weakly-supervised signals. However, large-scale vision-language pre-training (VLP) remains underexplored in 3D medical image analysis. Specifically, the insufficient investigation into multi-grained radiology semantics and their correlations across patients leads to underutilization of large-scale volume-report data. Considering intra-patient cross-modal semantic consistency and inter-patient semantic correlations, we propose a multi-task VLP method, MG-3D, pre-trained on large-scale data (47.1K), addressing the challenges by the following two aspects: 1) Establishing the correspondence between volume semantics and multi-grained medical knowledge of each patient with cross-modal global alignment and complementary modality-guided local reconstruction, ensuring intra-patient features of different modalities cohesively represent the same semantic content; 2) Correlating inter-patient visual semantics based on fine-grained report correlations across patients, and keeping sensitivity to global individual differences via contrastive learning, enhancing the discriminative feature representation. Furthermore, we delve into the scaling law to explore potential performance improvements. Comprehensive evaluations across nine uni- and cross-modal clinical tasks are carried out to assess model efficacy. Extensive experiments on both internal and external datasets demonstrate the superior transferability, scalability, and generalization of MG-3D, showcasing its potential in advancing feature representation for 3D medical image analysis. Code will be available: https://github.com/Xuefeng-Ni/MG-3D.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Anatomy-Grounded CT Vision-Language Representations with Organ-Hierarchical Report Knowledge

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Organ-hierarchical knowledge extracted from radiology reports improves CT vision-language pretraining for zero-shot abnormality diagnosis and retrieval.

  2. SimCroP: Radiograph Representation Learning with Similarity-driven Cross-granularity Pre-training

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SimCroP learns chest-CT representations by aligning each report sentence to its most similar visual patches and fusing whole-scan and word-patch features, reporting higher classification and segmentation scores than s...

  3. HSENet: Hybrid Spatial Encoding Network for 3D Medical Vision-Language Understanding

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HSENet improves 3D CT vision-language understanding by combining global and local 3D encoders with a centroid-based spatial token compressor, posting state-of-the-art results on CT-RATE and RadGenome-ChestCT.

Pith tools