Pith. sign in

REVIEW 3 cited by

Masked Contrastive Reconstruction for Cross-modal Medical Image-Report Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.15840 v2 pith:HLNDMTDA submitted 2023-12-26 cs.CV cs.AI

classification cs.CVcs.AI
keywords cross-modalmaskedtaskscontrastivemedicalreconstructionretrievaltask
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cross-modal medical image-report retrieval task plays a significant role in clinical diagnosis and various medical generative tasks. Eliminating heterogeneity between different modalities to enhance semantic consistency is the key challenge of this task. The current Vision-Language Pretraining (VLP) models, with cross-modal contrastive learning and masked reconstruction as joint training tasks, can effectively enhance the performance of cross-modal retrieval. This framework typically employs dual-stream inputs, using unmasked data for cross-modal contrastive learning and masked data for reconstruction. However, due to task competition and information interference caused by significant differences between the inputs of the two proxy tasks, the effectiveness of representation learning for intra-modal and cross-modal features is limited. In this paper, we propose an efficient VLP framework named Masked Contrastive and Reconstruction (MCR), which takes masked data as the sole input for both tasks. This enhances task connections, reducing information interference and competition between them, while also substantially decreasing the required GPU memory and training time. Moreover, we introduce a new modality alignment strategy named Mapping before Aggregation (MbA). Unlike previous methods, MbA maps different modalities to a common feature space before conducting local feature aggregation, thereby reducing the loss of fine-grained semantic information necessary for improved modality alignment. Qualitative and quantitative experiments conducted on the MIMIC-CXR dataset validate the effectiveness of our approach, demonstrating state-of-the-art performance in medical cross-modal retrieval tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Semantically Informed Salient Regions Guided Radiology Report Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SISRNet identifies medically salient regions in chest X-rays via cross-modal alignment and uses them to guide both masked image modeling and report generation, improving clinical accuracy metrics on IU-Xray and MIMIC-CXR.

  2. Prototype-Enhanced Confidence Modeling for Cross-Modal Medical Image-Report Retrieval

    cs.CV 2025-08 conditional novelty 5.0 of 10

    PECM combines multi-level prototypes with dual-stream confidence weighting and reports up to 10.17% retrieval gains on radiology datasets, including zero-shot transfer.

  3. Quality Versus Sparsity in Image Recovery by Dictionary Learning Using Iterative Shrinkage

    cs.CV 2025-08 reject novelty 3.0 of 10

    A dictionary learning paper whose abstract claims sparsity does not hurt recovery quality, but whose full text is an unrelated medical retrieval manuscript.

Pith tools