Pith. sign in

REVIEW 5 cited by

SAM2CLIP2SAM: Vision Language Model for Segmentation of 3D CT Scans for Covid-19 Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15728 v2 pith:KEMLUIYY submitted 2024-07-22 eess.IV cs.CV

classification eess.IVcs.CV
keywords segmentationcovid-19scansdetectionlungsmodelsegmentapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a new approach for effective segmentation of images that can be integrated into any model and methodology; the paradigm that we choose is classification of medical images (3-D chest CT scans) for Covid-19 detection. Our approach includes a combination of vision-language models that segment the CT scans, which are then fed to a deep neural architecture, named RACNet, for Covid-19 detection. In particular, a novel framework, named SAM2CLIP2SAM, is introduced for segmentation that leverages the strengths of both Segment Anything Model (SAM) and Contrastive Language-Image Pre-Training (CLIP) to accurately segment the right and left lungs in CT scans, subsequently feeding these segmented outputs into RACNet for classification of COVID-19 and non-COVID-19 cases. At first, SAM produces multiple part-based segmentation masks for each slice in the CT scan; then CLIP selects only the masks that are associated with the regions of interest (ROIs), i.e., the right and left lungs; finally SAM is given these ROIs as prompts and generates the final segmentation mask for the lungs. Experiments are presented across two Covid-19 annotated databases which illustrate the improved performance obtained when our method has been used for segmentation of the CT scans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors propose a new frame-level annotated violence detection dataset, DVD, with 500 videos and rich metadata, but it is not yet available and lacks validation experiments.

  2. Taming Domain Shift in Multi-source CT-Scan Classification via Input-Space Standardization

    eess.IV 2025-07 conditional novelty 4.0 of 10

    Input-space standardization via lung cropping and density-based slice sampling reduces inter-source feature variance by 75% and improves COVID-19 CT classification F1 by roughly 24 points across architectures.

  3. Multi-Source COVID-19 Detection via Variance Risk Extrapolation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A challenge entry combining VREx and Mixup in a two-stage training pipeline reports 0.96 macro F1 on a four-hospital COVID-19 CT validation set, but without baselines or statistical validation.

  4. Multi Source COVID-19 Detection via Kernel-Density-based Slice Sampling

    eess.IV 2025-07 conditional novelty 3.0 of 10

    EfficientNet-B7 with SSFL and KDS slice sampling achieves 94.68 F1 on a multi-source COVID-19 CT validation set, outperforming Swin Transformer-Base's 93.34, though the evaluation has class-imbalance and no-ablation issues.

  5. Advancing Lung Disease Diagnosis in 3D CT Scans

    eess.IV 2025-07 conditional novelty 2.0 of 10

    A 3D ResNeSt50 model with slice removal and weighted cross-entropy achieves a Macro F1 of 0.80 on the Fair Disease Diagnosis Challenge validation set.

Pith tools