Pith. sign in

REVIEW 11 cited by

Evaluating General Purpose Vision Foundation Models for Medical Image Analysis: An Experimental Study of DINOv2 on Radiology Benchmarks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02366 v4 pith:GHYW5ZEE submitted 2023-12-04 cs.CV cs.AI

classification cs.CVcs.AI
keywords dinov2analysisimagemodelsacrossdatafoundationgeneralizability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The integration of deep learning systems into healthcare has been hindered by the resource-intensive process of data annotation and the inability of these systems to generalize to different data distributions. Foundation models, which are models pre-trained on large datasets, have emerged as a solution to reduce reliance on annotated data and enhance model generalizability and robustness. DINOv2 is an open-source foundation model pre-trained with self-supervised learning on 142 million curated natural images that exhibits promising capabilities across various vision tasks. Nevertheless, a critical question remains unanswered regarding DINOv2's adaptability to radiological imaging, and whether its features are sufficiently general to benefit radiology image analysis. Therefore, this study comprehensively evaluates the performance DINOv2 for radiology, conducting over 200 evaluations across diverse modalities (X-ray, CT, and MRI). To measure the effectiveness and generalizability of DINOv2's feature representations, we analyze the model across medical image analysis tasks including disease classification and organ segmentation on both 2D and 3D images, and under different settings like kNN, few-shot learning, linear-probing, end-to-end fine-tuning, and parameter-efficient fine-tuning. Comparative analyses with established supervised, self-supervised, and weakly-supervised models reveal DINOv2's superior performance and cross-task generalizability. The findings contribute insights to potential avenues for optimizing pre-training strategies for medical imaging and enhancing the broader understanding of DINOv2's role in bridging the gap between natural and radiological image analysis. Our code is available at https://github.com/MohammedSB/DINOv2ForRadiology

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. In-Context Learning for Wound Classification with Small Multimodal Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Retrieval-based in-context learning, not zero-shot prompting, drives wound-classification gains in small multimodal models, with Qwen 3.5 27B reaching 0.872 accuracy on Kaggle and 0.678 on Medetec.

  2. Axial-Centric Cross-Plane Attention for 3D Medical Image Classification

    cs.CV 2026-02 conditional novelty 6.0 of 10

    An axial-centric cross-plane attention model with a frozen medical VFM reports top accuracy on five of six MedMNIST3D datasets.

  3. Brewing Stronger Features: Dual-Teacher Distillation for Multispectral Earth Observation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A dual-teacher contrastive distillation pipeline (multispectral EMA teacher + frozen DINOv3 optical teacher) yields a Swin-based Earth-observation model with state-of-the-art average results on segmentation, change de...

  4. VFM-Guided Semi-Supervised Detection Transformer under Source-Free Constraints for Remote Sensing Object Detection

    cs.CV 2025-08 conditional novelty 6.0 of 10

    VG-DETR combines a mean-teacher detector with DINOv2-guided pseudo-label mining and dual-level feature alignment, reporting 77.5% mAP on xView to DOTA and 70.6% on HRRSD to SSDD.

  5. SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications

    cs.CV 2025-07 conditional novelty 6.0 of 10

    General-purpose video foundation models, adapted with lightweight readout heads, reach state-of-the-art performance on three of five scientific video benchmarks.

  6. Simplifying DINO via Coding Rate Regularization

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Replacing DINO's complex anti-collapse machinery with an explicit coding rate regularizer yields simpler, more stable, and higher-performing self-supervised models.

  7. Is an Ultra Large Natural Image-Based Foundation Model Superior to a Retina-Specific Model for Detecting Ocular and Systemic Diseases?

    eess.IV 2025-02 conditional novelty 6.0 of 10

    A head-to-head benchmark shows DINOv2 generally outperforms RETFound on ocular disease detection, while RETFound is superior for predicting systemic disease incidence from retinal images.

  8. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A frozen 2D vision foundation model with lightweight LoRA adapters and attention-based slice fusion achieves state-of-the-art 3D medical image classification across 12 tasks with about 1M trainable parameters per task.

  9. MobQA: A Benchmark Dataset for Semantic Understanding of Human Mobility Data through Question Answering

    cs.CL 2025-08 unverdicted novelty 5.0 of 10

    LLMs handle factual lookups on mobility trajectories well but perform far worse on reasoning and explanation questions in the new 5,800-pair MobQA benchmark.

  10. CDPDNet: Integrating Text Guidance with Hybrid Vision Encoders for Medical Image Segmentation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A hybrid CNN-DINOv2-CLIP model with text-generated task prompts improves multi-organ and tumor segmentation on partially labeled CT datasets.

  11. Multi-Scale Feature Fusion with Image-Driven Spatial Integration for Left Atrium Segmentation from Cardiac MRI Images

    cs.CV 2025-02 conditional novelty 4.0 of 10

    A DINOv2 encoder with a UNet decoder, learned multi-scale feature weights, and input image integration improves left atrium segmentation over nnUNet on the LAScarQS 2022 dataset.

Pith tools