Pith. sign in

REVIEW 2 cited by

Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.14911 v2 pith:XNKKBGZO submitted 2025-03-19 cs.CV

classification cs.CV
keywords clinicalderm1mdatasetacrossdermatologymedicalmodelsskin
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of vision-language models has transformed medical AI, enabling unprecedented advances in diagnostic capability and clinical applications. However, progress in dermatology has lagged behind other medical domains due to the lack of standard image-text pairs. Existing dermatological datasets are limited in both scale and depth, offering only single-label annotations across a narrow range of diseases instead of rich textual descriptions, and lacking the crucial clinical context needed for real-world applications. To address these limitations, we present Derm1M, the first large-scale vision-language dataset for dermatology, comprising 1,029,761 image-text pairs. Built from diverse educational resources and structured around a standard ontology collaboratively developed by experts, Derm1M provides comprehensive coverage for over 390 skin conditions across four hierarchical levels and 130 clinical concepts with rich contextual information such as medical history, symptoms, and skin tone. To demonstrate Derm1M potential in advancing both AI research and clinical application, we pretrained a series of CLIP-like models, collectively called DermLIP, on this dataset. The DermLIP family significantly outperforms state-of-the-art foundation models on eight diverse datasets across multiple tasks, including zero-shot skin disease classification, clinical and artifacts concept identification, few-shot/full-shot learning, and cross-modal retrieval. Our dataset and code will be publicly available at https://github.com/SiyuanYan1/Derm1M upon acceptance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An End-to-End Deep Learning Framework for Arsenicosis Diagnosis Using Mobile-Captured Skin Images

    cs.CV 2025-09 conditional novelty 6.0 of 10

    On a newly curated 20-class dataset of 11,489 mobile skin images, a Swin Transformer reached 86% accuracy for classifying arsenicosis and similar dermatoses, with external validation claimed but underreported.

  2. Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning Qwen2-VL to predict 16 quantitative skin attributes yields embeddings that retrieve images matching a query in both appearance and the target attribute.

Pith tools