Pith. sign in

REVIEW 5 cited by

Towards a Visual-Language Foundation Model for Computational Pathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.12914 v2 pith:T7E24HWS submitted 2023-07-24 cs.CV cs.AI

classification cs.CVcs.AI
keywords histopathologymodelconchdiversepathologyvisual-languagearrayfoundation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The accelerated adoption of digital pathology and advances in deep learning have enabled the development of powerful models for various pathology tasks across a diverse array of diseases and patient cohorts. However, model training is often difficult due to label scarcity in the medical domain and the model's usage is limited by the specific task and disease for which it is trained. Additionally, most models in histopathology leverage only image data, a stark contrast to how humans teach each other and reason about histopathologic entities. We introduce CONtrastive learning from Captions for Histopathology (CONCH), a visual-language foundation model developed using diverse sources of histopathology images, biomedical text, and notably over 1.17 million image-caption pairs via task-agnostic pretraining. Evaluated on a suite of 13 diverse benchmarks, CONCH can be transferred to a wide range of downstream tasks involving either or both histopathology images and text, achieving state-of-the-art performance on histology image classification, segmentation, captioning, text-to-image and image-to-text retrieval. CONCH represents a substantial leap over concurrent visual-language pretrained systems for histopathology, with the potential to directly facilitate a wide array of machine learning-based workflows requiring minimal or no further supervised fine-tuning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation

    cs.CV 2025-04 conditional novelty 6.0 of 10

    MR-PLIP is a multi-resolution pathology vision-language model that aligns histology patches and generated text across 5x, 10x, 20x, and 40x magnifications and reports improved transfer to 26 downstream pathology benchmarks.

  2. ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A dual-scale vision-language MIL framework with LLM-generated descriptive prompts and prototype-guided feature aggregation improves few-shot whole slide image classification.

  3. Aligning Knowledge Concepts to Whole Slide Images for Precise Histopathology Image Analysis

    cs.CV 2024-11 conditional novelty 6.0 of 10

    ConcepPath aligns WSI patches to GPT-4-induced expert concepts plus learned concepts and uses a two-stage aggregation to outperform prior MIL methods on five histopathology classification tasks.

  4. SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A weakly supervised pre-training scheme that propagates bag labels to patches improves downstream MIL classification and survival prediction on WSI datasets, but the comparison baselines are not trained on the same ta...

  5. Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma

    cs.CV 2025-01 conditional novelty 4.0 of 10

    A frozen pathology vision-language model (PLIP) with a transformer-based multiple instance learning aggregator outperforms ImageNet VGG backbones at classifying Ewing sarcoma versus three similar sarcomas.

Pith tools