REVIEW 5 cited by
Towards a Visual-Language Foundation Model for Computational Pathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The accelerated adoption of digital pathology and advances in deep learning have enabled the development of powerful models for various pathology tasks across a diverse array of diseases and patient cohorts. However, model training is often difficult due to label scarcity in the medical domain and the model's usage is limited by the specific task and disease for which it is trained. Additionally, most models in histopathology leverage only image data, a stark contrast to how humans teach each other and reason about histopathologic entities. We introduce CONtrastive learning from Captions for Histopathology (CONCH), a visual-language foundation model developed using diverse sources of histopathology images, biomedical text, and notably over 1.17 million image-caption pairs via task-agnostic pretraining. Evaluated on a suite of 13 diverse benchmarks, CONCH can be transferred to a wide range of downstream tasks involving either or both histopathology images and text, achieving state-of-the-art performance on histology image classification, segmentation, captioning, text-to-image and image-to-text retrieval. CONCH represents a substantial leap over concurrent visual-language pretrained systems for histopathology, with the potential to directly facilitate a wide array of machine learning-based workflows requiring minimal or no further supervised fine-tuning.
Forward citations
Cited by 5 Pith papers
-
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
MR-PLIP is a multi-resolution pathology vision-language model that aligns histology patches and generated text across 5x, 10x, 20x, and 40x magnifications and reports improved transfer to 26 downstream pathology benchmarks.
-
ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image Classification
A dual-scale vision-language MIL framework with LLM-generated descriptive prompts and prototype-guided feature aggregation improves few-shot whole slide image classification.
-
Aligning Knowledge Concepts to Whole Slide Images for Precise Histopathology Image Analysis
ConcepPath aligns WSI patches to GPT-4-induced expert concepts plus learned concepts and uses a two-stage aggregation to outperform prior MIL methods on five histopathology classification tasks.
-
SimMIL: A Universal Weakly Supervised Pre-Training Framework for Multi-Instance Learning in Whole Slide Pathology Images
A weakly supervised pre-training scheme that propagates bag labels to patches improves downstream MIL classification and survival prediction on WSI datasets, but the comparison baselines are not trained on the same ta...
-
Exploring visual language models as a powerful tool in the diagnosis of Ewing Sarcoma
A frozen pathology vision-language model (PLIP) with a transformer-based multiple instance learning aggregator outperforms ImageNet VGG backbones at classifying Ewing sarcoma versus three similar sarcomas.
Discussion (0). Continue with ORCID to comment.