REVIEW 16 cited by
PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Foundation models in computational pathology promise to unlock the development of new clinical decision support systems and models for precision medicine. However, there is a mismatch between most clinical analysis, which is defined at the level of one or more whole slide images, and foundation models to date, which process the thousands of image tiles contained in a whole slide image separately. The requirement to train a network to aggregate information across a large number of tiles in multiple whole slide images limits these models' impact. In this work, we present a slide-level foundation model for H&E-stained histopathology, PRISM, that builds on Virchow tile embeddings and leverages clinical report text for pre-training. Using the tile embeddings, PRISM produces slide-level embeddings with the ability to generate clinical reports, resulting in several modes of use. Using text prompts, PRISM achieves zero-shot cancer detection and sub-typing performance approaching and surpassing that of a supervised aggregator model. Using the slide embeddings with linear classifiers, PRISM surpasses supervised aggregator models. Furthermore, we demonstrate that fine-tuning of the PRISM slide encoder yields label-efficient training for biomarker prediction, a task that typically suffers from low availability of training data; an aggregator initialized with PRISM and trained on as little as 10% of the training data can outperform a supervised baseline that uses all of the data.
Forward citations
Cited by 16 Pith papers
-
Do Multiple Instance Learning Models Transfer?
Pretrained multiple instance learning models transfer across organs and tasks in computational pathology, and pancancer pretraining can rival slide foundation models with far less data.
-
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.
-
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.
-
Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.
-
Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning
PathCTM uses adaptive continuous reasoning across scales to reduce patch processing in whole slide images by over 95% while preserving diagnostic AUC.
-
MOOZY: A Patient-First Foundation Model for Computational Pathology
Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.
-
Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
PathPT improves few-shot rare cancer subtyping by using zero-shot vision-language models to create tile-level pseudo-labels and learning prompt tokens with spatial context, outperforming standard MIL baselines when th...
-
WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis
A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.
-
Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis
TAPFM adapts pathology foundation models on a single GPU for WSI mutation prediction, outperforming fixed-feature and end-to-end fine-tuning baselines.
-
A Foundation Model for Spatial Proteomics
KRONOS, a self-supervised foundation model for spatial proteomics, outperforms existing vision models on cell phenotyping, retrieval, region classification, and label-efficient tasks across 11 cohorts.
-
The Butterfly Effect in Pathology: Exploring Security in Pathology Foundation Models
A label-free attack that perturbs only 0.1% of patches in a whole-slide image can shift the model's global representation and substantially degrade downstream pathology task accuracy.
-
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
Distilled, Apache-2.0-licensed GigaPath-Flash and GigaTIME-Flash models deliver most of the original models' accuracy at a fraction of the compute and memory.
-
Atlas 2 -- Foundation models for clinical deployment
Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.
-
EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision
EXAONE Path 2.0, a hierarchical vision transformer pretrained with slide-level supervision on 37k whole-slide images, reports the highest average AUROC across 10 pathology biomarker benchmarks.
-
Accelerating Data Processing and Benchmarking of AI Models for Pathology
The authors introduce Trident, a WSI processing package, Patho-Bench, a benchmarking library, and 42 curated pathology tasks to standardize foundation model evaluation.
-
MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
A GUI toolkit for secure local deployment of medical vision-language models, with a pathologist-scored evaluation showing these models still frequently fail on pathology questions.
Discussion (0). Continue with ORCID to comment.