Pith. sign in

REVIEW 16 cited by

PRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.10254 v2 pith:XOEPVTU4 submitted 2024-05-16 eess.IV cs.CVcs.LG

classification eess.IVcs.CVcs.LG
keywords prismmodelsslideclinicalembeddingsfoundationaggregatordata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Foundation models in computational pathology promise to unlock the development of new clinical decision support systems and models for precision medicine. However, there is a mismatch between most clinical analysis, which is defined at the level of one or more whole slide images, and foundation models to date, which process the thousands of image tiles contained in a whole slide image separately. The requirement to train a network to aggregate information across a large number of tiles in multiple whole slide images limits these models' impact. In this work, we present a slide-level foundation model for H&E-stained histopathology, PRISM, that builds on Virchow tile embeddings and leverages clinical report text for pre-training. Using the tile embeddings, PRISM produces slide-level embeddings with the ability to generate clinical reports, resulting in several modes of use. Using text prompts, PRISM achieves zero-shot cancer detection and sub-typing performance approaching and surpassing that of a supervised aggregator model. Using the slide embeddings with linear classifiers, PRISM surpasses supervised aggregator models. Furthermore, we demonstrate that fine-tuning of the PRISM slide encoder yields label-efficient training for biomarker prediction, a task that typically suffers from low availability of training data; an aggregator initialized with PRISM and trained on as little as 10% of the training data can outperform a supervised baseline that uses all of the data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Multiple Instance Learning Models Transfer?

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Pretrained multiple instance learning models transfer across organs and tasks in computational pathology, and pancancer pretraining can rival slide foundation models with far less data.

  2. From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.

  3. Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CRoMa scores each pathology image embedding by the margin between cross-site biological matches and same-site biological distractors, revealing distributional lower tails that pooled robustness scores hide.

  4. Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Distilling TITAN and CARE slide embeddings into MIL aggregators gives reusable pretrained weights that beat from-scratch training on most of 15 pathology tasks, with the largest gains in few-shot and linear-probing settings.

  5. Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous Reasoning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PathCTM uses adaptive continuous reasoning across scales to reduce patch processing in whole slide images by over 95% while preserving diagnostic AUC.

  6. MOOZY: A Patient-First Foundation Model for Computational Pathology

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.

  7. Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PathPT improves few-shot rare cancer subtyping by using zero-shot vision-language models to create tile-level pseudo-labels and learning prompt tokens with spatial context, outperforming standard MIL baselines when th...

  8. WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.

  9. Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis

    cs.CV 2025-06 conditional novelty 6.0 of 10

    TAPFM adapts pathology foundation models on a single GPU for WSI mutation prediction, outperforming fixed-feature and end-to-end fine-tuning baselines.

  10. A Foundation Model for Spatial Proteomics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    KRONOS, a self-supervised foundation model for spatial proteomics, outperforms existing vision models on cell phenotyping, retrieval, region classification, and label-efficient tasks across 11 cohorts.

  11. The Butterfly Effect in Pathology: Exploring Security in Pathology Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A label-free attack that perturbs only 0.1% of patches in a whole-slide image can shift the model's global representation and substantially degrade downstream pathology task accuracy.

  12. GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Distilled, Apache-2.0-licensed GigaPath-Flash and GigaTIME-Flash models deliver most of the original models' accuracy at a fraction of the compute and memory.

  13. Atlas 2 -- Foundation models for clinical deployment

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Atlas 2 and its distilled variants set new average state-of-the-art results across 80 pathology benchmarks, with larger robustness margins over prior models.

  14. EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision

    cs.CV 2025-07 reject novelty 5.0 of 10

    EXAONE Path 2.0, a hierarchical vision transformer pretrained with slide-level supervision on 37k whole-slide images, reports the highest average AUROC across 10 pathology biomarker benchmarks.

  15. Accelerating Data Processing and Benchmarking of AI Models for Pathology

    cs.CV 2025-02 conditional novelty 5.0 of 10

    The authors introduce Trident, a WSI processing package, Patho-Bench, a benchmarking library, and 42 curated pathology tasks to standardize foundation model evaluation.

  16. MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models

    cs.CV 2025-08 conditional novelty 4.0 of 10

    A GUI toolkit for secure local deployment of medical vision-language models, with a pathologist-scored evaluation showing these models still frequently fail on pathology questions.

Pith tools