Pith. sign in

REVIEW 22 cited by

Multimodal Whole Slide Foundation Model for Pathology

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.19666 v1 pith:G7TVLOQA submitted 2024-11-29 eess.IV cs.AIcs.CVcs.LGstat.AP

classification eess.IVcs.AIcs.CVcs.LGstat.AP
keywords clinicalpathologyslidefoundationtitanlearningmultimodalrare
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of computational pathology has been transformed with recent advances in foundation models that encode histopathology region-of-interests (ROIs) into versatile and transferable feature representations via self-supervised learning (SSL). However, translating these advancements to address complex clinical challenges at the patient and slide level remains constrained by limited clinical data in disease-specific cohorts, especially for rare clinical conditions. We propose TITAN, a multimodal whole slide foundation model pretrained using 335,645 WSIs via visual self-supervised learning and vision-language alignment with corresponding pathology reports and 423,122 synthetic captions generated from a multimodal generative AI copilot for pathology. Without any finetuning or requiring clinical labels, TITAN can extract general-purpose slide representations and generate pathology reports that generalize to resource-limited clinical scenarios such as rare disease retrieval and cancer prognosis. We evaluate TITAN on diverse clinical tasks and find that TITAN outperforms both ROI and slide foundation models across machine learning settings such as linear probing, few-shot and zero-shot classification, rare cancer retrieval and cross-modal retrieval, and pathology report generation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Do Multiple Instance Learning Models Transfer?

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Pretrained multiple instance learning models transfer across organs and tasks in computational pathology, and pancancer pretraining can rival slide foundation models with far less data.

  2. From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.

  3. Value-Monotonicity Matters: A Concordance Loss for Deep Survival Prediction

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A sigmoid concordance loss whose value approximates one minus the C-index stays coupled to ranking performance throughout training, while likelihood-based survival losses provably and empirically decouple from it.

  4. MOOZY: A Patient-First Foundation Model for Computational Pathology

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.

  5. Integrating Pathology and CT Imaging for Personalized Recurrence Risk Prediction in Renal Cancer

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Multimodal fusion of CT and pathology images improves recurrence risk prediction in kidney cancer, with the best model approaching the clinical Leibovich score.

  6. Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping

    cs.CV 2025-08 conditional novelty 6.0 of 10

    PathPT improves few-shot rare cancer subtyping by using zero-shot vision-language models to create tile-level pseudo-labels and learning prompt tokens with spatial context, outperforming standard MIL baselines when th...

  7. A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics

    q-bio.GN 2025-08 unverdicted novelty 6.0 of 10

    HESCAPE shows that cross-modal contrastive pretraining helps mutation classification but hurts gene expression prediction in spatial transcriptomics, implicating batch effects.

  8. Towards Robust Foundation Models for Digital Pathology

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PathoROB shows that all 20 evaluated pathology foundation models encode medical center information and that lower robustness correlates with larger downstream performance drops.

  9. WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.

  10. Single Domain Generalization for Multimodal Cross-Cancer Prognosis via Dirac Rebalancer and Distribution Entanglement

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Multimodal prognosis models often generalize worse than unimodal ones across cancer types; a sparse rebalancer plus distribution-entanglement module improves cross-cancer C-index from 0.5489 to 0.5625, though hyperpar...

  11. SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

    q-bio.QM 2025-07 conditional novelty 6.0 of 10

    A hierarchical multimodal model fusing morphology, expression, and spatial context that generates target-state cell morphologies from optimal-transport weak pairs, trained on a 25.9M-cell atlas and benchmarked against...

  12. A Foundation Model for Spatial Proteomics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    KRONOS, a self-supervised foundation model for spatial proteomics, outperforms existing vision models on cell phenotyping, retrieval, region classification, and label-efficient tasks across 11 cohorts.

  13. SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence

    cs.CV 2025-06 conditional novelty 6.0 of 10

    SurgVLM, a family of surgical vision-language models trained on 1.81M frames and 7.79M conversations, outperforms 14 commercial VLMs on a six-dataset surgical benchmark.

  14. The Butterfly Effect in Pathology: Exploring Security in Pathology Foundation Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A label-free attack that perturbs only 0.1% of patches in a whole-slide image can shift the model's global representation and substantially degrade downstream pathology task accuracy.

  15. Subspecialty-Specific Foundation Model for Intelligent Gastrointestinal Pathology

    eess.IV 2025-05 conditional novelty 6.0 of 10

    A GI-specific histopathology foundation model pretrained on 210,043 slides with a supervised ROI-mining second stage claims state-of-the-art results on 33 of 34 GI pathology tasks and 99.70 percent screening sensitivi...

  16. PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology

    cs.CV 2025-12 conditional novelty 5.0 of 10

    Splitting slide captions into random sentence subcaptions and aligning them with region features via text-conditioned attention improves whole-slide classification, retrieval, captioning and VQA in computational pathology.

  17. Exploiting Layer Normalization Fine-tuning in Visual Transformer Foundation Models for Classification

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    LayerNorm parameter shifts after fine-tuning encode domain-transition information; rescaling them via an FSR-dependent scalar lambda plus a cyclic step improves ViT classification under data scarcity and domain shift.

  18. EXAONE Path 2.0: Pathology Foundation Model with End-to-End Supervision

    cs.CV 2025-07 reject novelty 5.0 of 10

    EXAONE Path 2.0, a hierarchical vision transformer pretrained with slide-level supervision on 37k whole-slide images, reports the highest average AUROC across 10 pathology biomarker benchmarks.

  19. Accelerating Data Processing and Benchmarking of AI Models for Pathology

    cs.CV 2025-02 conditional novelty 5.0 of 10

    The authors introduce Trident, a WSI processing package, Patho-Bench, a benchmarking library, and 42 curated pathology tasks to standardize foundation model evaluation.

  20. Enhancing Zero-Shot Brain Tumor Subtype Classification via Fine-Grained Patch-Text Alignment

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    FG-PAN improves zero-shot brain tumor subtype classification by aligning refined visual patch features with LLM-generated fine-grained text prototypes.

  21. Leveraging Pathology Foundation Models for Panoptic Segmentation of Melanoma in H&E Images

    eess.IV 2025-07 conditional novelty 4.0 of 10

    A network combining a pathology foundation model (Virchow2) with an Efficient-UNet segments melanoma tissue types and won the PUMA challenge tissue segmentation task.

  22. Emerging AI Approaches for Cancer Spatial Omics

    q-bio.QM 2025-06 unverdicted novelty 2.0 of 10

    A review that groups AI methods for cancer spatial omics into data-driven, constraint-based, and mechanistic modeling paradigms, calling for more interpretable models and mouse-model-generated perturbational data.

Pith tools