REVIEW 37 cited by
Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Virchow2: Scaling Self-Supervised Mixed Magnification Models in Pathology
read the original abstract
Foundation models are rapidly being developed for computational pathology applications. However, it remains an open question which factors are most important for downstream performance with data scale and diversity, model size, and training algorithm all playing a role. In this work, we propose algorithmic modifications, tailored for pathology, and we present the result of scaling both data and model size, surpassing previous studies in both dimensions. We introduce three new models: Virchow2, a 632 million parameter vision transformer, Virchow2G, a 1.9 billion parameter vision transformer, and Virchow2G Mini, a 22 million parameter distillation of Virchow2G, each trained with 3.1 million histopathology whole slide images, with diverse tissues, originating institutions, and stains. We achieve state of the art performance on 12 tile-level tasks, as compared to the top performing competing models. Our results suggest that data diversity and domain-specific methods can outperform models that only scale in the number of parameters, but, on average, performance benefits from the combination of domain-specific methods, data scale, and model scale.
Forward citations
Cited by 37 Pith papers
-
STREAM: Stochastic Riemannian Flow Matching with Anisotropic Decoder for Digital Histopathology Image Generation
STREAM applies stochastic Riemannian flow matching on VFM-derived unit hypersphere latents with a novel anisotropic decoder to achieve SOTA reconstruction and generation on breast and colorectal cancer histopathology ...
-
Geometry-Aware State Space Model: A New Paradigm for Whole-Slide Image Representation
BatMIL uses hybrid hyperbolic-Euclidean geometry, an S4 state-space backbone, and chunk-level mixture-of-experts to outperform prior multiple-instance learning methods on seven whole-slide image datasets across six cancers.
-
A Generative Foundation Model for Multimodal Histopathology
MuPD is a pretrained generative foundation model using a diffusion transformer with cross-modal attention that synthesizes histopathology images from text or RNA data and outperforms task-specific models on generation...
-
LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models
Language-guided adaptive multi-teacher distillation yields an 87M pathology encoder that matches or exceeds GigaPath and UNI on WSI captioning, VQA, and MIL tasks.
-
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts
Multi-stage agglomerative distillation consolidates eight vision, vision-language, and slide-level pathology teachers into one backbone that ranks first on average across 96 downstream tasks.
-
Multi-Teacher Contrastive Distillation for Edge-Efficient Pathology Foundation Models
Multi-teacher contrastive distillation (MuCoDi) compresses Virchow2/UNI2/H-Optimus-1 into edge encoders that reach 71.0% external AUROC (vs 71.8% best teacher) and up to 605× Raspberry Pi speedup.
-
The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models
Mid-sized pathology foundation models match or beat billion-parameter ones on clinically realistic perturbations and distribution-shift tests, so scaling alone has largely saturated for robustness.
-
JASPR: Joint Spatial Representation learning of histology and spatial genomics for improved virtual genomic screening and clinical prognostication
JASPR integrates HE images and ST data through self-supervised cross-modal reconstruction with shared and modality-specific modules to improve virtual gene prediction and breast cancer prognostication.
-
A Multi-modal Agentic Co-pilot for Evidence Grounded Computational Pathology
PathPocket constructs a 4.55M-entity pathology hypergraph from 110k graded documents and deploys a multi-agent framework that outperforms prior systems on 200k cases while raising pathologist accuracy in user studies.
-
DaX: Learning General Pathology Representations Across Scales
DaX is a pathology vision foundation model that extends DINOv3 with continuous magnification training and cross-scale consistency, achieving top average performance on a benchmark of 161 tasks from 44 datasets coverin...
-
LRMIL: Efficient Low-Resolution Multiple Instance Learning via High-Resolution Knowledge Distillation for Whole Slide Image Classification
LRMIL employs a two-stage high-to-low resolution knowledge distillation strategy to train an efficient low-resolution MIL model for WSI classification that outperforms existing methods with lower computational cost.
-
Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology
Symb-xMIL is a post-hoc explanation framework that quantifies MIL model alignment with logical decision rules in histopathology to enable rule-based interpretability.
-
A Pathology Foundation Model for Gastric Cancer with Real-World Validation
GRACE, a gastric-specific pathology foundation model trained on multicenter HE-stained slides, outperforms pancancer models on 28 tasks and improves pathologist accuracy, speed, and agreement in a reader study while e...
-
Do Foundation Models See Biology? Evaluating Attention Coherence with Spatial Transcriptomics in Glioblastoma
A new evaluation framework aligns attention maps from five pathology foundation models with co-registered Visium spatial transcriptomics in glioblastoma, revealing five-fold stronger coherence with multi-gene pathways...
-
Spatial Transcriptomics-Guided Alignment Enhances Molecular Profiling in Pathology Foundation Model
STAMP uses a curated 1.8M-pair spatial transcriptomics atlas and pathway-informed alignment to augment pathology foundation models for molecular phenotype inference from H&E WSIs.
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.
-
CRISP -- Clustering-Based Redundancy-Reduced Instance Sampling for Pathology Case Representation and Retrieval
CRISP is a clustering-based sampling framework that builds case-level representations from multiple whole-slide images for improved pathology retrieval, matching or exceeding single-slide selection on two breast cance...
-
Unified Multi-Foundation-Model Slide Representation for Pan-Cancer Recognition and Text-Guided Tumor Localization
ASTRA unifies heterogeneous pathology foundation-model representations for pan-cancer classification and weakly supervised tumor localization using only slide-level structured annotations.
-
PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning
PC-MIL shows that anchoring supervision at a 2 mm scale and progressively mixing slide- and region-level labels improves cross-context accuracy in WSI cancer detection without reducing global performance.
-
MOOZY: A Patient-First Foundation Model for Computational Pathology
Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.
-
Enabling clinical use of foundation models for computational pathology
Novel robustness losses added during downstream training on foundation-model features from pathology slides improve both robustness to technical variation and classification accuracy.
-
Uncertainty Estimation in Pathology Foundation Models via Deep Mutual Learning
DICE ensembles frozen pathology foundation models, aligns them with deep mutual learning to make disagreement a reliable uncertainty proxy, and shows consensus-based localization on WSI tasks.
-
Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation
GLMP generates robust pathology embeddings by routing histology images through an intermediate textual representation produced by general-purpose MLLMs to mitigate batch effects.
-
Multi-FRuGaL: Multimodal Flexible Redundancy-aware Decomposed Gated Learning for Cancer Diagnosis and Prognosis
Multi-FRuGaL is a decomposition-aware gated fusion framework for multimodal cancer data that maintains performance under missing modalities and reports AUC gains on two head-and-neck cancer cohorts.
-
SlideCheck: Guiding Self-Supervised Pretraining of Pathology Foundation Models via Dataset Distributions
SlideCheck uses a dual-head MLP on frozen features plus MIL attention to score patches and filter pretraining subsets that approach full-data performance in self-supervised pathology ViT models.
-
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.
-
Retrieval-Guided Generation for Safer Histopathology Image Captioning
Retrieval-guided captioning from similar cases achieves higher semantic alignment (cosine similarity ~0.60 vs ~0.47) and fewer unsupported diagnoses than MedGemma on the ARCH dataset.
-
Weakly Supervised Multicenter Nancy Index Scoring in Ulcerative Colitis Using Foundation Models
Weakly supervised MIL with foundation models enables robust five-grade Nancy index prediction and neutrophilic activity assessment from slide-level labels in multicenter UC biopsies.
-
SSMamba: A Self-Supervised Hybrid State Space Model for Pathological Image Classification
SSMamba uses a two-stage self-supervised pretraining and fine-tuning pipeline with Mamba-based components to outperform prior pathological foundation models on ROI and WSI classification tasks.
-
Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts
Pathology foundation models deliver strong in-distribution prostate cancer grading performance but exhibit large drops under cross-site image appearance shifts while remaining relatively robust to label distribution shifts.
-
CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images
CellPrior-Net integrates hematoxylin channel prior into a lightweight CNN for nuclei detection and classification in H&E WSIs, claiming comparable accuracy to SOTA with significantly reduced inference time across 10.4...
-
Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge
MIDOG 2025 challenge shows top mitosis detection F1 of 0.740 and atypical figure balanced accuracy of 0.908 across diverse tumors, with clear drops in challenging regions and tumor-type variation.
-
Validation of an AI-based end-to-end model for prostate pathology using long-term archived routine samples
GleasonAI achieves quadratic-weighted kappa of 0.86 on ISUP grading of 10,366 long-term archived prostate biopsy cores, with performance stable over 17 years and a clear prognostic gradient for cancer-specific mortality.
-
Benchmarking Pathology Foundation Models for Breast Cancer Survival Prediction
H-optimus-1 achieves the strongest externally validated survival prediction from histopathology images, with second-generation PFMs outperforming first-generation counterparts and a compact distilled model offering ef...
-
OpenTME: An Open Dataset of AI-powered H&E Tumor Microenvironment Profiles from TCGA
OpenTME provides pre-computed TME profiles with over 4,500 quantitative readouts per slide from 3,634 TCGA H&E images using an AI pipeline based on pathology foundation models.
-
MMAP: A Multi-Magnification and Prototype-Aware Architecture for Predicting Spatial Gene Expression
MMAP uses multi-magnification patch features and slide-level prototype embeddings to predict spatial gene expression from H&E images and reports better MAE, MSE, and PCC than prior methods.
-
Transformer-Based Hematological Malignancy Prediction from Peripheral Blood Smears in a Real-World Cohort
cAItomorph applies a transformer aggregator on DinoBloom embeddings to predict eight coarse hematological malignancy classes from peripheral blood single-cell images, achieving 0.72 accuracy and reducing false discove...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.