Pith. sign in

REVIEW 3 major objections 6 minor 2 cited by

Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read The paper argues that foundation models in computational pathology are capable enough to transform diagnostics, but that progress is now gated by the absence of global benchmarks and standardized evaluation.

desk verdict A useful, honest pathology-FM review whose scoring scheme and abstract overreach need revision; the case studies alone make it worth a read. read the letter →

arxiv 2502.08333 v1 pith:JTLHVM2Y submitted 2025-02-12 cs.CV

classification cs.CV
keywords pathologyfoundationmodelscomputationalbenchmarkingclinicaladoptionvision-languagewholeslideimagesself-supervisedlearninggenerativeco-pilots
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper is a review that tries to establish a specific diagnosis of where computational pathology stands: foundation models, from vision-only encoders to generative co-pilots, have reached the point where they can mine subtle tissue cues, generate reports, and answer clinical questions, yet their capability claims cannot be trusted because each benchmarking study builds its own local benchmark. The authors argue that establishing global benchmarks is the crucial step for raising evaluation standards and enabling widespread clinical adoption. If this is right, the main bottleneck is no longer model scale or architecture but the evaluation infrastructure used to compare models fairly across institutions, tasks, and patient populations. The review shifts attention from building bigger models to building the measurement apparatus that would let the field know which models are actually better.

What carries the argument

The argument is carried by a taxonomy of pathology foundation models, organized into four architectural paradigms: image-only self-supervised models, multi-stain models, cross-modal image-text models, and generative multimodal co-pilots. On top of this, the paper applies a scoring rubric that rates each model on generality, multipurpose-ness, multimodality, and generativity, and it organizes evaluation into intrinsic evaluation of the model itself versus extrinsic evaluation of adapted downstream tasks. That intrinsic/extrinsic distinction is the load-bearing mechanism: it explains why performance gains are hard to attribute, because task-specific adaptation confounds the contribution of the foundation model, and it drives the conclusion that shared benchmarks are needed.

What would settle it

Take five leading pathology foundation models, such as UNI, Virchow2, RudolfV, CHIEF, and CONCH, and evaluate them on one fixed set of slide-level diagnostic tasks with the same linear-probing protocol, the same external cohorts, and the same metrics; if their observed performance differences largely disappear or reverse, the review's comparative rankings and its claim that model capability differences are real would be falsified. Alternatively, if clinical adoption proceeds widely without global benchmarks, the claim that benchmarks are crucial would be weakened.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that pathology foundation models have demonstrated strong predictive and generative capabilities, yet the evidence assembled from roughly forty models and twelve benchmarking studies supports a narrower conclusion: current evaluations are fragmented, with each benchmark constructed locally, so direct and fair comparison between models is limited. The authors therefore claim that global benchmarks, not further scaling alone, are necessary to enhance evaluation standards and foster clinical adoption. The review also advances a conceptual distinction between general and multipurpose models, where generality means breadth across tissues, resolutions, stains, and scanners, and multipurpose means the ability to perform many task types, and it documents how generative co-pilots such as report generators and interactive assistants are expanding the evaluation problem beyond accuracy to safety, interpretability, and workflow integration.

Load-bearing premise

The load-bearing premise is that the performance numbers reported across different source papers are comparable, even though each study uses its own evaluation protocol, cohorts, and metrics; the paper itself acknowledges that local benchmarks limit direct and fair comparison.

Editorial extensions

If this is right

  • If the field adopts global benchmarks, model rankings will become reproducible and differences in architecture, data, and adaptation methods can be separated.
  • Regulators and clinical buyers would get a common yardstick for deciding which foundation models are safe and effective enough for diagnostics.
  • Research groups without access to massive compute could still contribute by showing that their adapted models outperform existing baselines on the shared benchmark.
  • Benchmark results would likely change current leaderboards, because several models in the review outperform others only on their own local benchmarks.
  • Generative co-pilots would need evaluation suites that test not just accuracy but safety, interpretability, and human-in-the-loop performance before clinical adoption.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One consequence the authors leave implicit is that a shared benchmark should specify a fixed adaptation protocol, such as linear probing, or else rankings will still conflate model quality with fine-tuning choices.
  • A testable extension of the review's thesis: if a consortium re-evaluated the leading models on one external, multi-institution cohort with identical protocols, the advertised gaps, such as RudolfV outperforming UNI in 10 of 12 benchmarks, would shrink or reorder.
  • The review's emphasis on data diversity over data volume suggests that future foundation-model work may shift from collecting more slides to curating balanced cohorts across stains, scanners, populations, and rare diseases.
  • If global benchmarks materialize, they could also become a coordination point for reporting computational cost, carbon emissions, and failure modes, which the review identifies as underreported.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This manuscript is a narrative review of foundation models in computational pathology. It catalogs roughly 40 pathology foundation models, proposes a terminology for "general" versus "multipurpose" capabilities, and summarizes their architectures, training objectives, aggregation strategies, downstream tasks, and evaluation protocols. It also presents three clinical case studies (prostate cancer detection, MSI status prediction, and colon biopsy pre-screening), a discussion of adaptation methods, and a list of grand challenges. The paper's central thesis, stated in the abstract and repeated throughout, is that foundation models have demonstrated exceptional predictive and generative capabilities but that the main barrier to clinical adoption is the absence of standardized global benchmarks.

Significance. If taken as a field survey, the paper is a potentially useful reference: it collects a large number of models in one place, distinguishes image-only, vision-language, and multimodal architectures, and draws attention to real gaps such as the lack of intrinsic evaluation, the risk of data leakage in models pretrained on TCGA, and the difficulty of aggregating gigapixel WSIs. The three clinical case studies in Section 6 are a valuable counterweight to purely benchmark-driven narratives, and the paper explicitly acknowledges that many benchmarking studies are not directly comparable. The original scoring figures (Figures 3 and 4) and the clinical case studies are the most distinctive contributions. However, the paper does not propose a new benchmark, methodology, or formal analysis; its value depends on the accuracy and internal consistency of its synthesis. The main tension is that the abstract claims "exceptional" capabilities while the paper's own clinical-grade evidence shows that specialized models can outperform foundation models and that performance on clinically relevant tasks is often modest.

major comments (3)
  1. [Abstract; Sections 6.1-6.4 and 4.4.4] The abstract's claim that foundation models "have demonstrated exceptional predictive and generative capabilities" is contradicted by the paper's own clinical-grade evaluations. In Section 6.1, Virchow's prostate cancer detection AUC of 0.980 is significantly lower than Paige Prostate's 0.995 (P<0.05). In Section 6.2, MSI AUROC on the PAIP cohort ranges from 0.66 (Hibou-B) to 0.978 (Phikon), and for MMR classification UNI and Phikon reach AUROC 0.7136 versus ResNet50's 0.6709. In Section 6.3, colon biopsy pre-screening balanced accuracies are 65% (UNI), 52% (REMEDIS), and 39% (CTransPath), far below the 0.99 sensitivity target that specialized models approach. Section 6.4 concedes that "specialized models such as Paige Prostate currently deliver superior results in specific tasks," and Section 4.4.4 states that foundation models are "not yet transformative or ready for widespread clinical adoption." The paper needs to reconcile these statements with the abstract, or qualify the capability claim as benchmark-level success rather than clinically validated capability.
  2. [Figure 3 and Figure 4; Section 2] The original scoring scheme used in Figures 3 and 4 is not sufficiently specified to support the comparative statements it generates. The caption for Figure 3 defines the multimodality score as 0.2×WSIs + 0.5×Text + 0.3×Molecular, but the WSI counts range from roughly 3,000 to 3.1 million, so the term "WSIs" cannot enter unnormalized without overwhelming the binary text and molecular indicators; the required normalization is not given. Similarly, the "generality" and "multipurpose" scores average Boolean attributes with a "scaled" number of WSIs, but the scaling rule is not defined, and the generative score is assigned from an unexplained lookup table (1.0 for WSI-language assistant, 0.5 for report generation, 0.0 for none). Since these scores drive the model positioning in Figure 4 and related narrative comparisons, the authors should either remove the scores, define all transformations and weights explicitly, or add a sensitivity analysis showing that the qualitative conclusions do not change under reasonable alternative weights.
  3. [Section 2 and Section 4.4] Several headline comparisons aggregate performance numbers obtained under incompatible evaluation protocols. For example, Section 2 states that "RudolfV outperformed UNI in 10 out of 12 benchmarks and 27 out of 31 datasets," and Section 4.3.1 reports that Virchow2 "achieved state-of-the-art performance in cancer detection and subtyping... and ranks 1st Eva leaderboard." These claims mix linear-probing results, fine-tuned results, zero-shot results, and different cohorts and metrics, as the paper itself acknowledges in Section 4.4.4: "these benchmarking studies have constructed their own local benchmark limiting direct and fair comparison." The review should mark each such comparison as either protocol-matched or as a qualitative summary of heterogeneous studies, and it should avoid presenting ranked comparisons without a detailed protocol table.
minor comments (6)
  1. [Section 6.2] The heading "Microsatellite Stability (MSS) Screening in Colorectal Cancer" is mislabeled; the task is the prediction of microsatellite instability (MSI), while MSS denotes the stable phenotype that a screening test aims to rule out.
  2. [Section 2] The sentence "Analyzing WSIs of human tissue is central to computational pathology, which explains why early foundation model efforts focused solely on vision-based or image-only approaches" is repeated verbatim in the section; one copy should be removed.
  3. [Sections 4.2.1 and 4.3.5] The model name is inconsistent: "RodulfV" appears in Sections 4.2.1 and 4.3.5, while "RudolfV" is used elsewhere; please standardize the spelling throughout.
  4. [Section 5.1] The sentence describing Mallya et al. contains typographical artifacts ("6 dation models founCTransPath" and "MIL methds"); the text should be corrected to "six foundation models" and "MIL methods".
  5. [Section 4.2.2] The abbreviation "MSUK" on the zero-shot cross-modal retrieval line should be "MUSK".
  6. [Section 6.1] The claim that Virchow "show[s] better generalization" while reporting only a single AUC value is unsupported; the authors should either provide comparative generalization evidence or rephrase the sentence to describe what was actually measured.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the review summarizes external published results and does not derive predictions from its own inputs.

full rationale

This manuscript is a literature review, not a derivation or modeling paper. Its claims about foundation models in computational pathology are presented as summaries of externally published benchmark results (e.g., Virchow, UNI, RudolfV, Paige Prostate, MSIntuit CRC, CAIMAN, IGUANA), with citations to those original studies. The abstract's assertion that foundation models have demonstrated exceptional predictive and generative capabilities is an interpretive synthesis, not a quantity fitted or derived within the paper. Section 6's clinical-grade case studies explicitly report comparative numbers from external sources, including cases where foundation models underperform specialized models, and Section 6.4 concedes that specialized models currently deliver superior results in specific tasks. No parameter is fitted and then renamed as a prediction; no uniqueness theorem or ansatz is imported from the authors' own prior work to force a conclusion; and the review's central message about the need for global benchmarks is a stated recommendation, not a result constructed from its own definitions. Self-citations appear (e.g., Bilal, Tsang, et al. 2023 in the colon biopsy and MSI discussions), but they reference external validation studies with published performance metrics and do not carry the logical weight of the review's conclusions. The weakest methodological issue is comparability across heterogeneous evaluation protocols, which the paper itself acknowledges when stating that benchmarking studies 'have constructed their own local benchmark limiting direct and fair comparison'; this is a correctness or rigor concern, not circularity. There is therefore no load-bearing step that reduces to the paper's own inputs.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The central claim does not rest on a derivation, so the main ledger entries are the review's implicit assumptions: that the foundation model definition transfers to pathology, that benchmark numbers are comparable across studies, and that the authors' hand-assigned scores are meaningful. The scoring weights in Figures 3 and 4 are the clearest hand-chosen inputs.

free parameters (2)
  • Multimodality score weights = 0.2 for WSIs, 0.5 for text, 0.3 for molecular
    Hand-chosen weights used to compute the multimodality score in Figure 3. No justification or sensitivity analysis is provided.
  • Generative score lookup values = 1.0 for WSI-language assistant, 0.5 for report generation, 0.0 for none
    Arbitrary discretization used to color the scatter plot and cluster heatmap in Figures 3 and 4. The mapping is not validated against any external measure.
assumptions (3)
  • domain assumption The foundation model framing from Bommasani et al. applies to pathology models, with scale, self-supervision, and adaptability as the right defining criteria.
    Section 2.1 defines foundation models this way and uses it to include or exclude models from the review.
  • domain assumption Performance numbers from cited papers are comparable across heterogeneous cohorts, tasks, and adaptation protocols.
    Sections 4.2 through 4.4 compare AUROCs, task counts, and benchmark ranks across models with different evaluation setups, despite the paper noting that protocols vary.
  • ad hoc to paper Hand-assigned scores (generative lookup table, weighted multimodality score) are meaningful measures of model capability.
    Figure 3 uses these scores to cluster and rank models without sensitivity analysis or external validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact." pith.science (2026). https://pith.science/paper/JTLHVM2Y

@misc{pith2026250208333,
  author       = {Pith},
  title        = {Pith review of: Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JTLHVM2Y}},
  note         = {Machine review of arXiv:2502.08333}
}
read the original abstract

From self-supervised, vision-only models to contrastive visual-language frameworks, computational pathology has rapidly evolved in recent years. Generative AI "co-pilots" now demonstrate the ability to mine subtle, sub-visual tissue cues across the cellular-to-pathology spectrum, generate comprehensive reports, and respond to complex user queries. The scale of data has surged dramatically, growing from tens to millions of multi-gigapixel tissue images, while the number of trainable parameters in these models has risen to several billion. The critical question remains: how will this new wave of generative and multi-purpose AI transform clinical diagnostics? In this article, we explore the true potential of these innovations and their integration into clinical practice. We review the rapid progress of foundation models in pathology, clarify their applications and significance. More precisely, we examine the very definition of foundational models, identifying what makes them foundational, general, or multipurpose, and assess their impact on computational pathology. Additionally, we address the unique challenges associated with their development and evaluation. These models have demonstrated exceptional predictive and generative capabilities, but establishing global benchmarks is crucial to enhancing evaluation standards and fostering their widespread clinical adoption. In computational pathology, the broader impact of frontier AI ultimately depends on widespread adoption and societal acceptance. While direct public exposure is not strictly necessary, it remains a powerful tool for dispelling misconceptions, building trust, and securing regulatory support.

Figures

Figures reproduced from arXiv: 2502.08333 by the authors.

Figure 1
Figure 1. Three marked phases of Machine Learning evolution: from machine learning, deep learning to foundation models. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Pathology (Multimodal) Foundation Models – Data sources, compute-extensive training as number of GPUs needed is much larger ( [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Cluster heatmap of foundation models presenting multipurpose, generality, generative and multimodality scores. Multimodality score is sum [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Foundation models scatter plot positioning each model in terms of its generality score (x-axis), multipurpose score (y-axis), generative score [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Pathology Foundation Model Training: Vision-text and vision-only pipeline and aggregation methods – mSTAR’s multimodal aggregation and [PITH_FULL_IMAGE:figures/full_fig_p017_5.png]
Figure 6
Figure 6. Figure 6: Foundation model adapted to perform conventional, advanced, and unique computational pathology tasks [PITH_FULL_IMAGE:figures/full_fig_p029_6.png]
Figure 7
Figure 7. Figure 7: Understanding Pathology Foundation Models Evaluation and Adaptation [PITH_FULL_IMAGE:figures/full_fig_p034_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Robust Foundation Models for Digital Pathology

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PathoROB shows that all 20 evaluated pathology foundation models encode medical center information and that lower robustness correlates with larger downstream performance drops.

  2. PathBench: A comprehensive comparison benchmark for pathology foundation models towards precision oncology

    cs.CV 2025-05 reject novelty 6.0 of 10

    A large private multi-center benchmark of 19 pathology foundation models on 64 tasks finds Virchow2 and H-Optimus-1 best overall, with vision-language models lagging behind.

Reference graph

Works this paper leans on

14 extracted references · 4 canonical work pages · cited by 2 Pith papers

  1. [4]

    ‘HistGen: Histopathology Report Generation via Local-Global Feature Encoding and Cross-Modal Context Interaction’ . arXiv. http://arxiv.org/abs/2403.05396. Guo, Zishan, Renren Jin, Chuang Liu, Yufei Huang, Dan Shi, Supryadi, Linhao Yu, et al

  2. [6]

    Nature Medicine 29 (9): 2307–16

    ‘A Visual–Language Foundation Model for Pathology Image Analysis Using Medical Twitter’ . Nature Medicine 29 (9): 2307–16. https://doi.org/10.1038/s41591-023-02504-3. 52 Ilharco, Gabriel, Mitchell Wortsman, Ross Wightman, Cade Gordon, Nicholas Carlini, Rohan Taori, Achal Dave, et al. 2021. ‘OpenCLIP’ . Zenodo. https://doi.org/10.5281/zenodo.5143773. Ilse,...

  3. [7]

    ‘Multistain Pretraining for Slide Representation Learning in Pathology’ . arXiv. http://arxiv.org/abs/2408.02859. Javed, Sajid, Arif Mahmood, Muhammad Moazam Fraz, Navid Alemi Koohbanani, Ksenija Benes, Yee-Wah Tsang, Katherine Hewitt, David Epstein, David Snead, and Nasir Rajpoot. 2020. ‘Cellular Community Detection for Tissue Phenotyping in Colorectal C...

  4. [8]

    ‘Towards Graph Foundation Models: A Survey and Beyond’ . arXiv. http://arxiv.org/abs/2310.11829. Liu, Pei, Luping Ji, Jiaxiang Gou, Bo Fu, and Mao Ye. 2024. ‘Interpretable Vision- Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology’ . arXiv. http://arxiv.org/abs/2409.09369. Liu, Ze, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Z...

  5. [10]

    Renal digital pathology visual knowledge search platform based on language large model and book knowledge

    ‘BioGPT: Generative Pre-Trained Transformer for Biomedical Text Generation and Mining’ . Briefings in Bioinformatics 23 (6): bbac409. https://doi.org/10.1093/bib/bbac409. Lv, Xiaomin, Chong Lai, Liya Ding, Maode Lai, and Qingrong Sun. 2024. ‘Renal Digital Pathology Visual Knowledge Search Platform Based on Language Large Model and Book Knowledge’ . arXiv. ...

  6. [13]

    ‘DABS: A Domain-Agnostic Benchmark for Self-Supervised Learning’ . arXiv. http://arxiv.org/abs/2111.12062. Tamkin, Alex, Mike Wu, and Noah Goodman. 2021. ‘Viewmaker Networks: Learning Views for Unsupervised Representation Learning’ . arXiv. http://arxiv.org/abs/2010.07432. Touvron, Hugo, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Ba...

  7. [97]

    Bommasani, Rishi, Drew A

    https://doi.org/10.1016/S2589-7500(23)00148-6. Bommasani, Rishi, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, et al. 2022. ‘On the Opportunities and Risks of Foundation Models’ . arXiv. http://arxiv.org/abs/2108.07258. 47 Breen, Jack, Katie Allen, Kieran Zucker, Lucy Godson, Nicolas M. Orsi, and Nishant Rav...

  8. [108]

    57 Nechaev, Dmitry, Alexey Pchelnikov, and Ekaterina Ivanova

    https://doi.org/10.1053/j.semdp.2023.02.006. 57 Nechaev, Dmitry, Alexey Pchelnikov, and Ekaterina Ivanova. 2024. ‘Hibou: A Family of Foundational Vision Transformers for Pathology’ . arXiv. http://arxiv.org/abs/2406.05074. Neidlinger, Peter, Omar S. M. El Nahhas, Hannah Sophie Muti, Tim Lenz, Michael Hoffmeister, Hermann Brenner, Marko van Treeck, et al. ...

Show all 14 references
  1. [2017]

    ‘On the Expressive Power of Deep Neural Networks’ . arXiv. http://arxiv.org/abs/1606.05336. Ray, Partha Pratim. 2023. ‘Benchmarking, Ethical Alignment, and Evaluation Framework for Conversational AI: Advancing Responsible Development of ChatGPT’ . BenchCouncil Transactions on ...

  2. [2019]

    Annals of Oncology 30 (8): 1232–43

    ‘ESMO Recommendations on Microsatellite Instability Testing for Immunotherapy in Cancer, and Its Relationship with PD-1/PD-L1 Expression and Tumour Mutational Burden: A Systematic Review-Based Approach’ . Annals of Oncology 30 (8): 1232–43. https://doi.org/10.1093/annonc/mdz11...

  3. [2021]

    ‘Image BERT Pre-Training with Online Tokenizer’ . In . https://openreview.net/forum?id=ydopy-e6Dg. Zhou, Qifeng, Wenliang Zhong, Yuzhi Guo, Michael Xiao, Hehuan Ma, and Junzhou Huang. 2024. ‘PathM3: A Multimodal Multi-Task Multiple Instance Learning Framework for Whole Slide I...

  4. [2022]

    British Journal of Cancer, October

    ‘Role of AI and Digital Pathology for Colorectal Immuno-Oncology’ . British Journal of Cancer, October. https://doi.org/10.1038/s41416-022-01986-1. Bilal, Mohsin, Shan E Ahmed Raza, Ayesha Azam, Simon Graham, Mohammad Ilyas, Ian A Cree, David Snead, Fayyaz Minhas, and Nasir M ...

  5. [2023]

    ‘Evaluating Large Language Models: A Comprehensive Survey’ . arXiv. http://arxiv.org/abs/2310.19736. Gustafsson, Fredrik K., and Mattias Rantalainen. 2024. ‘Evaluating Computational Pathology Foundation Models for Prostate Cancer Grading under Distribution Shifts’ . arXiv. htt...

  6. [2024]

    ‘Eureka: Evaluating and Understanding Large Foundation Models’ . arXiv. http://arxiv.org/abs/2409.10566. Bengio, Yoshua, Aaron Courville, and Pascal Vincent. 2014. ‘Representation Learning: A Review and New Perspectives’ . arXiv. http://arxiv.org/abs/1206.5538. Beyer, Lucas, P...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.