Pith. sign in

REVIEW 6 cited by

Current Pathology Foundation Models are unrobust to Medical Center Differences

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.18055 v2 pith:PXMZMNKW submitted 2025-01-29 cs.LG cs.AI

Current Pathology Foundation Models are unrobust to Medical Center Differences

classification cs.LG cs.AI
keywords medicalcenterpathologyfeaturesrobustnessbiologicaldifferencesindex
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Pathology Foundation Models (FMs) hold great promise for healthcare. Before they can be used in clinical practice, it is essential to ensure they are robust to variations between medical centers. We measure whether pathology FMs focus on biological features like tissue and cancer type, or on the well known confounding medical center signatures introduced by staining procedure and other differences. We introduce the Robustness Index. This novel robustness metric reflects to what degree biological features dominate confounding features. Ten current publicly available pathology FMs are evaluated. We find that all current pathology foundation models evaluated represent the medical center to a strong degree. Significant differences in the robustness index are observed. Only one model so far has a robustness index greater than one, meaning biological features dominate confounding features, but only slightly. A quantitative approach to measure the influence of medical center differences on FM-based prediction performance is described. We analyze the impact of unrobustness on classification performance of downstream models, and find that cancer-type classification errors are not random, but specifically attributable to same-center confounders: images of other classes from the same medical center. We visualize FM embedding spaces, and find these are more strongly organized by medical centers than by biological factors. As a consequence, the medical center of origin is predicted more accurately than the tissue source and cancer type. The robustness index introduced here is provided with the aim of advancing progress towards clinical adoption of robust and reliable pathology FMs.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Good, the Bad, and the Brittle: Benchmarking Robustness and Generalisation of Histopathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    Mid-sized pathology foundation models match or beat billion-parameter ones on clinically realistic perturbations and distribution-shift tests, so scaling alone has largely saturated for robustness.

  2. Semantic-Anchored Evidential Fusion for Domain-Robust Whole-Slide Survival Analysis

    cs.CV 2026-06 unverdicted novelty 6.0

    SAEFS uses VQA-derived semantic anchors, dual-stream evidence extraction, and Dirichlet-based evidential fusion to achieve 10.2% higher average C-index in zero-shot cross-domain WSI survival analysis.

  3. When Are Multimodal Predictions Biologically Supported? A Diagnostic Evaluation Framework

    cs.LG 2026-05 unverdicted novelty 6.0

    DECAT classifies multimodal representations into four diagnostic scenarios using null-referenced metrics and a rule-based procedure to detect shared biology versus confounders without knowing the confounder identity.

  4. Enabling clinical use of foundation models for computational pathology

    cs.CV 2026-02 conditional novelty 6.0

    Novel robustness losses added during downstream training on foundation-model features from pathology slides improve both robustness to technical variation and classification accuracy.

  5. Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    GLMP generates robust pathology embeddings by routing histology images through an intermediate textual representation produced by general-purpose MLLMs to mitigate batch effects.

  6. Beyond the Failures: Rethinking Foundation Models in Pathology

    cs.AI 2025-10 unverdicted novelty 2.0

    Foundation models stumble in pathology due to conceptual mismatches with biological tissue, requiring explicitly designed models rather than adaptations of natural-image methods.