Pith. sign in

REVIEW 14 cited by

Phikon-v2, A large and public feature extractor for biomarker prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.09173 v1 pith:PSSPTBC7 submitted 2024-09-13 eess.IV cs.AIcs.CV

Phikon-v2, A large and public feature extractor for biomarker prediction

classification eess.IV cs.AIcs.CV
keywords modelphikon-v2datahistologymodelspredictionpubliclytasks
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Gathering histopathology slides from over 100 publicly available cohorts, we compile a diverse dataset of 460 million pathology tiles covering more than 30 cancer sites. Using this dataset, we train a large self-supervised vision transformer using DINOv2 and publicly release one iteration of this model for further experimentation, coined Phikon-v2. While trained on publicly available histology slides, Phikon-v2 surpasses our previously released model (Phikon) and performs on par with other histopathology foundation models (FM) trained on proprietary data. Our benchmarks include eight slide-level tasks with results reported on external validation cohorts avoiding any data contamination between pre-training and evaluation datasets. Our downstream training procedure follows a simple yet robust ensembling strategy yielding a +1.75 AUC increase across tasks and models compared to one-shot retraining (p<0.001). We compare Phikon (ViT-B) and Phikon-v2 (ViT-L) against 14 different histology feature extractors, making our evaluation the most comprehensive to date. Our result support evidences that DINOv2 handles joint model and data scaling better than iBOT. Also, we show that recent scaling efforts are overall beneficial to downstream performance in the context of biomarker prediction with GigaPath and H-Optimus-0 (two ViT-g with 1.1B parameters each) standing out. However, the statistical margins between the latest top-performing FMs remain mostly non-significant; some even underperform on specific indications or tasks such as MSI prediction - deposed by a 13x smaller model developed internally. While latest foundation models may exhibit limitations for clinical deployment, they nonetheless offer excellent grounds for the development of more specialized and cost-efficient histology encoders fueling AI-guided diagnostic tools.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Benchmarking Pathology Foundation Models for Spatial Domain Understanding

    cs.CV 2026-05 unverdicted novelty 7.0

    SpaPath-Bench evaluates spatial representation in 19 pathology foundation models via spatial domain identification on 42 paired WSI-ST slides using three agreement criteria across 83K runs.

  2. Topology-Driven Transferability Estimation for 3D Medical Vision Foundation Models

    cs.CV 2026-07 conditional novelty 6.5

    MST-based local boundary leakage and global topology divergence, fused by task complexity, rank SSL 3D medical encoders for segmentation without fine-tuning, beating prior TE metrics by 0.36 weighted Kendall τ at 56× speed.

  3. LaGuadia: Language-Guided Adaptive Distillation from Pathology Foundation Models

    cs.CV 2026-07 conditional novelty 6.0

    Language-guided adaptive multi-teacher distillation yields an 87M pathology encoder that matches or exceeds GigaPath and UNI on WSI captioning, VQA, and MIL tasks.

  4. Towards Autonomous and Auditable Medical Imaging Model Development

    cs.CV 2026-07 conditional novelty 6.0

    AMID, a verification-guided multi-agent MLE system for medical imaging, outperforms general MLE agents on 20 ReX-MLE challenges and approaches human challenge solutions on several tasks.

  5. DaX: Learning General Pathology Representations Across Scales

    eess.IV 2026-06 unverdicted novelty 6.0

    DaX is a pathology vision foundation model that extends DINOv3 with continuous magnification training and cross-scale consistency, achieving top average performance on a benchmark of 161 tasks from 44 datasets coverin...

  6. MOOZY: A Patient-First Foundation Model for Computational Pathology

    cs.CV 2026-03 conditional novelty 6.0

    Patient-level pretraining with a case transformer and multi-task public supervision yields transferable WSI embeddings that beat larger slide-centric models on held-out pathology tasks.

  7. Enabling clinical use of foundation models for computational pathology

    cs.CV 2026-02 conditional novelty 6.0

    Novel robustness losses added during downstream training on foundation-model features from pathology slides improve both robustness to technical variation and classification accuracy.

  8. APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms

    cs.CV 2026-06 unverdicted novelty 5.0

    APRIL-MedSeg is a new open-source modular toolbox that uses YAML configuration and component registries to unify multiple advanced paradigms for medical image segmentation.

  9. Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    GLMP generates robust pathology embeddings by routing histology images through an intermediate textual representation produced by general-purpose MLLMs to mitigate batch effects.

  10. Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction

    cs.CV 2026-05 unverdicted novelty 5.0

    A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.

  11. Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

    cs.CV 2026-05 unverdicted novelty 5.0

    FedHD is a federated learning framework for whole slide images that distills one-to-one synthetic features aligned via Gaussian mixtures and progressively integrates cross-site features through curriculum learning to ...

  12. Federated Distillation for Whole Slide Image via Gaussian-Mixture Feature Alignment and Curriculum Integration

    cs.CV 2026-05 unverdicted novelty 5.0

    FedHD performs federated distillation for whole slide images by generating one synthetic feature set per real slide via Gaussian-mixture alignment and adding them via curriculum integration, outperforming prior federa...

  13. CellPrior-Net: Prior-Guided Nuclei Detection and Classification for H&E Whole-Slide Images

    cs.MM 2026-07 unverdicted novelty 4.0

    CellPrior-Net integrates hematoxylin channel prior into a lightweight CNN for nuclei detection and classification in H&E WSIs, claiming comparable accuracy to SOTA with significantly reduced inference time across 10.4...

  14. APRIL-MedSeg: A Modular Medical Image Segmentation Toolbox Embracing Modern Paradigms

    cs.CV 2026-06 unverdicted novelty 4.0

    Presents APRIL-MedSeg, a modular YAML-configurable toolbox for 2D medical image segmentation integrating semi-supervised, domain adaptation, distillation, weakly supervised, text-guided, and foundation model paradigms...