REVIEW 4 major objections 5 minor 144 references
Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Aligning histology images with text at several magnifications—not a single resolution—makes pathology vision-language models generalize better, and the paper supports this with 34 million multi-resolution pairs.
desk verdict A broad multi-resolution pathology VLM with plausible gains, but the central claim lacks a single-resolution control and the SOTA claim is overbroad. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a hierarchy of four co-registered magnification levels organized as parent–child bags: one $5\times$ patch, four $10\times$ children, sixteen $20\times$ grandchildren, and sixty-four $40\times$ great-grandchildren, with $512\times512$ patches and a 70% tissue-coverage filter. On top of the bags sit two losses. CVTA (Eq. 1) treats the textual bag as a set of candidate keywords, selects the top $k_o=9$ words by cosine similarity to the UNI visual feature of each patch, and runs a contrastive loss with those as positives and the rest of the bag as negatives. MRTVA (Eqs. 3–4), built on the SimSiam formulation with projection and prediction heads and a stop-gradient, aligns the text-guided visual features produced by the multimodal encoder for each parent–child pair; alignments between grandparents and grandchildren are deliberately omitted because those patches share too little tissue content. The multimodal encoder is adapted from the mPLUG/ALBEF design, initialized from the latter six layers of QuiltNet's GPT-2/77 text encoder, and trained with ITC, ITM, MLM, and PLM as the baseline objective.
What would settle it
Train MR-PLIP identically but replace the Quilt-LLaVA captions in the text bags with pathologist-verified captions, and compare zero-shot weighted F1 on NCT-CRC and CAM16; the multi-resolution claim is falsified if the gap over single-resolution QuiltNet disappears or reverses. A second check is to test on held-out magnifications such as $15\times$ or $30\times$, which the training never saw—a model with genuinely resolution-invariant text-guided features should still classify accurately.
Extended reading notes
Core claim
On its own terms, the paper's central claim is that multi-resolution is the missing ingredient in pathology vision-language pre-training. The authors first show that existing models such as QuiltNet perform unevenly across magnifications—$20\times$ and $10\times$ are strong, $5\times$ and $40\times$ are weak—and that captions produced by Quilt-LLaVA change content as magnification changes. They then construct a visual bag in which each $5\times$ patch has children at $10\times$, $20\times$, and $40\times$, attach Quilt-LLaVA-generated descriptions to every patch, and train CVTA to pull each patch's UNI visual features toward its top nine keywords by cosine similarity while pushing away the rest of the text bag. A multimodal encoder fuses patch features with those keywords, and the MRTVA loss, a SimSiam-style symmetric objective with stop-gradient, aligns parent and child text-guided representations. The reported result is consistent gains in weighted F1 and balanced accuracy over seven prior VLMs and several vision-only foundation models across tile classification, WSI classification, segmentation, and retrieval.
Load-bearing premise
The load-bearing premise is that the Quilt-LLaVA-generated captions and the top nine keywords selected by cosine similarity between UNI visual features and QuiltNet word embeddings are semantically correct enough to serve as training targets, since the paper never validates them against pathologist-written annotations.
Editorial extensions
If this is right
- Pre-training on multiple resolutions can be added to existing pathology vision-language pipelines by re-running contrastive pre-training on multi-resolution patch bags, without changing the downstream task heads.
- Zero-shot classification becomes more reliable across magnifications, so the model does not need to know in advance whether a test patch came from a $5\times$, $10\times$, $20\times$, or $40\times$ scan.
- Selecting a small set of positive keywords per patch outperforms full-text alignment because it filters hallucinated or irrelevant words from auto-generated captions.
- Aligning text-guided visual features only between direct parent–child pairs is important; the ablation without the hierarchy shows that aligning all cross-resolution pairs hurts performance.
- Text-guided visual features from the multimodal encoder beat unimodal features in zero-shot classification, indicating that the selected keywords add discriminative signal rather than acting only as a regularizer.
Reading between the lines
- Because the keyword-selection step is a generic filter for noisy auto-generated captions, the same CVTA mechanism could plausibly improve cross-modal retrieval in other medical imaging domains, but the authors do not test that transfer.
- The paper does not separate the effect of multi-resolution alignment from the effect of having four times as many training pairs; an ablation that trains on 34 million single-resolution pairs would disentangle the two.
- If parent–child alignment is what makes the features scale-invariant, a testable prediction is that MR-PLIP should transfer to magnifications it never saw, such as $15\times$ or $30\times$, and that prediction is not currently in the paper.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MR-PLIP, a pathology vision-language model pre-trained on multi-resolution TCGA patches with captions generated by Quilt-LLaVA. It introduces two losses: CVTA, which selects top-k positive keywords per visual feature by cosine similarity and applies contrastive alignment, and MRTAV, which aligns text-guided visual features of parent and child patches across magnifications via a SimSiam-style loss. The model is initialized from UNI and QuiltNet encoders and trained with ITC/ITM/MLM/PLM objectives on 34 million image-text pairs at 5x, 10x, 20x, and 40x. The authors report zero-shot, linear-probe, weakly supervised, segmentation, and retrieval results on 26 datasets, claiming consistent improvements over prior VLMs and vision-only foundation models.
Significance. If the central claim holds, the paper would be a useful step toward multi-resolution vision-language pre-training for computational pathology, and the release of code would facilitate reproducibility. The evaluation breadth (26 datasets, multiple protocols, ablations over losses, captioning models, encoders, and keyword counts) is a genuine strength. However, the load-bearing attribution of the gains to multi-resolution alignment is currently not directly tested, because no single-resolution MR-PLIP control exists, and several hyperparameters appear to have been selected on the same datasets used for the headline zero-shot comparisons. The paper also shows point estimates without uncertainty quantification, which matters for a claim of 'significant margins.' The result is plausible and potentially valuable, but the experimental evidence needs additional controls and corrected claims before it can support the stated conclusions.
major comments (4)
- [Section 4.1, Supplementary Tables 5-6] The number of positive keywords (k0=9) and the choice of the magnification set {5x,10x,20x,40x} are justified using the same six datasets (CAM16, CPTAC, SICAP, DigestPath, Databiox, NCT-CRC) that appear in the headline zero-shot comparison (Table 1). Table 6 selects k0 by reading performance on these datasets, and the four-magnification set is motivated by fine-tuning experiments on the same seven benchmark datasets used for comparison in Fig. 3 and the supplementary. This is selection on the test data, and it can inflate the reported gains. Please either choose hyperparameters on a separate validation split or report results on held-out datasets that were not used for any model or hyperparameter selection.
- [Section 3.1-3.3, Table 5 and Table 3] The central claim that multi-resolution pre-training causes the reported improvements is not directly tested. The supplemental ablation in Table 5 compares pairs and triples of resolutions but contains no single-resolution row, and the Lbl-only row in Table 3 is still trained on multi-resolution data. A controlled comparison would be a single-resolution MR-PLIP variant trained at the best single magnification (e.g., 20x or 10x) with the same Lbl + LCVTA losses and the same zero-shot protocol; MRTAV would not apply at a single resolution. Without this control, the gains over PLIP, QuiltNet, CONCH, MI-Zero, and BioCLIP could equally be attributed to the UNI+QuiltNet initialization, the multi-modal encoder, or the larger effective training set rather than to cross-resolution alignment.
- [Abstract, Table 2] The abstract's claim that the fine-tuned model 'outperforms state-of-the-art counterparts across multiple datasets and tasks' is stronger than the data support. In Table 2, UNI achieves higher balanced accuracy and F1 on WILDS-CAM17 (0.983 vs. 0.975 and 0.980), and GigaPath achieves higher balanced accuracy on CAM16 (0.967 vs. 0.950). The claims should be qualified as 'most datasets' or 'on the majority of evaluated benchmarks,' and the CAM16 result should be discussed, since it is a standard WSI benchmark.
- [Section 3.2, Eq. (1), and Section 4.4] The training signal is partly self-bootstrapped: positive keywords for the CVTA loss are selected by cosine similarity between UNI visual features and QuiltNet text embeddings, and the trained model is initialized from the same model families. The paper does not validate the generated captions or selected keywords against pathologist annotations or any external ground truth, so it is unknown whether CVTA and MRTAV add multi-resolution knowledge or mainly reinforce existing biases of UNI and QuiltNet. Please provide a quantitative assessment of caption and keyword quality, for example a manual subset evaluation, agreement with tissue-type labels, or an analysis of which keywords are selected at each resolution.
minor comments (5)
- [Throughout, Section 3.3] The abbreviation for the second loss is inconsistent: the text uses 'MRTAV' in Section 3.3, while equations and tables use 'MRTVA'; please unify the notation.
- [Section 3, Figure references] Section 3 refers to the algorithm schematic as 'Fig. 6', but the workflow figure is earlier (Fig. 4); Section 3.3 also references 'Fig. 6 (h)' and '(i)' for the multimodal encoder. These cross-references appear to be off by one and should be corrected.
- [Section 4.4] The text says 'Table 13 displays the results of zero-shot tile-level classification,' but the zero-shot results are in Table 1 of the main text; the table numbering between the main text and the supplementary appendix is confusing and should be fixed.
- [Tables 1-3 and supplementary tables] All performance numbers are reported as point estimates without standard deviations or significance tests. Given that some differences are small (e.g., Table 3: 0.546 vs. 0.527 for the added losses on SICAP), error bars or confidence intervals over repeated runs would materially strengthen the claimed margins.
- [Abstract and Section 3.1] The phrase '34 million image-language pairs' is used interchangeably with '34 million patches' and '34 million image-text pairs'; since each patch receives a generated caption, the terminology should be made consistent, and the distinction between patches and pairs should be clarified.
Circularity Check
CVTA's positive keywords are selected by the same cosine similarity that Eq. (1) maximizes, making the text-guided alignment partly self-referential; final external benchmarks keep the central comparison honest.
-
self definitional
[Sec. 3.2, Eq. (1) (CVTA positive-keyword selection and loss)]
"To achieve this, for each visual feature represented by va, we identify ko < k positive keywords w+_b from the corresponding Ti,j. This identification is based on maximizing the cosine similarity between the word feature wb ∈ Ti,j and the visual feature va ∈ Vi,j, calculated as the dot product w⊤_b va / ||wb||2 ||va||2. ... We then apply a contrastive loss function that is minimized by training with these identified positive and negative keyword pairs."
The 'positive' keywords are not externally anchored: they are defined as the top-ko words under exactly the cosine similarity score that L_CVTA then maximizes. Eq. (1) pushes v_a toward w+_b and away from the rest, but w+_b was chosen because it already has the highest v·w score. The objective is therefore self-confirmatory: any top-scoring words become 'positive' by construction, and the loss can only sharpen the pre-existing UNI/QuiltNet alignment rather than independently establish the claimed 'accurate' visual-textual alignment. This makes the contribution of L_CVTA to the reported gains partly a self-distillation/bootstrap effect. The external benchmark evaluation is independent, so this is partial circularity in a pre-training component, not in the final test comparison.
full rationale
The paper's headline comparison is against external benchmarks (Table 1 and Figs. 3, 6) with held-out datasets and standard zero-shot/linear-probe protocols, so the final claim of improved downstream performance is not circular. The one genuinely self-referential construction is the CVTA loss (Eq. 1): the positive keywords are the model's own top-similarity words, so optimizing this loss only reinforces the encoders' existing ranking rather than adding a measured external text signal. The MRTAV loss is a SimSiam-style consistency objective, which is self-supervised but not circular in the target-label sense. The captions from Quilt-LLaVA and the ITC/ITM/MLM/PLM losses do provide external text input, though their semantic quality is unvalidated. The paper's ablations lack a single-resolution MR-PLIP control, so the causal attribution to multi-resolution padding is underdetermined; this is an experimental-control gap rather than a circularity. CPLIP [53] is a self-cited baseline but is not used to justify the method, so no load-bearing self-citation. Overall score 3: one pre-training objective reduces by construction to the model's own similarity, but the central empirical evaluation is independent and the main losses are not circular.
Assumptions & free parameters
free parameters (3)
- ko, number of positive keywords =
9
- Magnification set =
5x, 10x, 20x, 40x
- Initial contrastive temperature tau =
0.07 (learnable)
assumptions (4)
- domain assumption Quilt-LLaVA descriptions are accurate and resolution-specific enough to serve as text supervision.
- domain assumption Cosine similarity between UNI visual features and QuiltNet word embeddings selects semantically correct positive keywords.
- domain assumption TCGA pre-training transfers to the 26 evaluation datasets despite distribution shifts.
- domain assumption The parent-child hierarchy is a meaningful inductive bias for cross-resolution alignment.
Cite this review
Pith. "Pith review of Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation." pith.science (2026). https://pith.science/paper/A5BNCPGJ
@misc{pith2026250418856,
author = {Pith},
title = {Pith review of: Multi-Resolution Pathology-Language Pre-training Model with Text-Guided Visual Representation},
year = {2026},
howpublished = {\url{https://pith.science/paper/A5BNCPGJ}},
note = {Machine review of arXiv:2504.18856}
}
read the original abstract
In Computational Pathology (CPath), the introduction of Vision-Language Models (VLMs) has opened new avenues for research, focusing primarily on aligning image-text pairs at a single magnification level. However, this approach might not be sufficient for tasks like cancer subtype classification, tissue phenotyping, and survival analysis due to the limited level of detail that a single-resolution image can provide. Addressing this, we propose a novel multi-resolution paradigm leveraging Whole Slide Images (WSIs) to extract histology patches at multiple resolutions and generate corresponding textual descriptions through advanced CPath VLM. We introduce visual-textual alignment at multiple resolutions as well as cross-resolution alignment to establish more effective text-guided visual representations. Cross-resolution alignment using a multimodal encoder enhances the model's ability to capture context from multiple resolutions in histology images. Our model aims to capture a broader range of information, supported by novel loss functions, enriches feature representation, improves discriminative ability, and enhances generalization across different resolutions. Pre-trained on a comprehensive TCGA dataset with 34 million image-language pairs at various resolutions, our fine-tuned model outperforms state-of-the-art (SOTA) counterparts across multiple datasets and tasks, demonstrating its effectiveness in CPath. The code is available on GitHub at: https://github.com/BasitAlawode/MR-PLIP
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Computational pathology definitions, best practices, and recommendations for regulatory guidance: a white paper from the digital pathology association
Esther Abels, Liron Pantanowitz, Famke Aeffner, Mark D Zarella, Jeroen van der Laak, Marilyn M Bui, Venkata NP Vemuri, Anil V Parwani, Jeff Gibbs, Emmanuel Agosto- Arroyo, et al. Computational pathology definitions, best practices, and recommendations for regulatory guidance: a white paper from the digital pathology association. The Journal of pathology, ...
2019
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. 23
arXiv 2023
-
[3]
Multimodal biomedical ai
Juli ´an N Acosta, Guido J Falcone, Pranav Rajpurkar, and Eric J Topol. Multimodal biomedical ai. Nature Medicine, 28(9):1773–1784, 2022. 1
2022
-
[4]
A method for stochastic opti- mization
Kingma DP Ba J Adam et al. A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980, 1412, 2014. 6, 19
arXiv 2014
-
[5]
Social media: pathologists’ force multiplier
Timothy Craig Allen. Social media: pathologists’ force multiplier. Archives of Pathology and Laboratory Medicine, 138(8):1000–1001, 2014. 22
2014
-
[6]
Publicly available clinical bert embeddings
Emily Alsentzer, John R Murphy, Willie Boag, Wei-Hung Weng, Di Jin, Tristan Naumann, and Matthew McDermott. Publicly available clinical bert embeddings. arXiv preprint arXiv:1904.03323, 2019. 19
arXiv 1904
-
[7]
Bach: Grand challenge on breast cancer his- tology images
Guilherme Aresta, Teresa Ara ´ujo, Scotty Kwok, Sai Saketh Chennamsetty, Mohammed Safwan, Varghese Alex, Bahram Marami, Marcel Prastawa, Monica Chan, Michael Donovan, et al. Bach: Grand challenge on breast cancer his- tology images. Medical image analysis, 56:122–139, 2019. 2, 16
2019
-
[8]
Viable and necrotic tumor assessment from whole slide im- ages of osteosarcoma using machine-learning and deep- learning models
Harish Babu Arunachalam, Rashika Mishra, Ovidiu Daescu, Kevin Cederberg, Dinesh Rakheja, Anita Sen- gupta, David Leonard, Rami Hallac, and Patrick Leavey. Viable and necrotic tumor assessment from whole slide im- ages of osteosarcoma using machine-learning and deep- learning models. PloS one, 14(4):e0210706, 2019. 6, 7, 20
2019
Show all 144 references
-
[9]
Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging
Shekoofeh Azizi, Laura Culp, Jan Freyberg, Basil Mustafa, Sebastien Baur, Simon Kornblith, Ting Chen, Nenad Toma- sev, Jovana Mitrovi ´c, Patricia Strachan, et al. Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging. Nature Biome...
2023
-
[10]
From detection of individ- ual metastases to classification of lymph node status at the patient level: the camelyon17 challenge
Peter Bandi, Oscar Geessink, Quirine Manson, Mar- cory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, et al. From detection of individ- ual metastases to classification of lymph node status at the patient le...
2018
-
[11]
Unitopatho, a labeled histopathological dataset for colorectal polyps classification and adenoma dysplasia grading
Carlo Alberto Barbano, Daniele Perlo, Enzo Tartaglione, Attilio Fiandrotti, Luca Bertero, Paola Cassoni, and Marco Grangetto. Unitopatho, a labeled histopathological dataset for colorectal polyps classification and adenoma dysplasia grading. In 2021 IEEE International Conferen...
2021
-
[12]
A multi-scale superpixel classification approach to the detec- tion of regions of interest in whole slide histopathology im- ages
Babak Ehteshami Bejnordi, Geert Litjens, Meyke Hermsen, Nico Karssemeijer, and Jeroen AWM van der Laak. A multi-scale superpixel classification approach to the detec- tion of regions of interest in whole slide histopathology im- ages. In Medical Imaging 2015: Digital Pathology...
2015
-
[13]
Diagnos- tic assessment of deep learning algorithms for detection of lymph node metastases in women with breast cancer.Jama, 318(22):2199–2210, 2017
Babak Ehteshami Bejnordi, Mitko Veta, Paul Johannes Van Diest, Bram Van Ginneken, Nico Karssemeijer, Geert Litjens, Jeroen AWM Van Der Laak, Meyke Hermsen, Quirine F Manson, Maschenka Balkenhol, et al. Diagnos- tic assessment of deep learning algorithms for detection of lymph ...
2017
-
[14]
Computational pathology in 2030: a delphi study forecasting the role of ai in pathology within the next decade
M Alvaro Berb ´ıs, David S McClintock, Andrey By- chkov, Jeroen Van der Laak, Liron Pantanowitz, Jochen K Lennerz, Jerome Y Cheng, Brett Delahunt, Lars Egevad, Catarina Eloy, et al. Computational pathology in 2030: a delphi study forecasting the role of ai in pathology within ...
2023
-
[15]
Palm: Pre- training an autoencoding&autoregressive language model for context-conditioned generation
Bin Bi, Chenliang Li, Chen Wu, Ming Yan, Wei Wang, Songfang Huang, Fei Huang, and Luo Si. Palm: Pre- training an autoencoding&autoregressive language model for context-conditioned generation. arXiv preprint arXiv:2004.07159, 2020. 17
2004 arXiv
-
[16]
A histopatho- logical image dataset for grading breast invasive ductal carcinomas
Hamidreza Bolhasani, Elham Amjadi, Maryam Tabatabaeian, and Somayyeh Jafarali Jassbi. A histopatho- logical image dataset for grading breast invasive ductal carcinomas. Informatics in Medicine Unlocked, 19:100341,
-
[17]
Lung and colon cancer histopathological im- age dataset (lc25000)
Andrew A Borkowski, Marilyn M Bui, L Brannon Thomas, Catherine P Wilson, Lauren A DeLand, and Stephen M Mastorides. Lung and colon cancer histopathological im- age dataset (lc25000). arXiv preprint arXiv:1912.12142 ,
1912 arXiv
-
[18]
Bracs: A dataset for breast carcinoma subtyping in h&e histology images
Nadia Brancati, Anna Maria Anniciello, Pushpak Pati, Daniel Riccio, Giosu `e Scognamiglio, Guillaume Jaume, Giuseppe De Pietro, Maurizio Di Bonito, Antonio Foncu- bierta, Gerardo Botti, et al. Bracs: A dataset for breast carcinoma subtyping in h&e histology images. Database, 2...
2022
-
[19]
Integrative analysis of histological textures and lymphocyte infiltration in renal cell carcinoma using deep learning
Otso Brummer, Petri P ¨ol¨onen, Satu Mustjoki, and Oscar Br¨uck. Integrative analysis of histological textures and lymphocyte infiltration in renal cell carcinoma using deep learning. bioRxiv, pages 2022–08, 2022. 6, 7, 20
2022
-
[20]
Artificial intelligence for diagnosis and gleason grading of prostate cancer: the panda challenge
Wouter Bulten, Kimmo Kartasalo, Po-Hsuan Cameron Chen, Peter Str ¨om, Hans Pinckaers, Kunal Nagpal, Yuan- nan Cai, David F Steiner, Hester Van Boven, Robert Vink, et al. Artificial intelligence for diagnosis and gleason grading of prostate cancer: the panda challenge. Nature m...
2022
-
[21]
Clinical-grade computational pathology using weakly supervised deep learning on whole slide im- ages
Gabriele Campanella, Matthew G Hanna, Luke Genes- law, Allen Miraflor, Vitor Werneck Krauss Silva, Klaus J Busam, Edi Brogi, Victor E Reuter, David S Klimstra, and Thomas J Fuchs. Clinical-grade computational pathology using weakly supervised deep learning on whole slide im- a...
2019
-
[22]
Emerging properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 9650–9660, 2021. 3
2021
-
[23]
A new era in computational pathology: A survey on foundation and vision-language models
Dibaloke Chanda, Milan Aryal, Nasim Yahya Soltani, and Masoud Ganji. A new era in computational pathology: A survey on foundation and vision-language models. arXiv preprint arXiv:2408.14496, 2024. 3
2024 arXiv
-
[24]
Fast and scalable search of whole-slide images via self-supervised deep learning
Chengkuan Chen, Ming Y Lu, Drew FK Williamson, Tiffany Y Chen, Andrew J Schaumberg, and Faisal Mah- mood. Fast and scalable search of whole-slide images via self-supervised deep learning. Nature Biomedical Engi- neering, 6(12):1420–1434, 2022. 2, 15
2022
-
[25]
Scaling vision transformers to gigapixel images via hierarchical self-supervised learning
Richard J Chen, Chengkuan Chen, Yicong Li, Tiffany Y Chen, Andrew D Trister, Rahul G Krishnan, and Faisal Mahmood. Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Patter...
2022
-
[26]
Towards a general-purpose foundation model for computational pathology.Nature Medicine, 30(3):850–862,
Richard J Chen, Tong Ding, Ming Y Lu, Drew FK Williamson, Guillaume Jaume, Andrew H Song, Bowen Chen, Andrew Zhang, Daniel Shao, Muhammad Shaban, et al. Towards a general-purpose foundation model for computational pathology.Nature Medicine, 30(3):850–862,
-
[27]
A simple framework for contrastive learning of visual representations
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Ge- offrey Hinton. A simple framework for contrastive learning of visual representations. In International conference on machine learning, pages 1597–1607. PMLR, 2020. 6
2020
-
[28]
Exploring simple siamese representation learning
Xinlei Chen and Kaiming He. Exploring simple siamese representation learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 15750–15758, 2021. 5, 6
2021
-
[29]
An empirical study of train- ing self-supervised vision transformers
X Chen, S Xie, and K He. An empirical study of train- ing self-supervised vision transformers. in 2021 ieee. In CVF International Conference on Computer Vision (ICCV), pages 9620–9629. 6, 18
2021
-
[30]
Ai in computational pathol- ogy of cancer: Improving diagnostic workflows and clini- cal outcomes? Annual Review of Cancer Biology, 7:57–71,
Didem Cifci, Gregory P Veldhuizen, Sebastian Foersch, and Jakob Nikolas Kather. Ai in computational pathol- ogy of cancer: Improving diagnostic workflows and clini- cal outcomes? Annual Review of Cancer Biology, 7:57–71,
-
[31]
Whole-slide imaging: routine pathologic diagnosis
Toby C Cornish, Ryan E Swapp, and Keith J Kaplan. Whole-slide imaging: routine pathologic diagnosis. Ad- vances in anatomic pathology, 19(3):152–159, 2012. 1, 2, 15
2012
-
[32]
Accurate and reproducible invasive breast cancer detection in whole-slide images: A deep learning approach for quan- tifying tumor extent
Angel Cruz-Roa, Hannah Gilmore, Ajay Basavanhally, Michael Feldman, Shridar Ganesan, Natalie NC Shih, John Tomaszewski, Fabio A Gonz ´alez, and Anant Madabhushi. Accurate and reproducible invasive breast cancer detection in whole-slide images: A deep learning approach for quan...
2017
-
[33]
Artificial intelligence and computational pathology
Miao Cui and David Y Zhang. Artificial intelligence and computational pathology. Laboratory Investigation, 101 (4):412–422, 2021. 1
2021
-
[34]
Digestpath: A benchmark dataset with challenge review for the pathological detection and seg- mentation of digestive-system
Qian Da, Xiaodi Huang, Zhongyu Li, Yanfei Zuo, Chen- bin Zhang, Jingxin Liu, Wen Chen, Jiahui Li, Dou Xu, Zhiqiang Hu, et al. Digestpath: A benchmark dataset with challenge review for the pathological detection and seg- mentation of digestive-system. Medical Image Analysis , 8...
2022
-
[35]
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understanding. arXiv preprint arXiv:1810.04805, 2018. 17, 18
2018 arXiv
-
[36]
Multi-scale fully convolutional network for gland segmentation using three-class classification
Huijun Ding, Zhanpeng Pan, Qian Cen, Yang Li, and Shifeng Chen. Multi-scale fully convolutional network for gland segmentation using three-class classification. Neuro- computing, 380:150–161, 2020. 2, 3, 15
2020
-
[37]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv...
2010 arXiv
-
[38]
Deep learning in cancer pathology: a new gen- eration of clinical biomarkers
Amelie Echle, Niklas Timon Rindtorff, Titus Josef Brinker, Tom Luedde, Alexander Thomas Pearson, and Jakob Niko- las Kather. Deep learning in cancer pathology: a new gen- eration of clinical biomarkers. British journal of cancer , 124(4):686–696, 2021. 1
2021
-
[39]
The cptac data portal: a re- source for cancer proteomics research
Nathan J Edwards, Mauricio Oberti, Ratna R Thangudu, Shuang Cai, Peter B McGarvey, Shine Jacob, Subha Mad- havan, and Karen A Ketchum. The cptac data portal: a re- source for cancer proteomics research. Journal of proteome research, 14(6):2707–2713, 2015. 6, 7, 21
2015
-
[40]
Multiple instance cap- tioning: Learning representations from histopathology text- books and articles
Jevgenij Gamper and Nasir Rajpoot. Multiple instance cap- tioning: Learning representations from histopathology text- books and articles. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 16549–16559, 2021. 22
2021
-
[41]
Pannuke: an open pan- cancer histology dataset for nuclei instance segmentation and classification
Jevgenij Gamper, Navid Alemi Koohbanani, Ksenija Benet, Ali Khuram, and Nasir Rajpoot. Pannuke: an open pan- cancer histology dataset for nuclei instance segmentation and classification. In Digital Pathology: 15th European Congress, ECDP 2019, Warwick, UK, April 10–13, 2019, P...
2019
-
[42]
Hover-net: Simultaneous segmentation and classi- fication of nuclei in multi-tissue histology images
Simon Graham, Quoc Dang Vu, Shan E Ahmed Raza, Ayesha Azam, Yee Wah Tsang, Jin Tae Kwak, and Nasir Rajpoot. Hover-net: Simultaneous segmentation and classi- fication of nuclei in multi-tissue histology images. Medical image analysis, 58:101563, 2019. 22
2019
-
[43]
Bootstrap your own latent-a new approach to self-supervised learning
Jean-Bastien Grill, Florian Strub, Florent Altch ´e, Corentin Tallec, Pierre Richemond, Elena Buchatskaya, Carl Do- ersch, Bernardo Avila Pires, Zhaohan Guo, Mohammad Gheshlaghi Azar, et al. Bootstrap your own latent-a new approach to self-supervised learning. Advances in neur...
2020
-
[44]
Domain-specific language model pre- training for biomedical natural language processing
Yu Gu, Robert Tinn, Hao Cheng, Michael Lucas, Naoto Usuyama, Xiaodong Liu, Tristan Naumann, Jianfeng Gao, and Hoifung Poon. Domain-specific language model pre- training for biomedical natural language processing. ACM Transactions on Computing for Healthcare (HEALTH) , 3 (1):1–...
2021
-
[45]
Wsss4luad: Grand challenge on weakly- supervised tissue semantic segmentation for lung adenocar- cinoma
Chu Han, Xipeng Pan, Lixu Yan, Huan Lin, Bingbing Li, Su Yao, Shanshan Lv, Zhenwei Shi, Jinhai Mai, Ji- atai Lin, et al. Wsss4luad: Grand challenge on weakly- supervised tissue semantic segmentation for lung adenocar- cinoma. arXiv preprint arXiv:2204.06455, 2022. 6, 7, 20
2022 arXiv
-
[46]
Whole slide imaging: technology and appli- cations
Matthew G Hanna, Anil Parwani, and Sahussapont Joseph Sirintrapun. Whole slide imaging: technology and appli- cations. Advances in Anatomic Pathology, 27(4):251–259,
-
[47]
Multi-scale domain-adversarial multiple-instance cnn for cancer subtype classification with unannotated histopatho- logical images
Noriaki Hashimoto, Daisuke Fukushima, Ryoichi Koga, Yusuke Takagi, Kaho Ko, Kei Kohno, Masato Nakaguro, Shigeo Nakamura, Hidekata Hontani, and Ichiro Takeuchi. Multi-scale domain-adversarial multiple-instance cnn for cancer subtype classification with unannotated histopatho- l...
2020
-
[48]
Computational pathology: a survey review and the way forward
Mahdi S Hosseini, Babak Ehteshami Bejnordi, Vincent Quoc-Huy Trinh, Lyndon Chan, Danial Hasan, Xingwen Li, Stephen Yang, Taehyo Kim, Haochen Zhang, Theodore Wu, et al. Computational pathology: a survey review and the way forward. Journal of Pathology Informatics , page 100357,...
2024
-
[49]
A visual–language foundation model for pathology image analysis using medical twitter
Zhi Huang, Federico Bianchi, Mert Yuksekgonul, Thomas J Montine, and James Zou. A visual–language foundation model for pathology image analysis using medical twitter. Nature medicine, 29(9):2307–2316, 2023. 1, 2, 3, 5, 6, 16, 17, 19, 22, 24, 25, 26, 27
2023
-
[50]
Grand challenge on breast cancer histology images
BACH ICIAR. Grand challenge on breast cancer histology images. 2018, 2018. 6, 7, 20
2018
-
[51]
Quilt-1m: One million image-text pairs for histopathology
Wisdom Ikezogwo, Saygin Seyfioglu, Fatemeh Ghezloo, Dylan Geva, Fatwir Sheikh Mohammed, Pavan Kumar Anand, Ranjay Krishna, and Linda Shapiro. Quilt-1m: One million image-text pairs for histopathology. Advances in Neural Information Processing Systems, 36, 2024. 1, 2, 3, 4, 5, ...
2024
-
[52]
Attention-based deep multiple instance learning
Maximilian Ilse, Jakub Tomczak, and Max Welling. Attention-based deep multiple instance learning. In In- ternational conference on machine learning , pages 2127–
-
[53]
Cplip: Zero-shot learning for histopathology with comprehensive vision-language alignment
Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, and Mohammed Ben- namoun. Cplip: Zero-shot learning for histopathology with comprehensive vision-language alignment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat-...
2024
-
[54]
Mimic-iii, a freely accessible critical care database
Alistair EW Johnson, Tom J Pollard, Lu Shen, Li-wei H Lehman, Mengling Feng, Mohammad Ghassemi, Benjamin Moody, Peter Szolovits, Leo Anthony Celi, and Roger G Mark. Mimic-iii, a freely accessible critical care database. Scientific data, 3(1):1–9, 2016. 19
2016
-
[55]
Benchmarking self-supervised learn- ing on diverse pathology datasets
Mingu Kang, Heon Song, Seonwook Park, Donggeun Yoo, and S ´ergio Pereira. Benchmarking self-supervised learn- ing on diverse pathology datasets. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3344–3354, 2023. 3, 5, 6, 19, 22, 27
2023
-
[56]
Si- mil: Taming deep mil for self-interpretability in gigapixel histopathology
Saarthak Kapse, Pushpak Pati, Srijan Das, Jingwei Zhang, Chao Chen, Maria Vakalopoulou, Joel Saltz, Dimitris Samaras, Rajarsi R Gupta, and Prateek Prasanna. Si- mil: Taming deep mil for self-interpretability in gigapixel histopathology. In Proceedings of the IEEE/CVF Confer- e...
2024
-
[57]
100,000 histological images of human colorectal cancer and healthy tissue
Jakob Nikolas Kather, Niels Halama, and Alexander Marx. 100,000 histological images of human colorectal cancer and healthy tissue. Zenodo10, 5281, 2018. 6, 7
2018
-
[58]
Application of artificial intelligence in pathology: Trends and challenges
Inho Kim, Kyungmin Kang, Youngjae Song, and Tae-Jung Kim. Application of artificial intelligence in pathology: Trends and challenges. Diagnostics, 12(11):2794, 2022. 1
2022
-
[59]
Wilds: A benchmark of in-the-wild distri- bution shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distri- bution shifts. InInternational conference on machine learn- ing...
2021
-
[60]
Deep learning for the detection of anatomical tissue structures and neoplasms of the skin on scanned histopathological tissue sections
Katharina Kriegsmann, Frithjof Lobers, Christiane Zgorzelski, Joerg Kriegsmann, Charlotte Janssen, Rolf Ruedinger Meliss, Thomas Muley, Ulrich Sack, Georg Steinbuss, and Mark Kriegsmann. Deep learning for the detection of anatomical tissue structures and neoplasms of the skin ...
2022
-
[61]
Dual-stream multiple instance learning network for whole slide image classifica- tion with self-supervised contrastive learning
Bin Li, Yin Li, and Kevin W Eliceiri. Dual-stream multiple instance learning network for whole slide image classifica- tion with self-supervised contrastive learning. In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 14318–14328, 2021. 3
2021
-
[62]
mplug: Effective and efficient vision-language learning by cross-modal skip-connections
Chenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang, Ming Yan, Bin Bi, Jiabo Ye, Hehong Chen, Guohai Xu, Zheng Cao, et al. mplug: Effective and efficient vision-language learning by cross-modal skip-connections. arXiv preprint arXiv:2205.12005, 2022. 5, 6, 17, 18
2022 arXiv
-
[63]
Align and prompt: Video-and- language pre-training with entity prompts
Dongxu Li, Junnan Li, Hongdong Li, Juan Carlos Niebles, and Steven CH Hoi. Align and prompt: Video-and- language pre-training with entity prompts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 4953–4963, 2022. 18
2022
-
[64]
Generalizable whole slide image classification with fine- grained visual-semantic interaction
Hao Li, Ying Chen, Yifei Chen, Rongshan Yu, Wenxian Yang, Liansheng Wang, Bowen Ding, and Yuchen Han. Generalizable whole slide image classification with fine- grained visual-semantic interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[65]
A multi- resolution model for histopathology image classification and localization with multiple instance learning.Computers in biology and medicine, 131:104253, 2021
Jiayun Li, Wenyuan Li, Anthony Sisk, Huihui Ye, W Dean Wallace, William Speier, and Corey W Arnold. A multi- resolution model for histopathology image classification and localization with multiple instance learning.Computers in biology and medicine, 131:104253, 2021. 3
2021
-
[66]
Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi
Junnan Li, Ramprasaath R. Selvaraju, Akhilesh Deepak Gotmare, Shafiq Joty, Caiming Xiong, and Steven Hoi. Align before fuse: Vision and language representation learning with momentum distillation. In NeurIPS, 2021. 5, 6, 17, 18, 19
2021
-
[67]
Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation. In International Conference on Machine Learning , pages 12888–12900. PMLR, 2022. 18
2022
-
[68]
Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In In- ternational conference on machine learning, pages 19730– 19742. PMLR, 2023. 23
2023
-
[69]
Dynamic graph rep- resentation with knowledge-aware attention for histopathol- ogy whole slide image analysis
Jiawen Li, Yuxuan Chen, Hongbo Chu, Qiehe Sun, Tian Guan, Anjia Han, and Yonghong He. Dynamic graph rep- resentation with knowledge-aware attention for histopathol- ogy whole slide image analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...
2024
-
[70]
Self-supervised learning: Generative or contrastive
Xiao Liu, Fanjin Zhang, Zhenyu Hou, Li Mian, Zhaoyu Wang, Jing Zhang, and Jie Tang. Self-supervised learning: Generative or contrastive. IEEE transactions on knowledge and data engineering, 35(1):857–876, 2021. 1
2021
-
[71]
To- wards a visual-language foundation model for computa- tional pathology
Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Ivy Liang, Tong Ding, Guillaume Jaume, Igor Odintsov, Andrew Zhang, Long Phi Le, et al. To- wards a visual-language foundation model for computa- tional pathology. arXiv preprint arXiv:2307.12914, 2023. 1, 2, 5, 6, 7, ...
2023 arXiv
-
[72]
Visual lan- guage pretrained multiple instance zero-shot transfer for histopathology images
Ming Y Lu, Bowen Chen, Andrew Zhang, Drew FK Williamson, Richard J Chen, Tong Ding, Long Phi Le, Yung-Sung Chuang, and Faisal Mahmood. Visual lan- guage pretrained multiple instance zero-shot transfer for histopathology images. In Proceedings of the IEEE/CVF Conference on Comp...
2023
-
[73]
A mul- timodal generative ai copilot for human pathology
Ming Y Lu, Bowen Chen, Drew FK Williamson, Richard J Chen, Melissa Zhao, Aaron K Chow, Kenji Ikemura, Ahrong Kim, Dimitra Pouli, Ankush Patel, et al. A mul- timodal generative ai copilot for human pathology. Nature, pages 1–3, 2024. 3
2024
-
[74]
You’re on social media! so now what? Archives of Pathology & Lab- oratory Medicine, 140(5):393–393, 2016
Michael J Misialek and Timothy Craig Allen. You’re on social media! so now what? Archives of Pathology & Lab- oratory Medicine, 140(5):393–393, 2016. 22
2016
-
[75]
A threshold selection method from gray- level histograms
Nobuyuki Otsu. A threshold selection method from gray- level histograms. IEEE transactions on systems, man, and cybernetics, 9(1):62–66, 1979. 19
1979
-
[76]
Review of the current state of whole slide imaging in pathology
Liron Pantanowitz, Paul N Valenstein, Andrew J Evans, Keith J Kaplan, John D Pfeifer, David C Wilbur, Laura C Collins, and Terence J Colgan. Review of the current state of whole slide imaging in pathology. Journal of pathology informatics, 2(1):36, 2011. 1
2011
-
[77]
Hun- crc: annotated pathological slides to enhance deep learning applications in colorectal cancer screening
B ´alint ´Armin Pataki, Alex Olar, Dezs˝o Ribli, Adri´an Pesti, Endre Kontsek, Benedek Gy ¨ongy¨osi, ´Agnes Bilecz, Tekla Kov´acs, Krist ´of Attila Kov´acs, Zs ´ofia Kramer, et al. Hun- crc: annotated pathological slides to enhance deep learning applications in colorectal canc...
2022
-
[78]
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9,
-
[79]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In International conference on machine learning...
2021
-
[80]
The digital brain tumour atlas, an open histopathology resource
Thomas Roetzer-Pejrimovsky, Anna-Christina Moser, Baran Atli, Clemens Christian V ogel, Petra A Mercea, Ro- mana Prihoda, Ellen Gelpi, Christine Haberler, Romana H¨oftberger, Johannes A Hainfellner, et al. The digital brain tumour atlas, an open histopathology resource. Scient...
2022
-
[81]
Quilt-llava: Visual instruction tuning by extracting localized narratives from open-source histopathology videos
Mehmet Saygin Seyfioglu, Wisdom O Ikezogwo, Fatemeh Ghezloo, Ranjay Krishna, and Linda Shapiro. Quilt-llava: Visual instruction tuning by extracting localized narratives from open-source histopathology videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pa...
2024
-
[82]
Cancer diagnosis using artificial intelligence: A review
K Aditya Shastry and HA Sanjay. Cancer diagnosis using artificial intelligence: A review. Artificial Intelligence Re- view, pages 1–33, 2022. 1
2022
-
[83]
Tiager: Tumor-infiltrating lymphocyte scoring in breast cancer for the tiger challenge
Adam Shephard, Mostafa Jahanifar, Ruoyu Wang, Muham- mad Dawood, Simon Graham, Kastytis Sidlauskas, Syed Ali Khurram, Nasir Rajpoot, and Shan E Ahmed Raza. Tiager: Tumor-infiltrating lymphocyte scoring in breast cancer for the tiger challenge. arXiv preprint arXiv:2206.11943, ...
2022 arXiv
-
[84]
Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification
Jiangbo Shi, Chen Li, Tieliang Gong, Yefeng Zheng, and Huazhu Fu. Vila-mil: Dual-scale vision-language multiple instance learning for whole slide image classification. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 11248–11258, 2...
2024
-
[85]
Self-learning for weakly supervised glea- son grading of local patterns
Julio Silva-Rodriguez, Adri ´an Colomer, Jose Dolz, and Valery Naranjo. Self-learning for weakly supervised glea- son grading of local patterns. IEEE journal of biomedical and health informatics , 25(8):3094–3104, 2021. 6, 7, 20, 26
2021
-
[86]
Artificial intelligence for digital and computa- tional pathology
Andrew H Song, Guillaume Jaume, Drew FK Williamson, Ming Y Lu, Anurag Vaidya, Tiffany R Miller, and Faisal Mahmood. Artificial intelligence for digital and computa- tional pathology. Nature Reviews Bioengineering , 1(12): 930–949, 2023. 1, 3
2023
-
[87]
Mor- phological prototyping for unsupervised slide representa- tion learning in computational pathology
Andrew H Song, Richard J Chen, Tong Ding, Drew FK Williamson, Guillaume Jaume, and Faisal Mahmood. Mor- phological prototyping for unsupervised slide representa- tion learning in computational pathology. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- ter...
2024
-
[88]
Deep neural network models for computational histopathology: A survey
Chetan L Srinidhi, Ozan Ciga, and Anne L Martel. Deep neural network models for computational histopathology: A survey. Medical image analysis, 67:101813, 2021. 3
2021
-
[89]
Feature re-embedding: Towards foun- dation model-level performance in computational pathol- ogy
Wenhao Tang, Fengtao Zhou, Sheng Huang, Xiang Zhu, Yi Zhang, and Bo Liu. Feature re-embedding: Towards foun- dation model-level performance in computational pathol- ogy. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 11343–11352,
-
[90]
The digital brain tumour atlas, an open histopathology resource
Roetzer-Pejrimovsky Thomas, Moser Anna-Christina, Atli Baran, Clemens Christian V ogel, Petra A Mercea, Prihoda Romana, Ellen Gelpi, Haberler Christine, H ¨oftberger Ro- mana, Johannes A Hainfellner, et al. The digital brain tumour atlas, an open histopathology resource. Scien...
2022
-
[91]
Adaptive weighting multi-field-of-view cnn for semantic segmentation in pathology
Hiroki Tokunaga, Yuki Teramoto, Akihiko Yoshizawa, and Ryoma Bise. Adaptive weighting multi-field-of-view cnn for semantic segmentation in pathology. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12597–12606, 2019. 3
2019
-
[92]
Hooknet: Multi-resolution convolutional neural networks for seman- tic segmentation in histopathology whole-slide images
Mart Van Rijthoven, Maschenka Balkenhol, Karina Silin ¸a, Jeroen Van Der Laak, and Francesco Ciompi. Hooknet: Multi-resolution convolutional neural networks for seman- tic segmentation in histopathology whole-slide images. Medical image analysis, 68:101890, 2021. 2, 4
2021
-
[93]
Rotation equivariant cnns for dig- ital pathology
Bastiaan S Veeling, Jasper Linmans, Jim Winkens, Taco Cohen, and Max Welling. Rotation equivariant cnns for dig- ital pathology. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Pro...
2018
-
[94]
Computational pathology in cancer diagnosis, prog- nosis, and prediction–present day and prospects
Gregory Verghese, Jochen K Lennerz, Danny Ruta, Wen Ng, Selvam Thavaraj, Kalliopi P Siziopikou, Threnesan Naidoo, Swapnil Rane, Roberto Salgado, Sarah E Pinder, et al. Computational pathology in cancer diagnosis, prog- nosis, and prediction–present day and prospects. The Jour-...
2023
-
[95]
A foundation model for clinical-grade computational pathology and rare cancers detection
Eugene V orontsov, Alican Bozkurt, Adam Casson, George Shaikovski, Michal Zelechowski, Kristen Severson, Eric Zimmermann, James Hall, Neil Tenenholtz, Nicolo Fusi, et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nature Medicine ...
2024
-
[96]
Handcrafted histological transformer (h2t): Unsupervised representation of whole slide images
Quoc Dang Vu, Kashif Rajpoot, Shan E Ahmed Raza, and Nasir Rajpoot. Handcrafted histological transformer (h2t): Unsupervised representation of whole slide images. Medi- cal Image Analysis, 85:102743, 2023. 4
2023
-
[97]
Transformer-based unsupervised contrastive learning for histopathological image classification.Medical image anal- ysis, 81:102559, 2022
Xiyue Wang, Sen Yang, Jun Zhang, Minghui Wang, Jing Zhang, Wei Yang, Junzhou Huang, and Xiao Han. Transformer-based unsupervised contrastive learning for histopathological image classification.Medical image anal- ysis, 81:102559, 2022. 3, 6, 8, 19, 27
2022
-
[98]
A pathology foundation model for cancer diagnosis and prognosis prediction
Xiyue Wang, Junhan Zhao, Eliana Marostica, Wei Yuan, Jietian Jin, Jiayu Zhang, Ruijiang Li, Hongping Tang, Kan- ran Wang, Yu Li, et al. A pathology foundation model for cancer diagnosis and prognosis prediction. Nature, pages 1–9, 2024. 6, 8, 27
2024
-
[99]
A petri dish for histopathology image analysis
Jerry Wei, Arief Suriawinata, Bing Ren, Xiaoying Liu, Mikhail Lisovsky, Louis Vaickus, Charles Brown, Michael Baker, Naofumi Tomita, Lorenzo Torresani, et al. A petri dish for histopathology image analysis. In Artificial Intel- ligence in Medicine: 19th International Conferenc...
2021
-
[100]
The cancer genome atlas pan-cancer analysis project
John N Weinstein, Eric A Collisson, Gordon B Mills, Kenna R Shaw, Brad A Ozenberger, Kyle Ellrott, Ilya Shmulevich, Chris Sander, and Joshua M Stuart. The cancer genome atlas pan-cancer analysis project. Nature genetics, 45(10):1113–1120, 2013. 2, 4, 6, 17, 18
2013
-
[101]
A whole-slide founda- tion model for digital pathology from real-world data
Hanwen Xu, Naoto Usuyama, Jaspreet Bagga, Sheng Zhang, Rajesh Rao, Tristan Naumann, Cliff Wong, Zelalem Gero, Javier Gonz´alez, Yu Gu, et al. A whole-slide founda- tion model for digital pathology from real-world data. Na- ture, pages 1–8, 2024. 3, 6
2024
-
[102]
Hitea: Hierarchical temporal- aware video-language pre-training
Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, and Fei Huang. Hitea: Hierarchical temporal- aware video-language pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 15405–15416, 2023. 6, 17
2023
-
[103]
Prompting vision foundation models for pathology image analysis
Chong Yin, Siqi Liu, Kaiyang Zhou, Vincent Wai-Sun Wong, and Pong C Yuen. Prompting vision foundation models for pathology image analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11292–11301, 2024. 3
2024
-
[104]
Lungren, Tristan Naumann, Sheng Wang, and Hoifung Poon
Sheng Zhang, Yanbo Xu, Naoto Usuyama, Hanwen Xu, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, Andrea Tupini, Yu Wang, Matt Mazzola, Swadheen Shukla, Lars Liden, Jianfeng Gao, Matthew P. Lungren, Tristan Naumann, Sheng Wang, and Hoif...
2024
-
[105]
Pathologist-level interpretable whole-slide cancer diagnosis with deep learn- ing
Zizhao Zhang, Pingjun Chen, Mason McGough, Fuyong Xing, Chunbao Wang, Marilyn Bui, Yuanpu Xie, Manish Sapkota, Lei Cui, Jasreman Dhillon, et al. Pathologist-level interpretable whole-slide cancer diagnosis with deep learn- ing. Nature Machine Intelligence, 1(5):236–245, 2019. 1, 2, 15
2019
-
[106]
Algorithm 778: L-bfgs-b: Fortran subroutines for large-scale bound-constrained optimization
Ciyou Zhu, Richard H Byrd, Peihuang Lu, and Jorge No- cedal. Algorithm 778: L-bfgs-b: Fortran subroutines for large-scale bound-constrained optimization. ACM Trans- actions on mathematical software (TOMS), 23(4):550–560,
-
[107]
Can you describe the main features visible in this histopathology image?
Mengdan Zhu, Bing Ren, Ryland Richards, Matthew Suri- awinata, Naofumi Tomita, and Saeed Hassanpour. Develop- ment and evaluation of a deep neural network for histologic classification of renal cell carcinoma on biopsy and surgical resection slides. Scientific reports, 11(1):7...
2021
-
[110]
MR-PLIP Model: More Insights (Sec. 7)
-
[111]
Parent-Child Hierarchy (Sec. 8)
-
[112]
Multi-modal Encoder for Text-guided Visual Represen- tation (Sec. 9)
-
[113]
Pre-Training Baseline Objectives (Sec. 10)
-
[114]
Additional Training and Implementation Details (Sec. 11)
-
[115]
Evaluation Metrics (Sec. 12)
-
[116]
Histopathology Datasets (Sec. 13)
-
[117]
More Ablation Studies (Sec. 14)
-
[118]
Zero-Shot Experiments (Sec. 15)
-
[119]
Linear Probe Experiments (Sec. 16)
-
[120]
Weakly-Supervised WSI Classification Results (Sec. 17)
-
[121]
Fine-tune Evaluation (Sec. 18)
-
[122]
More Comparison with Recent MIL-based Methods (Sec. 19)
-
[123]
MR-PLIP Model: More Insights In clinical diagnostics, expert pathologists often analyze the WSI to predict the outcomes of diseases by inspecting it from multiple resolution levels. The multi-resolution anal- ysis of the WSIs assists expert pathologists to better analyze the t...
-
[124]
(3) of the main manuscript, we focus on maintaining the integrity of the parent-child relationship, as depicted in Fig 7
Parent-Child Hierarchy To clarify the approach used to preserve the hierarchical structure in our curated histopathology dataset, particu- larly when aligning text-guided visual features across dif- ferent resolution levels as outlined in Eq. (3) of the main manuscript, we foc...
-
[125]
Here, we explicitly adapt and merge their methodologies for the first time to suit the specific needs of the CPath domain, as demonstrated in Fig
Multi-modal Encoder for Text-guided Visual Representation Our multi-modal encoder’s architecture incorporates ele- ments from the frameworks presented in [62, 66, 102], which originally were not applied to Computational Pathol- ogy (CPath). Here, we explicitly adapt and merge ...
-
[126]
UNI (Vision Encoder) QuiltNet (Text Encoder) ITC Multi-Modal Encoder Text Decoder ITM MLM Prefix LM Invasive lobular carcinoma of the breast surrounding a Pacinian corpuscle
Pre-Training Baseline Objectives During the pre-training process, we perform four pre- training tasks: Image-Text Contrastive Learning ( LIT C), Image-Text Matching (LIT M ), Masked Language Mod- eling (LM LM), and Prefix Language Modeling ( LP LM). UNI (Vision Encoder) QuiltN...
-
[127]
Additional Training and Implementation Details The TCGA dataset is renowned for being one of the largest publicly available histopathology collections, covering a broad spectrum of cancer morphologies and subtypes across key organs [100]. We included WSIs from 21 major primary...
-
[128]
In alignment with SOTA methods [49, 51, 72, 104], we fine-tuned the baseline CLIP model [79] using a ViT- B/16-224 [37] as the image encoder and GPT-2/77 [78] as the text encoder
-
[129]
Recognizing the baseline CLIP’s training on out-of- domain paired data, we also fine-tuned the MR-PLIP model with a pathology domain-specific pre-trained PLIP [49] , using PLIP-ViT-B/32-224 as the image en- coder and GPT-2/347 as the text encoder
-
[130]
BioClinicalBert and PubMedBERT, both non-pathology text encoders, are trained on biomed- ical and clinical corpora, such as PubMed abstracts and MIMIC [54]
Following the approach of MI-zero and Biomed- CLIP [104], we fine-tuned the MR-PLIP model using BioClinicalBert/512 [6] and PubMedBERT/256 [44] as text encoders, alongside CTransPath/224 [97] as the im- age encoder. BioClinicalBert and PubMedBERT, both non-pathology text encod...
-
[131]
MR-PLIP was also fine-tuned using BioClinicalBert/512 as the text encoder and PLIP-ViT-B/32-224 as the in- domain image encoder
-
[132]
Furthermore, we fine-tuned MR-PLIP using CTransPath/224 as the in-domain image encoder and PLIP-GPT/347 as the in-domain text encoder
-
[133]
These experiments were conducted using inference time prompts similar to [71] for a fair comparison
Additionally, we fine-tuned MR-PLIP using ViT-B/16 as the in-domain image encoder from the QuiltNet and GPT-2/77 as the in-domain text encoder of the QuiltNet model [51]. These experiments were conducted using inference time prompts similar to [71] for a fair comparison. Throu...
-
[134]
Evaluation Metrics For evaluating performance across different CPath tasks [49, 71], we use a variety of metrics. These in- clude the weighted average F1 score, balanced accuracy, dice score, precision, recall, multi-class Panoptic Qual- ity (mPQ), Recall @1 (R @1), Recall @50...
-
[135]
To evaluate our MR-PLIP model on these tasks, we used 26 independent datasets
Downstream Histopathology Datasets We performed five different computational pathology (CPath) tasks, including tile-level classification, WSI-level classification, cross-modal retrieval , WSI segmentation , and nuclei segmentation. To evaluate our MR-PLIP model on these tasks...
-
[136]
Impact of Magnification Levels (Table 5) In Sec
More Ablation Studies 14.1. Impact of Magnification Levels (Table 5) In Sec. 7, we presented the argument that multi-resolution pre-training is advantageous, considering that downstream CPath tasks are executed at various magnifications. Our ex- perimental results indeed confi...
-
[137]
This is because UNI is pre-trained on unlabeled large histology images, and the QuiltNet text encoder is trained on 1M histology image-text pairs
The best results on six datasets are reported using UNI (ViT-L/16-224) as an image encoder and QuiltNet (GPT- 2/77) to initialize the text encoder. This is because UNI is pre-trained on unlabeled large histology images, and the QuiltNet text encoder is trained on 1M histology ...
-
[138]
Zero-Shot Experiments Zero-shot learning refers to the capability of models to ac- curately perform tasks on new, unseen data without direct training on those specific tasks, utilizing pre-learned repre- sentations from image-text pairs. 15.1. Zero-shot Segmentation Results (T...
-
[139]
Illustration of the zero-shot classification process used by SOTA VLMs, like PLIP [49] and MI-Zero [72]
73 0.48 0.32 0.54…..I An H&E image of Tumor Input Image Zero-Shot Classification used by SOTA VLMs (PLIP , MI-Zero) Vision Encoder An H&E image of Lymphocyte Stroma Epithelium Text Encoder Input Text Prompt: Tumor Figure 9. Illustration of the zero-shot classification process ...
-
[140]
Specifically, it involves training a simple linear classifier (e.g., logistic regression) on top of the features extracted from a pre-trained neural network
Linear Probe Experiments (Table 13) In the context of deep learning, linear probing refers to a technique used to evaluate the quality of features learned by a deep neural network. Specifically, it involves training a simple linear classifier (e.g., logistic regression) on top...
-
[141]
An H&E image of Tumor
92 0.49 0.22 0.42…..I T1 T2 T3 T5…. An H&E image of Tumor Lymphocyte Stroma Epithelium MR-PLIP (Text Encoder) Input Text Prompt: MR-PLIP (Multi-Modal Encoder) An H&E image of Tumor MR-PLIP (Vision Encoder) Text-Guided Visual Representation (𝓛𝑴𝑹𝑻𝑽𝑨) Zero-Shot Classification usi...
-
[142]
Weakly-Supervised WSI Classification Re- sults (Table 13) We performed weakly-supervised WSI classification to evaluate the text-guided visual representations learned by MR-PLIP across seven diverse WSI classification datasets. MR-PLIP was used to extract text-guided visual fe...
-
[143]
Specifi- cally, it involves fine-tuning the original network’s weights
Fine-tune Evaluation (Table 15) In the context of deep learning, fine-tuning refers to a tech- nique used to assess the adaptability and transfer potential of the learned weights by a deep neural network. Specifi- cally, it involves fine-tuning the original network’s weights. ...
-
[144]
For a fair comparison, we em- ployed the same ABMIL method for WSI-level feature ag- gregation and classification as discussed in Sec
More Comparisons with SOTA MIL-based Methods (Table 16) In this experiment, we compared the performance of our MR-PLIP model with recently proposed MIL-based meth- ods including FiVE [64], R2T [89], SI-MIL [56], ViLa-MIL [84], and PANTHER [87]. For a fair comparison, we em- pl...
-
[2024]
3, 4, 5, 6, 8, 19, 22, 27
-
[2136]
3, 4, 8, 27
PMLR, 2018. 3, 4, 8, 27
2018
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.