REVIEW 4 major objections 4 minor 34 references
The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population
T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read This paper releases a public dataset of 339 prostate biopsy whole-slide images from 185 patients in Erbil, Iraq, scanned with three scanners including a compact one, and graded independently by three pathologists.
desk verdict Genuinely new and useful dataset — first Middle Eastern prostate biopsy WSI set with triple grading and a compact scanner — but the 'consecutive representative' claim is unsupported as written and the manuscript needs internal consistency fixes before the deposit goes live. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the dataset itself: 339 WSIs in native scanner formats (.svs and .ndpi), organized so each slide has three scanner versions and linked to a single annotations table with independent Gleason/ISUP grades from three pathologists. The design work is carried by the pairing of multi-scanner capture (two high-throughput scanners and the Grundium Ocus40 compact scanner) with multi-pathologist independent reference grading; this combination is what makes cross-scanner and cross-rater analyses possible. A supporting mechanism is the decision to recut new H&E slides from archived FFPE blocks rather than scan faded original slides, which standardizes tissue quality across the colle
What would settle it
Compare the age, Gleason score distribution, and biopsy year of the 185 included cases against the 54 excluded cases; a statistically significant difference would show the remaining cohort is not a consecutive, representative series. Re-running the keyword search on the original archive and checking for missed prostate needle biopsy reports would also test the recall of the manual pipeline.
Extended reading notes
Core claim
The paper's central claim is that the PAR dataset is the first publicly available prostate biopsy whole-slide image collection from an underrepresented Middle Eastern population that is digitized with multiple scanners — including a compact, low-cost scanner — and labeled with slide-level Gleason scores and ISUP grades assigned independently by three pathologists. Each of the 339 slides exists in three scanned versions (one per scanner), and labels follow a fixed dictionary of ten Gleason classes and six ISUP grades. The authors position this as a direct response to three gaps in existing public datasets: single-reference grading, Western-only populations, and single-scanner capture. They fu
Load-bearing premise
The dataset is claimed to represent a consecutive series of routine prostate biopsies from 2013–2024, but 54 of 239 eligible cases were dropped because their tissue blocks were lost, and no evidence is given that the remaining 185 match the excluded cases in age, grade, or year.
Editorial extensions
If this is right
- If the dataset is representative, AI models trained on Western prostate biopsies can be tested for the first time on a Middle Eastern population, revealing whether grade distributions and tissue appearance shift across populations.
- The inclusion of compact-scanner images allows direct evaluation of whether low-cost scanners are adequate for AI validation in clinics that have not yet digitized.
- With three independent graders, the dataset supports inter-observer agreement studies and the construction of consensus labels under any explicit rule (majority, highest grade, etc.).
- Native-format 40x WSIs from three scanners enable color normalization and stain-robustness benchmarking that single-scanner datasets cannot offer.
Reading between the lines
- Because the slides were recut from archived blocks rather than scanned from the original diagnostic slides, the images may not reflect the exact tissue section that produced the clinical report; the labels come from the new sections, which could differ slightly in grade from the original clinical diagnosis.
- The dataset's value as a population-representative benchmark depends on the excluded cases; a reader could re-examine the hospital archive to test whether the keyword filter and the 54 lost blocks introduced selection bias.
- If the compact-scanner images show systematic color or focus shifts, they could serve as a natural testbed for stain normalization and domain adaptation methods rather than purely as additional data.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the PAR dataset: 339 glass slides from prostate core needle biopsies of 185 patients from PAR Hospital in Erbil, Iraq, digitized with three whole-slide scanners (Leica Aperio GT450, Hamamatsu NanoZoomer HT 2.0, and the compact Grundium Ocus40), yielding 1,017 WSI files in native formats. Slide-level Gleason scores and ISUP grades were assigned independently by three pathologists, with one pathologist grading all slides, one grading 337 slides, and one grading a stratified subset of 59. The authors claim this is the first public dataset with compact-scanner images and from an underrepresented Middle Eastern population, enabling cross-scanner, cross-population, and multi-pathologist validation studies. The manuscript describes curation, scanning, annotation, and technical validation steps.
Significance. If the dataset is released as described, it would be a genuinely valuable community resource. It addresses two documented gaps: public prostate biopsy WSIs from non-Western populations and WSIs from a compact scanner. The three-scanner design with independent multi-pathologist labels is well suited for color normalization and cross-scanner robustness studies. The authors are transparent about limitations (no pixel-level annotations, generalist vs. uropathologist differences) and the curation workflow is described in checkable detail. The dataset, once deposited, would support reproducibility research and help benchmark AI models across populations.
major comments (4)
- [Methods - Slide selection; Background & Summary] The central claim that the 185-patient series is a 'consecutive series representative of routine clinical practice between 2013-2024' is not supported by the evidence provided. Fifty-four of 239 eligible cases (23%) were excluded because FFPE blocks were lost, 'mostly from the early years of the collection.' No comparison is given between included and excluded cases on age, Gleason score, year, or tumor burden, and no audit is reported of the recall of the manual keyword search for 'pros' in .docx reports. If lost blocks or missed reports are grade- or year-dependent, the dataset is a convenience sample, not a consecutive series. Please provide any available comparison using the archived reports, or explicitly weaken the representativeness claim and discuss the potential bias.
- [Abstract vs. Data Records / Data Availability] There are load-bearing inconsistencies in the dataset description. The abstract states '1,017 whole slide images' and gives BioImage Archive accession S-BIAD2323; the body states '339 digitized prostate core needle biopsies' and gives accession 'TBA' with DOI 'TBA.' If 1,017 is the total number of WSI files (339 slides x 3 scanners), this should be stated explicitly and consistently throughout. The accession number is essential for a data paper and must be reconciled.
- [Table 1 and Reference standard protocol] The claim that the second pathologist (H.M.) graded 337 slides is inconsistent with Table 1: the non-missing Pat. II counts sum to 336 (208+2+7+8+25+86=336, with 3 missing). One of these is wrong. Additionally, Pat. III is described as a 'random subset of 59 slides stratified by ISUP grades assigned by the first pathologist,' but the Pat. III grade distribution (11, 5, 13, 3, 6, 21) is very different from the Pat. I distribution (164, 7, 38, 59, 30, 41) and does not appear proportionally stratified. Specify the stratification scheme or correct the description.
- [Methods - Slide selection; Table 1] The text refers to '185 prostate cancer needle biopsy cases,' yet Table 1 lists 164 benign slides in the Pat. I column. If the cohort includes patients whose biopsies were benign (or benign slides from cancer patients), the wording is inaccurate and should be changed to 'prostate needle biopsy cases.' Clarify whether all patients had a cancer diagnosis or whether some slides/patients are benign.
minor comments (4)
- [Figure 2 caption] The caption reports 0.2266 micrometers per pixel for Hamamatsu, while Methods reports 0.22 um/pixel. Use a consistent value.
- [Table 1] The 'Missing' row in the ISUP grade columns is ambiguous: it mixes ungraded slides for Pat. II (3) with the 280 slides not assessed by Pat. III. Consider labeling this row clearly, e.g., 'Not assessed by this pathologist.'
- [Technical Validation] The text says 'tissue was segmented using deep-learning–based algorithms' but Code Availability states 'No custom code was used.' Clarify whether existing tools were used and name them, or remove the claim about deep-learning segmentation.
- [Data Records] The filename scheme is described as 'c<slide_id><a|b|none>.<ext>,' but the example is not shown. A concrete example filename would help users.
Circularity Check
No circularity found: descriptive dataset paper with no fitted prediction or derivation chain.
full rationale
The paper is a dataset descriptor: it reports the collection, scanning, and grading of 339 prostate biopsy slides from 185 patients. There is no derivation, model fitting, or prediction that could collapse into its inputs. The central gap claim ('No public datasets currently include data digitized with compact scanners or originating from underrepresented populations') is supported by an enumerated survey of prior public datasets (PANDA, SPROB20, SICAPv2, UKK/WNS, DiagSet, Gleason 2019, TCGA-PRAD, EMPaCT, TMAZ, etc.), not by an assumption that contains the conclusion. The self-citations in the manuscript (refs 11 and 28, both by members of the same group) are contextual citations for AI-in-pathology capability and scanner-induced variation; they are not load-bearing justifications for the dataset's novelty or validity, and no uniqueness theorem or ansatz is imported from them. The skeptically identified representativeness limitation (23% of eligible blocks lost, manual keyword-filter recall unassessed, no included-vs-excluded comparison) is a legitimate data-quality and external-validity concern, but it is not a circularity issue: no quantity is defined in terms of a target, no fitted parameter is renamed as a prediction, and no conclusion is forced by construction. Internal inconsistencies in slide counts (339 vs 1,017 images across three scanner copies) and abstract/body wording are editorial issues, not circular reasoning. Therefore the appropriate circularity score is 0.
Assumptions & free parameters
assumptions (4)
- domain assumption Gleason scoring and ISUP grades (refs 1-3) are a valid reference standard for prostate biopsy severity.
- domain assumption Keyword search of the .docx archive for 'pros' exhaustively retrieves all prostate cases.
- domain assumption The 59-slide subset graded by the third pathologist is a usable third reference standard.
- domain assumption WSIs from three scanners with different pixel resolutions are directly usable for cross-scanner robustness without calibration metadata.
Cite this review
Pith. "Pith review of The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population." pith.science (2026). https://pith.science/paper/4EGZAJUE
@misc{pith2026251203854,
author = {Pith},
title = {Pith review of: The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population},
year = {2026},
howpublished = {\url{https://pith.science/paper/4EGZAJUE}},
note = {Machine review of arXiv:2512.03854}
}
read the original abstract
Artificial intelligence (AI) is increasingly used in digital pathology. Publicly available histopathology datasets remain scarce, and those that do exist predominantly represent Western populations. Consequently, the generalizability of AI models to populations from less digitized regions, such as the Middle East, is largely unknown. This motivates the public release of our dataset to support the development and validation of pathology AI models across globally diverse populations. We present 1,017 whole slide images by digitizing 339 glass slides of prostate core needle biopsies from a consecutive series of 185 patients collected in Erbil, Iraq. The slides are associated with Gleason scores and International Society of Urological Pathology grades assigned independently by three pathologists. Scanning was performed using two high-throughput scanners (Leica and Hamamatsu) and one compact scanner (Grundium). All slides were de-identified and are provided in their native formats without further conversion. The dataset enables grading concordance analyses, color normalization, and cross-scanner robustness evaluations. The PAR dataset has been deposited to the BioImage Archive under accession number S-BIAD2323 (https:// doi.org/10.6019/S-BIAD2323).
Reference graph
Works this paper leans on
-
[1]
Gleason, D. F. Histologic grading of prostate cancer: a perspective. Hum. Pathol. 23 , 273–279 (1992). https://doi.org/10.1016/0046-8177(92)90108-f
-
[2]
Epstein, J. I., Allsbrook, W. C. Jr., Amin, M. B., Egevad, L. L. & ISUP Grading Committee The 2005 International Society of Urological Pathology (ISUP) Consensus Conference on Gleason grading of prostatic carcinoma. Am. J. Surg. Pathol. 29 , 1228–1242 (2005). https://doi.org/10.1097/01.pas.0000173646.99337.b1
arXiv 2005
-
[3]
Epstein, J. I. et al. The 2014 International Society of Urological Pathology (ISUP) Consensus Conference on Gleason grading of prostatic carcinoma: definition of grading patterns and proposal for a new grading system. Am. J. Surg. Pathol. 40 , 244–252 (2016). https://doi.org/10.1097/PAS.0000000000000530
-
[4]
Melia, J. et al. A UK-based investigation of inter- and intra-observer reproducibility of Gleason grading of prostatic biopsies. Histopathology 48 , 644–654 (2006). https://doi.org/10.1111/j.1365-2559.2006.02393.x
arXiv 2006
-
[5]
Egevad, L. et al . Standardization of Gleason grading among 337 European pathologists. Histopathology 62 , 247–256 (2013). https://doi.org/10.1111/his.12008
-
[6]
Ozkan, T.A. et al. Interobserver variability in Gleason histological grading of prostate cancer. Scand. J. Urol. 50 , 420–424 (2016). https://doi.org/10.1080/21681805.2016.1206619
arXiv 2016
-
[7]
Pantanowitz, L. et al. Twenty years of digital pathology: An overview of the road travelled, what is on the horizon, and the emergence of vendor-neutral archives. J. Pathol. Inform. 9 , 40 (2018). https://doi.org/10.4103/jpi.jpi_69_18
-
[8]
Campanella, G. et al. Clinical-grade computational pathology using weakly supervised deep learning on whole-slide images. Nat. Med. 25 , 1301–1309 (2019). https://doi.org/10.1038/s41591-019-0508-1
Show all 34 references
-
[9]
Ström, P. et al. Artificial intelligence for diagnosis and grading of prostate cancer in biopsies: a population-based, diagnostic study. Lancet Oncol. 21 , 222–232 (2020). https://doi.org/10.1016/S1470-2045(19)30738-7
2020 doi
-
[10]
Bulten, W. et al. Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study. Lancet Oncol. 21 , 233–241 (2020). https://doi.org/10.1016/S1470-2045(19)30739-9
2020 doi
-
[11]
Mulliqi, N. et al. Foundation models: a panacea for artificial intelligence in pathology? Preprint at https://arxiv.org/abs/2502.21264 (2025)
2025
-
[12]
Bulten, W. et al. Artificial intelligence for diagnosis and Gleason grading of prostate cancer: the PANDA challenge. Nat. Med. 28 , 154–163 (2022). https://doi.org/10.1038/s41591-021-01620-2
2022 doi
-
[13]
Walhagen, P. et al. Spear Prostate Biopsy 2020 (SPROB20). AIDA Data Hub https://doi.org/10.23698/aida/sprob20 (2020)
2020 doi
-
[14]
SICAPv2 – Prostate whole slide images with Gleason grades annotations
Silva-Rodríguez, J. SICAPv2 – Prostate whole slide images with Gleason grades annotations. Mendeley Data . https://doi.org/10.17632/9xxm58dvs3.1 (2020)
2020 doi
-
[15]
A., Molina, R
Silva-Rodríguez, J., Colomer, A., Sales, M. A., Molina, R. & Naranjo, V. Going deeper through the Gleason scoring scale: an automatic end-to-end system for histology prostate grading and cribriform pattern detection. Comput. Methods Programs Biomed. 195 , 105637 (2020). https:...
2020
-
[16]
Tolkach, Y. et al. An international multi-institutional validation study of the algorithm for prostate cancer detection and Gleason grading. NPJ Precis. Oncol. 7 , 77 (2023) https://doi.org/10.1038/s41698-023-00424-6
2023 doi
-
[17]
Koziarski, M. et al. DiagSet: a dataset for prostate cancer histopathological image classification. Sci. Rep. 14 , 6780 (2024). https://doi.org/10.1038/s41598-024-52183-4
2024 doi
-
[18]
Huo, X. et al. A comprehensive AI model development framework for consistent Gleason grading. Commun. Med. 4 , 84 (2024). https://doi.org/10.1038/s43856-024-00502-1
2024 doi
-
[19]
The molecular taxonomy of primary prostate cancer
The Cancer Genome Atlas Research Network. The molecular taxonomy of primary prostate cancer. Cell 163 , 1011–1025 (2015). https://doi.org/10.1016/j.cell.2015.10.025
2015 doi
-
[20]
Heath, A. P. et al. The NCI Genomic Data Commons. Nat. Genet. 53 , 257–262 (2021). https://doi.org/10.1038/s41588-021-00791-5
2021 doi
-
[21]
Dataset EMPaCT TMA
Karkampouna, S., & Kruithof-de Julio, M. Dataset EMPaCT TMA . Zenodo . https://doi.org/10.5281/zenodo.10066853 (2023)
2023 doi
-
[22]
Zhong, Q. et al. A curated collection of tissue microarray images and clinical outcome data of prostate cancer patients. Sci. Data. 4 , 170014 (2017). https://doi.org/10.1038/sdata.2017.14
2017 doi
-
[23]
Arvaniti, E. et al. Replication Data for: Automated Gleason grading of prostate cancer tissue microarrays via deep learning. Harvard Dataverse. https://doi.org/10.7910/DVN/OCYCMP (2018)
2018 doi
-
[24]
Gleason 2019: Automatic Gleason grading of prostate cancer in digital histopathology
MICCAI 2019 Gleason Challenge Organizers. Gleason 2019: Automatic Gleason grading of prostate cancer in digital histopathology. MICCAI 2019 Conference - Challenge Website https://gleason2019.grand-challenge.org/ (2019; accessed 28 November 2025)
2019
-
[25]
Chen, R.J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 30 , 850–862 (2024). https://doi.org/10.1038/s41591-024-02857-3
2024 doi
-
[26]
Vorontsov, E. et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat. Med. 30 , 2924–2935 (2024). https://doi.org/10.1038/s41591-024-03141-0
2024 doi
-
[27]
Duenweg, S. R. et al. Whole slide imaging (WSI) scanner differences influence optical and computed properties of digitized prostate cancer histology. J. Pathol. Inform. 14 , 100321 (2023). https://doi.org/10.1016/j.jpi.2023.100321
2023
-
[28]
Ji, X. et al. Physical color calibration of digital pathology scanners for robust artificial intelligence–assisted cancer diagnosis. Mod. Pathol . 38 , 100715 (2025). https://doi.org/10.1016/j.modpat.2025.100715
2025
-
[29]
A benchmarking crisis in biomedical machine learning
Mahmood, F. A benchmarking crisis in biomedical machine learning. Nat. Med. 31 , 1060 (2025). https://doi.org/10.1038/s41591-025-03637-3
2025 doi
-
[30]
Marée, R. et al. Collaborative analysis of multi-gigapixel imaging data using Cytomine. Bioinformatics 32 , 1395–1401 (2016). https://doi.org/10.1093/bioinformatics/btw013
2016 doi
-
[31]
& Satyanarayanan, M
Goode, A., Gilbert, B., Harkes, J., Jukic, D. & Satyanarayanan, M. OpenSlide: A vendor-neutral software foundation for digital pathology. J. Pathol. Inform. 4 , 27 (2013). https://doi.org/10.4103/2153-3539.119005
2013
-
[32]
ASAP – Automated Slide Analysis Platform
Litjens, G. ASAP – Automated Slide Analysis Platform. GitHub https://computationalpathologygroup.github.io/ASAP (2017)
2017
-
[33]
Bankhead, P. et al. QuPath: Open source software for digital pathology image analysis. Sci. Rep. 7 , 16878 (2017). https://doi.org/10.1038/s41598-017-17204-5
2017 doi
-
[34]
Lu, M.Y. et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng. 5 , 555–570 (2021). https://doi.org/10.1038/s41551-020-00682-w Figure 1. Diagram illustrating dataset curation. The diagram outlines the process of identifying...
2021 doi
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.