Pith. sign in

REVIEW 4 major objections 4 minor 34 references

The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population

T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read This paper releases a public dataset of 339 prostate biopsy whole-slide images from 185 patients in Erbil, Iraq, scanned with three scanners including a compact one, and graded independently by three pathologists.

desk verdict Genuinely new and useful dataset — first Middle Eastern prostate biopsy WSI set with triple grading and a compact scanner — but the 'consecutive representative' claim is unsupported as written and the manuscript needs internal consistency fixes before the deposit goes live. read the letter →

arxiv 2512.03854 v2 pith:4EGZAJUE submitted 2025-12-03 cs.CV

classification cs.CV
keywords prostatecancerwholeslideimagingGleasongradingISUPgradedigitalpathologydatasetmulti-scannerMiddleEasternpopulationcompactscanner
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper aims to establish a new public resource: 339 digitized prostate core needle biopsy whole-slide images from 185 patients seen at a hospital in Erbil, Iraq, between 2013 and 2024. The slides were freshly recut from archived tissue blocks, scanned with two high-throughput scanners and one compact scanner, and given Gleason and ISUP grades independently by three pathologists. The authors argue that no public dataset currently combines an underrepresented Middle Eastern population, compact-scanner capture, multiple scanners, and multi-pathologist reference labels. If the dataset is representative, it would give AI pathology researchers a substrate for cross-population and cross-scanner validation that Western-only, single-scanner datasets cannot provide.

What carries the argument

The central object is the dataset itself: 339 WSIs in native scanner formats (.svs and .ndpi), organized so each slide has three scanner versions and linked to a single annotations table with independent Gleason/ISUP grades from three pathologists. The design work is carried by the pairing of multi-scanner capture (two high-throughput scanners and the Grundium Ocus40 compact scanner) with multi-pathologist independent reference grading; this combination is what makes cross-scanner and cross-rater analyses possible. A supporting mechanism is the decision to recut new H&E slides from archived FFPE blocks rather than scan faded original slides, which standardizes tissue quality across the colle

What would settle it

Compare the age, Gleason score distribution, and biopsy year of the 185 included cases against the 54 excluded cases; a statistically significant difference would show the remaining cohort is not a consecutive, representative series. Re-running the keyword search on the original archive and checking for missed prostate needle biopsy reports would also test the recall of the manual pipeline.

Watch

Extended reading notes

Core claim

The paper's central claim is that the PAR dataset is the first publicly available prostate biopsy whole-slide image collection from an underrepresented Middle Eastern population that is digitized with multiple scanners — including a compact, low-cost scanner — and labeled with slide-level Gleason scores and ISUP grades assigned independently by three pathologists. Each of the 339 slides exists in three scanned versions (one per scanner), and labels follow a fixed dictionary of ten Gleason classes and six ISUP grades. The authors position this as a direct response to three gaps in existing public datasets: single-reference grading, Western-only populations, and single-scanner capture. They fu

Load-bearing premise

The dataset is claimed to represent a consecutive series of routine prostate biopsies from 2013–2024, but 54 of 239 eligible cases were dropped because their tissue blocks were lost, and no evidence is given that the remaining 185 match the excluded cases in age, grade, or year.

Editorial extensions

If this is right

  • If the dataset is representative, AI models trained on Western prostate biopsies can be tested for the first time on a Middle Eastern population, revealing whether grade distributions and tissue appearance shift across populations.
  • The inclusion of compact-scanner images allows direct evaluation of whether low-cost scanners are adequate for AI validation in clinics that have not yet digitized.
  • With three independent graders, the dataset supports inter-observer agreement studies and the construction of consensus labels under any explicit rule (majority, highest grade, etc.).
  • Native-format 40x WSIs from three scanners enable color normalization and stain-robustness benchmarking that single-scanner datasets cannot offer.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the slides were recut from archived blocks rather than scanned from the original diagnostic slides, the images may not reflect the exact tissue section that produced the clinical report; the labels come from the new sections, which could differ slightly in grade from the original clinical diagnosis.
  • The dataset's value as a population-representative benchmark depends on the excluded cases; a reader could re-examine the hospital archive to test whether the keyword filter and the 54 lost blocks introduced selection bias.
  • If the compact-scanner images show systematic color or focus shifts, they could serve as a natural testbed for stain normalization and domain adaptation methods rather than purely as additional data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper introduces the PAR dataset: 339 glass slides from prostate core needle biopsies of 185 patients from PAR Hospital in Erbil, Iraq, digitized with three whole-slide scanners (Leica Aperio GT450, Hamamatsu NanoZoomer HT 2.0, and the compact Grundium Ocus40), yielding 1,017 WSI files in native formats. Slide-level Gleason scores and ISUP grades were assigned independently by three pathologists, with one pathologist grading all slides, one grading 337 slides, and one grading a stratified subset of 59. The authors claim this is the first public dataset with compact-scanner images and from an underrepresented Middle Eastern population, enabling cross-scanner, cross-population, and multi-pathologist validation studies. The manuscript describes curation, scanning, annotation, and technical validation steps.

Significance. If the dataset is released as described, it would be a genuinely valuable community resource. It addresses two documented gaps: public prostate biopsy WSIs from non-Western populations and WSIs from a compact scanner. The three-scanner design with independent multi-pathologist labels is well suited for color normalization and cross-scanner robustness studies. The authors are transparent about limitations (no pixel-level annotations, generalist vs. uropathologist differences) and the curation workflow is described in checkable detail. The dataset, once deposited, would support reproducibility research and help benchmark AI models across populations.

major comments (4)
  1. [Methods - Slide selection; Background & Summary] The central claim that the 185-patient series is a 'consecutive series representative of routine clinical practice between 2013-2024' is not supported by the evidence provided. Fifty-four of 239 eligible cases (23%) were excluded because FFPE blocks were lost, 'mostly from the early years of the collection.' No comparison is given between included and excluded cases on age, Gleason score, year, or tumor burden, and no audit is reported of the recall of the manual keyword search for 'pros' in .docx reports. If lost blocks or missed reports are grade- or year-dependent, the dataset is a convenience sample, not a consecutive series. Please provide any available comparison using the archived reports, or explicitly weaken the representativeness claim and discuss the potential bias.
  2. [Abstract vs. Data Records / Data Availability] There are load-bearing inconsistencies in the dataset description. The abstract states '1,017 whole slide images' and gives BioImage Archive accession S-BIAD2323; the body states '339 digitized prostate core needle biopsies' and gives accession 'TBA' with DOI 'TBA.' If 1,017 is the total number of WSI files (339 slides x 3 scanners), this should be stated explicitly and consistently throughout. The accession number is essential for a data paper and must be reconciled.
  3. [Table 1 and Reference standard protocol] The claim that the second pathologist (H.M.) graded 337 slides is inconsistent with Table 1: the non-missing Pat. II counts sum to 336 (208+2+7+8+25+86=336, with 3 missing). One of these is wrong. Additionally, Pat. III is described as a 'random subset of 59 slides stratified by ISUP grades assigned by the first pathologist,' but the Pat. III grade distribution (11, 5, 13, 3, 6, 21) is very different from the Pat. I distribution (164, 7, 38, 59, 30, 41) and does not appear proportionally stratified. Specify the stratification scheme or correct the description.
  4. [Methods - Slide selection; Table 1] The text refers to '185 prostate cancer needle biopsy cases,' yet Table 1 lists 164 benign slides in the Pat. I column. If the cohort includes patients whose biopsies were benign (or benign slides from cancer patients), the wording is inaccurate and should be changed to 'prostate needle biopsy cases.' Clarify whether all patients had a cancer diagnosis or whether some slides/patients are benign.
minor comments (4)
  1. [Figure 2 caption] The caption reports 0.2266 micrometers per pixel for Hamamatsu, while Methods reports 0.22 um/pixel. Use a consistent value.
  2. [Table 1] The 'Missing' row in the ISUP grade columns is ambiguous: it mixes ungraded slides for Pat. II (3) with the 280 slides not assessed by Pat. III. Consider labeling this row clearly, e.g., 'Not assessed by this pathologist.'
  3. [Technical Validation] The text says 'tissue was segmented using deep-learning–based algorithms' but Code Availability states 'No custom code was used.' Clarify whether existing tools were used and name them, or remove the claim about deep-learning segmentation.
  4. [Data Records] The filename scheme is described as 'c<slide_id><a|b|none>.<ext>,' but the example is not shown. A concrete example filename would help users.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: descriptive dataset paper with no fitted prediction or derivation chain.

full rationale

The paper is a dataset descriptor: it reports the collection, scanning, and grading of 339 prostate biopsy slides from 185 patients. There is no derivation, model fitting, or prediction that could collapse into its inputs. The central gap claim ('No public datasets currently include data digitized with compact scanners or originating from underrepresented populations') is supported by an enumerated survey of prior public datasets (PANDA, SPROB20, SICAPv2, UKK/WNS, DiagSet, Gleason 2019, TCGA-PRAD, EMPaCT, TMAZ, etc.), not by an assumption that contains the conclusion. The self-citations in the manuscript (refs 11 and 28, both by members of the same group) are contextual citations for AI-in-pathology capability and scanner-induced variation; they are not load-bearing justifications for the dataset's novelty or validity, and no uniqueness theorem or ansatz is imported from them. The skeptically identified representativeness limitation (23% of eligible blocks lost, manual keyword-filter recall unassessed, no included-vs-excluded comparison) is a legitimate data-quality and external-validity concern, but it is not a circularity issue: no quantity is defined in terms of a target, no fitted parameter is renamed as a prediction, and no conclusion is forced by construction. Internal inconsistencies in slide counts (339 vs 1,017 images across three scanner copies) and abstract/body wording are editorial issues, not circular reasoning. Therefore the appropriate circularity score is 0.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper is a data descriptor; nothing is derived and no numbers are fitted, so the ledger contains only domain assumptions about the reference standard, archive recall, sampling adequacy, and scanner comparability. The dataset's significance depends on these assumptions, particularly archive completeness and representativeness after block loss, rather than on theoretical postulates.

assumptions (4)
  • domain assumption Gleason scoring and ISUP grades (refs 1-3) are a valid reference standard for prostate biopsy severity.
    The entire label layer of the dataset uses GS/ISUP converted to six ordinal ISUP classes; the paper cites rather than re-derives this standard (Background & Summary).
  • domain assumption Keyword search of the .docx archive for 'pros' exhaustively retrieves all prostate cases.
    Slide selection begins with this string match on 30,056 reports (Methods - Slide selection); no recall check is reported, so cases whose reports lack the substring are invisible to the pipeline.
  • domain assumption The 59-slide subset graded by the third pathologist is a usable third reference standard.
    Sampling is described only as a random subset stratified by ISUP assigned by the first pathologist; no stratification targets or precision analysis are provided (Methods - Reference standard protocol).
  • domain assumption WSIs from three scanners with different pixel resolutions are directly usable for cross-scanner robustness without calibration metadata.
    The paper advertises cross-scanner robustness evaluations but provides no color/calibration targets or equivalence analysis; resolutions differ (0.22, 0.25, 0.26 um/pixel) (Methods - Whole slide image scanning).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population." pith.science (2026). https://pith.science/paper/4EGZAJUE

@misc{pith2026251203854,
  author       = {Pith},
  title        = {Pith review of: The PAR dataset: Prostate biopsy whole slide images from an underrepresented Middle Eastern population},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4EGZAJUE}},
  note         = {Machine review of arXiv:2512.03854}
}
read the original abstract

Artificial intelligence (AI) is increasingly used in digital pathology. Publicly available histopathology datasets remain scarce, and those that do exist predominantly represent Western populations. Consequently, the generalizability of AI models to populations from less digitized regions, such as the Middle East, is largely unknown. This motivates the public release of our dataset to support the development and validation of pathology AI models across globally diverse populations. We present 1,017 whole slide images by digitizing 339 glass slides of prostate core needle biopsies from a consecutive series of 185 patients collected in Erbil, Iraq. The slides are associated with Gleason scores and International Society of Urological Pathology grades assigned independently by three pathologists. Scanning was performed using two high-throughput scanners (Leica and Hamamatsu) and one compact scanner (Grundium). All slides were de-identified and are provided in their native formats without further conversion. The dataset enables grading concordance analyses, color normalization, and cross-scanner robustness evaluations. The PAR dataset has been deposited to the BioImage Archive under accession number S-BIAD2323 (https:// doi.org/10.6019/S-BIAD2323).

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

34 extracted references · 17 canonical work pages

  1. [1]

    Gleason, D. F. Histologic grading of prostate cancer: a perspective. Hum. Pathol. 23 , 273–279 (1992). https://doi.org/10.1016/0046-8177(92)90108-f

  2. [2]

    I., Allsbrook, W

    Epstein, J. I., Allsbrook, W. C. Jr., Amin, M. B., Egevad, L. L. & ISUP Grading Committee The 2005 International Society of Urological Pathology (ISUP) Consensus Conference on Gleason grading of prostatic carcinoma. Am. J. Surg. Pathol. 29 , 1228–1242 (2005). https://doi.org/10.1097/01.pas.0000173646.99337.b1

  3. [3]

    Epstein, J. I. et al. The 2014 International Society of Urological Pathology (ISUP) Consensus Conference on Gleason grading of prostatic carcinoma: definition of grading patterns and proposal for a new grading system. Am. J. Surg. Pathol. 40 , 244–252 (2016). https://doi.org/10.1097/PAS.0000000000000530

  4. [4]

    Melia, J. et al. A UK-based investigation of inter- and intra-observer reproducibility of Gleason grading of prostatic biopsies. Histopathology 48 , 644–654 (2006). https://doi.org/10.1111/j.1365-2559.2006.02393.x

  5. [5]

    Egevad, L. et al . Standardization of Gleason grading among 337 European pathologists. Histopathology 62 , 247–256 (2013). https://doi.org/10.1111/his.12008

  6. [6]

    Ozkan, T.A. et al. Interobserver variability in Gleason histological grading of prostate cancer. Scand. J. Urol. 50 , 420–424 (2016). https://doi.org/10.1080/21681805.2016.1206619

  7. [7]

    Pantanowitz, L. et al. Twenty years of digital pathology: An overview of the road travelled, what is on the horizon, and the emergence of vendor-neutral archives. J. Pathol. Inform. 9 , 40 (2018). https://doi.org/10.4103/jpi.jpi_69_18

  8. [8]

    Campanella, G. et al. Clinical-grade computational pathology using weakly supervised deep learning on whole-slide images. Nat. Med. 25 , 1301–1309 (2019). https://doi.org/10.1038/s41591-019-0508-1

Show all 34 references
  1. [9]

    Ström, P. et al. Artificial intelligence for diagnosis and grading of prostate cancer in biopsies: a population-based, diagnostic study. Lancet Oncol. 21 , 222–232 (2020). https://doi.org/10.1016/S1470-2045(19)30738-7

  2. [10]

    Bulten, W. et al. Automated deep-learning system for Gleason grading of prostate cancer using biopsies: a diagnostic study. Lancet Oncol. 21 , 233–241 (2020). https://doi.org/10.1016/S1470-2045(19)30739-9

  3. [11]

    Mulliqi, N. et al. Foundation models: a panacea for artificial intelligence in pathology? Preprint at https://arxiv.org/abs/2502.21264 (2025)

  4. [12]

    Bulten, W. et al. Artificial intelligence for diagnosis and Gleason grading of prostate cancer: the PANDA challenge. Nat. Med. 28 , 154–163 (2022). https://doi.org/10.1038/s41591-021-01620-2

  5. [13]

    Walhagen, P. et al. Spear Prostate Biopsy 2020 (SPROB20). AIDA Data Hub https://doi.org/10.23698/aida/sprob20 (2020)

  6. [14]

    SICAPv2 – Prostate whole slide images with Gleason grades annotations

    Silva-Rodríguez, J. SICAPv2 – Prostate whole slide images with Gleason grades annotations. Mendeley Data . https://doi.org/10.17632/9xxm58dvs3.1 (2020)

  7. [15]

    A., Molina, R

    Silva-Rodríguez, J., Colomer, A., Sales, M. A., Molina, R. & Naranjo, V. Going deeper through the Gleason scoring scale: an automatic end-to-end system for histology prostate grading and cribriform pattern detection. Comput. Methods Programs Biomed. 195 , 105637 (2020). https:...

  8. [16]

    Tolkach, Y. et al. An international multi-institutional validation study of the algorithm for prostate cancer detection and Gleason grading. NPJ Precis. Oncol. 7 , 77 (2023) https://doi.org/10.1038/s41698-023-00424-6

  9. [17]

    Koziarski, M. et al. DiagSet: a dataset for prostate cancer histopathological image classification. Sci. Rep. 14 , 6780 (2024). https://doi.org/10.1038/s41598-024-52183-4

  10. [18]

    Huo, X. et al. A comprehensive AI model development framework for consistent Gleason grading. Commun. Med. 4 , 84 (2024). https://doi.org/10.1038/s43856-024-00502-1

  11. [19]

    The molecular taxonomy of primary prostate cancer

    The Cancer Genome Atlas Research Network. The molecular taxonomy of primary prostate cancer. Cell 163 , 1011–1025 (2015). https://doi.org/10.1016/j.cell.2015.10.025

  12. [20]

    Heath, A. P. et al. The NCI Genomic Data Commons. Nat. Genet. 53 , 257–262 (2021). https://doi.org/10.1038/s41588-021-00791-5

  13. [21]

    Dataset EMPaCT TMA

    Karkampouna, S., & Kruithof-de Julio, M. Dataset EMPaCT TMA . Zenodo . https://doi.org/10.5281/zenodo.10066853 (2023)

  14. [22]

    Zhong, Q. et al. A curated collection of tissue microarray images and clinical outcome data of prostate cancer patients. Sci. Data. 4 , 170014 (2017). https://doi.org/10.1038/sdata.2017.14

  15. [23]

    Arvaniti, E. et al. Replication Data for: Automated Gleason grading of prostate cancer tissue microarrays via deep learning. Harvard Dataverse. https://doi.org/10.7910/DVN/OCYCMP (2018)

  16. [24]

    Gleason 2019: Automatic Gleason grading of prostate cancer in digital histopathology

    MICCAI 2019 Gleason Challenge Organizers. Gleason 2019: Automatic Gleason grading of prostate cancer in digital histopathology. MICCAI 2019 Conference - Challenge Website https://gleason2019.grand-challenge.org/ (2019; accessed 28 November 2025)

  17. [25]

    Chen, R.J. et al. Towards a general-purpose foundation model for computational pathology. Nat. Med. 30 , 850–862 (2024). https://doi.org/10.1038/s41591-024-02857-3

  18. [26]

    Vorontsov, E. et al. A foundation model for clinical-grade computational pathology and rare cancers detection. Nat. Med. 30 , 2924–2935 (2024). https://doi.org/10.1038/s41591-024-03141-0

  19. [27]

    Duenweg, S. R. et al. Whole slide imaging (WSI) scanner differences influence optical and computed properties of digitized prostate cancer histology. J. Pathol. Inform. 14 , 100321 (2023). https://doi.org/10.1016/j.jpi.2023.100321

  20. [28]

    Ji, X. et al. Physical color calibration of digital pathology scanners for robust artificial intelligence–assisted cancer diagnosis. Mod. Pathol . 38 , 100715 (2025). https://doi.org/10.1016/j.modpat.2025.100715

  21. [29]

    A benchmarking crisis in biomedical machine learning

    Mahmood, F. A benchmarking crisis in biomedical machine learning. Nat. Med. 31 , 1060 (2025). https://doi.org/10.1038/s41591-025-03637-3

  22. [30]

    Marée, R. et al. Collaborative analysis of multi-gigapixel imaging data using Cytomine. Bioinformatics 32 , 1395–1401 (2016). https://doi.org/10.1093/bioinformatics/btw013

  23. [31]

    & Satyanarayanan, M

    Goode, A., Gilbert, B., Harkes, J., Jukic, D. & Satyanarayanan, M. OpenSlide: A vendor-neutral software foundation for digital pathology. J. Pathol. Inform. 4 , 27 (2013). https://doi.org/10.4103/2153-3539.119005

  24. [32]

    ASAP – Automated Slide Analysis Platform

    Litjens, G. ASAP – Automated Slide Analysis Platform. GitHub https://computationalpathologygroup.github.io/ASAP (2017)

  25. [33]

    Bankhead, P. et al. QuPath: Open source software for digital pathology image analysis. Sci. Rep. 7 , 16878 (2017). https://doi.org/10.1038/s41598-017-17204-5

  26. [34]

    Lu, M.Y. et al. Data-efficient and weakly supervised computational pathology on whole-slide images. Nat. Biomed. Eng. 5 , 555–570 (2021). https://doi.org/10.1038/s41551-020-00682-w Figure 1. Diagram illustrating dataset curation. The diagram outlines the process of identifying...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.