{"id":"71d339da-3b43-4540-a1e0-cbe186b84880","arxiv_id":"2512.03854","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new public prostate-biopsy WSI dataset from 185 Iraqi patients with triple pathologist Gleason/ISUP labels and three-scanner digitization including a compact scanner.","lead":"The authors describe a new public dataset of 1,017 digitized prostate biopsy whole-slide images from 185 patients in Erbil, Iraq, each scanned with three different machines and graded independently by three pathologists. It exists to close a gap: most public pathology data comes from Western populations, so AI models are untested on Middle Eastern tissue and on cheaper compact scanners.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed representativeness of the 185-patient series is unverified: 23% of eligible blocks were lost and no comparison of included vs excluded cases is provided.","rationale":"The reader's weakest assumption correctly identifies the representativeness of the 185-patient series as load-bearing. The paper's stated purpose is to provide a dataset from an underrepresented population that enables cross-population and cross-scanner validation, and the Background explicitly claims the series is 'representative of the cases encountered in routine clinical practice between 2013-2024.' This claim rests on the cohort being selected without systematic bias. The Methods show two plausible bias sources: an unvalidated keyword search (recall unknown) and a 23% block-loss rate concentrated in early years. No evidence is given that the excluded cases are similar to the included ones. Without this comparison, the dataset may be a convenience sample, undermining the generalizability claims that motivate its release. This is a substantive scientific concern, not a textual error. The reader's verdict of CONDITIONAL is therefore appropriate: the dataset may be valuable, but its representativeness must be demonstrated before the generalization claims can be accepted. I agree with the reader, and no adjustment to the verdict is needed; the same conditional bar applies. The internal inconsistencies (abstract vs. body counts and DOI) are also present, but they are secondary because they can be corrected without changing the underlying data; the representativeness gap cannot be fixed by editing the manuscript alone.","tokens_in":12437,"tokens_out":3477,"duration_ms":32006,"concrete_test":"Retrieve the original pathology reports for the 54 excluded cases (they were identified in the archive, so reports should exist) and compare age at biopsy, reported Gleason score/ISUP grade, and year of biopsy against the 185 included cases using appropriate statistical tests (e.g., chi-square for grade, t-test/Wilcoxon for age/year). Additionally, audit the recall of the 'pros' keyword filter: take a random sample of 100 reports that did not match the keyword, manually classify whether they are prostate-related, and estimate the number of missed prostate biopsies. If the excluded group differs significantly from the included group in grade distribution or temporal spread, the consecutive/representative claim is falsified. If the keyword filter misses a substantial fraction of prostate cases, the filtering step may itself introduce selection bias.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that this is a 'consecutive series representative of routine clinical practice between 2013–2024' rests on the cohort being an unbiased sample of prostate needle biopsies at PAR Hospital. The curation pipeline (Methods - Slide selection) introduces at least two potential selection mechanisms: (1) a manual keyword search for 'pros' in .docx reports of unknown recall, and (2) exclusion of 54 of 239 eligible cases (23%) because FFPE blocks were lost, 'mostly from the early years of the collection.' The paper provides no comparison of the included 185 cases with the excluded 54 on age, reported Gleason score, year, or tumor burden, and no audit of the keyword filter's sensitivity. If the lost blocks or missed reports skew toward particular grades, years, or patient demographics, the dataset becomes a convenience sample rather than a consecutive series. This directly affects the paper's motivation—cross-population generalization—because models trained on a biased subset may not represent the broader Iraqi or Middle Eastern population that the dataset is intended to inform. The issue is not merely a documentation gap; it is a load-bearing assumption about the dataset's external validity. The internal inconsistencies in WSI counts and accession status (abstract vs. body) are concerning but fixable; the representativeness question requires additional data that may not exist, and cannot be resolved by text cleanup alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the PAR dataset: 339 glass slides from prostate core needle biopsies of 185 patients from PAR Hospital in Erbil, Iraq, digitized with three whole-slide scanners (Leica Aperio GT450, Hamamatsu NanoZoomer HT 2.0, and the compact Grundium Ocus40), yielding 1,017 WSI files in native formats. Slide-level Gleason scores and ISUP grades were assigned independently by three pathologists, with one pathologist grading all slides, one grading 337 slides, and one grading a stratified subset of 59. The authors claim this is the first public dataset with compact-scanner images and from an underrepresented Middle Eastern population, enabling cross-scanner, cross-population, and multi-pathologist validation studies. The manuscript describes curation, scanning, annotation, and technical validation steps.","tokens_in":12593,"tokens_out":4239,"duration_ms":37773,"significance":"If the dataset is released as described, it would be a genuinely valuable community resource. It addresses two documented gaps: public prostate biopsy WSIs from non-Western populations and WSIs from a compact scanner. The three-scanner design with independent multi-pathologist labels is well suited for color normalization and cross-scanner robustness studies. The authors are transparent about limitations (no pixel-level annotations, generalist vs. uropathologist differences) and the curation workflow is described in checkable detail. The dataset, once deposited, would support reproducibility research and help benchmark AI models across populations.","major_comments":[{"comment":"The central claim that the 185-patient series is a 'consecutive series representative of routine clinical practice between 2013-2024' is not supported by the evidence provided. Fifty-four of 239 eligible cases (23%) were excluded because FFPE blocks were lost, 'mostly from the early years of the collection.' No comparison is given between included and excluded cases on age, Gleason score, year, or tumor burden, and no audit is reported of the recall of the manual keyword search for 'pros' in .docx reports. If lost blocks or missed reports are grade- or year-dependent, the dataset is a convenience sample, not a consecutive series. Please provide any available comparison using the archived reports, or explicitly weaken the representativeness claim and discuss the potential bias.","section":"Methods - Slide selection; Background & Summary"},{"comment":"There are load-bearing inconsistencies in the dataset description. The abstract states '1,017 whole slide images' and gives BioImage Archive accession S-BIAD2323; the body states '339 digitized prostate core needle biopsies' and gives accession 'TBA' with DOI 'TBA.' If 1,017 is the total number of WSI files (339 slides x 3 scanners), this should be stated explicitly and consistently throughout. The accession number is essential for a data paper and must be reconciled.","section":"Abstract vs. Data Records / Data Availability"},{"comment":"The claim that the second pathologist (H.M.) graded 337 slides is inconsistent with Table 1: the non-missing Pat. II counts sum to 336 (208+2+7+8+25+86=336, with 3 missing). One of these is wrong. Additionally, Pat. III is described as a 'random subset of 59 slides stratified by ISUP grades assigned by the first pathologist,' but the Pat. III grade distribution (11, 5, 13, 3, 6, 21) is very different from the Pat. I distribution (164, 7, 38, 59, 30, 41) and does not appear proportionally stratified. Specify the stratification scheme or correct the description.","section":"Table 1 and Reference standard protocol"},{"comment":"The text refers to '185 prostate cancer needle biopsy cases,' yet Table 1 lists 164 benign slides in the Pat. I column. If the cohort includes patients whose biopsies were benign (or benign slides from cancer patients), the wording is inaccurate and should be changed to 'prostate needle biopsy cases.' Clarify whether all patients had a cancer diagnosis or whether some slides/patients are benign.","section":"Methods - Slide selection; Table 1"}],"minor_comments":[{"comment":"The caption reports 0.2266 micrometers per pixel for Hamamatsu, while Methods reports 0.22 um/pixel. Use a consistent value.","section":"Figure 2 caption"},{"comment":"The 'Missing' row in the ISUP grade columns is ambiguous: it mixes ungraded slides for Pat. II (3) with the 280 slides not assessed by Pat. III. Consider labeling this row clearly, e.g., 'Not assessed by this pathologist.'","section":"Table 1"},{"comment":"The text says 'tissue was segmented using deep-learning–based algorithms' but Code Availability states 'No custom code was used.' Clarify whether existing tools were used and name them, or remove the claim about deep-learning segmentation.","section":"Technical Validation"},{"comment":"The filename scheme is described as 'c<slide_id><a|b|none>.<ext>,' but the example is not shown. A concrete example filename would help users.","section":"Data Records"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially valuable and the description is mostly workable, but the representativeness claim is the key load-bearing assertion and it is currently unsupported by the missing-case analysis. If the authors cannot access the excluded cases' data, they should reframe the paper around a documented consecutive-block-availability cohort. The abstract/body inconsistencies (1,017 vs 339; S-BIAD2323 vs TBA) must be fixed before publication, as a data paper is defined by its accession. I believe this is within scope for a data journal and can be brought to an acceptable level with a major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: this is the first public prostate biopsy WSI resource from an underrepresented Middle Eastern population, and the first public set to include a compact scanner (Grundium Ocus40) alongside two high-throughput scanners, with three independent pathologist label sets. The paper deserves a serious look.\n\nWhat it does well: the survey of 12 existing public prostate pathology datasets is concrete and shows the gap is real. The curation workflow is describable and mostly reproducible: 30,056 archived cases, keyword filter, addenda merging, block retrieval, recutting all slides, scanning at 40x in native formats. The decision to recut fresh H&E slides rather than scan faded archival slides is sensible for image quality. The authors are transparent about what is not there — no pixel-level annotations, no consensus label, only a random stratified subset graded by the uropathologist. That is honest and appropriate for a data descriptor.\n\nThe soft spots are real and proportionately sized. The biggest is the representativeness claim. The paper calls this a 'consecutive series ... representative of routine clinical practice between 2013–2024,' but 54 of 239 eligible cases (23%) were dropped because FFPE blocks were lost, mostly from the early years. No comparison of included vs excluded cases on age, year, grade, or tumor burden is provided, and the recall of the keyword 'pros' search on .docx reports is unassessed. That means the series could be a convenience sample biased toward later years or particular grades. The gap this dataset fills — cross-population generalizability — depends on the sample being reasonably representative, so this is not a minor omission. It can be fixed in part by reframing the claim as 'a cohort of available blocks from a consecutive archive' and, if possible, reporting what is known about the lost cases.\n\nSecond, the manuscript contradicts itself in load-bearing places. The arXiv abstract says 1,017 WSIs (339 × 3) and a deposited BioImage Archive accession with DOI; the full-text abstract and Data Records say 339 WSIs and TBA for accession and DOI. Table 1's pathologist II column sums to 339, not 337 as stated in the text; the missing-count rows in Table 1 need careful checking. None of this is fatal to the resource, but a data descriptor that cannot be used as advertised is not yet a dependable citation.\n\nWho this is for: researchers working on prostate cancer grading, cross-scanner robustness, stain normalization, or AI validation in non-Western populations. They will want this dataset once the deposit is confirmed and the metadata is harmonized. A referee should engage with it — the resource is important enough to deserve careful review — but the authors should be required to resolve the inconsistencies and soften the representativeness language before publication.","headline":"Genuinely new and useful dataset — first Middle Eastern prostate biopsy WSI set with triple grading and a compact scanner — but the 'consecutive representative' claim is unsupported as written and the manuscript needs internal consistency fixes before the deposit goes live.","tokens_in":13273,"tokens_out":2621,"would_cite":true,"duration_ms":23778,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper releases a public dataset of 339 prostate biopsy whole-slide images from 185 patients in Erbil, Iraq, scanned with three scanners including a compact one, and graded independently by three pathologists.","keywords":["prostate cancer","whole slide imaging","Gleason grading","ISUP grade","digital pathology dataset","multi-scanner","Middle Eastern population","compact scanner"],"falsifier":"Compare the age, Gleason score distribution, and biopsy year of the 185 included cases against the 54 excluded cases; a statistically significant difference would show the remaining cohort is not a consecutive, representative series. Re-running the keyword search on the original archive and checking for missed prostate needle biopsy reports would also test the recall of the manual pipeline.","tokens_in":12215,"feed_emoji":"🔬","tokens_out":5153,"duration_ms":45093,"temperature":0.7,"pith_summary":"This paper aims to establish a new public resource: 339 digitized prostate core needle biopsy whole-slide images from 185 patients seen at a hospital in Erbil, Iraq, between 2013 and 2024. The slides were freshly recut from archived tissue blocks, scanned with two high-throughput scanners and one compact scanner, and given Gleason and ISUP grades independently by three pathologists. The authors argue that no public dataset currently combines an underrepresented Middle Eastern population, compact-scanner capture, multiple scanners, and multi-pathologist reference labels. If the dataset is representative, it would give AI pathology researchers a substrate for cross-population and cross-scanner validation that Western-only, single-scanner datasets cannot provide.","feed_headline":"First public prostate WSI dataset from Middle East spans 3 scanners","feed_subtitle":"Three pathologists graded 339 slides; a compact scanner opens low-cost AI validation for under-digitized regions.","key_machinery":"The central object is the dataset itself: 339 WSIs in native scanner formats (.svs and .ndpi), organized so each slide has three scanner versions and linked to a single annotations table with independent Gleason/ISUP grades from three pathologists. The design work is carried by the pairing of multi-scanner capture (two high-throughput scanners and the Grundium Ocus40 compact scanner) with multi-pathologist independent reference grading; this combination is what makes cross-scanner and cross-rater analyses possible. A supporting mechanism is the decision to recut new H&E slides from archived FFPE blocks rather than scan faded original slides, which standardizes tissue quality across the colle","core_discovery":"The paper's central claim is that the PAR dataset is the first publicly available prostate biopsy whole-slide image collection from an underrepresented Middle Eastern population that is digitized with multiple scanners — including a compact, low-cost scanner — and labeled with slide-level Gleason scores and ISUP grades assigned independently by three pathologists. Each of the 339 slides exists in three scanned versions (one per scanner), and labels follow a fixed dictionary of ten Gleason classes and six ISUP grades. The authors position this as a direct response to three gaps in existing public datasets: single-reference grading, Western-only populations, and single-scanner capture. They fu","pith_inferences":["Because the slides were recut from archived blocks rather than scanned from the original diagnostic slides, the images may not reflect the exact tissue section that produced the clinical report; the labels come from the new sections, which could differ slightly in grade from the original clinical diagnosis.","The dataset's value as a population-representative benchmark depends on the excluded cases; a reader could re-examine the hospital archive to test whether the keyword filter and the 54 lost blocks introduced selection bias.","If the compact-scanner images show systematic color or focus shifts, they could serve as a natural testbed for stain normalization and domain adaptation methods rather than purely as additional data."],"forward_implications":["If the dataset is representative, AI models trained on Western prostate biopsies can be tested for the first time on a Middle Eastern population, revealing whether grade distributions and tissue appearance shift across populations.","The inclusion of compact-scanner images allows direct evaluation of whether low-cost scanners are adequate for AI validation in clinics that have not yet digitized.","With three independent graders, the dataset supports inter-observer agreement studies and the construction of consensus labels under any explicit rule (majority, highest grade, etc.).","Native-format 40x WSIs from three scanners enable color normalization and stain-robustness benchmarking that single-scanner datasets cannot offer."],"fun_headline_variants":["Middle East's first public prostate WSI set, 3 scanners each","Triple-graded prostate slides from Iraq: first public Middle East WSI set","First prostate WSI dataset from Iraq includes compact scanner scans","New public prostate WSI dataset from Middle East uses 3 scanners","3 pathologists, 3 scanners: first public prostate WSI set from Iraq"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The dataset is claimed to represent a consecutive series of routine prostate biopsies from 2013–2024, but 54 of 239 eligible cases were dropped because their tissue blocks were lost, and no evidence is given that the remaining 185 match the excluded cases in age, grade, or year.","fun_headline_variants_meta":{"raw":{"variants":["Middle East's first public prostate WSI set, 3 scanners each","Triple-graded prostate slides from Iraq: first public Middle East WSI set","First prostate WSI dataset from Iraq includes compact scanner scans","New public prostate WSI dataset from Middle East uses 3 scanners","3 pathologists, 3 scanners: first public prostate WSI set from Iraq"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000729,"raw_usage":{"total_tokens":3093,"prompt_tokens":725,"completion_tokens":2368,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":2272}},"tokens_in":469,"tokens_out":2368,"duration_ms":14336,"temperature":1.0,"reasoning_tokens":2272,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T18:41:01.944597+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the age, Gleason score distribution, and biopsy year of the 185 included cases against the 54 excluded cases; a statistically significant difference would show the remaining cohort is not a consecutive, representative series. Re-running the keyword search on the original archive and checking for missed prostate needle biopsy reports would also test the recall of the manual pipeline.","supporting_citations":[],"review_version":1}