{"id":"0dbb5b74-a369-488b-8ed6-4f0a926db720","arxiv_id":"2508.03762","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"A prespecified protocol to test whether an AI system can match expert readers for prostate cancer detection on MRI across a global, multiethnic population.","lead":"This paper presents a protocol for a global, multiethnic study to validate an AI model, PI-CAI-2B, for prostate cancer detection on MRI against standard-of-care diagnoses. The study will test whether the AI can achieve diagnostic interchangeability across 22,481 MRI exams from 46 cities in 22 countries.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Non-inferiority margin of 0.05 comes from the PI-CAI observer study; the abstract provides no evidence that reader variability in the external multiethnic screening and diagnostic cohorts matches that estimate, making the interchangeability test potentially mis-calibrated.","rationale":"The paper is a study protocol and does not report results, so the scientific claim is not yet testable; the reader's UNVERDICTED assessment remains appropriate. Within the protocol, the most load-bearing design choice is the non-inferiority margin of 0.05, because it directly determines whether diagnostic interchangeability can be claimed. The abstract justifies this margin solely through reader estimates from the PI-CAI observer study, and no information is provided to show that those estimates generalize to the external multiethnic screening and diagnostic populations. This is not an internal inconsistency but a potentially invalid transportability assumption. If the external reader variability is greater than that in PI-CAI, the margin is too strict and a truly interchangeable AI could be rejected; if it is smaller, the margin is too lenient and an inferior AI could be accepted. The proposed concrete test uses existing PI-CAI data and published cohort characteristics to evaluate whether the margin holds under case-mix differences; this would settle whether the concern actually lands. Until that analysis is performed or the full protocol demonstrates margin calibration, the protocol's central hypothesis is unverified but not contradicted. Therefore, the verdict remains UNVERDICTED, and the concern should be weighed when the full protocol is reviewed.","tokens_in":895,"tokens_out":7010,"duration_ms":81779,"concrete_test":"Re-analyze the PI-CAI observer study dataset to assess margin transportability: for each external cohort (STHLM3-MRI, IP1-PROSTAGRAM, PRIME), obtain the published PI-RADS score distribution and cancer prevalence, then resample the 400 PI-CAI observer cases to match each cohort's case-mix. For each resampled set, compute the pooled inter-reader agreement proportion with the reference standard and construct a bootstrap 95% confidence interval around the observed variability. If the upper bound of the confidence interval for any external cohort exceeds 0.05, the prespecified margin is not conservative for that setting and the protocol should include a margin calibration step or a sensitivity analysis before the primary hypothesis is tested.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The primary endpoint is the proportion of AI assessments agreeing with a composite reference standard, and the protocol declares an absolute non-inferiority margin of 0.05 based on reader estimates from the PI-CAI observer study (62 radiologists, 400 cases). This margin is the sole quantitative basis for the interchangeability claim. For the claim to be valid, the inter-reader variability measured in that study must be representative of the variability in the external cohorts (STHLM3-MRI, IP1-PROSTAGRAM, PRIME; 2010 cases, 20 cities, 12 countries). Neither the abstract nor the available text provides evidence for this transportability. Reader variability is known to depend on disease prevalence (screening versus primary diagnosis), PI-RADS distribution, image quality, and reader expertise; these are likely to differ markedly across the external settings. If the external reader variability exceeds the PI-CAI estimate, the 0.05 margin understates the clinically acceptable difference and the study may falsely reject an interchangeable AI; if the external variability is smaller, the margin overstates the difference and the study may falsely accept a non-interchangeable AI. No sensitivity analysis or margin calibration to the external cohorts is described or referenced. Because the study is confirmatory and the conclusion hinges entirely on this threshold, this unverified transportability assumption is the most load-bearing element of the protocol.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This abstract-only manuscript describes a confirmatory, intercontinental study protocol for validating the PI-CAI-2B AI system for detecting Gleason grade group ≥2 prostate cancer on MRI. The plan includes training and internal testing on 20,471 examinations from 26 cities in 14 countries, and external testing on 2,010 examinations from population-based screening (STHLM3-MRI, IP1-PROSTAGRAM) and primary diagnostic (PRIME) settings across 20 cities in 12 countries. The primary endpoint is the proportion of AI assessments in agreement with the standard of care (histopathology or expert radiologist consensus) at a PI-RADS ≥3 (primary diagnosis) or ≥4 (screening) cut-off, with a prespecified non-inferiority margin of 0.05 derived from reader variability in the PI-CAI observer study. Secondary endpoints include AUROC stratified by imaging quality, age, and ethnicity. The protocol emphasizes prespecification and external validation, but the abstract alone provides no results and only limited details on the analysis plan.","tokens_in":1181,"tokens_out":2892,"duration_ms":34210,"significance":"If successfully executed, this study would provide a large, multiethnic, externally validated assessment of an AI tool in both screening and primary diagnostic settings, with a prespecified statistical analysis plan. The use of separate external trials (STHLM3-MRI, IP1-PROSTAGRAM, PRIME) and independent centers is a genuine strength for testing generalizability, and the focus on interchangeability (rather than superiority) is clinically relevant. However, the confirmatory interpretation hinges entirely on the non-inferiority margin, and the abstract provides no evidence that the reader-variability estimate used to set this margin is transportable to the external populations. This is a load-bearing concern that must be addressed before the protocol can be considered methodologically sound.","major_comments":[{"comment":"The non-inferiority margin of 0.05 is justified by reader estimates from the PI-CAI observer study (62 radiologists, 400 cases). The abstract provides no evidence that this estimate of reader variability is representative of the external testing cohorts (STHLM3-MRI, IP1-PROSTAGRAM, PRIME; 2,010 cases, 20 cities, 12 countries), which differ in disease prevalence (screening vs. primary diagnosis), PI-RADS distribution, image quality, and reader expertise. If external reader variability exceeds the PI-CAI estimate, the margin is too narrow and could falsely reject an interchangeable AI; if it is smaller, the margin is too wide and could falsely accept a non-interchangeable AI. The protocol should either justify the transportability of the margin, pre-specify sensitivity analyses across plausible margins, or calibrate the margin using reader data from the external sites.","section":"Abstract (primary endpoint)"},{"comment":"The primary endpoint is described as 'the proportion of AI-based assessments in agreement with the standard of care diagnoses,' but the abstract does not define the pairwise comparison for interchangeability. It is unclear whether the analysis uses a two-sided confidence interval for the difference in proportions, a paired test, or a specific equivalence/non-inferiority test statistic, and how the PI-RADS ≥3/≥4 cut-off is integrated into the endpoint. The protocol must operationalize the endpoint and the non-inferiority test explicitly to be confirmatory.","section":"Abstract (primary endpoint)"},{"comment":"The abstract states that 20,471 examinations are used for 'training and internal testing' and 2,010 for external testing, but it does not state whether the AI model is fully locked before any external validation occurs, or whether the external data are used in any model selection or threshold adjustment. To maintain the confirmatory nature of the study, the protocol should specify that the model was frozen at internal testing completion and that external data are accessed only once for the final evaluation.","section":"Abstract (methods)"}],"minor_comments":[{"comment":"Please expand the trial acronyms STHLM3-MRI, IP1-PROSTAGRAM, and PRIME at first mention, and include citations to the original trial protocols or publications.","section":"Abstract (background)"},{"comment":"The composite reference standard is described as histopathology 'if available, or at least two expert urogenital radiologists in consensus.' The hierarchy and adjudication rules for cases where histopathology is unavailable or discordant with radiology should be specified, including how missing data are handled.","section":"Abstract (primary endpoint)"},{"comment":"The stratification of AUROC by ethnicity is a valuable bias assessment, but the abstract does not specify the ethnicity categories or whether they are self-reported, registry-based, or imputed; these details should be provided in the full protocol.","section":"Abstract (secondary endpoints)"},{"comment":"The phrase 'reader estimates derived from the PI-CAI observer study' is vague; the protocol should state the exact quantity (e.g., the standard deviation of reader agreements or the width of a 95% limits-of-agreement interval) and how it maps to an absolute 0.05 margin.","section":"Abstract (margin)"},{"comment":"As an abstract-only protocol, there are no results to report; if this is intended for a journal that expects protocol papers, please confirm that the full protocol is available as supplementary material or under a repository, so that the prespecified statistical plan can be audited.","section":"Abstract (results)"}],"recommendation":"major_revision","confidential_remarks":"This is an abstract-only submission, so the referee cannot verify whether the full protocol already addresses the margin-transportability concern or the endpoint definition. The stress-test concern about the non-inferiority margin is legitimate and load-bearing; the authors should be required to provide the full protocol or to revise the abstract to include the missing validation of the margin. The choice of external cohorts is strong, and the prespecified analysis plan is a plus, but the confirmatory claim is not yet defensible from the submitted text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a protocol for validating PI-CAI-2B, an existing AI model, against standard of care for detecting GG>=2 prostate cancer in 2,010 external cases across screening and primary diagnostic settings. What is actually new is the scale and diversity: 20 cities in 12 countries, including population-based screening trials and a primary care trial. The design has real strengths. The external cohorts are independent of training data, the reference standard is histopathology or consensus radiology (not AI), and the primary endpoint is prespecified. The secondary analysis stratified by imaging quality, age, and ethnicity is a sensible way to probe for bias. This is a serious, confirmatory protocol, not a fishing expedition.\n\nThe soft spot is exactly the stress-test concern: the absolute margin of 0.05 for interchangeability rests on reader variability from the PI-CAI observer study (62 radiologists, 400 cases). The abstract gives no evidence that this estimate transports to the external screening and primary care populations, where disease prevalence, PI-RADS distribution, reader expertise, and image quality will differ. If the margin is mis-calibrated, the interchangeability test could be too lenient or too strict. This is not a fatal flaw in the abstract alone — a full protocol might justify the margin with sensitivity analyses or external reader data — but as presented, it is the weakest link.\n\nTwo other things are worth saying. First, this is an abstract-only submission, so I cannot verify the training data handling, the model development, or the actual statistical analysis plan. The reader's \"unverdictable\" score is appropriate. Second, the protocol is explicitly an extension of the PI-CAI program; novelty is in the breadth of validation, not in new methods. That is fine for a confirmatory study, but it should be framed as such.\n\nWho is this for? Anyone running or evaluating AI validation studies, especially for prostate MRI, and anyone interested in the standards for claiming AI interchangeability with human readers. The study deserves a serious referee — the design is largely sound, the question is important, and the external cohort is genuinely valuable. My recommendation: send the full protocol to peer review, and make the margin calibration the first question for the authors.","headline":"A well-designed confirmatory protocol for a large-scale, multiethnic external validation of an existing prostate MRI AI, but the non-inferiority margin derived from the PI-CAI observer study is a load-bearing assumption that the abstract does not defend.","tokens_in":1871,"tokens_out":1645,"would_cite":false,"duration_ms":21780,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A prespecified protocol will test whether the PI-CAI-2B AI system is diagnostically interchangeable with the standard of care in detecting Gleason grade group ≥2 prostate cancer on MRI, within an absolute margin of 0.05.","keywords":["prostate cancer MRI","artificial intelligence","diagnostic interchangeability","non-inferiority margin","Gleason grade group","screening","multiethnic","study protocol"],"falsifier":"If, in the 2,010 external cases, the observed proportion of AI agreement with the standard of care falls more than 0.05 below the standard-of-care agreement rate at the prespecified PI-RADS cutoff, the primary hypothesis of interchangeability is rejected. A reader could also look for whether the AI's AUROC drops significantly in any single country or ethnic stratum, which would undercut the global claim.","tokens_in":745,"feed_emoji":"🩺","tokens_out":3352,"duration_ms":34298,"temperature":0.7,"pith_summary":"This paper is a study protocol, not a result. It prespecifies a confirmatory, intercontinental test of whether the PI-CAI-2B artificial intelligence system can replace the standard of care in detecting clinically significant prostate cancer on MRI. The primary endpoint is the proportion of AI assessments that agree with expert histopathology or consensus radiology reads, and the hypothesis is that the AI is within an absolute margin of 0.05 of that standard in external cohorts. The study enrolls 22,481 MRI examinations from screening and primary diagnostic settings across Europe, the Americas, Asia, and Australia. If the margin holds, the AI would be considered diagnostically interchangeable with human readers for both screening and primary diagnosis.","feed_headline":"Prostate-cancer AI heads to a 0.05 interchangeability test","feed_subtitle":"22,481 scans across 22 countries will judge whether the AI matches radiologists at screening and primary diagnosis.","key_machinery":"The central object is the prespecified non-inferiority margin of 0.05, together with the PI-CAI-2B model. The margin is sized from reader variability measured in the PI-CAI observer study (62 radiologists reading 400 cases), so that 'interchangeability' is defined as the AI performing no worse than a reasonable human reader. The external cohorts (STHLM3-MRI, IP1-PROSTAGRAM, PRIME) provide the screening and primary-diagnosis settings where the claim must hold.","core_discovery":"The paper's central claim is a prespecified hypothesis: PI-CAI-2B achieves diagnostic interchangeability with the standard of care in detecting Gleason grade group ≥2 prostate cancer on MRI, defined as agreement within an absolute margin of 0.05 at the PI-RADS ≥3 (primary diagnosis) or ≥4 (screening) cutoff. The discovery, for now, is the design of a test that could establish this: 2,010 external test cases from population-based screening and primary diagnostic trials, with the margin derived from a 62-radiologist observer study that estimated reader variability. The authors assert that passing this threshold would justify using the AI as an independent reader in organized screening and primary diagnosis in a global, multiethnic population.","pith_inferences":["The 0.05 margin is a single global threshold, but acceptable interchangeability might need to be stricter in screening settings where lower disease prevalence means the same absolute error yields more false positives per cancer detected.","If the model passes, a natural next test is a prospective randomized comparison to biopsy decision-making, since retrospective agreement with histopathology does not directly measure patient outcomes.","The study's separation of training from external testing centers is a stronger form of validation than random splitting and could serve as a model for other AI diagnostic approvals."],"forward_implications":["If the margin holds, PI-CAI-2B could act as a first-line reader in screening programs, reducing radiologist workload.","The same threshold would apply across ethnic groups and geographic regions, supporting global deployment.","Passing the primary endpoint would establish a template for approving AI as an independent diagnostic reader using a reader-variability-derived margin.","Secondary AUROC stratification may reveal whether the model's performance varies by imaging quality, age, or ethnicity."],"supporting_citations":[],"fun_headline_variants":["AI vs radiologists: 0.05 margin test on 22k scans","Global prostate MRI AI faces 0.05 interchangeability bar","22,481 scans to test if AI matches radiologists on prostate MRI","Interchangeability test: AI vs standard of care on 2k external scans","Global multiethnic MRI study: AI heads to 0.05 interchangeability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 0.05 margin, estimated from a 62-radiologist observer study, is assumed to represent how much disagreement between expert readers would be clinically acceptable in the external multiethnic centers where the AI will be tested.","fun_headline_variants_meta":{"raw":{"variants":["AI vs radiologists: 0.05 margin test on 22k scans","Global prostate MRI AI faces 0.05 interchangeability bar","22,481 scans to test if AI matches radiologists on prostate MRI","Interchangeability test: AI vs standard of care on 2k external scans","Global multiethnic MRI study: AI heads to 0.05 interchangeability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000971,"raw_usage":{"total_tokens":4199,"prompt_tokens":1089,"completion_tokens":3110,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":705,"completion_tokens_details":{"reasoning_tokens":3011}},"tokens_in":705,"tokens_out":3110,"duration_ms":22302,"temperature":1.0,"reasoning_tokens":3011,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:56:30.828019+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If, in the 2,010 external cases, the observed proportion of AI agreement with the standard of care falls more than 0.05 below the standard-of-care agreement rate at the prespecified PI-RADS cutoff, the primary hypothesis of interchangeability is rejected. A reader could also look for whether the AI's AUROC drops significantly in any single country or ethnic stratum, which would undercut the global claim.","supporting_citations":[],"review_version":1}