{"id":"0864692e-0722-4e36-8323-bb091305e6d2","arxiv_id":"2508.13821","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A deep learning model segments ICA and MCA vascular territories directly on cerebral DSA images, outperforming atlas registration in Dice score and qualitative success rate.","lead":"Researchers trained an AI model to outline brain regions fed by the carotid and middle cerebral arteries directly on X-ray angiography images from stroke procedures. The model matched manually corrected territory maps better than the existing atlas-based method, and ran in seconds rather than minutes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference labels are human-adjusted autoTICI atlas masks, so the reported superiority over atlas registration may largely reflect imitation of human edits rather than anatomical truth; independent ground truth is needed.","rationale":"The reader's weakest assumption is exactly the load-bearing concern identified here: the reference standard shares provenance with the atlas baseline, so the comparison may be biased by the atlas prior. I agree with this assessment. The paper's own Discussion flags the annotation-bias limitation, and the suggested CTA co-registration would be the appropriate independent reference. Because the reader already returned a CONDITIONAL verdict, my stress-test does not move the verdict; the central direction is plausible and the practical autoTICI use case remains reasonable, but the exact accuracy claims are not fully supported without an independent anatomical reference. The internal inconsistencies in success rate, ASD, and Table 1's impossible IQR are real and should be corrected, but they are secondary to the reference-standard concern.","tokens_in":10717,"tokens_out":3036,"duration_ms":34345,"concrete_test":"Select a blinded subset of ~30 acquisitions stratified by occlusion location and view. Have two neuroradiologists independently delineate ICA and MCA territories from scratch on the DSA MinIPs, without access to the autoTICI atlas, the model outputs, or the atlas-registration outputs. Measure inter-rater agreement (e.g., DSC between the two raters), then recompute model and atlas DSC, ASD, and success rates against these from-scratch labels. If the model's advantage over atlas registration largely disappears or inter-rater agreement is low, the reference standard cannot support the accuracy claim as stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a deep network can directly segment ICA/MCA vascular territories on DSA more accurately than atlas registration. This requires the reference standard to be a valid measure of true vascular territory. However, the reference labels were 'derived from autoTICI atlases' and 'manually adjusted for size, rotation, and scaling' (Methods, Section 2), and the comparator is the same autoTICI atlas registration pipeline. Thus the training labels and the baseline share the same atlas prior. The model can learn to reproduce the human corrections applied to the atlas, while the baseline receives no such correction; the measured DSC/ASD advantage may therefore reflect imitation of human edits rather than recovery of independently established anatomy. The Discussion acknowledges this: 'the manual annotation of reference standards could have introduced bias... the results may not entirely reflect the true vascular territories' and recommends CTA co-registered references for future work. This does not invalidate the practical claim that the learned model could replace atlas registration inside the autoTICI pipeline, but it weakens the stronger claim of 'accurate segmentations' of true vascular territories. Secondary inconsistencies—abstract success rate 85% vs body per-patient 80%, abstract ASD 13.8 vs table 14, and the impossible atlas ICA DSC IQR of 0.82 [0.62-0.80] in Table 1—further reduce confidence in the exact numbers as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a deep learning approach, based on nnUNet, to segment internal carotid artery (ICA) and middle cerebral artery (MCA) vascular territories directly from 2D minimal intensity projection (MinIP) cerebral DSA images acquired during endovascular thrombectomy. The model was trained on 1,224 acquisitions from 361 patients whose reference labels were autoTICI atlas segmentations manually adjusted for size, rotation, and scaling. The authors compare the model with an atlas registration baseline (autoTICI) using Dice similarity coefficient, Jaccard index, average surface distance, and Hausdorff distance, and also report a qualitative Likert-scale success-rate comparison on a held-out cohort. The paper reports significantly better metric values for the deep learning model, a higher qualitative success rate, and much lower computational time, and it concludes that direct segmentation can replace atlas registration for vascular territory visualization in DSA.","tokens_in":10927,"tokens_out":2632,"duration_ms":29858,"significance":"If the reported results hold, the paper would offer a practically useful and much faster alternative to atlas-based vascular territory mapping in DSA, with potential value for the autoTICI reperfusion scoring pipeline and for broader visualization during X-ray-guided procedures. The study has several strengths: patient-level stratified splits, a multicenter registry cohort, use of a strong segmentation baseline (nnUNet), public code, and an external qualitative assessment by two raters. The claimed computational speedup (seconds versus over two minutes) is clinically relevant. However, the central comparison is weakened by the shared provenance of the reference labels and the atlas baseline: the labels are manually corrected autoTICI atlas masks and the comparator is the automatic autoTICI atlas registration. This does not invalidate the practical claim that the learned model could replace atlas registration inside the autoTICI pipeline, but it limits the strength of the anatomical-accuracy claim.","major_comments":[{"comment":"The reference standard labels are human-adjusted masks derived from the autoTICI atlas, and the comparator is the automatic autoTICI atlas registration. This shared provenance means the reported DSC, ASD, and success-rate advantages may partly measure the model's ability to mimic human corrections to the same atlas prior, rather than recovery of independently established vascular anatomy. The Discussion acknowledges this possibility, but the manuscript's central claim of 'accurate segmentations' of true vascular territories rests on the reference standard. Please provide a sensitivity analysis on a subset with an independent reference, for example territories delineated on CTA and co-registered to DSA, or at least re-annotate a random sample with a protocol that does not start from the atlas. Without such evidence, the conclusions should be explicitly limited to 'superior to atlas registration within the autoTICI framework'.","section":"Methods, Section 2; Discussion, Section 4"},{"comment":"Several numbers reported in the abstract and body are mutually inconsistent. The abstract states a success rate of 85% for the segmentation model, whereas the per-patient success rate reported in Section 3.2 is 80%; the abstract reports ICA ASD of 13.8 versus 47.3, while Table 1 gives medians of 14 and 47; and Table 1 reports the atlas ICA DSC as 0.82 [0.62-0.80], which is impossible because the upper IQR bound is below the median. Please correct these inconsistencies and state unambiguously which values are means, medians, per-view, or per-patient.","section":"Abstract; Results, Section 3.2; Table 1"},{"comment":"The term 'external test set' in the abstract is misleading. The qualitative comparison was performed on 564 out of 660 patients from a previous study in the same MR CLEAN Registry, excluding patients who overlapped with training or test sets. This is an external cohort in the sense of not being used for training, but it is not an independent acquisition protocol or institution. Please reword to 'held-out cohort' or 'external to training' and clarify the relationship to the registry.","section":"Methods, Section 3.2; Abstract"},{"comment":"The qualitative Likert assessment is central to the success-rate claim, but the methods do not state whether the two raters were blinded to the method producing each segmentation, nor is inter-rater agreement (e.g., Cohen's kappa) reported. If the raters were aware of which output came from the atlas versus the deep learning model, bias in scoring is possible. Please provide blinding details and inter-rater reliability.","section":"Results, Section 3.2 (success rate comparison)"}],"minor_comments":[{"comment":"There is a typo in the text: 'significantly better DCS' should read 'DSC'.","section":"Results, Section 3.2 (Model performance)"},{"comment":"The phrase 'non-statistical significant different performance' should be reworded, for example to 'no statistically significant difference in performance'.","section":"Methods, Section 2 (Model training)"},{"comment":"The sentence 'if very few or no vessels are depicted, the might be little use for specific vessel territory assessment' contains a grammatical error ('the might be') and should be revised.","section":"Discussion, Section 4"},{"comment":"The symbols and asterisks in Table 2 are not fully explained: only some phase-specific comparisons are marked with '***', but the text does not state what the significance markers refer to (presumably comparison with the arterial or non-contrast phase). Please clarify in the table caption or text.","section":"Table 2"},{"comment":"Reference 6 is listed as an arXiv preprint although the journal version of the nnU-Net paper is already cited as reference 5; please check whether both are necessary and cite the peer-reviewed version consistently.","section":"General / References"}],"recommendation":"major_revision","confidential_remarks":"The central technical contribution is plausible and the practical claim about replacing atlas registration within autoTICI may well survive further scrutiny. The main risk is the reference-standard circularity, which I see as fixable with additional analysis rather than fatal. The inconsistent numbers in the abstract and Table 1 must be corrected before acceptance. I therefore recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead the DSA territory paper. The core result is plausible: a standard nnUNet trained on DSA MinIPs can segment ICA and MCA vascular territories, and it beats the autoTICI atlas registration both in overlap metrics and qualitative scoring. The speed gain is real and clinically relevant (4s vs 141s). The task framing is new—prior DSA segmentation papers go after vessels or aneurysms, not territories—and the evaluation has genuine care: patient-level splits, an external qualitative test, and public code.\n\nThe weak spot is the reference standard. The training labels are autoTICI atlases that were manually adjusted, and the comparator is the unadjusted autoTICI atlas. That shared provenance means the model's superiority partly reflects learning the human edits rather than recovering independent anatomy. The authors admit as much in the Discussion and suggest CTA-based references for future work. This does not sink the practical claim—if you want to replace the atlas inside autoTICI, a model trained on corrected atlases is a reasonable approach—but it does mean the words \"accurate segmentations\" in the conclusions are stronger than the evidence supports. The qualitative success-rate comparison (80% vs 66%) is less affected by this issue, and the fact that the atlas's 66% matches previously reported 64% is a good consistency check.\n\nThe reporting has avoidable sloppiness. Abstract says 85% success and ASD 13.8; the body says 80% and Table 1 says 14. Table 1 also lists an atlas ICA DSC IQR of 0.82 [0.62-0.80], which is impossible. These are minor and easily fixed, but they undermine confidence in the headline numbers.\n\nBottom line: the direction is likely right and the clinical motivation is concrete. Who is this for? Neurointerventional imaging researchers working on automated TICI scoring or real-time anatomical overlays. It deserves a serious referee, but the revision needs to confront the reference-standard provenance head-on and reconcile the numbers. The core experiment is worth redoing with a CTA-derived or independent ground truth before anyone leans too hard on the absolute accuracy.","headline":"The core direction is likely right — a nnUNet can segment ICA/MCA territories on DSA and beat atlas registration — but the reference labels inherit the atlas's provenance, so the absolute accuracy claims are over-stated and the headline numbers need reconciling.","tokens_in":11531,"tokens_out":2082,"would_cite":false,"duration_ms":22503,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A trained network maps brain territories invisible on DSA scans","keywords":["digital subtraction angiography","vascular territory segmentation","deep learning","nnU-Net","ischemic stroke","atlas registration","minimal intensity projection","autoTICI"],"falsifier":"Segment the same DSA images with reference labels derived from co-registered CT angiography rather than from a 2D atlas, and compare the model's Dice and surface distance against both label sources; if the advantage over atlas registration shrinks or vanishes against the CTA-based labels, the claimed superiority is an artifact of label bias.","tokens_in":10529,"feed_emoji":"🧠","tokens_out":4939,"duration_ms":49167,"temperature":0.7,"pith_summary":"The paper tries to show that a deep learning model can directly segment the vascular territories of the internal carotid and middle cerebral arteries from digital subtraction angiography (DSA), even though these territories have no visible borders in the images. Training on 1,224 acquisitions from 361 stroke patients, the model beat the standard atlas-registration approach on overlap, surface distance, per-case success rate, and runtime. If the result holds, physicians performing thrombectomy could see which brain regions are at risk without waiting for atlas registration or extra imaging.","feed_headline":"Neural net maps invisible brain territories on DSA","feed_subtitle":"Model beats atlas registration on overlap, speed, and success rate for stroke angiography.","key_machinery":"The load-bearing mechanism is the nnUNet, a self-configuring U-Net architecture adapted to the data, operating on 2D minimal-intensity projections built from each DSA sequence. The training labels come from atlas-registered territory masks that were manually adjusted for size, rotation, and scaling; the ICA mask is reconstructed as the union of the anterior cerebral artery (ACA) and MCA labels. Morphological erosion, connected-component analysis, and dilation clean the predicted masks before evaluation.","core_discovery":"The central claim is that an nnUNet trained on minimal-intensity projections of cerebral DSA, with labels that are manually adjusted versions of atlas-based segmentations, can predict the internal carotid artery (ICA) and middle cerebral artery (MCA) territories directly from the image. Compared with a conventional atlas registration method, the segmentation model achieved higher Dice similarity (0.96 vs 0.82 for the ICA territory), lower average surface distance, and a higher qualitative success rate (85% versus 66% on the external test set), while running in about 4 seconds instead of 141 seconds. The model relies most on the capillary phase of contrast passage; non-contrast frames produce substantially worse segmentations.","pith_inferences":["The measured margin over atlas registration may partly reflect the model learning the human editors' corrections of the same atlas labels, so a comparison against labels from an independent modality would be needed to gauge the true anatomical accuracy.","The model's strong dependence on the capillary phase suggests it is reading contrast fill of the parenchyma rather than bone or vessel silhouettes, which is consistent with learning a proxy for perfusion territory.","A testable extension would be to feed the model phase-selected capillary minimal-intensity projections only, which the phase-dependency results predict should improve segmentation quality in pre-EVT ICA-occlusion cases.","If the approach transfers to other anatomical structures, it would offer a general route to soft-tissue visualization during X-ray-guided procedures, where such anatomy is otherwise invisible."],"forward_implications":["Replacing atlas registration in the existing reperfusion-scoring pipeline would raise the overall per-patient success rate from 66% to about 80% while cutting runtime from more than two minutes to under ten seconds.","Because the model reads all phases but performs best on capillary-phase frames, future versions could select the capillary phase when available or be trained phase-specifically to improve robustness in ICA occlusions.","The same direct-segmentation approach could be extended to other territories with unclear borders, such as the posterior cerebral artery or functional regions like Broca's and Wernicke's areas, and to other X-ray-guided interventions.","Pre-EVT images, especially with proximal ICA occlusions, remain harder than post-EVT images, so larger, more balanced training sets would be needed before relying on the model in the most occluded cases."],"supporting_citations":[{"why":"Supplies the autoTICI atlas registration method that serves as the comparison baseline and the source of the manually adjusted territory labels.","marker":"[16]"},{"why":"Provides the external test set and the previously reported atlas success rate (64%) against which the segmentation model's success rate is compared.","marker":"[14]"},{"why":"Defines the nnUNet segmentation architecture with its self-configuring training pipeline used for the model.","marker":"[5]"},{"why":"Presents the earlier automated design principle for the nnUNet that underlies the model's adaptivity.","marker":"[6]"}],"fun_headline_variants":["AI maps hidden brain territories on stroke DSA","Neural net segments brain vascular zones in seconds","Deep learning beats atlas for DSA territory mapping","nnUNet reveals invisible brain regions in DSA","Rapid AI segmentation of brain territories on DSA"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reference standard is assumed to be true vascular territory, but it comes from manually adjusted atlas projections, so if the atlas inherited a systematic bias, every comparison metric inherits it too.","fun_headline_variants_meta":{"raw":{"variants":["AI maps hidden brain territories on stroke DSA","Neural net segments brain vascular zones in seconds","Deep learning beats atlas for DSA territory mapping","nnUNet reveals invisible brain regions in DSA","Rapid AI segmentation of brain territories on DSA"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000427,"raw_usage":{"total_tokens":2192,"prompt_tokens":959,"completion_tokens":1233,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":1160}},"tokens_in":575,"tokens_out":1233,"duration_ms":9401,"temperature":1.0,"reasoning_tokens":1160,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:10:42.198005+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Segment the same DSA images with reference labels derived from co-registered CT angiography rather than from a 2D atlas, and compare the model's Dice and surface distance against both label sources; if the advantage over atlas registration shrinks or vanishes against the CTA-based labels, the claimed superiority is an artifact of label bias.","supporting_citations":[{"cited_title":"IEEE transactions on medical imaging40(9), 2380–2391 (2021)","cited_arxiv_id":null,"evidence_quote":"Supplies the autoTICI atlas registration method that serves as the comparison baseline and the source of the manually adjusted territory labels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the external test set and the previously reported atlas success rate (64%) against which the segmentation model's success rate is compared."}],"review_version":2}