{"id":"942c4ece-0b7d-4573-b472-144d7044f4f2","arxiv_id":"2502.01023","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A pipeline combining MFAT vesselness, MIP-based seeds, and geometry-guided region growing achieves the highest Dice scores for vessel segmentation on chi-separation brain maps.","lead":"Chi-separation MRI maps iron and myelin in the brain, but blood vessels create artifacts. This paper introduces a three-step vessel segmentation method that reports higher Dice scores than existing approaches and changes ROI analyses in 16 of 27 brain regions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dice superiority may be inflated by per-subject tuning of the anisotropy threshold on the same subjects used for evaluation.","rationale":"The reader's weakest assumption concerned the stability of the anisotropy measure in Eq. 7 across subjects, resolutions, and χ-separation algorithms. Our stress test identifies the same underlying risk but locates it more precisely in the evaluation protocol: the anisotropy threshold is adjusted per subject on the same subjects used for Dice computation, which can compensate for instability in the measure and inflate the reported advantage. This is partially aligned with the reader's concern, since the need for per-subject adjustment is direct evidence that Eq. 7 is not stable. The proposed concrete test—holding the threshold fixed or tuning only on a development set—would settle whether the reported superiority is genuine or an artifact of hyperparameter flexibility. We do not recommend changing the reader's CONDITIONAL verdict: the method is clearly described, the qualitative results are plausible, and the concern is addressable with additional evaluation. The paper's own acknowledgment of subject-wise parameter tuning challenges supports the need for this condition before the central claim is accepted as evidence of a fully automatic, generalizable tool.","tokens_in":18805,"tokens_out":6584,"duration_ms":73409,"concrete_test":"Fix Aniso_Thresh = 0.0012 for all subjects and all four χ-separation algorithms, with no per-subject adjustment, and recompute Dice against the manual labels. Also evaluate a pre-registered variant where the threshold is chosen once on a development subject (one 3T, one 7T) and applied to the remaining validation subjects. If the fixed or held-out Dice remains at least 5 percentage points above the Frangi and GRE-based baselines, the superiority claim survives; if it drops to baseline levels, the reported advantage depends on per-subject tuning. Additionally, report how many subjects required a threshold increase to exclude deep gray matter and the magnitude of that increase.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of highest Dice and robust performance rests on the anisotropy threshold (Aniso_Thresh, Eq. 7) being both discriminative and stable. The paper states this threshold was 'adjusted for each subject' starting from 1.2e-3 and increased when deep gray matter regions were not properly excluded (Methods 2.2, Discussion). Table 1 and Table 2 report Dice on the same subjects (three 3T, three 7T) that were used for this per-subject adjustment. If threshold selection uses knowledge of each test subject's false positives, the comparison against Frangi (parameters fixed per dataset via ROC) and the GRE-based method (no adjustable parameters) is not an out-of-the-box evaluation. The reported threshold range spans 0.0012–0.0108 for χ_para with MEDI/iLSQR, a factor of ~9, indicating the anisotropy measure is not automatically stable across algorithms or subjects without manual intervention. Additionally, the manual labels exclude calcifications, meninges, and registration artifacts (Supp. Fig. 3), which the Discussion acknowledges as failure modes; these exclusions remove known difficult cases from the Dice computation. Therefore, the reported 'highest Dice, effectively excluding non-vessel structures' may not generalize to automatic application on new data, which is needed to support the claimed utility in population-averaged ROI analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a three-step vessel segmentation pipeline for chi-separation maps: (1) seed generation from R2* and the product of chi_para and |chi_dia| using MFAT filtering and MIP-based back-projection, (2) vessel-geometry-guided region growing with intensity limits and novel directional/intensity/anisotropy criteria, and (3) removal of non-vessel structures by connected-component anisotropy thresholding (Eq. 7). The method is evaluated against Frangi filter and a GRE-based vein-segmentation method on six manually labeled subjects (three 3T, three 7T) and across four chi-separation algorithms, reporting higher Dice scores (Tables 1 and 2). Two demonstrations are provided: improved RMSE/PSNR/SSIM when evaluating chi-sepnet-R2* against chi-sep-COSMOS after vessel masking (Table 3), and statistically significant effects of vessel exclusion in population-averaged ROI analysis of 106 subjects (Table 4). Code is made available as part of the chi-separation toolbox.","tokens_in":19087,"tokens_out":3450,"duration_ms":34882,"significance":"If the reported performance holds, this is practically valuable: it addresses a known artifact source in chi-separation analysis and provides a tool with modest computational cost and public code. The idea of using the product of chi_para and |chi_dia| for small-vessel seeds and the anisotropy-based connected-component refinement to suppress globus pallidus and optic radiation is sensible and clearly described. The paper also contributes two application-level validations (reconstruction-quality evaluation and ROI analysis), which increase the utility of the method beyond a purely algorithmic comparison. The central claims are supported by consistent DSC improvements, but the evaluation design has issues that need addressing before the superiority claim is fully convincing.","major_comments":[{"comment":"The anisotropy threshold Aniso_Thresh was adjusted for each subject, starting from 1.2e-3 and increased when deep gray matter regions were not properly excluded. Because the DSC values in Tables 1 and 2 are computed on the same subjects used for this per-subject adjustment, the proposed method is not being compared under the same 'out-of-the-box' conditions as the Frangi filter (whose parameters were optimized per dataset via ROC) or the GRE-based method (which has no adjustable parameters). The reported threshold range spans nearly an order of magnitude (0.0012-0.0108 for chi_para with MEDI/iLSQR), indicating that the anisotropy measure is not automatically stable. Please report DSC with a single fixed threshold (e.g., the starting value) for all subjects, or use a cross-validation/held-out protocol, and provide a sensitivity analysis of the threshold. Without this, the claim of 'highest Dice' may be inflated by peaking.","section":"Section 2.2, Eq. (7); Section 4 (Discussion)"},{"comment":"The manual ground-truth masks explicitly exclude calcifications, meninges, and artifacts caused by mis-registration between R2* and R2. These are precisely among the structures that the proposed method is designed to exclude as non-vessel (the paper acknowledges calcification and meninges as failure modes in the Discussion). Excluding them from the labeled voxels removes known difficult cases from the Dice computation, which may overstate the method's ability to 'effectively exclude non-vessel structures' (Abstract, Conclusion). Please justify this exclusion or evaluate on a subset that includes these structures; at minimum, quantify the number and spatial extent of these excluded regions and discuss how their inclusion would affect DSC.","section":"Section 2.3, Supplementary Figure 3"},{"comment":"The robustness claim across four chi-separation algorithms is weakened by per-algorithm threshold adjustment. The Discussion states that higher Aniso_Thresh values were needed for chi-sep-MEDI and chi-sep-iLSQR maps (0.0072-0.0108 for chi_para) and that this may have led to exclusion of small vessels and slightly reduced DSC. With the threshold varied per algorithm and per subject, the similar DSCs in Table 2 partly reflect this tuning rather than automatic robustness. Please report the DSC obtained with a single fixed threshold (e.g., 1.2e-3) applied to all algorithms, and show how the threshold would need to be chosen in practice for a new dataset.","section":"Section 2.4, Table 2; Section 4 (Discussion)"}],"minor_comments":[{"comment":"The vesselness vMFAT in Eq. (3) is stated to be computed from the susceptibility map, but in Step 1 vMFAT is computed from R2* and from chi_para * |chi_dia|. Please clarify explicitly which input image is used for the vMFAT entering Eq. (3) for each of the two masks (chi_para and |chi_dia|).","section":"Eq. (3)"},{"comment":"The notation for anisotropy is inconsistent: Eq. (6) defines Ani(q) as a per-voxel value, while Eq. (7) computes the mean over a connected component. A subscript such as Ani_CC or an explicit definition would improve readability.","section":"Eq. (7)"},{"comment":"For some ROIs (e.g., red nucleus, subthalamic nucleus) the p-value is reported as '-' with no explanation. Please state why no p-value is given, presumably because the vessel proportion is zero or the paired difference has zero variance.","section":"Table 4"},{"comment":"The ROC curves for Frangi filter parameter optimization are not fully described: it is unclear which parameter was swept and what the 'optimum' criterion means in terms of the ROC operating point. Please define the swept parameter(s) and the selection rule.","section":"Supplementary Figure 5"},{"comment":"The phrase 'highest Dice score coefficient' should be 'highest Dice similarity coefficient' for consistency with the standard terminology.","section":"Abstract and throughout"},{"comment":"The code is said to be available as part of the chi-separation toolbox, but no version or specific repository path is given. Please provide a stable link or a DOI to facilitate reproducibility.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering contribution to a niche but active area. The main novelty is the combination of chi-separation-specific inputs with a region-growing/anisotropy refinement; the individual components (MFAT, inverse Hamming filter, MIP back-projection) are drawn from prior work. The evaluation is the crucial part, and its current form (per-subject threshold tuning on the evaluation subjects, manual labels that exclude known difficult structures) makes the superiority claim hard to assess. I would encourage the editor to require a fixed-threshold or cross-validated evaluation and an analysis of the excluded manual-label regions before accepting. The application demonstrations are useful and, if the evaluation issues are fixed, the paper would be a reasonable methods contribution. The writing is generally clear; the main technical description is reproducible from the text and appendices."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, clearly described vessel segmentation method for chi-separation, and the applications are sensible. But the headline result—highest Dice—is not fully out-of-the-box, because the anisotropy threshold was tuned per subject on the same six subjects used for evaluation. That doesn't sink the paper, but it does mean the Dice advantage over fixed-parameter baselines is inflated.\n\nWhat's actually new: the combination of MFAT vesselness on high-pass-filtered R2*, MIP-based small-vessel seeds, and region growing with added intensity-similarity and anisotropy criteria, followed by connected-component anisotropy filtering. I don't see that exact pipeline in the cited vessel segmentation literature. The method is described in enough detail to reproduce, and the code is in the chi-separation toolbox. That's a real contribution to the subfield.\n\nThe evaluation is better than average for a methods paper: manual labels on six subjects (3T and 7T), comparison with Frangi and a GRE-based method, robustness across four chi-separation algorithms, and two applications (chi-sepnet evaluation and 106-subject ROI analysis) that give the mask a purpose. The Dice gains are consistent, and the figures show the method does exclude globus pallidus and optic radiation better than the baselines. The citation pattern covers the relevant literature; self-citations to the group's chi-separation work are appropriate.\n\nThe soft spots are real but mostly acknowledged. The per-subject tuning of Aniso_Thresh (Methods 2.2) is the main one. Since the same six subjects are used for tuning and evaluation, Tables 1 and 2 are effectively in-sample. The reported threshold range spans a factor of roughly 9 (0.0012–0.0108 for chi_para with MEDI/iLSQR), which suggests the anisotropy measure is not automatically stable across subjects or algorithms without manual adjustment. The ground truth also excluded calcifications, meninges, and registration artifacts (Supp. Fig. 3), which are exactly the known failure modes, so the Dice is computed on the easier cases. The sample size is small, but that's common for manual-label studies. On the plus side, the Discussion states the tuning issue and the failure modes explicitly; the authors are not hiding it.\n\nThe central claim should be read as 'with per-subject tuning, this beats fixed-parameter baselines.' That's useful, but the paper would be stronger with a fixed-threshold run and per-subject-tuned baselines for comparison.\n\nWho is this for? Researchers doing chi-separation or QSM-based iron/myelin quantification who need a vessel mask and are willing to tune a threshold. It deserves a serious referee. I'd send it out with a request for the fixed-threshold analysis, not desk-reject it.","headline":"Solid vessel segmentation adaptation for chi-separation, but the reported Dice advantage is weakened by per-subject threshold tuning on the evaluation subjects.","tokens_in":19546,"tokens_out":3594,"would_cite":true,"duration_ms":31501,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A three-step, geometry-guided vessel segmentation pipeline outperforms Frangi and GRE-based methods on χ-separation maps, with the highest Dice scores.","keywords":["χ-separation","vessel segmentation","quantitative susceptibility mapping (QSM)","MFAT vesselness","anisotropy-based cleanup","region growing","iron and myelin mapping","ROI analysis"],"falsifier":"Compute the per-connected-component anisotropy of Eq. 7 for manually labelled vessels, globus pallidus, and optic radiation across many subjects; if the vessel distribution overlaps substantially with the other two, no single threshold can separate them and the reported Dice advantage will not generalize. A complementary test is to run the pipeline on images containing confirmed calcifications, which the paper itself notes can mimic vessels by being hyperintense and tube-like.","tokens_in":18651,"feed_emoji":"🧠","tokens_out":10739,"duration_ms":96851,"temperature":0.7,"pith_summary":"$\\chi$-separation produces paramagnetic and diamagnetic susceptibility maps intended to reflect iron and myelin in the brain, but blood vessels generate artifacts that distort local values. This paper aims to establish that a dedicated vessel segmentation step, built from the same maps that $\\chi$-separation already produces, can remove vessels while sparing deep gray matter and myelinated fiber tracts. The proposed three-step pipeline seeds from high-pass-filtered $R_2^*$ and from the product of the two susceptibility maps, grows regions along vessel geometry using direction, intensity, and anisotropy criteria, and deletes connected components whose mean anisotropy is too low to be vessels. The paper reports that this mask outperforms Frangi and GRE-based vessel segmentation in Dice score, improves error metrics when evaluating a neural-network $\\chi$-separation reconstruction, and changes population-averaged ROI susceptibility values significantly in 16 of 27 tested regions.","feed_headline":"Geometry-guided mask clears vessels from χ-separation maps","feed_subtitle":"Vessel masking lifts Dice scores to 76.7% and shifts ROI susceptibility values in 16 of 27 brain regions.","key_machinery":"The load-bearing object is the anisotropy measure $Ani(q)=|\\lambda_2(q)\\lambda_3(q)|$ formed from the two larger-magnitude Hessian eigenvalues of the susceptibility map; it is used as a soft term inside the region-growing acceptance rule (Eq. 3) and as a per-connected-component mean threshold (Eq. 7) that removes non-vessel structures. The seed generation relies on an inverse Hamming high-pass filter on $R_2^*$ for large vessels and a maximum-intensity projection of $\\chi_{\\mathrm{para}}\\cdot|\\chi_{\\mathrm{dia}}|$ for small vessels, both feeding the multi-scale fractional anisotropy tensor (MFAT) vesselness map, a Hessian-eigenvalue measure of how tubular a voxel is. The region-growing step then follows vessel direction by comparing the principal Hessian eigenvectors of neighbouring voxels while requiring intensity coherence and anisotropy, which together stop the mask from bleeding into bright but non-tubular structures.","core_discovery":"The central discovery is that vessel segmentation for $\\chi$-separation can be treated as a geometric filtering problem rather than a pure intensity-thresholding problem. Vessels are seeded from two complementary signals—large vessels from a high-pass-filtered $R_2^*$ map and small vessels from a maximum-intensity projection of $\\chi_{\\mathrm{para}}\\cdot|\\chi_{\\mathrm{dia}}|$—then propagated by a region-growing rule that requires directionality similarity to the seed's Hessian eigenvector, intensity similarity, and a vesselness-anisotropy term. The final cleanup uses the mean of $|\\lambda_2\\lambda_3|$ over each connected component (Eq. 7) to discard non-vessel structures such as globus pallidus and optic radiation. In the reported experiments the method achieves Dice scores of $76.7\\pm4.2\\%$ ($\\chi_{\\mathrm{para}}$) and $68.7\\pm7.9\\%$ ($|\\chi_{\\mathrm{dia}}|$) at 3T and $76.9\\pm2.7\\%$ and $72.6\\pm5.7\\%$ at 7T, all above the two comparison methods, and it stays within a few points of its best score across four different $\\chi$-separation algorithms.","pith_inferences":["The anisotropy threshold in Eq. 7 is the pivot on which the method stands; a natural next experiment is to measure the anisotropy distributions of vessels versus globus pallidus and optic radiation across a large population to see how cleanly the two can be separated.","Because the inputs are generic susceptibility-style maps, the same seed-and-grow machinery could transfer to QSM or SWI vessel segmentation; the paper names this as future work, but it is not demonstrated here.","The masks produced with individually tuned hyperparameters could serve as training labels for a deep segmentation network, potentially removing the per-subject tuning the paper reports, though such a network is not tested in this work.","A testable extension is to validate the pipeline against whole-brain manual labels, since the current manual ground truth covers 12 central slices in three planes per subject rather than the full volume."],"forward_implications":["Vessel masking becomes a practical preprocessing step for $\\chi$-separation studies, since it changes reconstruction-quality metrics and ROI means enough to alter conclusions.","The reported Dice advantage implies that combining geometry-guided region growing with anisotropy cleanup captures more true vessels and fewer false positives than Hessian-only or multi-contrast heuristic segmentation.","The mask's stability across $\\chi$-sep-COSMOS, $\\chi$-sep-MEDI, $\\chi$-sep-iLSQR, and $\\chi$-sepnet-$R_2^*$ means a single segmentation protocol can be used regardless of which reconstruction algorithm generated the maps.","The modest resource footprint (2 GB RAM and 4 minutes on 3T data) makes the method practical for large-cohort analyses such as the 106-subject template study reported here.","Significant ROI shifts in 16 of 27 regions, including caudate and genu of corpus callosum, imply that previously reported atlas values without vessel masking may be biased in vessel-rich structures."],"supporting_citations":[{"why":"Provides the $\\chi$-separation method that generates the paramagnetic and diamagnetic susceptibility maps the segmentation uses as inputs.","marker":"[9]"},{"why":"Supplies the multi-scale fractional anisotropy tensor (MFAT) vesselness filter used to compute vesselness maps for seeds and region growing.","marker":"[41]"},{"why":"Defines the GRE-based vessel segmentation baseline and the inverse Hamming high-pass filter parameters used for large-vessel seed generation.","marker":"[44]"},{"why":"Defines the Hessian-based Frangi vesselness filter that serves as the first comparison baseline.","marker":"[38]"},{"why":"Provides the region-growing condition that the paper extends with intensity similarity and anisotropy terms.","marker":"[48]"},{"why":"Supplies the $\\chi$-sepnet-$R_2^*$ reconstruction network whose quantitative evaluation is improved by applying the vessel mask.","marker":"[49]"},{"why":"Provides the $\\chi$-separation atlas and the 106-subject template dataset used for population-averaged ROI analysis.","marker":"[27]"},{"why":"Supplies the maximum-intensity-projection strategy used to enhance small vessel visibility in the seed step.","marker":"[37]"},{"why":"Provides the background-suppressed high-pass filtering approach that suppresses non-vessel structures in the $R_2^*$ map.","marker":"[45]"},{"why":"Defines the Dice similarity coefficient used to quantify segmentation agreement with manual ground truth.","marker":"[62]"}],"fun_headline_variants":["Geometry-guided mask cleans vessels from χ-separation maps","Vesselness filtering improves vessel masks for χ-separation","χ-separation vessel segmentation via Hessian-eigenvector region growing","Vessel-free analysis of χ-separation maps boosts ROI accuracy","Dice 76.7%: geometric vessel mask for χ-separation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that vessels are the only structures whose connected components are both bright in the seed maps and highly anisotropic, so that the Eq. 7 anisotropy threshold can delete globus pallidus and optic radiation without also deleting small vessels.","fun_headline_variants_meta":{"raw":{"variants":["Geometry-guided mask cleans vessels from χ-separation maps","Vesselness filtering improves vessel masks for χ-separation","χ-separation vessel segmentation via Hessian-eigenvector region growing","Vessel-free analysis of χ-separation maps boosts ROI accuracy","Dice 76.7%: geometric vessel mask for χ-separation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000385,"raw_usage":{"total_tokens":2126,"prompt_tokens":1123,"completion_tokens":1003,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":739,"completion_tokens_details":{"reasoning_tokens":916}},"tokens_in":739,"tokens_out":1003,"duration_ms":10337,"temperature":1.0,"reasoning_tokens":916,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T16:51:12.199767+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the per-connected-component anisotropy of Eq. 7 for manually labelled vessels, globus pallidus, and optic radiation across many subjects; if the vessel distribution overlaps substantially with the other two, no single threshold can separate them and the reported Dice advantage will not generalize. A complementary test is to run the pipeline on images containing confirmed calcifications, which the paper itself notes can mimic vessels by being hyperintense and tube-like.","supporting_citations":[{"cited_title":"Evidence of demyelination in mild cognitive impairment and dementia using a direct and specific magnetic resonance imaging measure of myelin content","cited_arxiv_id":null,"evidence_quote":"Provides the $\\chi$-separation method that generates the paramagnetic and diamagnetic susceptibility maps the segmentation uses as inputs."},{"cited_title":"Evaluation of SWI in Children with Sickle Cell Disease","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-scale fractional anisotropy tensor (MFAT) vesselness filter used to compute vesselness maps for seeds and region growing."},{"cited_title":"A novel gradient echo data based vein segmentation algorithm and its application for the detection of regional cerebral differences in venous susceptibility","cited_arxiv_id":null,"evidence_quote":"Defines the GRE-based vessel segmentation baseline and the inverse Hamming high-pass filter parameters used for large-vessel seed generation."},{"cited_title":"Quantification of cerebral veins in patients with acute migraine with aura: A fully automated quantification algorithm using susceptibility-weighted imaging","cited_arxiv_id":null,"evidence_quote":"Defines the Hessian-based Frangi vesselness filter that serves as the first comparison baseline."},{"cited_title":"2D and 3D Vascular Structures Enhancement via Multiscale Fractional Anisotropy Tensor","cited_arxiv_id":"1902.00550","evidence_quote":"Provides the region-growing condition that the paper extends with intensity similarity and anisotropy terms."},{"cited_title":"2D and 3D Vascular Structures Enhancement Via Improved Vesselness Filter and Vessel Enhancing Diffusion","cited_arxiv_id":null,"evidence_quote":"Supplies the $\\chi$-sepnet-$R_2^*$ reconstruction network whose quantitative evaluation is improved by applying the vessel mask."},{"cited_title":"So You Want to Image Myelin Using MRI: Magnetic Susceptibility Source Separation for Myelin Imaging","cited_arxiv_id":null,"evidence_quote":"Provides the $\\chi$-separation atlas and the 106-subject template dataset used for population-averaged ROI analysis."},{"cited_title":"Improved susceptibility‐weighted imaging for high contrast and resolution thalamic nuclei mapping at 7T","cited_arxiv_id":null,"evidence_quote":"Supplies the maximum-intensity-projection strategy used to enhance small vessel visibility in the seed step."}],"review_version":1}