{"id":"b79b7c63-7ab8-4827-9df3-c58b254d1a35","arxiv_id":"2507.05451","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"HA2HA uses odd and even plane-wave angle subsets as noisy pairs to train a U-Net that denoises blood-flow RF signals, improving power and color Doppler image quality across species and contrast conditions.","lead":"A self-supervised deep learning framework, HA2HA, denoises ultrasound microvascular images by training on paired half-angle subsets of the same radio-frequency blood flow data, without clean labels. It reports over 15 dB gains in contrast-to-noise ratio and SNR on pig kidney, human liver, and human kidney data, and also improves color Doppler images.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The N2N target in Eq. (3) is not established: odd/even half-angle compounding plus independent SVD filtering makes Y1 and Y2 share a biased, not identical, signal, so the >15 dB gains may partly reflect suppression of half-angle artifacts rather than unbiased denoising.","rationale":"The reader's weakest assumption correctly identifies Eq. (3) as the load-bearing premise, and the paper's own Section II-B admits that the 'noise' includes structured, signal-dependent components. My stress-test sharpens the same concern: the two observations are not even noisy versions of the same X. Independent SVD clutter filtering of two half-angle compounds is a nonlinear, data-adaptive operation, so the retained blood-flow components in Y1 and Y2 can differ systematically. This makes the training target a biased estimate of the true vascular signal, not X plus zero-mean noise. That bias is especially problematic because the network is trained on half-angle data but evaluated on full-angle compounded data; the network may learn to remove half-angle-specific structure that has a different character in the full-angle test distribution. The empirical results are internally consistent and the multi-dataset evaluation is a genuine strength, but they do not test the core N2N assumption. The proposed Field II simulation with a known ground-truth signal would directly measure whether the learned mapping is unbiased. Given that the paper is otherwise clearly written and the reported gains are consistent across datasets, a CONDITIONAL verdict remains appropriate; the missing validation of Eq. (3) is a concrete, testable condition rather than a demonstrated failure.","tokens_in":13753,"tokens_out":5858,"duration_ms":74491,"concrete_test":"Run a Field II simulation with a known flow phantom: generate 10-angle RF frames, split into odd/even groups, compound and SVD-filter each group exactly as in Section II-F, and train HA2HA on those paired inputs. Then compare the HA2HA output on full-angle SVD-filtered test frames against the known clean flow signal X. If the mean output in vessel regions is biased away from X by more than about 1 dB while background noise is reduced, Eq. (3) is violated and the reported gains are inflated by structured artifact suppression rather than unbiased denoising.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II-B defines Y1 = X + N1 and Y2 = X + N2, and Eq. (3) requires E[Y2|Y1] = X. The construction in Section II-F violates the premises needed for this equality. Y1 and Y2 are produced by coherently compounding two disjoint 5-angle subsets and then applying independent SVD clutter filters. Coherent half-angle compounding changes the point spread function and sidelobe structure, while SVD clutter filtering is data-adaptive and selects different subspaces in each subset. Thus the signal components in Y1 and Y2 are not the same X: the target Y2 is X_true plus a structured, flow-dependent bias, not X_true plus zero-mean noise. The paper explicitly counts angle-dependent sidelobes and clutter variations as noise (Section II-B), but those are not independent of Y1 and are not zero-mean conditional on it. Under MAE training, the minimizer is driven toward the conditional median of Y2 given Y1; if that conditional median differs from X_true by a flow-dependent amount, the >15 dB CNR/SNR gains and lowest BNP values in Tables I and II may in part reflect suppression of half-angle artifacts rather than unbiased recovery of the vascular signal. Because inference is applied to full-angle compounded data, a network trained to remove half-angle-specific structure could over- or under-suppress on the test distribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HA2HA, a self-supervised deep learning framework for denoising ultrasound microvascular imaging (UMI) radio-frequency (RF) data. Training pairs are formed by splitting the plane-wave steering angles into odd and even subsets, independently compounding and SVD-clutter-filtering each subset, and using one subset as input and the other as target in a Noise2Noise-style loss with symmetric, consistency, and L1 regularization terms. A single model trained on 90 frames of contrast-free pig kidney RF data is tested on contrast-free and contrast-enhanced pig kidney, human liver, and human kidney with CKD, as well as on pig kidney data acquired at varying transmit duty cycles. The paper reports improvements exceeding 15 dB in CNR and SNR over conventional, angular processing (AP), and spatiotemporal non-local means (ST-NLM) methods, and shows qualitative improvements in color Doppler imaging derived from the denoised RF signals.","tokens_in":14055,"tokens_out":5640,"duration_ms":65166,"significance":"If the reported results are robust, the work is significant: it offers a label-free, RF-domain denoising method that generalizes across species, contrast conditions, and anatomical regions, with potential clinical translation. The use of complementary angular subset pairs is a novel adaptation of Noise2Noise for UMI, and the demonstration of downstream color Doppler improvement is a useful addition. The evaluation is broad, and the supplementary materials include videos and figures that support the qualitative findings. However, the central quantitative claims rest on single measurements without uncertainty quantification, and the theoretical justification relies on a noise-independence assumption that is likely violated by the proposed data construction. These issues need to be addressed before the conclusions can be fully relied upon.","major_comments":[{"comment":"The central assumption E[Y2|Y1] = X is not established for the proposed pairing scheme. In the construction, Y1 and Y2 are formed by coherently compounding disjoint angle subsets and independently applying SVD clutter filters. Coherent half-angle compounding changes the point-spread function and sidelobe structure, and independent SVD filtering can select different subspaces in each subset; hence the clean signal components in Y1 and Y2 are not identical X, and the structured interference (sidelobes, clutter variations) counted as \"noise\" in Section II-B is signal-dependent and not zero-mean conditional on Y1. Under MAE optimization, the network output tends toward the conditional median of Y2 given Y1, which may differ from X by a flow-dependent bias. The reported >15 dB gains may therefore partly reflect suppression of half-angle artifacts rather than unbiased denoising. To substantiate the claim, I recommend a simulation or phantom study with known ground truth X, using the same odd/even pairing, to measure the bias of the HA2HA output relative to X; additionally, comparing a network trained with the proposed pairs against one trained with the full-angle compounded data as the target would help separate artifact suppression from denoising.","section":"Section II-B and II-F, Eq. (3)"},{"comment":"All quantitative results are single measurements per dataset and condition. For each dataset, the power Doppler image appears to be computed from a single acquisition, with manually selected ROIs (Fig. 7d1 and Supplementary Fig. S3) and per-method dynamic-range optimization. No error bars, repeated acquisitions, or statistical tests are reported for CNR, SNR, or BNP. The claim of an \"improvement exceeding 15 dB\" and \"consistently the best\" performance is therefore not substantiated with uncertainty. The authors should report mean ± standard deviation over multiple frames, subjects, or independent ROI selections, and provide a statistical comparison (e.g., paired tests or confidence intervals) against the best baseline.","section":"Tables I and II, Section III"},{"comment":"The duty-cycle experiment is presented as a systematic SNR variation, but it is unclear how many independent measurements underlie each curve in Fig. 8. The text states that the four DC levels were acquired sequentially within each frame to share the same imaging section and motion status, which is a repeated-measures design, but it is not stated whether the quantitative values are computed from a single frame, an average over frames, or a summary over multiple acquisitions. Without this information and without variance estimates, the conclusion that HA2HA is best for DC ≥ 0.2 and that AP outperforms at DC = 0.1 is not reliable. Please specify the number of independent samples, report variability, and justify the absence of error bars.","section":"Section III-E and Fig. 8"}],"minor_comments":[{"comment":"The text says the training data construction is \"described in detail in Section 2.6,\" but the manuscript uses Roman numeral section numbering; this should be Section II-F.","section":"Section II-F"},{"comment":"The normalization factor (2 + λc) in the combined loss is not explained; please clarify why this particular denominator is chosen and how the three loss terms are weighted relative to each other.","section":"Eq. (4)"},{"comment":"The phrase \"The single frame MB image is shown\" should be rephrased as \"The single-frame blood flow image is shown\" for grammatical clarity.","section":"Section III-B"},{"comment":"The ROI definitions used for the quantitative metrics appear only in the Fig. 7 caption and Supplementary Fig. S3; they should be stated in the main text near Section II-H, since they are essential for interpreting the reported CNR, SNR, and BNP values.","section":"Fig. 7 and Section II-H"},{"comment":"The claim that HA2HA effectively suppresses spurious velocity artifacts in color Doppler is supported only by visual inspection and line profiles; consider adding a quantitative metric, such as velocity variance or the fraction of colored pixels in a noise-only region, to the CDI evaluation.","section":"Section III-F"}],"recommendation":"major_revision","confidential_remarks":"The manuscript comes from a group with strong expertise in ultrasound microvascular imaging, and the HA2HA idea is interesting and potentially publishable. The main risks are (i) the Noise2Noise assumption violation, which could be addressed with a phantom or simulation experiment measuring bias, and (ii) the absence of statistical rigor in the single-image evaluations. Both are fixable within the scope of a revision. I would not reject based on the current evidence, but the revision should include the additional validation requested in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the half-angle pairing idea is new and worth testing seriously, but the paper overstates what the evidence supports. The method is clearly described and the cross-species, cross-contrast generalization is encouraging. If the >15 dB CNR/SNR gains hold up with proper uncertainty, this is a practical tool for UMI.\n\nWhat's actually new: constructing paired noisy observations from complementary odd/even subsets of steered plane-wave angles, then compounding and SVD-filtering each subset separately, is a legitimate adaptation of N2N to UMI. It avoids the temporal-misalignment problem of frame pairs and the simulation bias of added noise. Direct RF-domain denoising is also a sensible choice because it preserves phase for color Doppler. The manuscript is honest about the training being on a single species and organ, and the limitations section is reasonable.\n\nSoft spots, in order of severity. First, the theoretical foundation is shaky. Eq. (3) requires E[Y2|Y1] = X, which demands zero-mean noise independent of Y1. The paper's own definition of 'noise' includes angle-dependent sidelobes and clutter variations. Those are structured, signal-dependent, and not obviously zero-mean. Coherent half-angle compounding changes the point spread function, and SVD clutter filtering is data-adaptive and picks different subspaces per subset. So Y1 and Y2 do not share the same X; the target is X plus a flow-dependent bias. Training under MAE then moves toward the conditional median of Y2 given Y1, which need not be the true clean signal. The >15 dB gains may partly represent suppression of half-angle artifacts rather than unbiased denoising. That doesn't kill the method—artifact suppression is clinically useful—but it changes what the claim is, and the paper should test the assumption, e.g., by simulating known ground-truth flow or comparing against full-angle data as a target.\n\nSecond, the validation is thin. Each condition is one image, ROIs are manually selected, and dynamic range is optimized per method. There are no error bars, no statistics, no multi-subject analysis. The 15 dB claim is therefore a point estimate, not a robust effect. No code or data is released, which makes independent verification hard.\n\nThird, minor: lambda_c and lambda_1 are tuned on the same kind of data as the test sets, which is mild fitting, not fatal. The self-citations to AP and ST-NLM baselines are fine; those are the relevant comparisons.\n\nOverall, this deserves a serious referee. The idea is novel and practically motivated, and the generalization results are promising if real. But the manuscript needs a test of the N2N assumption and a real validation study before the >15 dB claim can be taken at face value. I'd send it to peer review with a request for major revision.","headline":"A genuinely useful pairing idea for self-supervised RF denoising in UMI, but the N2N justification is load-bearing and unproven, and the validation is single-shot.","tokens_in":14665,"tokens_out":2363,"would_cite":false,"duration_ms":25948,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A self-supervised framework trains on paired half-angle ultrasound RF frames and, without any clean labels, denoises unseen UMI data enough to exceed 15 dB gains in CNR and SNR.","keywords":["self-supervised denoising","ultrasound microvascular imaging","plane-wave compounding","radio-frequency denoising","Noise2Noise","power Doppler imaging","color Doppler imaging","SVD clutter filtering"],"falsifier":"Train or test HA2HA on simulated plane-wave RF data with a known ground-truth vascular signal and controlled additive electronic noise plus angle-dependent sidelobe interference; if the network output systematically deviates from the known clean signal by a structured, angle-correlated component rather than by zero-mean residual, then the conditional-zero-mean assumption is violated and the denoised images are biased.","tokens_in":13578,"feed_emoji":"🩸","tokens_out":6482,"duration_ms":71143,"temperature":0.7,"pith_summary":"HA2HA is a self-supervised denoising framework for ultrasound microvascular imaging (UMI). It builds training pairs by splitting the steered plane-wave angles of beamformed radio-frequency (RF) blood-flow data into odd and even half-angle groups, compounding and clutter-filtering each group separately, so that the vascular signal is shared while noise differs. The paper claims that a network trained on just 90 in-vivo contrast-free pig kidney frames with a Noise2Noise-style loss removes enough noise in unseen contrast-free and contrast-enhanced pig kidney, human liver, and human kidney datasets to raise contrast-to-noise ratio (CNR) and signal-to-noise ratio (SNR) by more than 15 dB and to lower background noise power below that of the compared methods. Because denoising is done directly on RF data rather than on envelope or image-domain data, the cleaned signal also improves downstream color Doppler imaging by suppressing spurious velocity estimates in vessel-free regions. If correct, this gives a label-free, generalizable route to higher-quality microvascular imaging without contrast agents or clean ground truth.","feed_headline":"No-label denoiser lifts ultrasound vascular clarity by 15 dB","feed_subtitle":"Trained once on contrast-free pig kidney data, it cleans unseen pig and human blood-flow images and color Doppler.","key_machinery":"The load-bearing object is the paired-input construction: for each acquisition, the steering angles are split into odd and even groups; each group is coherently compounded and SVD clutter-filtered to produce two RF volumes that share the same vascular signal but carry independent noise. The network is trained with a composite Noise2Noise loss made of a forward term, a reverse term, and a consistency term, with an L1 weight penalty. The theoretical engine is the Noise2Noise identity that a zero-mean, signal-independent noise leaves the conditional expectation of one noisy observation equal to the clean signal, which is what makes a clean-target-free regression valid. The encoder-decoder backbone processes full RF frames at inference, so the denoised RF signal retains phase information needed for color Doppler imaging.","core_discovery":"On the paper's own terms, the central claim is that complementary angular subsets of a plane-wave acquisition provide statistically valid noisy pairs for self-supervised denoising of UMI blood-flow RF data. Because the same vessel geometry appears in both half-angle compounds but the electronic noise, angle-dependent sidelobes, and residual clutter differ, the network can learn the mapping from one noisy observation to the other and thereby converge to the underlying clean vascular signal. The authors demonstrate empirically that a model trained solely on 90 contrast-free pig kidney frames transfers without fine-tuning to contrast-enhanced pig kidney and to human liver and kidney data, producing power Doppler images with more than 15 dB higher CNR and SNR and the lowest background noise power among conventional, angular-processing, and spatiotemporal non-local means baselines, while HA2HA-denoised RF data also yields cleaner color Doppler velocity maps.","pith_inferences":["If the independence and zero-mean noise assumption degrades with stronger angle-dependent sidelobes or tissue motion, the learned target would be a biased version of the true vascular signal; a useful test would be to run HA2HA on simulated plane-wave RF with known ground truth and measure residual bias.","Because training used a single scanner and probe configuration, the reported cross-dataset generalization leaves open whether the same half-angle construction transfers to other frequencies, probes, or imaging protocols; a multi-scanner study would settle this.","The same odd/even angle-splitting recipe could be applied to raw channel data or IQ data rather than beamformed RF, potentially enabling system-level denoising before beamforming and reducing computational cost.","If the 15 dB gains reproduce in controlled phantom or clinical studies, the method could enable reduced acoustic output in contrast-free microvascular imaging, at least down to the transmit-energy regime where it still beats the baselines."],"forward_implications":["A single HA2HA model, trained once on contrast-free pig kidney RF data, can be applied to unseen contrast-enhanced and human datasets without fine-tuning, so no labeled or clean reference data is needed for new sites.","Denoising at the RF level rather than at the envelope or image level preserves phase information, so downstream Doppler processing such as color Doppler imaging inherits the noise suppression.","The method's gains exceed 15 dB in CNR and SNR on the tested pig kidney, human liver, and human kidney volumes, with background noise power lower than the conventional, angular processing, and ST-NLM baselines.","Because the training pairs come from a single angular sweep rather than temporally adjacent frames, the approach avoids motion-related misalignment between noisy pairs.","At the lowest tested transmit energy (DC = 0.1), the method no longer outperforms angular processing, indicating a practical operating range for the self-supervised strategy."],"supporting_citations":[{"why":"Defines the angular processing (AP) baseline that HA2HA is compared against and that also splits steering angles into odd and even groups.","marker":"[14]"},{"why":"Supplies the spatiotemporal non-local means (ST-NLM) post-filtering baseline against which HA2HA's CNR, SNR, and background noise power are measured.","marker":"[20]"},{"why":"Introduces spatiotemporal SVD-based clutter filtering used to isolate blood-flow RF signals before HA2HA training and inference.","marker":"[1]"},{"why":"Provides block-wise adaptive local clutter filtering for ultrasound small-vessel imaging, supporting the SVD clutter-filtering pipeline HA2HA relies on.","marker":"[2]"},{"why":"Applies the Noise2Noise principle to consecutive RF frames in ultrasound, the direct precedent that HA2HA adapts to angular subsets.","marker":"[38]"}],"fun_headline_variants":["Self-supervised denoising pairs half-angles for cleaner ultrasound","15 dB gain in ultrasound microvascular clarity without labels","Ultrasound denoiser learns from half-angle pairs, boosts CNR 15 dB","Trained on pig kidney, cleans human ultrasound: 15 dB better","Half-angle self-supervision cleans ultrasound RF for 15 dB gain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole method depends on the assumption that the two half-angle observations are the true vascular signal plus independent noise that averages to zero, so that one observation is an unbiased guide to the other.","fun_headline_variants_meta":{"raw":{"variants":["Self-supervised denoising pairs half-angles for cleaner ultrasound","15 dB gain in ultrasound microvascular clarity without labels","Ultrasound denoiser learns from half-angle pairs, boosts CNR 15 dB","Trained on pig kidney, cleans human ultrasound: 15 dB better","Half-angle self-supervision cleans ultrasound RF for 15 dB gain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000792,"raw_usage":{"total_tokens":3495,"prompt_tokens":959,"completion_tokens":2536,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":2451}},"tokens_in":575,"tokens_out":2536,"duration_ms":18962,"temperature":1.0,"reasoning_tokens":2451,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:26:01.614719+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train or test HA2HA on simulated plane-wave RF data with a known ground-truth vascular signal and controlled additive electronic noise plus angle-dependent sidelobe interference; if the network output systematically deviates from the known clean signal by a structured, angle-correlated component rather than by zero-mean residual, then the conditional-zero-mean assumption is violated and the denoised images are biased.","supporting_citations":[{"cited_title":"Simultaneous noise suppression and incoherent artifact reduction in ultrafast ultrasound vascular imaging,","cited_arxiv_id":null,"evidence_quote":"Defines the angular processing (AP) baseline that HA2HA is compared against and that also splits steering angles into odd and even groups."},{"cited_title":"Improved ultrafast power Doppler imaging by using spatiotemporal non -local means filtering,","cited_arxiv_id":null,"evidence_quote":"Supplies the spatiotemporal non-local means (ST-NLM) post-filtering baseline against which HA2HA's CNR, SNR, and background noise power are measured."},{"cited_title":"Spatiotemporal clutter filtering of ultrafast ultrasound data highly increases Doppler and fUltrasound sensitivity,","cited_arxiv_id":null,"evidence_quote":"Introduces spatiotemporal SVD-based clutter filtering used to isolate blood-flow RF signals before HA2HA training and inference."},{"cited_title":"Ultrasound small vessel imaging with block-Wise adaptive local clutter filtering,","cited_arxiv_id":null,"evidence_quote":"Provides block-wise adaptive local clutter filtering for ultrasound small-vessel imaging, supporting the SVD clutter-filtering pipeline HA2HA relies on."},{"cited_title":"Hidden Bethe states in a partially integrable model","cited_arxiv_id":"2205.03425","evidence_quote":"Applies the Noise2Noise principle to consecutive RF frames in ultrasound, the direct precedent that HA2HA adapts to angular subsets."}],"review_version":1}