{"id":"d23f55b0-9ba1-4528-90c6-0fa0b1536d61","arxiv_id":"2607.24579","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":6,"one_line_summary":"Bottleneck, Wasserstein, persistence-landscape, and persistence-image measures are more robust to Gaussian noise and Gaussian/ML denoising of synthetic 3D porous-media images than generator-count or average-lifespan statistics.","lead":"The paper compares how common persistent-homology summaries hold up when synthetic 3D porous-media images are noised and then denoised. It finds distance- and vector-based summaries stay more stable than raw generator counts or average lifespans, which matters for anyone using topology on noisy scans.","discovery_kind":"extension","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The robustness ranking is partly built into the vectorizations: landscapes keep only the top 100 envelopes and persistence images downweight short lifetimes, whereas ΔN and ΔL count every noise-born generator.","rationale":"The reader's stated weakest assumption concerns whether i.i.d. Gaussian synthetic noise transfers to real µCT noise. That remains a genuine external-validity limitation, but it is not the most direct threat to the comparative synthetic-data claim. The reader did separately note the free PI weights and landscape truncation, so agreement is partial: this critique turns that acknowledged parameter dependence into the load-bearing comparability issue.\n\nThe paper is transparent about truncating landscapes at 100 envelopes and about the PI weighting choice, and the discussion even speculates that a larger weighting power might further suppress short intervals. This is not an internal contradiction or evidence of carelessness. It is, however, central to interpreting the ranking: the supposedly less robust statistics are precisely the summaries that retain all noise-born generators, while several supposedly robust summaries suppress them. The authors' narrower recommendation—to discount small intervals and distrust persistence statistics in noisy applications—may remain well motivated. The broader conclusion that the selected vectorizations and distances are generally better indicators of denoising quality should be conditioned on an equal-weighting ablation or an independent downstream quality criterion. Because the reader already gave a CONDITIONAL verdict, I would leave that verdict unchanged while adding this as an explicit condition.","tokens_in":27849,"tokens_out":1136,"duration_ms":226225,"concrete_test":"Run a matched-weight ablation on representative 128³ volumes from each morphology over the same Gaussian-σ grid: after computing each original/denoised PD, filter both PDs to lifespan ≥t for several common thresholds t, then recompute all six measures using no additional lifespan weighting and untruncated—or systematically increasing-K—landscapes. Include t=0 with unweighted PIs. If the L2/diagram measures remain lower and flatter and select the same σ across t and K, the hierarchy is substantive; if the ordering changes materially, it is an artifact of unequal low-persistence weighting.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The measures are not compared at a common level of feature weighting. Sec. 5.1.5 restricts the landscape calculation to the first 100 upper envelopes, which the authors themselves note has a denoising effect. Sec. 3.2 constructs persistence images with m=1 and θ=Lmax, so essentially every nonmaximal interval is linearly downweighted by lifespan, strongly suppressing the short intervals created by noise. Bottleneck distance likewise ignores the multiplicity of small generators through its minimax definition. By contrast, ΔN_i in Eq. (1) counts every generator and ΔL_i in Eq. (2) averages every lifespan. Since the added i.i.d. noise predominantly creates enormous numbers of short-lived generators (Table 1), the compared measures include the perturbation to very different degrees by construction. The lower or flatter normalized curves may therefore reflect the summaries' built-in filtering rather than establish that they are better indicators of overall noising/denoising quality. The different normalizing denominators also do not make error magnitudes or curve slopes commensurate across measures. This does not invalidate the within-measure recovery results, or the more limited conclusion that these summaries are appropriate when only dominant features matter, but it leaves the broad robustness hierarchy under-identified.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript studies how six topological summaries of 3D grayscale porous-media images respond to a controlled noising/denoising pipeline. Three synthetic dataset families (Fourier-series, overlapping-sphere PuMA, Worley/cellular) are generated with known ground truth, corrupted with i.i.d. Gaussian voxel noise (σ=5.76, calibrated from blank regions of a coquina µCT scan), and denoised either by Gaussian convolution with varying bandwidth σ or by a 3D U-Net trained on 450 volumes per morphology (288/72/90 split). Six normalized measures are compared between original and denoised images: generator-count difference ΔN_i (Eq. 1), average-lifespan difference ΔL_i (Eq. 2), normalized bottleneck and Wasserstein-2 distances, and normalized L2 distances between persistence landscapes (top 100 envelopes) and persistence images. The authors find a consistent optimal Gaussian bandwidth near σ≈0.5, which coincides with the minimizer of ∥f_o−f_d∥_∞ (Fig. 7), and conclude that L2/distance-based measures (bottleneck, Wasserstein, landscapes, and to a lesser degree images) are more robust indicators of denoising quality than the persistence statistics ΔN and ΔL, under both denoising paradigms.","tokens_in":23355,"tokens_out":5744,"duration_ms":203045,"significance":"If the conclusions hold, the paper gives practitioners of TDA on µCT porous-media data useful, concrete guidance: which PH summaries can be trusted after denoising, and a practical observation that the optimal Gaussian bandwidth is well tracked by the (computationally cheap) sup-norm of the image difference. The experimental design has real strengths: known ground truth on three morphologically distinct synthetic families, averages over 10 realizations (Gaussian path) and 90 hold-out volumes (ML path), and an explicit numerical verification of the bottleneck stability bound (Fig. 7) that also yields the falsifiable observation that most measures are simultaneously minimized near the ∥·∥_∞ minimizer. The within-measure results — existence and location of an optimal denoising level, and the unreliability of ΔL at high σ — appear solid. The significance is tempered by the fact that the headline cross-measure robustness hierarchy is partly under-identified (Major Comment 1), and by the restriction to a single, spatially uncorrelated noise model, which limits transfer to real scans.","major_comments":[{"comment":"The cross-measure robustness hierarchy — the paper's central comparative claim — is confounded by heterogeneous treatment of short-lived generators. The compared measures do not operate at a common level of feature weighting: (a) the landscape measure PL_i uses only the first 100 upper envelopes (Sec. 5.1.5), which the authors themselves note 'has a denoising effect' (Sec. 6); (b) the persistence image uses m=1, θ=L_max in Eq. (6), so every nonmaximal interval is linearly downweighted by lifespan; (c) the bottleneck distance is insensitive by its minimax definition to the multiplicity of near-diagonal points; whereas (d) ΔN_i (Eq. (1)) and ΔL_i (Eq. (2)) count every generator with equal weight. Since the added i.i.d. noise predominantly creates enormous numbers of short-lived generators (Table 1: e.g., Cellular N_1 grows from ~12k to ~1.3M), the compared measures include the perturbation","section":"Secs. 3.2, 5.1.5–5.1.6, 6 (Eqs. (1)–(2), Def. 4/Eq. (6))"},{"comment":"All curves are stated to be averages over 10 realizations (Sec. 2), but no measure of variability (error bars, shaded bands, standard deviation) is shown anywhere in the Gaussian-denoising results. This matters specifically where the paper draws conclusions from fine structure: the oscillatory behavior and secondary minima of ΔL for PuMA at σ≈5 and σ≈7 (Fig. 9b), the 'minor oscillatory behavior' acknowledged for PI (Fig. 13), and the precision of the optimal band σ∈[0.5,0.7]. With n=10, realization-to-realization spread could plausibly explain these features; if it does not, showing the spread would substantially strengthen the claims. Please add variability indicators, at least for Figs. 9 and 13.","section":"Sec. 5.1, Figs. 8–13"},{"comment":"The ML-based path reports only four of the six measures, excluding bottleneck and Wasserstein distances 'due to prohibitive computation times' (Sec. 5.2). The Discussion nonetheless states that 'the measure robustness hierarchy identified under Gaussian denoising ... carries over to the ML setting.' That claim is only partially supported: two of the measures identified as most robust in Sec. 5.1 are exactly the ones missing, so the ML evidence bears only on PL/PI vs. ΔN/ΔL. Either temper the sentence to scope the carried-over hierarchy to the four measured quantities, or provide BD/W_2 on a subsample (e.g., 10 of the 90 test volumes, possibly at reduced resolution) to justify the stronger statement.","section":"Sec. 5.2 vs. Sec. 6"},{"comment":"The noising step clips voxel values to [0,255], but the Cellular dataset's histogram is strongly skewed toward dark values (Fig. 4c), so for a large fraction of its voxels the additive N(0,5.76) noise is one-sided after clipping — i.e., the effective noise on the dataset that drives most of the 'less robust' verdicts for ΔN/ΔL (broad ML distributions in Figs. 14c–15c, anomalous i=2 optimum σ≈2 in Fig. 8c, ~100× generator explosion in Table 1) is not the i.i.d. Gaussian the paper assumes. Please quantify the clipped fraction per dataset and discuss how much of Cellular's outlier status is attributable to clipping-induced noise asymmetry rather than to morphology per se. This also bears on the transferability of the σ≈0.5 regime and the rankings to real µCT noise, which may be spatially correlated or intensity-dependent; the Discussion's one-sentence limitation should be expanded according","section":"Sec. 4.1 (noise model), Table 1, Figs. 8c, 14c, 15c"}],"minor_comments":[{"comment":"Step (ii) wraps out-of-range values around modulo 256 (overflow 255→0, underflow 0→255). Since convolution with a normalized Gaussian kernel is a convex combination of its inputs, convolved values cannot leave [min,max] of the (already clipped) noisy image, so wrap-around should never trigger; if it does trigger in practice (e.g., due to the integer truncation in step (i)), it would create spurious extreme gradients (white voxels becoming black) with large topological consequences. Please clarify, or replace with clipping for consistency with the noising step.","section":"Sec. 4.1, denoising normalization"},{"comment":"Implementation details of the persistence image are incomplete: the grid resolution M×N (Def. 5) is never stated, and the units of the kernel variance σ=0.5 (birth–lifespan grayscale units?) are ambiguous. These materially affect the measure; please report them. Similarly, the common mesh used for the landscape L2 computation (Sec. 5.1.5) is unspecified.","section":"Sec. 3.2, Def. 4 / Sec. 5.1.6"},{"comment":"The weight function is written in terms of |x−y|, but the persistence surface is constructed on the transformed (birth, lifespan) coordinates where lifespan is the second coordinate; as written it is unclear whether |x−y| refers to pre- or post-transformation coordinates. Please make the notation consistent.","section":"Eq. (6)"},{"comment":"Unlike the other normalized measures, ΔL_i can exceed 1 (the Cellular distributions in Fig. 15c extend to ~3), so its values are not on the same [0,1] scale the Discussion implicitly assumes when comparing measure magnitudes. Worth one sentence when the measure is introduced (Eq. (2)).","section":"Secs. 5.2.2, Fig. 15"},{"comment":"Caption reads 'over 90×3 testing datasets' while Figs. 15–17 say 'over 90 testing datasets'; please make the captions consistent (presumably 90 volumes × 3 homological dimensions).","section":"Fig. 14 caption"},{"comment":"The Fourier coefficient decay 1/(1+i^{1/2}+j^{1/2}+k^{1/2}) is unusually slow (exponent 1/2), giving substantial high-frequency content; a one-line motivation (or a note on how the choice affects the multiscale feature distribution) would help readers judge representativeness.","section":"Sec. 2.1"},{"comment":"No code or data availability statement is given. Given that the generators (PuMA, Porespy/PyFastNoiseSIMD) and the PH code (Cubicle) are public, releasing the noising/denoising scripts and U-Net configuration would make the study reproducible and is standard for the venue.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The work is a competent, well-organized empirical study and a reasonable fit for the journal's applied computational-geometry scope. My main reservation, detailed in Major Comment 1, is that the headline robustness ranking conflates the measures with the short-lifespan filtering embedded in their computation; the authors' own closing recommendation shows they are aware of the issue, so I expect a revision can resolve it through reframing plus a modest control experiment rather than new methodology. No concerns about novelty claims or citation practice."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is empirical and narrow: on three synthetic porous-media ensembles with known ground truth, normalized bottleneck, W2, persistence landscapes, and persistence images recover more cleanly under Gaussian convolution and a 3D U-Net than generator-count and average-lifespan statistics. That ranking, plus the dataset-dependent blow-ups of ΔN and ΔL on the Cellular volumes, is the actual new result. It is incremental relative to Turkeš et al. and the existing porous-media denoising literature, but it is cleanly executed.\n\nWhat they do well: fixed i.i.d. Gaussian noise, averages over 10 realizations (Gaussian path) and 90 hold-outs (ML path), explicit check that bottleneck stays under the L∞ stability bound (Fig. 7), and honest reporting that average lifespan oscillates and that Cellular is harder. The synthetic generators (Fourier, PuMA spheres, Worley) are reasonable stand-ins for the morphologies they claim. Citations are appropriate; no invented entities.\n\nThe stress-test concern is real and should be stated plainly. Landscapes are truncated to the top 100 envelopes (authors note this itself denoises), persistence images use m=1 and θ=Lmax so short intervals are linearly down-weighted, and bottleneck ignores multiplicity by definition. ΔN and ΔL count every noise-born generator. The hierarchy therefore partly reflects built-in filtering rather than a pure head-to-head of “robustness.” Different normalizers also make curve heights non-commensurate. This does not kill the within-measure recovery plots or the practical advice to prefer vector/diagram summaries when dominant features matter; it does mean the broad claim needs that caveat. Other soft spots are secondary: single uncorrelated noise model (σ=5.76 from one coquina blank), synthetic-only evaluation, free parameters in PI/landscapes/U-Net, and no code/data release. ML section drops bottleneck/Wasserstein for cost.\n\nWho it is for: people already computing sublevel PH on noisy 3D greyscale volumes who need a concrete ranking of summaries after denoising. Not a theory paper and not a general denoising advance. I would send it to referees; the experiments are solid enough to deserve that time, with the filtering caveat and transfer limits made explicit. I would cite the ranking if I were writing on PH of porous media, with the caveat attached.","headline":"Useful controlled ranking of PH summaries under matched noising/denoising on synthetic porous volumes; the hierarchy is partly baked into how the summaries treat short intervals.","tokens_in":24022,"tokens_out":614,"would_cite":true,"duration_ms":10426,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["55N31","68U10","62H35"],"pacs":[],"model":"grok-4.5","headline":"Vectorized and distance-based persistent-homology measures stay reliable under noise and denoising of 3D porous images; raw generator counts and average lifespans do not.","keywords":["persistent homology","image denoising","porous media","bottleneck distance","Wasserstein distance","persistence landscapes","persistence images","topological data analysis"],"falsifier":"Repeat the same measure comparison on real µCT volumes whose noise is known to be spatially correlated or intensity-dependent; if generator-count and lifespan measures suddenly become as stable as the vectorized distances, the claimed ranking fails.","tokens_in":23883,"feed_emoji":"🔬","tokens_out":778,"duration_ms":13352,"temperature":0.7,"pith_summary":"Noise in 3D grayscale images of porous media floods persistent homology with millions of short-lived generators, making computation expensive and any analysis that uses generator counts unreliable. This paper tests how well common topological summaries survive a controlled noising-then-denoising cycle on three families of synthetic porous volumes. After adding spatially uncorrelated Gaussian noise, the authors smooth with either a Gaussian filter or a trained 3-D U-Net and compare the cleaned images with the known originals. They find that normalized bottleneck and Wasserstein distances, persistence landscapes, and persistence images recover a clear minimum near a modest smoothing strength and stay close to the original topology across a useful range of parameters; normalized generator counts and average lifespans swing wildly and are dataset-dependent. The practical message is that practitioners who must denoise before computing persistent homology should trust the vectorized and diagram-distance summaries and treat raw persistence statistics with caution.","feed_headline":"Which PH measures survive noise in 3D porous images","feed_subtitle":"Vectorized distances stay reliable; raw generator counts and lifespans do not","key_machinery":"A controlled original-noisy-denoised pipeline on three synthetic porous datasets, evaluated by six normalized topological measures (generator count, average lifespan, bottleneck, Wasserstein-2, persistence-landscape L2, persistence-image L2).","core_discovery":"On synthetic 3-D porous-media volumes, L2-based vectorizations (persistence landscapes and images) and diagram distances (bottleneck, Wasserstein-2) are consistently more robust indicators of noising/denoising quality than scalar persistence statistics (normalized generator count and average lifespan), under both Gaussian convolution and U-Net denoising.","pith_inferences":["If the ranking holds for real correlated noise, software libraries could expose a default “robust PH score” based on landscapes or images instead of Betti curves.","The observation that restricting to the first 100 landscapes already acts as a soft denoiser suggests a cheap pre-filter before full barcode computation.","Edge-preserving or topology-aware denoisers could be scored by the same pipeline to decide whether they improve on simple Gaussian smoothing for PH fidelity."],"forward_implications":["Denoising pipelines for sub/super-levelset PH should monitor landscape or image L2 distance (or bottleneck/Wasserstein) rather than raw Betti numbers.","A modest Gaussian bandwidth near the noise standard deviation simultaneously minimizes most robust measures across homology dimensions.","U-Net denoising preserves landscape and image structure more reliably than generator counts or average lifespans across heterogeneous pore geometries.","Persistence statistics that count or average short intervals should be discounted or heavily weighted when noise is present.","The same robustness hierarchy appears under both classical smoothing and learned denoising, suggesting it is intrinsic to the measures rather than the denoiser."],"fun_headline_variants":["L2 PH vectorizations beat scalar stats under 3D noise","Bottleneck and Wasserstein hold up after porous-image denoising","Persistence images stay reliable when generator counts fail","Diagram distances track denoising quality better than lifespans","Which PH measures remain stable on noisy 3D porous volumes"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Real experimental noise is adequately captured by spatially uncorrelated Gaussian noise of fixed strength, so the robustness ranking found on these synthetic trials transfers to actual scans.","fun_headline_variants_meta":{"raw":{"variants":["L2 PH vectorizations beat scalar stats under 3D noise","Bottleneck and Wasserstein hold up after porous-image denoising","Persistence images stay reliable when generator counts fail","Diagram distances track denoising quality better than lifespans","Which PH measures remain stable on noisy 3D porous volumes"]},"model":"grok-4.5","effort":"low","cost_usd":0.00286,"raw_usage":{"total_tokens":963,"prompt_tokens":675,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":28604000,"prompt_tokens_details":{"text_tokens":675,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":218,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":675,"tokens_out":70,"duration_ms":4686,"temperature":1.0,"reasoning_tokens":218,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T11:20:23.585685+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Repeat the same measure comparison on real µCT volumes whose noise is known to be spatially correlated or intensity-dependent; if generator-count and lifespan measures suddenly become as stable as the vectorized distances, the claimed ranking fails.","supporting_citations":[],"review_version":1}