{"id":"14c2d4d9-da03-4266-b86a-2ad5652ecbe7","arxiv_id":"2606.04879","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"bootESA constructs empirical confidence intervals and one-sample tests for elastic shape distance of image contours via an m-out-of-N bootstrap that accounts for ESD non-differentiability.","lead":"The paper introduces bootESA, a bootstrap hypothesis test for whether an estimated 2D image contour matches a hypothesized shape under elastic shape distance. It matters for quantifying symmetry and shape uncertainty in noisy scientific images such as inertial confinement fusion neutron reconstructions.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged Assumptions 1–2; the bootstrap justification and simulations hold under the paper's stated conditions.","rationale":"The paper's central claim is the construction of a one-sample ESD test via m-out-of-N bootstrap that accounts for the non-smoothness of the elastic metric. The mathematical scaffolding (consistency of the mean image, continuous mapping through the SRVF under Assumptions 1–2, directional differentiability of the ESD, and the m-out-of-N remedy) is standard and correctly applied. Simulations (Section 4) and the real-data illustration (Section 5) support the practical recommendations. The free parameters ε and p are acknowledged and explored; data-on-request and missing code are reproducibility limitations already noted by the reader. Because the only genuine soft spot is precisely the one the reader already flagged, no adjustment to the CONDITIONAL verdict is warranted.","tokens_in":19877,"tokens_out":530,"duration_ms":5753,"concrete_test":"Re-run the Type-I error experiment of Table 3 at p=0.65 with N=500, forcing a controlled topology change (e.g., additive high-frequency noise that splits the 0.65-level set into two components on ~10% of bootstrap replicates) and record whether the empirical rejection rate under H0 remains ≤α or inflates; if it stays controlled, the method is robust even when Assumption 2 is mildly violated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the load-bearing point: continuous differentiability of the image-to-SRVF map (and therefore validity of the ordinary or m-out-of-N bootstrap for the ESD functional) rests on Assumptions 1–2 in Section 3.2. Those assumptions are standard for percentile contours of smooth images and are checked in the ICF simulations (Figure 4, Table 3). The non-differentiability of the ESD itself is handled by the m-out-of-N construction (Section 3.2.1), whose practical choice of ε is explored and shown to recover Type-I error near the nominal level for the recommended ε=0.8. No deeper internal inconsistency or hidden failure mode of the central claim (Eqs. 9–12) is apparent once those assumptions are granted.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes bootESA, a one-sample hypothesis test for 2D contours under elastic shape analysis. After reconstructing a mean image from multiple observations (e.g., pinhole sub-images), a percentile contour is extracted, converted to its square-root velocity function, and compared to a hypothesized shape via the elastic shape distance (ESD). Because the ESD is non-differentiable at zero, ordinary bootstrap can be invalid; the authors therefore construct empirical confidence intervals and p-values with an m-out-of-N bootstrap (Eqs. 9–12, Sections 3.1–3.2). Validity rests on two smoothness/topology assumptions that make the image-to-SRVF map continuously differentiable. Simulations examine Type I error under different m = N^ε and percentiles (Table 3) and power against indented circles, ellipses and polygons (Figure 6); two real ICF neutron images illustrate the procedure.","tokens_in":20081,"tokens_out":1057,"duration_ms":10541,"significance":"Formal one-sample inference for ESA contours has been largely missing; existing ESA tests are multi-class permutation or energy tests. A bootstrap procedure that explicitly accounts for the non-smoothness of the ESD is therefore a useful methodological contribution, especially for imaging applications (ICF, TDA-style density contours) where a target shape is of scientific interest. The paper supplies concrete Type I error and power evidence under the stated assumptions and demonstrates that a practical choice ε ≈ 0.8 recovers near-nominal size for mid-percentiles. The framework is generalizable beyond ICF once a consistent mean image and bootstrap of that mean are available.","major_comments":[{"comment":"Section 3.2.1 and Table 3: the practical selection of ε (hence m = N^ε) is left largely empirical. Table 3 shows that m = N yields Type I error near 0 for most percentiles, while ε = 0.6–0.7 can inflate it well above α = 0.1; ε = 0.8 works for p ≤ 0.85 but still under-covers at p = 0.95. The supplemental stability plots (Figure 11) are helpful but do not constitute a data-driven rule that a practitioner can apply without knowing the true distribution. A clearer, reproducible recommendation (or a diagnostic that flags when ordinary bootstrap is safe) is needed for the method to be usable outside the authors’ simulations.","section":null},{"comment":"Assumptions 1–2 (Section 3.2) are load-bearing for continuous differentiability of the image-to-SRVF map and therefore for both ordinary and m-out-of-N bootstrap validity. Figure 4 shows that topology can break for N ≲ 30; the real-data examples use N = 55 and m = N “to be conservative,” yet no quantitative check of smoothness or single-component topology is reported for those images. The paper should either supply a practical diagnostic or quantify how mild violations affect coverage, otherwise the real-data conclusions rest on unverified assumptions.","section":null},{"comment":"Section 5 and Figure 8: the real-data analysis tests every percentile contour separately at α = 0.05 without multiplicity adjustment, while the text notes that a single p is chosen a priori in practice. Presenting a battery of unadjusted tests invites over-interpretation of which percentiles “are circular.” Either restrict the main analysis to one pre-specified percentile or apply a multiple-testing correction and discuss the hierarchical dependence across percentiles.","section":null}],"minor_comments":[{"comment":"Notation for the bootstrap scale is inconsistent: √m appears in Eqs. 9–10 while the test statistic later uses √N (Section 3.2.1); a single, explicit statement of the scaling used for R_obs versus R*_m would help.","section":null},{"comment":"Table 1 caption and surrounding text refer to “spherical harmonic fit” while the method is Legendre polynomial decomposition; the terminology should be aligned.","section":null},{"comment":"Figure 5c and the corresponding violin plots in the supplement would be clearer if the asymptotic scaling (√m or √N) were applied before plotting, so that the visual comparison matches the quantities used in the hypothesis test.","section":null},{"comment":"A few typographical issues: “bootESAfor short,” “ak-sample,” “them-out-of-N,” and missing spaces after periods appear in the introduction and Section 3.","section":null},{"comment":"The claim of being “the first ESA-based one-sample test” is plausible given the cited multi-class work, but a brief comparison with existing Procrustes or landmark-based one-sample bootstrap procedures would better situate the contribution.","section":null}],"recommendation":"major_revision","confidential_remarks":"The methodological core is sound under the stated assumptions and the simulations are informative; the main obstacles to acceptance are the underspecified choice of m and the lack of diagnostics for Assumptions 1–2 on the real data. These are fixable within a revision. Scope is appropriate for a statistics/methodology journal with an imaging application."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is the first one-sample hypothesis test for elastic shape distance on contours. Prior ESA inference (Strait 2017, Zhang 2021) was multi-sample; they correctly flag non-differentiability of ESD at the origin and use m-out-of-N bootstrap to build empirical CIs and a test (Eqs. 9–12). That is the real contribution.\n\nWhat works: the setup is clean. They map images to percentile contours to SRVFs to ESD, justify the ordinary bootstrap failing under directional differentiability, and show via simulation (Table 3, Figures 5–6) that m = N^0.8 roughly recovers Type I error near 0.1 for mid-percentiles while ordinary bootstrap is overly conservative. Power against indented circles, ellipses, and polygons is high once N is moderate. The ICF examples are illustrative and the supplemental KDE manifold example shows the method is not locked to neutron imaging. Math is standard once you accept the non-smooth functional; citations cover the ESA and bootstrap literature without padding.\n\nSoft spots are real but proportionate. Assumptions 1–2 (smooth non-degenerate pixel distribution, fixed single closed topology under small perturbations) are load-bearing for continuous mapping and bootstrap validity; they check them for ICF (Figure 4) and note failures only at tiny N. Free parameters remain: ε for m = N^ε, percentile p, and B. They explore ε and recommend conservative m = N when images are smooth, but do not automate selection. Real data are on request and no code ships, so reproducibility is limited. Real-data results are not a formal validation. None of this breaks the central claim under the stated conditions.\n\nThis is for people already working in ESA or needing shape diagnostics on image contours (ICF symmetry, TDA-style density contours). A serious referee should see it; the gap is genuine and the simulations support the practical advice. I would engage and expect revision mainly around parameter guidance and code/data release.","headline":"Solid first one-sample ESD test via m-out-of-N bootstrap; usable for ICF and generalizable, with free parameters and data-on-request as the main soft spots.","tokens_in":20683,"tokens_out":518,"would_cite":true,"duration_ms":5949,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G09","62H35","62G10"],"pacs":[],"model":"grok-4.5","headline":"A one-sample bootstrap test for whether an image contour matches a hypothesized shape, using elastic shape distance.","keywords":["elastic shape analysis","elastic shape distance","bootstrap confidence intervals","one-sample hypothesis test","percentile contours","inertial confinement fusion","m-out-of-N bootstrap","non-smooth functionals"],"falsifier":"On synthetic images whose true source is known to be a circle, run the test at the recommended m = N^0.8 and check whether the observed Type I error approaches the nominal level (e.g., 0.1) as sample size grows; systematic under- or over-rejection would falsify the calibration claim.","tokens_in":20768,"feed_emoji":"○","tokens_out":621,"duration_ms":5498,"temperature":0.7,"pith_summary":"Elastic shape analysis can measure how two curves differ after removing rotation, scale, translation, and reparameterization, but it has lacked a simple one-sample hypothesis test for contours extracted from images. This paper supplies that test: construct an empirical confidence interval for the elastic shape distance between a hypothesized true contour and the contour estimated from data, then reject if the hypothesized shape falls outside the interval. The interval is built by an m-out-of-N bootstrap that handles the fact that the elastic distance is not differentiable at zero. Simulations show controlled Type I error once the bootstrap sample size is chosen appropriately, and power against common deviations such as indents, ellipses, and polygons. Real neutron images from inertial confinement fusion experiments illustrate when reconstructed source shapes can or cannot be treated as circular.","feed_headline":"Bootstrap test asks if an image contour matches a target shape","feed_subtitle":"Elastic shape distance plus m-out-of-N bootstrap yields the first one-sample ESA test for contours","key_machinery":"bootESA: map images to percentile contours, convert to square-root velocity functions, form the elastic shape distance, then build a rescaled bootstrap distribution of those distances (with m = N^ε when the ordinary bootstrap is invalid) to produce the confidence set and p-value.","core_discovery":"The first elastic-shape-analysis one-sample test for image contours: empirical confidence intervals for the elastic shape distance between a proposed true shape and an estimated shape, obtained from an m-out-of-N bootstrap that accounts for the non-differentiability of that distance.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Bootstrap one-sample test for elastic shape of image contours","m-out-of-N bootstrap CIs for elastic shape distance of contours","ESA hypothesis test checks if contour matches proposed true shape","Non-smooth bootstrap tests 2D contours via elastic shape analysis","Empirical ESD intervals for one-sample contour shape testing"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The mean image must be smooth enough that the chosen percentile contour stays a single closed curve under small data changes; if topology flips or the intensity map is too rough, the bootstrap justification fails.","fun_headline_variants_meta":{"raw":{"variants":["Bootstrap one-sample test for elastic shape of image contours","m-out-of-N bootstrap CIs for elastic shape distance of contours","ESA hypothesis test checks if contour matches proposed true shape","Non-smooth bootstrap tests 2D contours via elastic shape analysis","Empirical ESD intervals for one-sample contour shape testing"]},"model":"grok-4.5","effort":"low","cost_usd":0.004034,"raw_usage":{"total_tokens":1205,"prompt_tokens":703,"num_sources_used":0,"completion_tokens":88,"cost_in_usd_ticks":40340000,"prompt_tokens_details":{"text_tokens":703,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":414,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":703,"tokens_out":88,"duration_ms":4269,"temperature":1.0,"reasoning_tokens":414,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T15:08:38.005228+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On synthetic images whose true source is known to be a circle, run the test at the recommended m = N^0.8 and check whether the observed Type I error approaches the nominal level (e.g., 0.1) as sample size grows; systematic under- or over-rejection would falsify the calibration claim.","supporting_citations":[],"review_version":2}