{"id":"750242b1-56e4-4a69-9d01-efe0840d0245","arxiv_id":"2607.18633","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using low-discrepancy sequences to seed stochastic gravitational-wave template banks achieves the same coverage with 12–27% fewer proposal points and nearly unchanged bank size.","lead":"To catch gravitational waves, search pipelines compare detector data with large libraries of predicted waveforms. This paper shows that seeding those libraries with evenly spread 'low-discrepancy' points instead of random ones reaches the same coverage with 12–27% fewer candidates, saving memory and computation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed proposal-set reduction is a post hoc operating-point choice, not a measured property of LDS sampling","rationale":"The reader's weakest_assumption focused on metric accuracy near boundaries (Fig. A1 validating only at central points). While that is a legitimate limitation for absolute recovery fractions, it is not the most load-bearing issue for the central relative claim: both random and LDS banks use the same metric, so metric inaccuracies affect both sides of the comparison roughly equally. The decisive weakness is the experimental protocol for choosing |U|. The paper's own text in Section 4 reveals that LDS |U| was tuned until recovery became 'comparable' — i.e., the savings were selected post hoc. This makes the headline percentages unverifiable from the reported data. The reader's rationale does mention 'proposal-set sizes were selected post hoc' in passing, but their explicit weakest_assumption is the metric accuracy, so I mark partial disagreement on that point. The paper otherwise has genuine value: the top-down algorithm is clear, the slab-wise coarse filter is benchmarked and shown to preserve the exact metric cut, and the observed ~1% change in final template count is consistent with metric-volume intuition. The central claim may well be true; the paper just does not currently demonstrate it in a falsifiable way. The fix is straightforward — a saturation-curve study — so the verdict remains CONDITIONAL, not REJECT. Hence verdict_should_be is UNCHANGED relative to the reader's CONDITIONAL.","tokens_in":16076,"tokens_out":5879,"duration_ms":54808,"concrete_test":"Compute R0.97 as a function of |U| for both random-uniform and LDS proposal sets, with at least 20 realizations per |U|. For Set I use |U| = 200, 250, 300, 350, 400 (×10^3); for Set II use 15, 16, 17.6, 18.5, 20 (×10^6). Fit a smooth saturation curve to each method's mean recovery and determine the minimum |U| needed to reach a fixed target: R0.97 ≥ 96% for Set I and R0.97 ≥ 94% for Set II. The claim is supported only if the LDS-required |U| is lower than the random-required |U| by the reported margins (27.5% and 12%) with non-overlapping bootstrap confidence intervals. If the curves cross or overlap, the reported reduction is an artifact of the chosen operating points.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — 27.5% fewer proposals in 2D and 12% fewer in 3D at comparable recovery — rests entirely on the chosen values of |U| for each method. Section 4 states: 'For LDS sampling, |U| is then adjusted until an average recovery fraction comparable to that of random uniform sampling is achieved.' This is circular: the reported savings are whatever the adjustment produces, with no pre-registered target, equivalence margin, or saturation curve. The 2D random baseline |U|=400K is asserted without showing how recovery degrades at 350K, 300K, etc.; if random sampling at 300K already gives 96.39%, the 27.5% saving is not real. In 3D, the LDS recovery is actually lower (94.10±0.06 vs. 94.22±0.04), and the text concedes 200–300K additional points are needed to match the random baseline. Thus the '12% fewer' is a lower-recovery operating point, not an equivalent-coverage comparison. The final template count changes by only ~1%, but this is consistent with many choices of |U| and does not validate the comparison. Without a protocol that fixes a target recovery and measures the minimum |U| required for each method, the headline claim is unfalsifiable.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes using low-discrepancy sequences (Halton, Sobol, and scrambled variants) as the proposal set for the top-down stochastic template-bank algorithm of Ref. [22], replacing uniform random sampling. In two dimensions (θ0, θ3) and three dimensions (θ0, θ3, θ3s) using IMRPhenomD and the aLIGOZeroDetHighPower PSD with Mmin=0.97, the authors report that LDS requires 27.5% and 12% fewer proposal points than uniform random sampling while achieving 'comparable' recovery fractions (R0.97) on a common set of 10^5 injections, with final template counts differing by about 1%. The paper also introduces a slab-wise coarse-filtering step based on the oriented bounding box of the local metric ellipsoid, with a synthetic benchmark showing candidate-count reductions for anisotropic metrics. The evaluation uses 10 independent realizations per configuration.","tokens_in":16405,"tokens_out":7200,"duration_ms":72860,"significance":"If established, the LDS proposal-set reduction would be a simple, drop-in improvement for stochastic and hybrid template-bank generation, reducing memory and nearest-neighbour costs with no change in final bank size. The paper's strengths are the external random-sampling baseline, common injection set, 10 realizations, and the separate analysis of proposal-set size versus final template count. The slab-wise filter is a neat and potentially useful algorithmic addition. However, the headline percentages are not currently supported by a controlled protocol: the LDS proposal-set sizes were selected post hoc until recovery was 'comparable,' with no saturation curves for random sampling or equivalence margins, and in 3D the LDS recovery is actually slightly lower than the random baseline.","major_comments":[{"comment":"The headline reductions (27.5% in 2D, 12% in 3D) are the result of an after-the-fact choice of |U| for LDS: 'For LDS sampling, |U| is then adjusted until an average recovery fraction comparable to that of random uniform sampling is achieved.' The random baseline sizes (400K, 20000K) are fixed, but no recovery-vs-|U| saturation curves are shown for random sampling. Without such curves one cannot exclude that random sampling at 290K (2D) or 17600K (3D) already achieves the same R0.97, which would eliminate the claimed savings. Please report R0.97 as a function of |U| for each method and dimension, with error bars, and use a pre-specified equivalence margin (e.g., two one-sided tests) to define 'comparable.' This is essential to the abstract's central quantitative claim.","section":"§4, Table 4 and text following"},{"comment":"The Sobol and Sobol-S recovery fractions (94.10±0.06 and 94.13±0.07) are lower than the random baseline (94.22±0.04). The text labels these 'comparable,' but the difference is in the direction adverse to the claim, and with σ≈0.05–0.07 it is about 1.5–2σ. The proposed remedy of adding 200–300K points to 17600K would reduce the headline saving from 12% to roughly 10.5–11%, so the numbers should be recomputed. Please either run LDS at the larger |U| and report the result or provide a formal equivalence test with a pre-specified margin.","section":"§4, Table 4 Set II (3D)"},{"comment":"The sentence 'We therefore determine the required size of U by studying the bank coverage as a function of |U| and selecting the point beyond which the improvement in coverage is negligible' describes a procedure that is never shown. The manuscript does not present the coverage-versus-|U| curves on which the choices 400K/20000K (random) and 290K/17600K (LDS) are based. Without these data the selection is not reproducible and the reported percentage reductions are not testable. Please include those curves for all methods and intermediate |U| values, or release the relevant data/code.","section":"§4, paragraph on determining |U|"}],"minor_comments":[{"comment":"The synthetic slab-filter benchmark uses uniform points in the unit hypercube and random orthogonal metrics with condition number 10^4. The connection to the actual parameter-space anisotropy encountered in Set I and Set II is not established; a sentence reporting the typical κ(g) values would help the reader judge the relevance.","section":"§3.3, Table 2"},{"comment":"The metric is validated only at the central point of each parameter space, while the text says 'validity of this approximation is assessed.' Please add a caveat, or include boundary checks, since the covering claim relies on the metric remaining accurate throughout the domain.","section":"Fig. A1 caption"},{"comment":"The wall-clock speed-ups (4.5% in 2D, 7% in 3D) are based on a single representative run per method. With differences this small, please report the spread over multiple runs or clearly label these as indicative.","section":"§4, CPU-time paragraph"},{"comment":"The precise percentages '27.5%' and '12%' are not warranted by the post hoc protocol; consider reporting ranges (e.g., '10–30%') or softening the quantitative claim until the saturation curves are provided.","section":"Abstract and §1"},{"comment":"In Eq. (3.10) 'V nan max' appears to be a typo for the volume of the n-ball times a_max^n. The Fig. A1 caption has awkward grammar ('we illustrate the validity of this approximation is assessed').","section":"Eq. (3.10) and general typos"}],"recommendation":"major_revision","confidential_remarks":"This is a practically useful paper with a clean evaluation framework, but the central quantitative claim is not yet supported by the current protocol. The post hoc selection of |U| and the negative 3D recovery difference are load-bearing. The slab-wise coarse filter is a nice contribution and could be emphasized more; if the LDS comparison cannot be tightened, the authors might consider making the filter the primary result. No concerns about citation practice or scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a clean, honest engineering study: it applies low-discrepancy sequences (Halton/Sobol) to the proposal set in top-down stochastic template-bank placement for CBC searches, and adds a slab-wise pre-filter that reduces candidate counts for anisotropic metrics. The core idea is not deep—Manca & Vallisneri already showed LDS helps covering in flat spaces—but the extension to a curved CBC metric space with IMRPhenomD and the recovery-fraction measurements are new, and the slab-filter benchmark is genuinely useful. The paper deserves a serious referee.\n\nWhat it does well: the 2D result is clean (27.5% fewer proposals at comparable or slightly better recovery), the 10-realization/100k-injection protocol is solid, and the authors are refreshingly candid that the final template count changes only ~1% because bank size is metric-volume-limited. The runtime savings (5–7%) are modest and the paper says so. The slab-wise filter derivation (Eqs. 3.10–3.11) is neat, and the local benchmark shows a 20–30% speedup for n>3.\n\nThe soft spot: the headline proposal-set reduction is partly an artifact of post hoc adjustment. Section 4 states that for LDS, |U| is adjusted 'until an average recovery fraction comparable to random uniform sampling is achieved.' Without a pre-specified equivalence margin or a saturation curve for both methods, the 27.5% and 12% numbers are operating points, not measured properties. In 3D the LDS recovery is actually slightly lower (94.10±0.06 vs 94.22±0.04), and the text concedes an extra 200–300K points would be needed to match the random baseline. So the '12% fewer' is a lower-recovery point, not an equal-coverage comparison. This doesn't sink the paper—the general space-filling advantage of LDS is real, and Fig. 1 shows it—but it means the quantitative claim should be softened to '10–30% fewer proposals, with recovery differences within a few tenths of a percent.' The metric validation only at central points (Fig. A1) is a minor concern, not a fatal one; the recovery fractions are measured with injections throughout the space, so serious boundary errors would likely show up there. No code or data release limits immediate adoption, but the algorithm is described well enough to reproduce.\n\nMy take: solid incremental work. I'd accept it for peer review with a request to tighten the |U| selection protocol, report recovery as a function of |U| for both methods, and release code/data. I'd bring it to a reading group focused on GW search infrastructure, but it won't change how I build banks tomorrow.\n\nBest,\n[You]","headline":"Honest, modest engineering result: LDS proposal sets for stochastic template banks are tested carefully, but the headline savings are partly an artifact of post hoc |U| choice and the real gain is small.","tokens_in":16893,"tokens_out":2087,"would_cite":true,"duration_ms":24494,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["04.30.-w"],"model":"deepseek-v4-flash","headline":"This paper shows that replacing uniform random sampling with low-discrepancy sequences when proposing candidate templates cuts the required proposal-set size by 27.5% in two dimensions and 12% in three, with essentially no loss in detection","keywords":["gravitational waves","template banks","stochastic placement","low-discrepancy sequences","Halton sequence","Sobol sequence","matched filtering","fitting factor"],"falsifier":"Compute the exact match (not the metric approximation) between waveform pairs at the boundaries of the 2D and 3D parameter ranges used here, and compare the exact mismatch ellipsoids against the metric ellipsoids of Eq. (A.8); if the metric systematically underestimates mismatch near boundaries, the LDS banks would recover a smaller fraction of boundary injections than the claimed R0.97, and the equal-coverage claim would fail for real signals. Alternatively, measure the recovery fraction for a 4D aligned-spin bank with LDS vs random proposals at the same |U|; if LDS does not reduce |U| at equ","tokens_in":15947,"feed_emoji":"🌌","tokens_out":6643,"duration_ms":93612,"temperature":0.7,"pith_summary":"Gravitational-wave searches work by matching detector data against a bank of theoretical waveforms scattered across a parameter space of masses and spins. Building that bank with the standard stochastic recipe requires an oversampled set of random proposal points, and a large fraction of those points are wasted. This paper argues that drawing the proposals from low-discrepancy sequences—Halton and Sobol sets that fill space more evenly than random points—lets the same covering algorithm reach equal recovery fractions with about 27.5% fewer proposals in two dimensions and 12% fewer in three. The final template count barely moves (~1%), because the number of templates is set by the metric volume of the search space, not by how the proposals were drawn. If the result holds, it is a drop-in efficiency upgrade for stochastic and hybrid bank construction: lower memory use, cheaper nearest-neighbour searches, and less bookkeeping, with larger gains expected in higher dimensions.","feed_headline":"Quasirandom proposals trim GW template-bank needs by 27.5%","feed_subtitle":"Extra-even point sets match detection rates with fewer candidates, easing memory and search costs for LIGO-era banks.","key_machinery":"The central objects are low-discrepancy sequences—deterministic point sets such as Halton and Sobol that fill the unit cube more uniformly than pseudo-random points—and the top-down stochastic placement algorithm that consumes them. Proposals are drawn once from the chosen sequence over the chirp-time coordinates (θ0, θ3, θ3s), mapped to physical masses and spins, and filtered by physical constraints. A KD-tree is built once over the surviving proposals; at each iteration a random active proposal is promoted to the template list, a conservative Euclidean ball (scaled by the inverse square root of the smallest metric eigenvalue) retrieves candidates, a slab-wise filter prunes those outside th","core_discovery":"The central claim is that low-discrepancy sequences (LDS) are a drop-in replacement for uniform random sampling in the proposal stage of top-down stochastic template placement. Using the IMRPhenomD waveform model and the standard metric-based mismatch ellipsoid, the authors construct 2D and 3D banks over chirp-time coordinates and show that Halton (2D) and Sobol (3D) proposal sets achieve statistically indistinguishable recovery fractions—96.4% and 94.1% respectively—with 27.5% and 12% fewer proposal points than random sampling. The final template count is reduced by only about 1%, which they interpret as confirming that bank size is governed by metric volume and minimal match, not by propos","pith_inferences":["An immediate testable extension is to plug LDS proposal sets into hybrid geometric-random placement (which currently starts from uniformly sampled U); the same 10–30% reduction in |U| should carry over, and the paper notes this as a natural next step.","The 27.5%/12% savings are measured for one waveform family (IMRPhenomD) and one PSD; the generic claim—that better space-filling at the proposal stage reduces waste—should transfer to precessing and eccentric waveform families, where the parameter space is higher-dimensional and the metric is less reliable, but the size of the gain needs re-measuring.","The paper implicitly predicts that the savings increase with dimension: at fixed covering radius, the covering-radius advantage of low-discrepancy sets over random points grows in higher dimensions, so a 4D or 5D intrinsic space should show a larger reduction in |U| than 12%.","The slab-wise filter's speed-up is independent of anisotropy but assumes a locally uniform proposal density; in regions where the physical constraint boundaries leave holes in U, the filter's candidate-count ratios may differ—an effect the controlled benchmark in Table 2 does not capture."],"forward_implications":["For 2D banks over (θ0, θ3), LDS proposal sets deliver the same recovery fraction (≈96.4%) with 27.5% fewer proposals than uniform random sampling; for 3D banks over (θ0, θ3, θ3s), the saving is ≈12% at ≈94.1% recovery.","Final template counts change by only ~1%, confirming that bank size is set by metric volume and minimal match, not by the sampling strategy.","Because memory and nearest-neighbour overhead scale with the proposal set, LDS sampling reduces memory usage and bookkeeping without touching the covering logic—applicable to stochastic and hybrid placement alike.","The slab-wise coarse filter reduces local covering-step time by 20–30% for dimensions > 3, independent of metric anisotropy, making higher-dimensional banks cheaper to build.","The benefit is expected to grow with dimensionality, where adequate coverage demands proposal sets far larger than the final bank."],"fun_headline_variants":["Quasirandom sampling needs 27.5% fewer GW template-bank proposals","Low-discrepancy sequences cut proposal count in GW banks by 27.5%","GW template banks: quasirandom proposals slash 27.5% off prep","Quasirandom proposals: 27.5% fewer tries for same GW coverage","Low-discrepancy sampling: 27.5% fewer proposals for GW banks"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The whole placement pipeline assumes that the local quadratic metric g_ij (built from the Fisher matrix) accurately predicts waveform mismatch everywhere in the search space, including boundaries and strongly anisotropic regions—the paper validates it only at a representative central point in Fig. A1.","fun_headline_variants_meta":{"raw":{"variants":["Quasirandom sampling needs 27.5% fewer GW template-bank proposals","Low-discrepancy sequences cut proposal count in GW banks by 27.5%","GW template banks: quasirandom proposals slash 27.5% off prep","Quasirandom proposals: 27.5% fewer tries for same GW coverage","Low-discrepancy sampling: 27.5% fewer proposals for GW banks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000665,"raw_usage":{"total_tokens":2895,"prompt_tokens":791,"completion_tokens":2104,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":535,"completion_tokens_details":{"reasoning_tokens":1995}},"tokens_in":535,"tokens_out":2104,"duration_ms":106966,"temperature":1.0,"reasoning_tokens":1995,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:46:53.794608+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the exact match (not the metric approximation) between waveform pairs at the boundaries of the 2D and 3D parameter ranges used here, and compare the exact mismatch ellipsoids against the metric ellipsoids of Eq. (A.8); if the metric systematically underestimates mismatch near boundaries, the LDS banks would recover a smaller fraction of boundary injections than the claimed R0.97, and the equal-coverage claim would fail for real signals. Alternatively, measure the recovery fraction for a 4D aligned-spin bank with LDS vs random proposals at the same |U|; if LDS does not reduce |U| at equ","supporting_citations":[],"review_version":1}