{"id":"6842fc2c-75b2-47a3-8dee-89af51b7c636","arxiv_id":"2608.09203","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"A Poisson-Fisher degeneracy index (β, γ) predicts which endpoint of a rare-event search interval is vulnerable to model misspecification and whether a goodness-of-fit test can detect the distortion.","lead":"This paper introduces a two-number diagnostic that predicts whether a model error will quietly break the stated confidence of rare-event searches, and which end of the reported interval fails. It offers experimentalists a cheap screening step before expensive toy Monte Carlo calibration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Boundary use of Eq. (7) is not well-defined when a shape parameter is unidentified at μ=0, because the tangent matrix has a zero column; the reported β may depend on an arbitrary reference branch, so the Fig. 9 endpoint prediction needs an explicit convention.","rationale":"The reader's weakest assumption already identifies the boundary as the soft spot; my pass sharpens it into a concrete singularity of Eq. (7) at μ=0. The qualitative core claim — sign of β identifies the threatened endpoint and γ orders residual detectability — is supported by the exact counting benchmark and the 0νββ line, and the paper's own limitations section concedes that Eq. (8) is not a finite-sample calibration. The remaining load-bearing gap is that the boundary points in Fig. 9 rely on a projection that is not uniquely defined when a secondary shape parameter is unidentified; this can be fixed by specifying a branch convention or by restricting boundary claims to identifiable models. This does not overturn the reader's conditional verdict, so no verdict change is needed.","tokens_in":17101,"tokens_out":7481,"duration_ms":78982,"concrete_test":"Take the dark-matter tail deformation at true μ=0 used in Fig. 9, fix the deformation and all count settings, and recompute β by Eq. (7) for mχ ∈ {10, 30, 50, 100, 500} GeV, removing the zero ∂ν/∂mχ column from T (or applying the paper's candidate-branch rule). If the sign of β or the ordering of the DM-tail curve changes with the chosen branch, the boundary index is not uniquely defined and the lower-endpoint prediction at μ=0 needs an explicit convention before it can be used for triage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"At the boundary μ=0, the spectral analysis has an unidentified WIMP mass mχ. Since ν_i = b_i + μ s_i(mχ), ∂ν/∂mχ = 0 when μ=0, so the 'complete fitted tangent space' T of Eq. (7) has a zero column and F is singular. The text's solution — evaluate separately in each candidate local tangent space and use the best-fitting branch — does not define a unique projection, because under the null all mχ branches give identical expected counts, while the μ-direction derivative s_i(mχ) depends on the chosen branch. Hence β at the boundary is convention-dependent unless the authors specify how the branch is selected. This is not a merely technical point: Fig. 9 left uses boundary β values for the dark-matter tail as evidence for the endpoint-prediction claim, and the screening workflow is most valuable exactly at μ=0, where lower-endpoint coverage is the whole-interval claim. The counting and 0νββ examples escape this because they have no unidentified secondary shape, but the DM-tail boundary validation is not a clean test of Eq. (7). The paper honestly disclaims Eq. (8) as a local reference and reports the β=4.2 quantitative failure; the remaining gap is the boundary definition of β itself.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how controlled model misspecification changes the frequentist coverage of intervals for a non-negative signal strength in low-count searches, using an exact Poisson counting benchmark, a dark-matter recoil spectrum, and a 0νββ peak search. It introduces the Poisson–Fisher degeneracy index I_PF = (β, γ): after projecting a candidate expected-count deformation onto the full fitted tangent space, β is the signed fitted signal shift in profiled standard-error units and γ is the Poisson–Fisher norm of the residual. The authors claim that locally the sign of β identifies which interval endpoint is threatened and that γ orders the detectability of the deformation by a saturated-Poisson goodness-of-fit test. They validate the qualitative predictions with toy Monte Carlo, show that calibration under the nominal simulator does not protect against misspecification, and demonstrate in an exactly collinear benchmark that promoting the deformation to a constrained nuisance restores coverage at the evaluated grid points at a quantifiable sensitivity cost.","tokens_in":17353,"tokens_out":9323,"duration_ms":96203,"significance":"If the screening diagnostic works as claimed, it would give experimental analyses a cheap, principled way to triage candidate model deformations before expensive toy calibration: large-|β|, small-γ directions are the dangerous ones and should be promoted to constrained nuisances. The paper's strengths include an explicit and simple formula (Eq. 7) for the index; a transparent controlled-severity protocol with stated sample sizes and binomial errors; exact Poisson summation in the counting benchmark; independent toy validation of the direction and rough severity of endpoint failures; and unusually honest reporting of the limitations of Eq. (8), including the large-β failure. The central qualitative claims—signed endpoint vulnerability and the importance of the full fitted tangent space—are supported by the clean counting and 0νββ examples as well as by the interior spectral tests. The main unresolved issue is the definition of the index at the boundary in settings with an unidentified secondary shape parameter, which affects part of the validation.","major_comments":[{"comment":"The definition of β (and γ) at μ=0 for the spectral dark-matter analysis is not well-posed. At μ=0 the WIMP mass mχ is unidentified: the mχ column of the tangent matrix T vanishes, so F is singular, and the statement in Section 4 that \"the projection is evaluated separately in each candidate local tangent space and the best-fitting branch is used\" does not select a unique branch, because under the null all mχ branches give identical expected counts. The DM-tail points in Fig. 9 left, and the corresponding residual in Fig. 9 right, are therefore convention-dependent unless a branch-selection rule is specified (for example, a fixed reference mass, or the branch that maximizes |β|). This matters because the screening workflow is most valuable precisely at μ=0 and because the DM-tail point is used as boundary validation. Please specify the convention and recompute the affected points; if no unique convention is intended, state that boundary β is defined only up to branch and restrict the empirical boundary validation to settings with identified secondary parameters (counting and 0νββ).","section":"Section 4, Eq. (7); Section 7, Fig. 9"},{"comment":"The paper honestly reports the failure of Eq. (8) at β=4.2 (predicted coverage ≈0.005, measured 0.20), but the assessment plot in Fig. 9 left still draws the Gaussian reference curve through the full range of plotted points. This curve is not a predictive relation in the large-deformation regime, and its presence could visually overstate the quantitative support for the index. Since the paper's central claim is that β is a screening coordinate rather than a finite-sample calibration, please either remove the dashed curve from the coverage-vs-β plot, restrict it to a clearly marked local regime, or add an explicit caption note that the curve is not a calibration and is known to fail at large β. The qualitative ordering claim is acceptable if framed in this way.","section":"Section 7, first paragraph; Eq. (8)"}],"minor_comments":[{"comment":"The spectral model description would benefit from explicitly listing which parameters are fitted (μ, mχ and, if any, background scale or shape parameters), so that the \"complete fitted tangent space\" in Eq. (7) is unambiguous.","section":"Section 2.1 and Fig. 1"},{"comment":"Reference [1] contains the typo \"tonne−Years\" and the misspelling \"lux-zeplin\"; these should be corrected to \"tonne-years\" and \"LUX-ZEPLIN\".","section":"Reference [1]"},{"comment":"The legend entry \"CLs / Bayesian upper limit: upper edge\" may confuse readers because both constructions are pure upper limits whose lower-endpoint coverage is trivially 1 at all severities; please clarify this in the caption.","section":"Section 6.3, Fig. 3"},{"comment":"The sentence \"the right-panel tail and line residuals use the same reference points\" is ambiguous: please specify that the DM-tail residual is evaluated at the boundary branch used for β, and the line residuals at the Q-value, to avoid confusion with the interior halo point.","section":"Section 7, right panel"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid and unusually honest contribution, and the central qualitative claims are likely correct. The main technical gap is the branch ambiguity for the index at μ=0 when a secondary shape parameter is unidentified; this is fixable with a stated convention or with a restriction of the boundary validation to identified-parameter settings. The paper is better suited to a statistics-for-physics venue than a pure phenomenology journal, but that is the editor's call. No concerns about novelty or citation behavior."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this paper gives the field a cheap pre-toy triage tool for coverage risk under model misspecification, and it validates the qualitative claims carefully. The main caveats are the missing code, the acknowledged quantitative failure of Eq. (8) at large deformation, and a boundary identifiability issue that makes the dark-matter validation at mu=0 convention-dependent.\n\nThe new construction is I_PF = (beta, gamma): project a specified deformation onto the full fitted tangent space, take the signed standardized shift of the signal parameter, and pair it with the orthogonal Poisson–Fisher residual. That combination is not in the cited literature, and it is practically motivated. The paper deserves credit for the endpoint-separated stress-test protocol and for demonstrating the directionality: positive signal-like contamination threatens the lower endpoint, negative signal bias can make upper limits anti-conservative, and the constrained-nuisance remedy restores coverage at the evaluated points at a measurable sensitivity cost.\n\nThe paper is also honest. It states explicitly that Eq. (8) is a local reference, not a finite-sample calibration, and reports the beta=4.2 case where the formula gives 0.005 but the measured coverage is 0.20. Its own limitations section lists the restricted dimensionality and the grid-points-only nuisance calibration. That honesty makes the read easier.\n\nThe soft spots are real but not disqualifying. No code or data are shipped, which undercuts the \"screening tool\" pitch. The quantitative prediction is only reliable in the small-deformation regime, and the paper's own validation says so. The stress-test concern about the boundary is legitimate: at mu=0 in the dark-matter spectral analysis, the WIMP mass column of the tangent matrix is zero, F is singular, and the text's \"best-fitting branch\" convention is not well defined because all mass branches give identical expected counts under the null. The boundary beta values in Fig. 9 left therefore depend on an arbitrary branch choice unless the paper specifies a rule. The counting and 0nu-beta-beta examples are unaffected because they have no unidentified shape parameter, so the qualitative screening claim likely survives once the convention is fixed.\n\nWho is this for? Experimental statisticians in rare-event searches who need to decide whether a candidate deformation can be ignored or must be propagated into a nuisance. This deserves peer review; the right outcome is a major revision asking for code and data, an explicit branch-selection convention at non-identifiable boundaries, and a more modest presentation of the quantitative coverage-prediction claims. I would not desk-reject.","headline":"A genuinely useful screening diagnostic for coverage risk under model misspecification, with a real boundary-identifiability gap that needs a convention; deserves a serious referee, not a desk reject.","tokens_in":17895,"tokens_out":2878,"would_cite":true,"duration_ms":84547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two numbers predict silent coverage failures in rare-event searches.","keywords":["confidence intervals","coverage","model misspecification","Poisson–Fisher degeneracy index","rare-event searches","dark matter","neutrinoless double-beta decay","goodness-of-fit"],"falsifier":"A decisive test is a single-bin Poisson experiment with fixed background and an exactly signal-collinear positive excess scanned over amplitude: the index predicts $\\beta>0$, $\\gamma=0$, monotone decrease of lower-endpoint coverage in $\\beta$, and goodness-of-fit rejection at the 5\\% floor for every amplitude. If toy coverage at $\\mu=0$ is non-monotone in $\\beta$, or if the goodness-of-fit rejection probability rises above the floor for any of these deformations, the sign and detectability ordering is falsified; the paper's own reported miss at $\\beta=4.2$ (predicted 0.005, measured 0.20) already shows that the Gaussian reference can fail without falsifying the qualitative ordering.","tokens_in":16823,"feed_emoji":"🔭","tokens_out":9617,"duration_ms":82539,"temperature":0.7,"pith_summary":"This paper claims that rare-event search intervals can silently miss their nominal frequentist coverage: the interval fails while the accompanying goodness-of-fit test stays at its nominal false-positive rate. It introduces the Poisson–Fisher degeneracy index, a pair $(\\beta,\\gamma)$ computed from expected bin counts and local model derivatives, and argues that the sign of $\\beta$ identifies which endpoint of the confidence interval is threatened, while $\\gamma$ orders how detectable the model error is by a saturated-Poisson goodness-of-fit test at fixed 5\\% type-I error. The claim is tested on controlled model deformations in an exact Poisson counting experiment, a dark-matter recoil spectrum, and a neutrinoless-double-$\\beta$-decay peak search, where positive signal-like bias degrades discovery-side coverage and negative signal bias from overestimated efficiency makes the upper endpoint undercover. If the claim is right, analyses can triage candidate model deformations before toy calibration, promoting large-$|\\beta|$, small-$\\gamma$ deformations to constrained nuisances with auxiliary information. The paper states that Eq. (8) is only a local reference, not a finite-sample coverage calibration.","feed_headline":"Two numbers predict silent coverage failures in rare-event searches","feed_subtitle":"A new diagnostic maps each model deformation to the threatened interval endpoint and the chance a goodness-of-fit test detects it.","key_machinery":"The load-bearing object is Eq. (7), the Poisson–Fisher degeneracy index. With the Poisson–Fisher inner product $\\langle u,v\\rangle_W=\\sum_i u_i v_i/\\nu_{0,i}$ and the tangent matrix $T$ whose columns are the local derivatives $\\partial\\nu/\\partial\\vartheta_a$, the fitted displacement is $\\Delta\\hat\\vartheta=F^{-1}T^T W\\,\\delta\\nu$ with $F=T^T W T$; then $\\beta=\\Delta\\hat\\mu/\\sqrt{(F^{-1})_{\\mu\\mu}}$ and $\\gamma^2=(\\delta\\nu-T\\Delta\\hat\\vartheta)^T W(\\delta\\nu-T\\Delta\\hat\\vartheta)$. The paper connects $\\beta$ to endpoint coverage through the local Gaussian reference $c_L\\simeq\\Phi(z_{1-\\alpha/2}-\\beta)$ and $c_U\\simeq\\Phi(z_{1-\\alpha/2}+\\beta)$, and uses $\\gamma$ to order the rejection probability of the saturated-Poisson deviance test. A constrained-nuisance extension adds auxiliary Fisher information $F_{\\mathrm{aux}}$ to the total information and a tension term $\\Delta\\hat\\vartheta^T F_{\\mathrm{aux}}\\Delta\\hat\\vartheta$ to the residual. Projection onto the full fitted tangent space is what makes the diagnostic nuisance-aware: a deformation parallel to a background or mass derivative is absorbed there, changing both the predicted damage and the residual available to a goodness-of-fit test.","core_discovery":"The central discovery is the Poisson–Fisher degeneracy index $\\mathcal{I}_{\\mathrm{PF}}(\\delta\\nu;\\vartheta_0)=(\\beta,\\gamma)$, a two-coordinate screening diagnostic for a candidate model deformation. Given a deformation $\\delta\\nu$ of the nominal expected-count vector, one projects it onto the complete fitted tangent space using the local Poisson–Fisher metric $W=\\mathrm{diag}(1/\\nu_{0,i})$; $\\beta$ is the resulting signed shift in the parameter of interest in profiled standard-error units, and $\\gamma$ is the Poisson–Fisher norm of the residual outside all fitted directions. The paper shows empirically that the sign of $\\beta$ tells which interval endpoint loses coverage—positive $\\beta$ threatens the lower (discovery) endpoint, negative $\\beta$ threatens the upper (exclusion) endpoint—while $\\gamma$ tracks the power of a saturated-Poisson deviance test. Across the counting, dark-matter, and $0\\nu\\beta\\beta$ settings, a deformation with large $|\\beta|$ and small $\\gamma$ is the dangerous case: the interval can move while the fitted spectrum looks fine. The paper also shows that calibration under the nominal simulator does not protect against misspecification of that simulator, and that modelling an exactly collinear deformation as a constrained nuisance restores coverage at evaluated points, at a sensitivity cost quantified by a factor up to 1.6.","pith_inferences":["A natural extension is to compute the same projection with an observed-information metric for unbinned or learned likelihoods, giving $\\beta$ and $\\gamma$ for simulation-based inference pipelines without toy calibration.","Because $\\gamma$ orders the power of a residual check, it could guide test construction: binning and region-of-interest choices that increase $\\gamma$ for a candidate deformation family should also increase the sensitivity of a goodness-of-fit test.","The large-$|\\beta|$, small-$\\gamma$ corner is effectively an identifiability statement: near collinearity, no finite data-driven interval exists without external information, so the practical boundary is set by the strength of auxiliary constraints rather than by sample size.","One testable extension is to turn the index into a calibration target, choosing deformation severities that reach a specified $\\beta$ and using toy coverage to set analysis-specific 'large' and 'small' cutoffs, which the paper leaves open."],"forward_implications":["Robustness studies should report lower-endpoint, upper-endpoint, and whole-interval coverage separately; whole-interval coverage at $\\mu=0$ tests only the lower endpoint and cannot validate a pure upper limit.","Positive signal-like contamination degrades discovery-side coverage while making upper limits conservative, whereas negative signal bias from overestimated efficiency can make the upper endpoint undercover, so both deformation signs should be tested.","Calibration under the nominal simulator, including a simulation-based method's internal coverage check, does not detect misspecification of that simulator relative to the data-generating process; a separate data-based model check is needed.","A plausible deformation with large $|\\beta|$ and small $\\gamma$ should be promoted to a nuisance constrained by auxiliary data or a declared envelope; in the exactly collinear benchmark this restores coverage at evaluated grid points at up to a factor-1.6 sensitivity cost.","The index can serve as pre-toy triage, but Eq. (8) is only a local reference; final coverage statements still require toy calibration, especially at boundaries and for nonlocal deformations."],"supporting_citations":[{"why":"Supplies the unified Neyman construction used as the exact counting reference and implementation check.","marker":"[11]"},{"why":"Provides the asymptotic profile-likelihood intervals and boundary formulas that the local Gaussian reference in Eq. (8) extends.","marker":"[9]"},{"why":"Establishes the nonstandard likelihood-ratio limit at the boundary, motivating why coverage at $\\mu=0$ needs a separate diagnostic.","marker":"[8]"},{"why":"Covers maximum-likelihood asymptotics when parameters are unidentified under the null, the situation at the boundary for $m_\\chi$.","marker":"[10]"},{"why":"Gives the pseudo-true displacement and sandwich covariance for misspecified models, which the paper distinguishes from its own projected diagnostic.","marker":"[29]"},{"why":"Defines the astrophysical halo uncertainty that generates the halo-model deformation family in the spectral stress test.","marker":"[25]"}],"fun_headline_variants":["Two numbers flag silent coverage failures in rare-event searches","New diagnostic maps silent coverage failures in rare-event searches","Degeneracy index predicts silent coverage failures in rare-event searches","Silent coverage failures in rare-event searches now predictable","A diagnostic for silent coverage failures in rare-event searches"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The diagnostic assumes that a first-order local projection computed at a reference point remains informative for coverage behavior at the boundary $\\mu=0$ and for nonlocal deformations, where profile-likelihood asymptotics are nonstandard and parameters such as the WIMP mass are unidentified.","fun_headline_variants_meta":{"raw":{"variants":["Two numbers flag silent coverage failures in rare-event searches","New diagnostic maps silent coverage failures in rare-event searches","Degeneracy index predicts silent coverage failures in rare-event searches","Silent coverage failures in rare-event searches now predictable","A diagnostic for silent coverage failures in rare-event searches"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00072,"raw_usage":{"total_tokens":3336,"prompt_tokens":1152,"completion_tokens":2184,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":768,"completion_tokens_details":{"reasoning_tokens":2106}},"tokens_in":768,"tokens_out":2184,"duration_ms":15027,"temperature":1.0,"reasoning_tokens":2106,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T21:23:38.994858+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test is a single-bin Poisson experiment with fixed background and an exactly signal-collinear positive excess scanned over amplitude: the index predicts $\\beta>0$, $\\gamma=0$, monotone decrease of lower-endpoint coverage in $\\beta$, and goodness-of-fit rejection at the 5\\% floor for every amplitude. If toy coverage at $\\mu=0$ is non-monotone in $\\beta$, or if the goodness-of-fit rejection probability rises above the floor for any of these deformations, the sign and detectability ordering is falsified; the paper's own reported miss at $\\beta=4.2$ (predicted 0.005, measured 0.20) already shows that the Gaussian reference can fail without falsifying the qualitative ordering.","supporting_citations":[],"review_version":1}