{"id":"b123740a-4833-41f7-bd22-adf005df744d","arxiv_id":"2508.13012","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new Wald-type test statistic built from a ridge-regularized estimator gives uniformly shorter valid confidence intervals than the partial-conditioning solution for marginal inference in the two-normal-means problem.","lead":"This paper builds confidence intervals for one of two normal means when the two means are known to be close, using a regularized estimator and a possibility-based inferential model. The new interval is always at least as short as a recently proposed partial-conditioning interval, with no loss of finite-sample validity.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Lemma 1 is misstated as written and its proof omits the slope bound needed for Proposition 2; the underlying inequality is true and easily repaired, so the reader's conditional verdict stands.","rationale":"The reader's weak-assumption diagnosis is accurate: Proposition 2 depends on Lemma 1, and the proof of Lemma 1 as written is incomplete and the notation around z_alpha is inconsistent. My stress-test confirms this and adds the concrete observation that the printed lemma is literally false at gamma = 0 under conventional quantile notation. However, the intended inequality is true and the missing slope bound h'(mu) < 1 follows directly from the paper's own Eq. (25), so the flaw is a presentational/proof-writing gap rather than a substantive mathematical error. The rest of the construction -- regularized estimator, noncentral chi-square test statistic, supremum marginalization, and length comparison -- is internally coherent and the validity argument is sound. Consequently the appropriate action is to require the authors to correct Lemma 1's statement and complete the proof, i.e., a conditional acceptance, which is exactly the reader's verdict. I see no reason to move to accept or reject, so the verdict remains unchanged.","tokens_in":8508,"tokens_out":8635,"duration_ms":82557,"concrete_test":"Evaluate the printed Lemma 1 at alpha = 0.05, gamma = 0: compute sqrt(Q_{0.05}(0)) = z_{0.975} = 1.96, z_{0.05} = -1.645, and observe 1.96 - (-1.645) > 0, which falsifies the statement under the standard z_alpha notation. Then verify the corrected inequality: for h(mu) = sqrt(Q_alpha(mu^2)), use Eq. (25) to confirm h'(mu) = (phi(h-mu) - phi(h+mu))/(phi(h-mu) + phi(h+mu)) is in (0,1) for all mu > 0, so by the mean value theorem h(mu) - h(0) <= mu; with z_alpha replaced by z_{(1+alpha)/2} the lemma holds, and substituting into Eq. (26) gives L2 <= L1.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of the paper, Proposition 2, reduces entirely to Lemma 1, which states sqrt(Q_alpha(gamma)) - z_alpha <= sqrt(gamma) for all gamma >= 0. As printed, this is false under the standard convention that z_alpha is the alpha-quantile of N(0,1). At gamma = 0, sqrt(Q_alpha(0)) = z_{(1+alpha)/2}, so the left-hand side is z_{(1+alpha)/2} - z_alpha, which is positive for every alpha in (0,1); for alpha = 0.05 it equals 3.605. The proof in Section 4.2 implicitly redefines z_alpha as z_{(1+alpha)/2} (the constant h(0) in Eq. (24)), which is inconsistent with Eq. (8) and with the use of z_{1-alpha/2} in Eq. (26). Even under the intended notation, the proof establishes only h'(mu) > 0, while the desired inequality h(mu) - h(0) <= mu requires h'(mu) <= 1. This slope bound is immediate from the displayed formula (25), since h'(mu) = (phi(h-mu) - phi(h+mu))/(phi(h-mu) + phi(h+mu)) lies strictly between 0 and 1, but the paper never states it. The gap is therefore real but cosmetic: the lemma becomes correct once z_alpha is replaced by z_{(1+alpha)/2} and the one-line slope bound is added. Proposition 2 then follows exactly as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This short note studies the two-normal-means problem with a known upper bound B on |θ2 − θ1|. The authors construct a possibilistic inferential model (IM) based on a ridge-type regularized maximum likelihood estimator, first re-deriving the partial-conditioning interval of Yang et al. (2023) from a Wald-type statistic T1, and then proposing a new interval C2 based on an uncentered statistic T2. The main analytical contribution is Proposition 2, which states that for any fixed λ ≥ 0, α ∈ (0,1), and B > 0, the length L2(λ; α, B) of the new interval is no larger than the length L1(λ; α, B) of the partial-conditioning interval. Proposition 1 asserts that L2 has a finite minimizer in λ, so the improved interval can be tuned numerically. The authors also provide a small numerical illustration comparing interval lengths.","tokens_in":8837,"tokens_out":1466,"duration_ms":13971,"significance":"If the main result holds, this is a modest but clean contribution to the IM literature and to finite-sample inference under Hölder-type constraints. The paper is honest about its narrow scope (n = 2) and identifies a concrete open problem for the general many-normal-means case. The construction is transparent: the comparison of L2 and L1 is parameter-free in the sense that the same λ, α, and B appear on both sides, and the proof relies on a new quantile inequality for the noncentral chi-square distribution. The proof of Proposition 2 does rest on a lemma whose printed statement and proof are not fully consistent, but the underlying inequality is correct and the gap is easily repaired. The paper also gives explicit formulas for the intervals, which will be useful for practitioners seeking a finite-sample-valid interval that exploits the constraint.","major_comments":[{"comment":"As printed, Lemma 1 is misstated under the paper's own convention for z_α (defined in Eq. (8) as the βth quantile of N(0,1)). Taking γ = 0, the left-hand side is sqrt(Q_α(0)) − z_α = z_{(1+α)/2} − z_α, which is positive for every α ∈ (0,1); for α = 0.05 it equals about 3.6, so the stated inequality fails. The proof, Eq. (24), implicitly redefines z_α as the constant h(0) = z_{(1+α)/2}, which is inconsistent with Eq. (8) and with the use of z_{1−α/2} in Eq. (26). The statement should read sqrt(Q_α(γ)) − z_{(1+α)/2} ≤ sqrt(γ) (with the obvious replacement in Eq. (26)), or the notation should be changed throughout.","section":"Lemma 1 and Proposition 2 (pp. 8–9)"},{"comment":"The proof only establishes h'(μ) > 0, which gives h(μ) > h(0) = z_{(1+α)/2}. The desired inequality h(μ) − h(0) ≤ μ requires the additional slope bound h'(μ) ≤ 1. This bound is immediate from the displayed formula (25), since h'(μ) = (φ(h−μ) − φ(h+μ))/(φ(h−μ) + φ(h+μ)) lies strictly between 0 and 1 for μ > 0, but the paper never states or uses it. Without this step, Proposition 2 does not follow from the written proof, even under the corrected notation.","section":"Lemma 1 proof, Eq. (25) (p. 9)"},{"comment":"The argument that L2' is positive for all sufficiently large λ is only sketched. The proof states that 2λ(1+λ)(1+2λ) is cubic while 2[λ²+(1+λ)²] is quadratic, but this comparison alone does not establish the sign of the whole bracket in Eq. (23), because Q_{1−α}(g(λ,B)) and Q'_{1−α}(g(λ,B)) also depend on λ through g(λ,B). Since g(λ,B) → B²/2 and both Q and Q' are positive and continuous on [0, B²/2], the desired conclusion does follow, but the proof should spell out the boundedness of Q and Q' over the relevant compact range.","section":"Proposition 1 proof, Eq. (23) (p. 7)"}],"minor_comments":[{"comment":"The notation in Eq. (11) defines a normal random variable with mean λ(θ1 − θ2), and Eq. (12) centers by that mean; it may help readers to see the centering term written explicitly as λ(θ1 − θ2) in both equations.","section":"p. 4, Eq. (11)–(12)"},{"comment":"The caption says 'α is fixed at 0.05 while B varies'; this is fine, but the figure would be easier to read if the standard interval length were labeled explicitly as 2z_{0.975} in the legend text.","section":"p. 9, Fig. 2 caption"},{"comment":"The word 'illustarte' in the first paragraph is a typo for 'illustrate'.","section":"p. 10, Discussion"},{"comment":"The validity inequality is stated as P_{Y|θ}{π_Y(θ) ≤ α} ≤ α, but the subsequent derivation in Eq. (3) uses the right-tail probability; for completeness, the authors may want to note that the p-value construction satisfies the inequality via the probability integral transform even when the test statistic is not continuous.","section":"p. 2, Eq. (2)"},{"comment":"The notation F(·; k, γ) is standard, but the first occurrence of the noncentrality parameter 0 is written as 'χ²(1,0)'; this is clear enough, yet a brief definition of the chi-square noncentrality parameter convention would improve accessibility.","section":"p. 3, Eq. (7)"}],"recommendation":"minor_revision","confidential_remarks":"The paper's main claim is correct and the proof gap is cosmetic, but the discrepancy in Lemma 1 is exactly the sort of technical inconsistency that editors and readers will notice. I recommend minor revision rather than major revision because the fix is a one-line slope bound and a notational correction; no part of the central construction needs to change. The paper is well within scope for a statistics methodology journal and is appropriately cautious about the limitations of the n = 2 setting."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I've read the Liu–Williams note on the two-normal-means problem. The central result is real: the uncentered Wald statistic in Eq. (18) gives a possibilistic IM interval that, for every λ, α, B, is no wider than the partial-conditioning interval of Yang et al. (2023). That uniform dominance is a genuine, if modest, contribution to the IM subfield. The earlier rederivation of the partial-conditioning solution via ridge regularization is a nice pedagogical bridge, and the authors are appropriately careful about scope—no overclaiming beyond the two-mean case.\n\nThe construction is coherent. The marginalization over θ1 works because the noncentral chi-square contour increases in the noncentrality parameter, and the formulas are explicit enough to reproduce without code. The numerical illustrations are helpful, though not essential.\n\nThe soft spot is Lemma 1. As printed, sqrt(Q_α(γ)) - z_α ≤ sqrt(γ) is false under the usual convention that z_α is the α-quantile of the standard normal. At γ=0, the left side is z_{(1+α)/2} - z_α, which is positive. The proof in Section 4.2 implicitly works with h(0) = z_{(1+α)/2}, so the lemma is misstated, not fundamentally wrong. The proof also only establishes h'(μ) > 0; the inequality h(μ)-h(0) ≤ μ actually requires h'(μ) ≤ 1. That slope bound is immediate from the displayed formula for h'(μ)—the ratio is strictly between 0 and 1—so this is a one-line repair. Once the lemma is restated with z_{(1+α)/2} and the slope bound added, Proposition 2 follows exactly as intended. The notation mismatch between z_α and z_{1-α/2} in the proof of Proposition 2 is part of the same issue. These are real but cosmetic; a referee will spot them quickly.\n\nThe citation pattern is reasonable: the paper builds on Yang et al. and the Martin–Liu IM line without overselling. No data or code, but none is needed for an analytic note of this kind.\n\nBottom line: this is a solid short note for the IM audience. The claim is likely correct, the proof has a fixable gap, and the contribution is modest but honest. I'd send it to peer review and expect it to come back with a corrected Lemma 1 and a cleaner proof. I wouldn't cite it in my own work unless I had a specific need for that noncentral chi-square quantile inequality, but I'd point a student working on regularized confidence sets to it.","headline":"A clean, modest result in the inferential-models program, held up by a misstated lemma that is easily corrected; worth a serious referee.","tokens_in":9339,"tokens_out":7044,"would_cite":false,"duration_ms":63846,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62F10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Ridge regularization of the two-normal-means problem yields valid confidence intervals for one mean that are never wider than the previous best partial-conditioning intervals, and strictly narrower once the penalty weight is optimized.","keywords":["confidence intervals","inferential models","regularization","statistical inference","two normal means","partial conditioning","noncentral chi-square distribution","possibility theory"],"falsifier":"Compute $\\sqrt{Q_{1-\\alpha}(\\gamma)} - z_{1-\\alpha/2} - \\sqrt{\\gamma}$ on a dense grid of $\\gamma \\geq 0$ for fixed $\\alpha$ such as 0.05, 0.1, and 0.2: any positive value is a counterexample to Lemma 1 and therefore to Proposition 2. Equivalently, evaluate the two length formulas $L_2$ and $L_1$ on a grid of $(\\lambda, B, \\alpha)$ and check whether $L_2 > L_1$ anywhere; the claim is settled by whichever sign appears.","tokens_in":8297,"feed_emoji":"📏","tokens_out":20963,"duration_ms":180253,"temperature":0.7,"pith_summary":"This paper takes the simplest hard case of the classic many-normal-means problem --- two independent unit-variance normal observations whose means are known to be at most $B$ apart --- and asks how to make a valid confidence interval for the second mean while using the first observation. The construction starts from a ridge-regularized estimator that pulls the two means toward each other and then inverts an uncentered Wald statistic for the focal mean inside a possibilistic inferential model, whose plausibility contours are p-values and whose $\\alpha$-cuts are finite-sample-valid confidence sets. The main analytic result is that, for every penalty weight $\\lambda \\geq 0$, confidence level $\\alpha \\in (0,1)$, and bound $B>0$, the new interval is no longer than the partial-conditioning interval of Yang et al. (2023), which was itself the best available construction; after optimizing the penalty weight --- a numerical search that the paper proves to be well-posed --- the improvement is strict for every $B>0$. If the result holds, then 'regularize, then drop the centering term' is a general recipe that converts prior knowledge about the closeness of means into shorter interval estimates at no cost in finite-sample coverage, and it supplies a template for the full many-normal-means problem.","feed_headline":"New interval for a normal mean is never wider than the old best","feed_subtitle":"A ridge penalty and uncentered Wald statistic convert known closeness of the two means into shorter, still-valid intervals.","key_machinery":"The load-bearing object is the uncentered Wald-type statistic (18) built from the focal coordinate of the ridge-regularized estimator: unlike the partial-conditioning statistic, it does not subtract the unknown mean difference $\\lambda(\\theta_1-\\theta_2)$, so its distribution keeps a noncentrality bounded by $\\lambda^2 B^2/(\\lambda^2+(1+\\lambda)^2)$. The comparison is carried by Lemma 1, a quantile inequality for the noncentral chi-square distribution --- $\\sqrt{Q_{1-\\alpha}(\\gamma)} \\leq z_{1-\\alpha/2} + \\sqrt{\\gamma}$ for all $\\gamma\\geq 0$ --- which says that the square root of the noncentral quantile never exceeds the standard normal quantile plus the square root of the noncentrality. Multiplying that inequality by the statistic's scale factor $\\sqrt{\\lambda^2+(1+\\lambda)^2}$ converts $\\sqrt{\\gamma}$ into exactly the $\\lambda B$ term that appears in the partial-conditioning length $L_1$, so the dominance $L_2 \\leq L_1$ holds term-by-term for every $\\lambda$, $\\alpha$, and $B$. The whole procedure lives inside a possibilistic inferential model (IM): a p-value-based framework in which each parameter value receives a plausibility equal to the tail probability of a test statistic, with validity inherited from the probability integral transform and marginal inference obtained by taking the supremum over nuisance parameters. Proposition 1 completes the machinery by showing $L_2$ attains its minimum at some positive finite $\\lambda$, making numerical tuning feasible.","core_discovery":"Fixing the two-normal-means problem with $|\\theta_2 - \\theta_1| \\leq B$, the paper studies the regularized estimator $\\hat\\theta_2(\\lambda) = (\\lambda y_1 + (1+\\lambda)y_2)/(1+2\\lambda)$ and the uncentered Wald statistic $T_2 = [\\lambda(Y_1-\\theta_2) + (1+\\lambda)(Y_2-\\theta_2)]^2 / (\\lambda^2 + (1+\\lambda)^2)$, which under the data-generating distribution is noncentral chi-square with noncentrality $\\lambda^2(\\theta_1-\\theta_2)^2 / (\\lambda^2+(1+\\lambda)^2) \\leq \\lambda^2 B^2/(\\lambda^2+(1+\\lambda)^2)$. Taking the supremum of the resulting p-value over $\\theta_1 \\in [\\theta_2-B, \\theta_2+B]$ yields a valid marginal possibility contour for $\\theta_2$, whose $\\alpha$-cut is the explicit interval (21) of length $L_2(\\lambda;\\alpha,B) = 2\\sqrt{Q_{1-\\alpha}(\\lambda^2 B^2/(\\lambda^2+(1+\\lambda)^2))\\,(\\lambda^2+(1+\\lambda)^2)}/(1+2\\lambda)$, where $Q$ is the noncentral chi-square quantile. By the quantile inequality $\\sqrt{Q_{1-\\alpha}(\\gamma)} \\leq z_{1-\\alpha/2} + \\sqrt{\\gamma}$ for all $\\gamma \\geq 0$, this length is at most the length $L_1$ of the partial-conditioning interval for every $\\lambda \\geq 0$, $\\alpha \\in (0,1)$, and $B>0$; since $L_2$ is provably minimized at a finite positive penalty weight, the optimally tuned new interval is strictly shorter than the partial-conditioning interval for all $B>0$ and strictly dominates the textbook marginal interval that ignores $y_1$.","pith_inferences":["Editor's inference: the mechanism is generic --- whenever a known bound on parameter differences enters through a ridge penalty, the Wald statistic's noncentrality has the form $\\lambda^2 B^2/(\\lambda^2 + \\text{scale}^2)$, and the same quantile inequality converts that bound directly into an interval-length budget, so analogous dominance results should hold for other constrained-estimation problem","Editor's inference: in the full many-means generalization the noncentrality will accumulate across neighbors of the focal mean, so the per-mean gain likely shrinks as local degree grows; a concrete testable prediction is that dominance persists but the optimal penalty weight and the size of the improvement depend on the Hölder bound configuration.","Editor's inference: the interval length used as the efficiency criterion is independent of both data and parameters in this two-means case; in larger problems the length will depend on both, so a worst-case or average-case length criterion will be needed, and the choice of criterion will affect which penalty weight is optimal --- a difficulty the paper itself flags in its discussion."],"forward_implications":["The new interval achieves nominal finite-sample coverage and, for every fixed $\\lambda \\geq 0$, $\\alpha \\in (0,1)$, and $B>0$, is no wider than the partial-conditioning interval; after numerical optimization of the penalty weight it is strictly shorter for every $B>0$.","The centered version of the same regularized statistic reproduces the Yang et al. (2023) partial-conditioning interval exactly, so the regularization viewpoint yields a simpler derivation and pinpoints the centering term as the sole difference between the old and new constructions.","Because the textbook interval that ignores $y_1$ arises as the $\\lambda = 0$ (or $B \\to \\infty$) limit, the tuned regularized interval also dominates the standard marginal inference that uses only the focal observation.","Proposition 1 guarantees that the interval length has a finite minimizing penalty weight, so the efficiency tuning is a provably well-posed one-dimensional numerical search.","The authors conjecture that the same construction extends to the full many-normal-means problem with Hölder constraints, using one ridge penalty per adjacent pair and a modified Wald statistic, with a similar efficiency gain."],"supporting_citations":[{"why":"Supplies the partial-conditioning IM solution for the many-normal-means problem with Hölder constraints; its marginal interval length L1 is the baseline that the new interval must dominate.","marker":"Yang et al. (2023)"},{"why":"Foundational IM papers establishing conditional and marginal inferential-model validity; their marginalization-by-supremum argument underwrites the finite-sample coverage of the new interval.","marker":"Martin and Liu (2015a,b)"},{"why":"The possibilistic IM construction used in the paper, in which the possibility contour is a p-value function of a test statistic and validity follows from the probability integral transform.","marker":"Martin (2022a,b,c)"},{"why":"Source of the 'possibility contour' formulation and of the supremum-based marginalization operation used in equation (4).","marker":"Liu and Martin (2024)"},{"why":"Shows relative-likelihood IMs are asymptotically efficient under classical regularity conditions and motivates regularization as the route to efficiency in over-parameterized problems.","marker":"Martin and Williams (2025)"},{"why":"The probability integral transform (Theorem 2.1.10 and Exercise 2.10) cited to justify validity of the p-value-based contour in equation (3).","marker":"Casella and Berger (2002)"}],"fun_headline_variants":["Ridge penalty shortens interval for a normal mean","Shorter valid intervals via regularization in two-means problem","Never wider: regularized interval beats partial conditioning","Two-means trick: penalized Wald yields tighter intervals","Regularization tightens normal mean interval without losing validity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is the quantile inequality $\\sqrt{Q_{1-\\alpha}(\\gamma)} \\leq z_{1-\\alpha/2} + \\sqrt{\\gamma}$ for noncentral chi-square quantiles --- whose proof in Section 4.2 shows only that the quantile function is increasing, not the slope bound it also needs --- and if that inequality failed for any $\\alpha$ and $\\gamma$, the new interval could be wider than the partial-conditioning interval it claims to improve.","fun_headline_variants_meta":{"raw":{"variants":["Ridge penalty shortens interval for a normal mean","Shorter valid intervals via regularization in two-means problem","Never wider: regularized interval beats partial conditioning","Two-means trick: penalized Wald yields tighter intervals","Regularization tightens normal mean interval without losing validity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1461,"prompt_tokens":1087,"completion_tokens":374,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":703,"completion_tokens_details":{"reasoning_tokens":297}},"tokens_in":703,"tokens_out":374,"duration_ms":3878,"temperature":1.0,"reasoning_tokens":297,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:16:32.602635+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute $\\sqrt{Q_{1-\\alpha}(\\gamma)} - z_{1-\\alpha/2} - \\sqrt{\\gamma}$ on a dense grid of $\\gamma \\geq 0$ for fixed $\\alpha$ such as 0.05, 0.1, and 0.2: any positive value is a counterexample to Lemma 1 and therefore to Proposition 2. Equivalently, evaluate the two length formulas $L_2$ and $L_1$ on a grid of $(\\lambda, B, \\alpha)$ and check whether $L_2 > L_1$ anywhere; the claim is settled by whichever sign appears.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the partial-conditioning IM solution for the many-normal-means problem with Hölder constraints; its marginal interval length L1 is the baseline that the new interval must dominate."},{"cited_title":"and Martin, R","cited_arxiv_id":null,"evidence_quote":"Source of the 'possibility contour' formulation and of the supremum-based marginalization operation used in equation (4)."},{"cited_title":"and Berger, R","cited_arxiv_id":null,"evidence_quote":"The probability integral transform (Theorem 2.1.10 and Exercise 2.10) cited to justify validity of the p-value-based contour in equation (3)."}],"review_version":2}