{"id":"3ce227df-e3e1-4a15-a652-ccd63211c063","arxiv_id":"2507.19633","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":8.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Score-based confidence regions in linear mixed models are shown to have approximately correct coverage uniformly over parameters, including boundary cases where variance components are zero or correlations are plus or minus one.","lead":"This paper proves finite-sample and asymptotic bounds showing that score-based statistics in linear mixed models stay close to normal even when random effect variances are zero or correlations are at their extremes. This matters because standard Wald and likelihood-ratio confidence intervals fail near such boundaries, while the proposed score regions keep near-nominal coverage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; the uniform-coverage claims follow from the stated bounds, pending the deferred proof of Lemma 4.","rationale":"The reader's weakest_assumption points to the design replication conditions (Theorem 3(ii) and n_min requirements). These are real and are indeed load-bearing for the finite-sample bounds, but the authors state them explicitly and discuss their necessity, so I do not see them as undermining the central claim. The more delicate point is that the main theorems rest on Lemma 4 and Theorem 4, both proved only in the supplement. For covariance parameters, A(v, psi) is indefinite, so the key normal approximation must be valid beyond the positive-semidefinite case common in single-variance-component settings. If that lemma has an unstated condition, the uniform coverage results would fail. My reading of the main text found no contradiction or obvious counterexample; the logic from the bounds to uniform coverage is sound, including the use of Lemma 2 and fixed-r Cramer-Wold. The simulation evidence is consistent with the theory, though Monte Carlo error bars are absent. Overall, the paper's claims are conditionally correct given the deferred proofs, and the reader's ACCEPT with moderate confidence remains appropriate. I therefore recommend no change to the verdict, while noting that a direct verification of Lemma 4 and Theorem 4 is the most useful next step.","tokens_in":15886,"tokens_out":22897,"duration_ms":275439,"concrete_test":"Fetch the supplementary material and re-derive Lemma 4 from the Zhang et al. (2025) quadratic-form bound, checking that it covers indefinite A(v, psi) and that the constant 0.14 is valid under only a^2 < 1/8; then numerically maximize a(v, psi)^2 over a grid of psi and v for n1 = n2 = 50 in model (19) and compare with Theorem 4's bound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"I read the paper as making a conditional claim: if sup_v a_n(v, psi_n) tends to 0 (Theorems 2-4), then score-based confidence regions have asymptotically correct uniform coverage. The conditions under which this smallness is established (Theorem 3(ii), Theorem 4's n_min >= 2 or 3) are explicit design-replication assumptions, and the paper acknowledges the exclusion of zero-column clusters in Section 4.2. Those are scope limitations, not internal inconsistencies. The one step I cannot check from the main text is Lemma 4, the quantitative density bound underlying all results, which is proved only in the supplement. In particular, A(v, psi) is typically indefinite for covariance parameters (H_j has off-diagonal ones), so the cited normal approximation for quadratic forms must handle indefinite matrices with a^2 < 1/8. If Lemma 4, or the psi-independent bound in Theorem 4, contains a hidden regularity condition, the crossed-effects uniformity would collapse. I found no concrete error in the main text, so this is a verification risk rather than a demonstrated flaw.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops finite-sample and asymptotic distributional approximations for score-based inference in Gaussian linear mixed models, with emphasis on uniform coverage near the boundary of the random-effects covariance parameter space, where variances may vanish or correlations may approach ±1. The main engine is Lemma 4, a finite-sample bound on the sup-norm distance between the density of a standardized linear combination of the score and the standard normal density, controlled by the ratio a(v, psi) of spectral to Frobenius norms of A(v, psi). Lemma 5 converts this into a uniform bound; Theorem 2 derives asymptotic normality of arbitrary linear combinations of the standardized score under sup_v a_n(v, psi_n) -> 0, including with unknown fixed effects. Theorems 3 and 4 give design-specific bounds for independent clusters and crossed random effects, and Corollaries 2-4 translate these into asymptotically correct uniform coverage for chi-square confidence regions, including at the boundary. Simulations illustrate near-nominal coverage for score regions and failures of Wald and likelihood-ratio regions in the same settings.","tokens_in":16060,"tokens_out":16055,"duration_ms":181873,"significance":"If the deferred proofs are correct, this is a substantial contribution: it provides a general treatment of uniform inference for covariance parameters at the singular boundary of the parameter set, with explicit finite-sample bounds and clear sufficient conditions. The design-specific bounds (m^{-1/2} for independent clusters, functions of n_min for crossed effects) are informative and lead to practically implementable confidence regions. The paper's strengths include explicit constants, a clean separation between the generic approximation and design-specific analyses, treatment of crossed random effects with diverging dimensions, reproducible simulation code, and falsifiable predictions about coverage. The main reservations are verification risks rather than demonstrated errors: the central normal approximation is imported from a co-authored prior paper whose proof is not in the reviewed material, and the extension to unknown beta requires a nontrivial argument that the main text does not supply.","major_comments":[{"comment":"Lemma 4 is the sole finite-sample engine of the paper, but its proof is deferred to a supplement that was not included in the reviewed material, and the lemma is attributed to a normal approximation for quadratic forms due to Zhang et al. (2025), which was developed for a single variance component. In the present setting A(v, psi) is generally indefinite because the H_j contain off-diagonal ones, so the cited result must handle indefinite quadratic forms under the condition a(tilde v, psi)^2 < 1/8. Please state the precise conditions and either prove Lemma 4 or reproduce the relevant theorem from the cited paper; without this, the main results cannot be independently checked.","section":"Section 4.1, Lemma 4"},{"comment":"The claim that u_n^T W_n^S(theta_n) converges to N(0,1) when beta is unknown is not immediate from Lemma 5. The score for beta is exactly normal, but it is uncorrelated with, not generally independent of, the quadratic score for psi; in a Gaussian vector a linear form and a quadratic form can have zero covariance yet be dependent. The proof must show how this dependence is controlled in the linear combination u_n^T W_n^S(theta_n). The remark that the result is 'suggested, but not implied' by the marginal normality of the beta-score confirms that a gap remains in the main text.","section":"Section 4.1, Theorem 2, second assertion"},{"comment":"The finite-sample bound (15) is valid only when a(tilde v, psi)^2 < 1/8. The design-specific bounds in Theorem 3 (18) and Theorem 4 are often larger than this threshold for small m or small n_min; for example, Theorem 4 with n_min = 3 gives tilde a^2 <= 1. The paper should state explicitly that the finite-sample statements are conditional on the right-hand side being below 1/8, and that the asymptotic corollaries rely on this threshold being eventually satisfied.","section":"Sections 4.2 and 4.3, finite-sample applicability"}],"minor_comments":[{"comment":"In the displayed definition of q_{1-alpha}(psi), the inequality appears reversed: it should be min{t : F(t) >= 1-alpha}, not min{t : F(1-alpha) >= t}.","section":"Section 2, Lemma 2"},{"comment":"In the display defining the crossed-effects projections, P2 = Z^{(2)}Z^{(2)T}/n_2 should be divided by n_1, not n_2, to make P2 a projection matrix; the subsequent expression for Sigma confirms this.","section":"Section 4.3"},{"comment":"The notation P_n subset of R^{r_n x r_n} in the theorem statement should be P_n subset of R^{r_n}; the parameter vector psi_n is in R^{r_n}, not in a matrix space.","section":"Theorem 2"},{"comment":"The symbol P_j is used both for the n_j x n_j projection matrix 1_{n_j}1_{n_j}^T/n_j and for the n x n Kronecker-product projection; please use distinct symbols or otherwise clarify to avoid confusion.","section":"Section 4.3"},{"comment":"The phrase 'fine-sample coverage probabilities' should read 'finite-sample coverage probabilities'.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The key inequality in Lemma 4 is borrowed from a paper co-authored by the first author; this is not circular because that paper is published and its result is used as a black box, but it does mean that the main theorem's correctness is partly inherited from a prior submission. The absence of the supplement from the reviewed material is the main obstacle to independent verification. I would be comfortable with acceptance if the supplement supplies the cited proofs in full and if the unknown-beta argument in Theorem 2 is made explicit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the paper. The headline: this is the first uniform theory I've seen for score-based confidence regions in linear mixed models that covers the singular boundary and crossed random effects. If the supplement checks out, the main theorems are genuine advances. The core mechanism—a finite-sample bound on the density of standardized score statistics, uniform in the parameter—is new, and it doesn't reduce to Zhang et al. (2025) or Ekvall and Bottai (2022), though it builds on both.\n\nThe paper is honest about what it does and doesn't do. Section 4.2 flags the zero-column cluster exclusion explicitly. Theorem 4's n_min ≥ 2 or 3 conditions are stated plainly. The simulations, while not error-barred, show the pattern clearly: score-based regions hold near-nominal coverage across boundary and interior, while Wald and likelihood-ratio regions drift. The code is linked.\n\nThe soft spots are real but proportionate. All proofs are deferred to a supplement I haven't seen, and Lemma 4 is the load-bearing step. The bound requires a(v, psi)^2 < 1/8, and as the stress test notes, A(v, psi) is typically indefinite for covariance parameters. I can't rule out a hidden regularity condition in the crossed-effect argument without reading the supplement. That's a verification risk, not a demonstrated flaw. The design-replication assumptions are restrictive—Theorem 3(ii) rules out random effects appearing via zero columns in any cluster—but the authors acknowledge this and it's a scope limitation, not an internal inconsistency. The profile-score region's poor performance when p is large isn't swept under the rug.\n\nVerdict: this deserves refereeing. The central claim is well-specified, the literature treatment is fair, and the gap it fills is real. Whether the bounds hold as stated depends on a supplement that should be reviewed carefully. I'd want the referee to check Lemma 4 and the spectral arguments in Theorem 4.","headline":"This paper makes a genuine advance in uniform inference for variance parameters in linear mixed models, including crossed random effects; the main risk is the deferred proof of the key lemma.","tokens_in":16591,"tokens_out":2220,"would_cite":true,"duration_ms":25234,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F12","62F25","62J10"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proves that score-based chi-squared confidence regions for random-effect variances and covariances have uniformly correct coverage in linear mixed models, including on the singular boundary where variances are zero or…","keywords":["linear mixed models","random effects","score statistic","uniform inference","boundary of parameter set","crossed random effects","finite-sample bounds","confidence regions"],"falsifier":"Simulate a crossed random-effects model with $r=3$ fixed factors but with one factor's size kept at $n_{\\min}=1$ while the others grow, compute the restricted score confidence region at parameters where a random-effect variance is zero, and check whether the actual coverage error stays as small as the paper's uniform bound predicts; if it does, the paper's sufficient condition $n_{\\min}\\ge3$ is shown to be unnecessary, and if it does not, the condition is doing real work. A more direct check is to evaluate $a(v,\\psi)$ numerically for such a design and verify whether the Theorem 4 bound is violated.","tokens_in":15663,"feed_emoji":"📊","tokens_out":8416,"duration_ms":86238,"temperature":0.7,"pith_summary":"The paper proves that, for linear mixed models, the score statistic standardized by expected Fisher information is approximately normal with finite-sample error bounds that hold uniformly over the parameter set, including points where the random-effect covariance matrix is singular—variance zero or correlation ±1. That uniformity is the key advance: existing pointwise asymptotic results fail near the boundary because Wald and likelihood ratio statistics change their asymptotic distributions as the parameter approaches the boundary, while the score statistic does not. From the bounds the authors construct chi-squared confidence regions for variance and covariance parameters whose coverage error is controlled explicitly, so one can judge whether a sample is large enough and obtain asymptotically correct uniform coverage. The same mechanism covers crossed random effects, where the number of independent observations is not the sample size, and where no comparable boundary theory existed. Simulations show these regions stay near nominal coverage in finite samples in settings where Wald and likelihood ratio regions drift.","feed_headline":"Score statistic gives uniform inference at the mixed-model boundary","feed_subtitle":"Chi-squared regions for random-effect variances stay near nominal even at variance zero or correlation ±1.","key_machinery":"The load-bearing object is the ratio $a(v,\\psi)$, defined as the spectral norm divided by the Frobenius norm of $A(v,\\psi)=\\Sigma(\\psi)^{-1/2}\\Sigma(v)\\Sigma(\\psi)^{-1/2}$; it measures how far a direction $v$ in parameter space is from being an identifiable, well-conditioned direction. Lemma 4, built on a normal approximation for quadratic forms, turns smallness of $a(v,\\psi)$ into an explicit bound on $|g(t;v,\\psi)-\\phi(t)|$, the difference between the density of $v^T W^S(\\psi)$ and the standard normal density. Lemma 5 converts the pointwise bound into a uniform one, and Theorem 2 converts that into asymptotic normality of every linear combination of the standardized score. For crossed random effects the analysis uses the fact that $\\Sigma(\\psi)$ is a linear combination of fixed projection matrices, so the eigenvectors of $\\Sigma$ do not depend on $\\psi$; this permits a spectral argument that yields the $\\psi$-free bound involving $n_{\\min}$ and $r$.","core_discovery":"On the paper's own terms, the central discovery is that the distribution of any linear combination of the standardized score $W^S(\\psi) = I(\\psi)^{-1/2}S(\\psi)$ is close to standard normal whenever the quantity $a(v,\\psi)=\\|A(v,\\psi)\\|/\\|A(v,\\psi)\\|_F$ is small, where $A(v,\\psi)=\\Sigma(\\psi)^{-1/2}\\Sigma(v)\\Sigma(\\psi)^{-1/2}$ and $\\Sigma(v)=\\sum_j v_j \\partial\\Sigma/\\partial\\psi_j$. Theorem 2 formalizes this as asymptotic normality, uniformly over unit vectors $v_n$, and extends it to unknown fixed effects $\\beta$ when $X$ has full column rank. In independent-cluster settings Theorem 3 gives the bound $a(v,\\psi)\\le c_3 m^{-1/2}(1+\\|\\psi_{-r}\\|/\\psi_r)$, so the approximation improves as the number of clusters grows even at the boundary. In crossed random effects settings Theorem 4 gives a bound depending only on the smallest factor size $n_{\\min}$ and the number of factors $r$—not on $\\psi$—so Corollary 4 yields $\\sup_{\\psi}|P_\\psi\\{\\psi\\in \\tilde{C}^S_n(\\alpha)\\}-(1-\\alpha)|\\to0$ over the whole parameter set when $r\\ge3$ is fixed and $n_{\\min}\\to\\infty$. The paper thus claims that score-based chi-squared confidence regions solve the boundary problem for these models, and the experiments support that claim for the restricted likelihood version.","pith_inferences":["We infer that the same $a(v,\\psi)$ ratio could serve as a practical diagnostic: before reporting a score-based confidence region, one could compute its maximum over the boundary of the fitted region and flag cases where the normal approximation is likely poor.","We infer that the mechanism is not fundamentally tied to normality: the score is a quadratic form, and the paper's remarks suggest the bound could be reworked for non-normal responses by replacing the normality-based moment conditions with estimates of third and fourth moments, perhaps via resampling.","We infer that allocating experimental units across all crossed factors matters more than total sample size for variance-component inference, a point with direct implications for designing multi-factor studies.","We infer that extending this style of bound to generalized linear mixed models would hit the same difficulty the paper identifies: without the linearity and joint normality, the spectral decomposition of $\\Sigma$ is lost, so the crossed-random-effects proof would need a different route."],"forward_implications":["If the central claim is right, the chi-squared quantile gives confidence regions for random-effect variances and covariances with asymptotically correct uniform coverage, including on the boundary where a variance is zero or a correlation is ±1.","For independent clusters, coverage improves at rate roughly $m^{-1/2}$ in the number of clusters, so the finite-sample bounds can be used to say how many clusters are enough before trusting a chi-squared region.","For crossed random effects, the effective sample size for variance-component inference is governed by the smallest factor size $n_{\\min}$, not the total number of observations, and uniform coverage holds over the entire parameter set when $r$ is fixed and $n_{\\min}\\to\\infty$.","The results also hold when the number of random-effect parameters grows with the sample size, and the restricted-likelihood version remains valid with a growing number of fixed-effect predictors.","Wald and likelihood-ratio regions do not have this property, so score-based regions are the default choice when boundary parameters are a live possibility."],"supporting_citations":[{"why":"Provides the normal approximation for quadratic forms that Lemma 4 is built on, and gives the single-variance-component case the paper generalizes.","marker":"Zhang et al. (2025)"},{"why":"Prior confidence regions near singular information and boundary points; the paper extends its ideas and contrasts its settings.","marker":"Ekvall and Bottai (2022)"},{"why":"Source of the uniform-inference lemma that turns convergence along parameter sequences into uniform coverage.","marker":"Mikusheva (2007)"},{"why":"Documents likelihood inference for small variance components and boundary behavior, motivating the score-statistic approach.","marker":"Stern and Welsh (2000)"},{"why":"Provides increasing-dimension asymptotics for two-way crossed mixed models and a decomposition idea used for crossed random effects.","marker":"Lyu et al. (2024)"},{"why":"Establishes pointwise asymptotic normality for crossed random effects, which the paper shows is insufficient near the boundary.","marker":"Jiang et al. (2024)"},{"why":"Gives the inverse-of-sum-of-matrices result used in the spectral decomposition part of the proof of Theorem 4.","marker":"Henderson and Searle (1981)"},{"why":"Classic treatment of nonstandard asymptotics at the boundary, the failure mode the uniform results address.","marker":"Self and Liang (1987)"}],"fun_headline_variants":["Score-based regions stay valid at variance boundaries","Uniform inference for mixed models on the boundary","Score-based confidence regions hit nominal coverage at boundary","Near-exact coverage for random-effect variances at zero","Uniform inference on the boundary of mixed models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole argument leans on the design matrices giving every random effect enough independent replication: in the clustered setting each cluster's within-cluster design matrix must be sufficiently well conditioned, and in the crossed settings every crossed factor must have at least two (or three for the restricted score) levels, with the smallest factor size growing for the asymptotic results; if that fails, the key ratio $a(v,\\psi)$ need not be small and the uniform normality bound does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Score-based regions stay valid at variance boundaries","Uniform inference for mixed models on the boundary","Score-based confidence regions hit nominal coverage at boundary","Near-exact coverage for random-effect variances at zero","Uniform inference on the boundary of mixed models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000479,"raw_usage":{"total_tokens":2416,"prompt_tokens":1037,"completion_tokens":1379,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":653,"completion_tokens_details":{"reasoning_tokens":1310}},"tokens_in":653,"tokens_out":1379,"duration_ms":9866,"temperature":1.0,"reasoning_tokens":1310,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:12:32.526586+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate a crossed random-effects model with $r=3$ fixed factors but with one factor's size kept at $n_{\\min}=1$ while the others grow, compute the restricted score confidence region at parameters where a random-effect variance is zero, and check whether the actual coverage error stays as small as the paper's uniform bound predicts; if it does, the paper's sufficient condition $n_{\\min}\\ge3$ is shown to be unnecessary, and if it does not, the condition is doing real work. A more direct check is to evaluate $a(v,\\psi)$ numerically for such a design and verify whether the Theorem 4 bound is violated.","supporting_citations":[{"cited_title":"O., and Molstad, A","cited_arxiv_id":null,"evidence_quote":"Provides the normal approximation for quadratic forms that Lemma 4 is built on, and gives the single-variance-component case the paper generalizes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior confidence regions near singular information and boundary points; the paper extends its ideas and contrasts its settings."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the uniform-inference lemma that turns convergence along parameter sequences into uniform coverage."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents likelihood inference for small variance components and boundary behavior, motivating the score-statistic approach."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides increasing-dimension asymptotics for two-way crossed mixed models and a decomposition idea used for crossed random effects."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the inverse-of-sum-of-matrices result used in the spectral decomposition part of the proof of Theorem 4."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classic treatment of nonstandard asymptotics at the boundary, the failure mode the uniform results address."}],"review_version":1}