{"id":"b2240904-5efd-4659-a3f0-37bde016f873","arxiv_id":"2509.00255","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A randomized split likelihood ratio test provides finite-sample valid inference for variance components at the boundary, including heritability near 1, with computational shortcuts for structured models.","lead":"This paper shows how to build confidence intervals and tests for variance components that stay valid in finite samples even when some variance components are zero or a proportion like heritability is close to one. It adapts the split likelihood ratio framework from universal inference and adds algorithms that make the method much faster for common model structures.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Sec. 3.4's approximate diagonalization replaces known K_m with \\tilde{K}_m in the likelihood; this misspecification is not covered by Theorem 1, so the finite-sample uniform-validity claim does not extend to the approximate algorithm shown in Fig. 3 (middle), leaving the central claim overbroad with","rationale":"The reader's verdict is CONDITIONAL, and the identified weakest assumption flags the approximate diagonalization as a misspecification not covered by the theory. My concern is the same point, sharpened: the approximate diagonalization in Sec. 3.4 is not merely a minor numerical detail but a change to the likelihood that breaks the core e-value property on which universal inference rests. The central claim of finite-sample uniform validity holds for the exact likelihood, but the paper's presentation of the approximated method in simulations and its silence about the lack of a guarantee could mislead readers into believing the fast method is valid.\n\nI do not see an internal inconsistency in the exact theory; Theorem 1 is correctly proved and the application to variance components is sound when K_m are known and used exactly. The issue is scope: the paper should either restrict the claim to exact K_m and clearly state that the approximate diagonalization is heuristic, or prove a finite-sample bound on the misspecification error. Therefore the verdict remains CONDITIONAL, as the reader concluded. My proposed simulation is a direct check of whether the approximate method actually fails in a challenging case; if it does not fail in that case, the concern would be weakened, but the theoretical gap remains.agreement_with_reader is 'agree' because the reader's weakest assumption already identifies approximate diagonalization as a second misspecification; my stress-test independently confirms it as the most load-bearing concern for the central claim.","tokens_in":16213,"tokens_out":11915,"duration_ms":150623,"concrete_test":"Simulate the Sec. 3.4 setting with poor approximation quality (e.g., c = 1, so the trailing eigenvalues are not small), generating Y from the true K_m and computing 95% randomized split-LRT confidence intervals for h^2_1 using \\tilde{K}_m. Vary true h^2_1 and h^2_2 over a grid including boundary and near-unity values; estimate coverage with at least 10,000 replications. If any configuration shows coverage below 0.95 beyond Monte Carlo error, the approximate method is not uniformly valid. Alternatively, analytically bound E_{θ*}[T_n] - 1 under this misspecified likelihood; if the bound exceeds α for some θ*, the failure is proven.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Universal inference (Theorem 1) requires the numerator and denominator in the split likelihood ratio to be true likelihoods under the data-generating distribution. The proof in Appendix A uses the Markov inequality and the identity E_{θ*}[L(\\hat{θ}_1)/L(θ*)] = 1, which holds only when the likelihood is correctly specified. Section 3.4 proposes replacing K_m by jointly diagonalizable approximations \\tilde{K}_m when the K_m are 'approximately' diagonalizable. The middle panel of Fig. 3 runs the randomized split LRT with these approximate covariance matrices and reports size and power nearly unchanged. However, no theorem establishes that the resulting statistic retains the e-value property E_{θ*}[T_n] ≤ 1. With \\tilde{K}_m, the likelihood is misspecified; the density ratio is no longer a true likelihood ratio, and the expectation can exceed 1, breaking uniform finite-sample coverage. The paper does not explicitly state that this approximate implementation lacks the theoretical guarantee, so readers may incorrectly infer that the fast algorithms (a central contribution) inherit universal validity. This is load-bearing because the abstract and Section 6 claim uniform finite-sample validity without qualification, while the approximate diagonalization is presented as a useful algorithmic device in simulations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies universal inference—specifically the split likelihood-ratio test and its randomized version—to Gaussian variance components models with parameterization θ=(h²,τ²). The authors show that tests and confidence intervals for variance components and their proportions are uniformly valid in finite samples, including settings where nuisance components are near their boundary or where the proportion of variability is near unity. They develop faster algorithms for models with diagonalizable covariance structure, including the case of crossed random effects, and they propose an approximate diagonalization scheme for near-block-diagonal settings. The main theorem (Theorem 1) is known from the universal inference literature, but the paper provides a self-contained proof and adapts the method to the variance components setting. Simulations and a crossed-random-effects data example illustrate the methods.","tokens_in":16548,"tokens_out":8058,"duration_ms":115259,"significance":"If the stated guarantees hold, the paper makes a useful contribution: it offers finite-sample valid inference for boundary and near-unity variance components, where standard asymptotic/score-based methods can be unreliable. The computational gains from diagonalization are concrete and potentially important for large n. The paper is honest about conservativeness, gives reproducible code, and includes a clear proof of the underlying split-LRT validity. Its main limitation is that the validity theorem requires the Gaussian model and the covariance matrices K_m to be correctly specified; the approximate diagonalization used in one simulation and the centering step in the data example are not covered by that theorem. These gaps are fixable by qualification or additional theory, but they currently make the abstract's and Section 6's blanket uniformity claims overbroad.","major_comments":[{"comment":"The approximate diagonalization is not covered by Theorem 1. Replacing K_m by \\tilde{K}_m in Eq. (4) means the numerator and denominator are densities from a misspecified model, so the key identity E_{θ*}[L_{\\hatθ_1}/L_{θ*} | Y^(1)] = 1 in Appendix A fails and E[T_n] ≤ 1 is not guaranteed. The middle panel of Fig. 3 therefore demonstrates only empirical size/power, not finite-sample uniform validity. The abstract and Section 6 state uniform finite-sample validity without this qualification. Please either provide conditions/proofs for the approximate version or clearly label it as heuristic and restrict all validity claims to the correctly specified model.","section":"Sec. 3.4 / Fig. 3 (middle)"},{"comment":"The data example centers the response by the sample mean and then analyzes the centered vector as if model (11) held with mean zero and full-rank covariance. Under the stated Gaussian model this creates a misspecified or singular likelihood, so the finite-sample guarantee of Theorem 1 cannot be directly invoked for the reported intervals. Please either use the known-X/unknown-β extension mentioned in the introduction or explicitly state that the data example is approximate and outside the formal guarantees.","section":"Sec. 5 / model (11)"}],"minor_comments":[{"comment":"The conditional log-likelihood appears to omit the n^(0) log τ² term. Later profiling formulas include the τ² dependence, so this is presumably a typographical omission, but it should be corrected for clarity.","section":"Eq. (5)"},{"comment":"Equation (8) is stated for general joint diagonalizability, but Theorem 2 is proved only under the stronger condition K_m K_ℓ = 0. The crossed random effects construction in Sec. 3.3 yields Λ_m Λ_ℓ ≠ 0, so the reader cannot infer joint diagonalizability from Theorem 2 there. Since O is constructed explicitly, this is not an algorithmic gap, but the text should either state the standard commuting-symmetric-matrices fact or clarify that Theorem 2 is only a sufficient condition.","section":"Sec. 3.2"},{"comment":"The k-fold averaging of test statistics is used without formal justification. Please add a sentence explaining why the averaged statistic retains the e-value property, or cite the relevant result in Wasserman et al. (2020).","section":"Sec. 5"},{"comment":"'Mote Carlo' should be 'Monte Carlo'.","section":"Fig. 3 caption"},{"comment":"Since the formal theorem assumes a correctly specified Gaussian model with known K_m, I suggest adding an explicit qualifier such as 'under correct specification of the covariance model' when stating the finite-sample uniformity guarantee.","section":"Abstract / Sec. 6"}],"recommendation":"major_revision","confidential_remarks":"The central theorem is known and correctly proved; the contribution is primarily the application and the computational algorithms. The approximate-diagonalization gap is fixable by qualification, but it is important because the abstract and conclusions currently overstate the scope of the validity guarantee. I would not reject; I would like to see the claims carefully re-scoped."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The paper applies split likelihood ratio tests (Wasserman et al. 2020) with the randomized threshold (Ramdas and Manole 2023) to variance components models, focusing on boundary cases like heritability near 1. The core theoretical novelty is modest: Theorem 1 is a restatement of existing results, proven cleanly. The real contributions are the case-specific likelihood simplifications, the algorithms exploiting shared eigenvectors and crossed random effects, and the demonstration of practical speedups in simulations. These are genuinely useful for practitioners who need finite-sample valid intervals in mixed models. The paper is well-written, the derivations in Sections 3.1–3.3 check out, and the code is available.\n\nThe stress-test concern is on point. Section 3.4 proposes approximate diagonalization, replacing the true K_m with approximations that are jointly diagonalizable. But universal inference only guarantees coverage when the likelihood is correctly specified. With approximated covariance matrices, the test statistic is not an e-value in general, so the finite-sample coverage claim does not extend to that implementation. The middle panel of Fig. 3 is not backed by any theorem. The paper never flags this, so a reader could easily infer that the fast algorithm inherits the validity guarantee. That is a real exposition problem, not a fatal one: the exact algorithms in Sections 3.1–3.3 are valid, and the fix is to add a sentence saying the approximate version is a heuristic. Also, the data example centers the response using an estimated mean and treats it as known, a minor misspecification only addressed indirectly through diagnostics. There is also a dangling '’c.f. 4'’ reference in Section 5, trivial to fix.\n\nOverall, this is a solid application of universal inference to an important problem. The main theorem is known, but the boundary and near-unity heritability setting is new, and the computational contributions are substantial. I would send it to a serious referee. The paper needs a revision that clearly separates theoretical guarantees from algorithmic heuristics, but after that it’s a good, useful paper.","headline":"A solid application of known universal inference machinery to the boundary/near-unity variance components problem, with real algorithmic contributions; the main caveat is that the approximate diagonalization in Sec. 3.4 is used without flagging that it falls outside the finite-sample coverage guarantee.","tokens_in":16977,"tokens_out":2970,"would_cite":true,"duration_ms":36501,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F25","62J10"],"pacs":[],"model":"deepseek-v4-flash","headline":"Split-data likelihood tests produce finite-sample valid confidence intervals for variance components at the boundary.","keywords":["universal inference","variance components","split likelihood ratio test","finite-sample validity","heritability","boundary parameters","crossed random effects","confidence intervals"],"falsifier":"Simulate Y from the stated Gaussian model with M=2, K_1 and K_2 fixed and known, true h^2_1 = 0 and h^2_2 = 0.99, n = 300, and construct the randomized split likelihood ratio confidence interval for h^2_1 over 10,000 replicates; if the empirical coverage falls below the nominal 95% level by more than simulation error, the finite-sample uniform validity claim is false.","tokens_in":16158,"feed_emoji":"🧬","tokens_out":5167,"duration_ms":56485,"temperature":0.7,"pith_summary":"Variance components models attribute trait variability to genetic, environmental, or other sources, but classical confidence intervals break down exactly when a variance component is near zero or one, or when other components are. This paper shows that a universal inference scheme, the randomized split likelihood ratio test, fixes the problem. Splitting the data, fitting parameters on one half, and testing on the other gives confidence intervals whose coverage is at least nominal for every finite sample, with no asymptotic approximation. The paper also develops fast implementations for models with diagonalizable covariance structure, including crossed random effects, where naive computation would scale as n^3.","feed_headline":"Split-likelihood tests give valid CIs for heritability near one","feed_subtitle":"The method stays honest in finite samples even when other variance components sit at zero or one.","key_machinery":"The split likelihood ratio test: partition Y into Y(0) and Y(1) independently of the data, compute estimates from Y(1), evaluate the conditional likelihood of Y(0) given Y(1) at those estimates under the null and the alternative, and reject when the ratio T_n exceeds U/alpha with U ~ Uniform(0,1). Its validity comes from the conditional expectation of the ratio equaling exactly one under the null, so Markov's inequality gives a finite-sample bound. Computation exploits the covariance structure tau^2 Psi(h^2): when the K_m share eigenvectors or are mutually orthogonal, Psi becomes diagonal, reducing likelihood and derivative evaluations from O(n^3) to O(n); crossed random effects are mapped o","core_discovery":"The paper establishes uniformly valid finite-sample confidence intervals for variance components and for proportions of variability such as heritability, including when the proportion is near one. For the Gaussian model Y ~ N(0, tau^2(sum_m h^2_m K_m + (1 - sum_m h^2_m) I)), the split likelihood ratio statistic compares the conditional likelihood of one data half at estimates from the other half under the null and alternative hypotheses. Markov's inequality bounds the rejection probability by alpha in any finite sample, and randomizing the threshold with a uniform variable preserves validity while increasing power. This yields, to the paper's knowledge, the first method that is uniformly val","pith_inferences":["The finite-sample guarantee is purchased from the likelihood being exactly right; applying the same split-test recipe to non-Gaussian data or to K_m estimated from the same data would require a separate argument.","Approximate diagonalization replaces the true covariance structure with a close but misspecified one; the simulation suggests the coverage loss is small, but the paper's theory does not formally cover that trade-off.","The choice of how to split the data, uniformly at random versus balancing factor levels, likely affects power more than coverage; the paper leaves this as an open question, and comparing interval widths across split designs would test it directly.","If the conservativeness can be reduced, the method could become a practical default replacement for Wald and profile intervals in mixed-model software, not just a safeguard for boundary cases."],"forward_implications":["Tests and confidence intervals for one variance component remain valid when other variance components sit at zero or one, where score-based and asymptotic intervals under-cover.","Heritability and other proportions of variability can receive finite-sample confidence intervals even when the estimate is near one, a case previously not handled.","The same split-likelihood recipe applies to any correctly specified parametric model, so boundary problems in other mixed models inherit the finite-sample guarantee.","The randomized threshold reduces conservativeness relative to the plain split likelihood ratio test without sacrificing the coverage bound.","For models with shared eigenvectors, orthogonal K_m, or crossed random effects, the proposed algorithms make the method practical for n in the thousands."],"supporting_citations":[{"why":"Establishes the split likelihood ratio test and its finite-sample validity; Theorem 1 of this paper restates and proves it for completeness.","marker":"Wasserman et al., 2020"},{"why":"Introduces the randomized test that rejects when T_n > U/alpha, which the paper uses to gain power while preserving uniform validity.","marker":"Ramdas and Manole, 2023"},{"why":"Provides the crossed-random-effects formulation and uniform-inference context that support the paper's efficient algorithms.","marker":"Ekvall and Bottai, 2025"},{"why":"Supplies the fast score-based confidence interval for a single variance component that the paper extends and shows is invalid with nuisance components near the boundary.","marker":"Zhang et al., 2025"}],"fun_headline_variants":["Split-likelihood yields valid CIs for heritability near one","Boundary-proof CIs for variance components in finite samples","Uniformly valid CIs for heritability at the edge","New method gives honest CIs when variance shares hit limits","Finite-sample valid inference for variance components at boundary"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The finite-sample coverage guarantee assumes the Gaussian model is exactly right: the mean is zero or known, the K_m are known, and the response is multivariate normal; coverage is not claimed if the covariance structure is estimated from data, if the response is non-Gaussian, or if K_m is replaced by an approximation.","fun_headline_variants_meta":{"raw":{"variants":["Split-likelihood yields valid CIs for heritability near one","Boundary-proof CIs for variance components in finite samples","Uniformly valid CIs for heritability at the edge","New method gives honest CIs when variance shares hit limits","Finite-sample valid inference for variance components at boundary"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000391,"raw_usage":{"total_tokens":1870,"prompt_tokens":693,"completion_tokens":1177,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":437,"completion_tokens_details":{"reasoning_tokens":1095}},"tokens_in":437,"tokens_out":1177,"duration_ms":9434,"temperature":1.0,"reasoning_tokens":1095,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:47:52.087641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate Y from the stated Gaussian model with M=2, K_1 and K_2 fixed and known, true h^2_1 = 0 and h^2_2 = 0.99, n = 300, and construct the randomized split likelihood ratio confidence interval for h^2_1 over 10,000 replicates; if the empirical coverage falls below the nominal 95% level by more than simulation error, the finite-sample uniform validity claim is false.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the split likelihood ratio test and its finite-sample validity; Theorem 1 of this paper restates and proves it for completeness."},{"cited_title":"Uniform inference in linear mixed models","cited_arxiv_id":"2507.19633","evidence_quote":"Provides the crossed-random-effects formulation and uniform-inference context that support the paper's efficient algorithms."},{"cited_title":"O., and Molstad, A","cited_arxiv_id":null,"evidence_quote":"Supplies the fast score-based confidence interval for a single variance component that the paper extends and shows is invalid with nuisance components near the boundary."}],"review_version":1}