{"id":"468f78a6-1261-44d9-bc94-351e845b502b","arxiv_id":"2607.18836","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":1,"one_line_summary":"The pooling factor predicts where PSIS-LOO will fail, and Rao–Blackwellised integration over random effects corrects those folds, reproducing exact leave-one-out refits at zero refit cost.","lead":"For hierarchical Bayesian models, PSIS-LOO — the usual cheap shortcut for leave-one-out cross-validation — fails on observations in small groups. This paper predicts those failures from the model's pooling structure and ships an estimator (RB-LOO) that matches exact refits without refitting the model.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"C3's 'refit only flagged folds' guarantee rests on residual base-khat operating point, which is only summarized by AUC 0.93; no threshold-level sensitivity/specificity is reported, so missed RB-LOO failures could break the exact-recovery claim.","rationale":"The reader's weakest assumption identifies the same underlying issue: RB-LOO's accuracy depends on base-posterior stability. My concern extends this to the operational reliability of the residual base-khat flag, which is the mechanism by which the paper claims to recover exact LOO while refitting only a few folds. The paper reports AUC but not the confusion matrix at the 0.7 threshold, so the safety net's false-negative rate is unknown. This is a genuine soft spot because the central claim—'two-level triage recovers the exact answer'—depends on the flag not missing folds where RB-LOO errs. The authors are appropriately honest about RB-LOO's limitations and they do provide E-basestress as a falsification test, which is strong independent support. However, the evidence is still one simulation family, and the operating point is uncharacterized. A concrete diagnostic-validation study would settle whether the exact-recovery claim holds beyond the specific designs tested. I do not think the paper should be rejected; the methods and empirical benchmarks are sound and well-scoped. But the exactness guarantee should be conditioned on the residual diagnostic's sensitivity/specificity being validated, hence CONDITIONAL rather than unconditional ACCEPT.","tokens_in":12854,"tokens_out":9386,"duration_ms":84836,"concrete_test":"Re-analyze the E-basestress data (Section 5.5) and report the full 2x2 table of residual base-khat>0.7 vs exact-refit divergence>0.25 nats, at the standard threshold and at the Youden-optimal threshold. Then run the same protocol with a higher-dimensional base (e.g., add 10 uncorrelated fixed effects) and J=4; if any divergence>0.25 nat occurs with base-khat<0.7, the exact-recovery claim needs qualification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central two-level triage (C3) claims that residual base-khat>0.7 flags exactly the folds where RB-LOO diverges from an exact refit, so refitting only flagged folds recovers exact LOO at a few percent of the cost. The load-bearing element is the reliability of this residual diagnostic, yet Section 5.5 reports only AUC 0.93 (base leverage 0.90) for separating divergences >0.25 nats. AUC is rank-based and does not characterize the operating point at the 0.7 threshold; with only 65 flagged folds among 2340, the threshold could silently miss a non-negligible share of folds whose RB-LOO error exceeds the tolerance. If so, the 'refitting only 3%' result is an in-sample artifact of the E-basestress design rather than a guarantee, and the model-selection verdict of Section 5.4 could shift in other weakly-identified hierarchies. The paper's own honesty note says the base leverage is theoretically correct but operationally close to residual base-khat; however, E-basestress covers only Gaussian LMMs with J=4-6 and does not vary base dimension, group count, or likelihood family. Thus the assumption that base-posterior instability is either absent or detectable is the least secured condition supporting the exact-recovery claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies why PSIS-LOO fails for hierarchical models and proposes a two-part remedy: a weight-free triage based on the Gelman–Pardoe pooling factor and the associated structural leverage (C1), and an integrated importance-sampling estimator, RB-LOO, that marginalises the random-effect block before forming base-level importance weights (C2). A Schur-complement decomposition (C3) separates case-deletion influence into a vertical/pooling component, claimed to govern PSIS-LOO failures, and a horizontal/variance-component component, claimed to govern residual RB-LOO failures. The empirical sections validate the predictor against exact refits in Gaussian LMMs and logistic GLMMs, benchmark RB-LOO against moment matching on singleton-heavy designs, and use the overdispersed epilepsy data to show that RB-LOO changes a model-selection verdict. The paper is explicit about the prior art for marginalisation, about pooled AUC optimism, and about the limited regime in which RB-LOO itself is stressed.","tokens_in":13197,"tokens_out":5795,"duration_ms":57348,"significance":"If the main claims hold, the paper provides a practical and conceptually useful advance: a principled, closed-form predictor of where PSIS-LOO will fail, and a cheap marginalisation-based repair that is well matched to the failure mode. The algebraic propositions are clean and the reproduction scripts are a genuine strength. The paper also makes honest, explicit limitation statements that help bound the claims. The least secured part is the two-level triage claim in C3: the operational threshold for sending folds to exact refit is validated only by a rank-based AUC, not by threshold-level sensitivity/specificity, and the stress-test design covers only Gaussian LMMs with few groups. Because the abstract's 'refitting only the few folds that need it' guarantee rests on this operating point, this needs additional evidence or a more carefully qualified statement.","major_comments":[{"comment":"The protocol in C3 refits folds with residual base-khat > 0.7, yet the supporting evidence is summarized only by pooled AUC 0.93 for separating |Delta(elpd)| > 0.25 divergences. AUC is rank-based and does not characterize sensitivity/specificity at the 0.7 cutoff. With 65 flagged folds among 2340, the paper should report the confusion matrix at the threshold, the missed-divergence count and magnitude, and the hybrid RMSE under a range of thresholds. The statement that the two-level triage 'recovers the exact answer' is an operating-point claim, and the current reporting does not establish it.","section":"Section 5.5 (E-basestress) and Section 4 (refit flag)"},{"comment":"The stress test of RB-LOO is limited to Gaussian LMMs with J=4-6 groups, singletons, and a weakly identified sigma_u; no variation in likelihood family, base dimension, or group structure is exercised. Since the abstract generalizes to 'the few folds that need it' and Section 5.4 uses the result to support a decision-changing claim, the authors should either add stress-test configurations covering broader hierarchical settings (e.g., logistic GLMMs, more base parameters) or explicitly restrict the C3 guarantee to the demonstrated regime. This is load-bearing for the exact-recovery claim.","section":"Section 5.5 and Abstract (generality of C3)"},{"comment":"In the Gaussian LMM case, the leverage reduces to group size plus a fitted scalar sigma_u, so the map is available after a variance-component estimate, not literally before seeing outcomes. The paper's Section 7 honesty note is accurate, but the abstract and Section 4 use 'a-priori' and 'design-time' in ways that could mislead readers. Please carry the Section 7 qualification into the abstract and introduction.","section":"Section 4 and Abstract ('design-time' wording)"}],"minor_comments":[{"comment":"The text reports only AUC for the residual base-khat flag. A calibration plot or a table of sensitivity/specificity for khat thresholds near 0.7 would greatly improve the practical usability of the triage.","section":"Section 5.5"},{"comment":"'rbloopackage' appears without a space/backtick in the code/software paragraph; check formatting. Consider adding a versioned release URL for the package.","section":"Reproducibility section"},{"comment":"The statement that the marginalised estimator is 'the standing recommendation of the loo documentation' is helpful, but citing the specific documentation version would make the provenance checkable.","section":"Section 2"},{"comment":"The paper's own limitation notes are commendable. In particular, the acknowledgment that pooled AUCs are optimistic and that per-replicate Spearman is primary is exactly the right way to report dependent folds; the same standard should be applied to the AUC-0.93 threshold analysis in Section 5.5.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is a strong, well-written applied paper with reproducible scripts and honest limitations. My main concern is narrow but load-bearing: the C3 'refit only flagged folds' protocol needs threshold-level operating characteristics and broader stress testing before the abstract-level guarantee is credible. I would be happy to accept after those additions. The paper is well within the scope of a statistics methodology journal and the contribution is likely to be useful to practitioners."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this paper is worth engaging with. The new thing is the pooling-factor-to-PSIS-reliability map (C1) and the packaged observation-level closed forms for RB-LOO, plus the head-to-head showing marginalisation beats moment matching on singleton-heavy logistic GLMMs. The marginalisation itself is prior art, and the paper says so clearly. What impressed me is the honesty infrastructure: exact refits as gold standard, per-replicate MCSE, explicit statements about what the experiments can and cannot falsify, and a limitations section that names the estimand-sharing issue on singleton folds. That is how you do simulation work.\n\nThe core empirical claims hold up in the tested regimes. On the overdispersed epilepsy data, RB-LOO reproduces an 82-minute exact refit with RMSE 0.04 while PSIS-LOO and moment matching are over-optimistic and even biased on sub-threshold folds. The model-selection flip (z=4.9 to z=1.0) is dramatic but the paper backs it with exact refits, so it's credible.\n\nSoft spots, in proportion. The two-level triage (C3) is the least secured claim. The 'refit only 3% of folds and still match exact LOO' result comes from a single regime: few-groups Gaussian LMMs with weak base identifiability. The residual base-khat flag is summarized only by AUC 0.93, not by a confusion matrix at the 0.7 threshold. AUC is rank-based; it doesn't tell you how many diverging folds get sent to refit or missed. With only 65 flagged folds among 2340, a missed divergence could shift the aggregate RMSE. The stress-test note is right about this. It's not a fatal flaw—the combined RMSE 0.067 is reported directly—but the operating point deserves a table, and the experiment should be extended to logistic/Poisson blocks with small J before calling it a general guarantee.\n\nAlso the 'design-time' framing is a bit strong even for Gaussian LMMs: the leverage ranking depends on the fitted variance ratio, so it's post-fit unless you treat the ratio as known. The paper's own honesty section acknowledges this, but the abstract oversells it.\n\nWho gets value: anyone doing hierarchical model comparison with loo, and methodologists working on PSIS repairs. It deserves a serious referee. I'd send it out, with a request to add threshold-level diagnostics and broaden the few-groups stress test.","headline":"A genuinely useful and unusually honest methods paper; the core triage works in the tested regimes, and the RB-LOO cure is right, but the 'refit only flagged folds' guarantee is narrower than the abstract suggests.","tokens_in":13648,"tokens_out":3400,"would_cite":true,"duration_ms":30137,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62G09","62J12"],"pacs":[],"model":"deepseek-v4-flash","headline":"In hierarchical models, the observations where cheap leave-one-out cross-validation collapses can be predicted in advance from the pooling factor, and a marginalised estimator corrects them without refits.","keywords":["cross-validation","leave-one-out","PSIS-LOO","hierarchical models","partial pooling","importance sampling","pooling factor","structural leverage"],"falsifier":"Run RB-LOO and exact leave-one-out refits on a Gaussian LMM with only 4–6 groups and singleton observations; if the residual base-khat diagnostic fails to separate the folds where RB-LOO's elpd deviates from exact refits by more than 0.25 nats (the paper reports AUC 0.93), the two-level triage claim is false.","tokens_in":12759,"feed_emoji":"📊","tokens_out":13157,"duration_ms":112785,"temperature":0.7,"pith_summary":"The paper sets out to show that the standard cheap approximation to leave-one-out cross-validation (PSIS-LOO) fails in hierarchical models in a predictable, structured way: exactly on the folds where a random-effect coordinate is driven by the data and its group is small. It argues that the pooling factor—the share of a group's posterior precision that comes from the prior—and an associated per-observation structural leverage identify these failing folds without forming any importance weights, and in Gaussian linear mixed models the map is determined by group sizes alone before outcomes are seen. The proposed cure, RB-LOO, marginalises the random-effect block and importance-samples only the fixed-effect and variance parameters, in closed form for Gaussian blocks and one-dimensional quadrature for Bernoulli, binomial and Poisson blocks. If the paper is right, the most troublesome failures of routine cross-validation can be flagged in advance and corrected with zero refits, and model-comparison verdicts that look decisive under the standard tool can be shown to be an artefact of the approximation.","feed_headline":"Partial pooling predicts where cross-validation fails","feed_subtitle":"A pooling-factor score flags the failing folds; a marginalised estimator matches exact refits at zero cost.","key_machinery":"The central object is a geometric split of each observation's case-deletion influence, obtained as a Schur complement of the Fisher information: a vertical term living in the random-effect direction, which sums to the pooling factor and predicts where PSIS-LOO fails, and a horizontal term living in the base (fixed-effect/variance) direction, which predicts where the replacement estimator is itself strained. The estimator that carries the cure, RB-LOO, is a variance-reduced importance-sampling scheme: instead of reweighting the random effect, it marginalises it out of the leave-one-out predictive—analytically via a rank-1 Gaussian downdate, or by one-dimensional quadrature for scalar Bernoull","core_discovery":"The paper's core claim: in hierarchical models, the folds where PSIS-LOO fails are exactly those where deleting one case moves a data-driven random-effect coordinate, and the pattern is visible before any importance weights. The pooling factor and structural leverage separate failing folds (khat>0.7) with AUC 0.96 in Gaussian LMMs from group sizes alone, and AUC 0.81 in logistic GLMMs from one fit. RB-LOO marginalises the random-effect block and importance-samples only base parameters (analytic Gaussian downdate; 1-D quadrature for Bernoulli/binomial/Poisson). On 97 failing overdispersion folds it matches exact refits (elpd RMSE 0.041 vs 0.638) and flips a model-selection z from 4.9 to 1.0;","pith_inferences":["Beyond the paper, the pooling-factor-to-reliability map could be used at study design time in hierarchical settings: given planned group sizes and a prior guess at the variance component, one could flag in advance the observations whose deletion would make cross-validation unstable.","Beyond the paper, the decision flip from a per-observation latent model versus a marginal alternative suggests that latent-vs-marginal model comparisons may carry systematic PSIS-LOO bias that does not cancel in the difference; testing this on other latent structures (mixtures, factor models) is a natural extension.","Beyond the paper, the base-fiber split is a geometric claim that likely applies to other case-deletion diagnostics, not just PSIS-LOO; it gives a template for predicting when any leave-one-out approximation that reweights the full posterior will have heavy-tailed weights.","Beyond the paper, combining the residual base-khat flag with an automated exact-refit fallback would produce a fully adaptive workflow that never silently trusts the approximation; the paper reports the flag's accuracy but leaves the automation as a design choice."],"forward_implications":["Model-selection verdicts in hierarchical and overdispersed data can be artifacts of the approximation: the paper shows a decisive-looking z=4.9 difference becoming z=1.0 under exact refits, with RB-LOO reproducing the exact refit.","In Gaussian LMMs the failure map is available before outcomes are collected, so group sizes alone can warn which observations will need special handling.","The drop-in implementation lets a standard workflow flag the risky folds, marginalise them, and refit only the residual few percent, recovering refit-quality LOO at a small fraction of the cost.","On singleton-driven failures RB-LOO is substantially more accurate than the standard moment-matching repair—about 3x pooled on logistic GLMMs—and corrects failures that moment matching leaves in place on real count data.","The method applies to single random-intercept groupings; crossed or multiple grouping factors are explicitly out of scope and fall back to plain PSIS-LOO."],"fun_headline_variants":["Pooling factor predicts where cross-validation fails","Predict PSIS-LOO failures before weighting","Marginalised estimator matches exact refits cheaply","Design-time map flags failing cross-validation folds","RB-LOO: exact accuracy without costly refits"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that deleting one observation leaves the posterior of the fixed effects and variance components essentially unchanged, so that reweighting the full-data posterior over only those base parameters is safe; when a single deletion moves the base posterior appreciably—few groups, weak identification—RB-LOO itself needs an exact refit.","fun_headline_variants_meta":{"raw":{"variants":["Pooling factor predicts where cross-validation fails","Predict PSIS-LOO failures before weighting","Marginalised estimator matches exact refits cheaply","Design-time map flags failing cross-validation folds","RB-LOO: exact accuracy without costly refits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000192,"raw_usage":{"total_tokens":1292,"prompt_tokens":964,"completion_tokens":328,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":258}},"tokens_in":708,"tokens_out":328,"duration_ms":3799,"temperature":1.0,"reasoning_tokens":258,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T14:10:28.709967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run RB-LOO and exact leave-one-out refits on a Gaussian LMM with only 4–6 groups and singleton observations; if the residual base-khat diagnostic fails to separate the folds where RB-LOO's elpd deviates from exact refits by more than 0.25 nats (the paper reports AUC 0.93), the two-level triage claim is false.","supporting_citations":[],"review_version":1}