{"id":"5c26fd2e-b034-4da5-a16d-6cd8535f9aa1","arxiv_id":"2506.02303","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A two-stage Bayesian downscaling model estimates county-level opioid use disorder risk from state-level survey counts, but its validation uses data generated by the model itself and shows large county-level errors.","lead":"Researchers built a two-stage Bayesian model that turns state-level drug survey numbers into county-level estimates of opioid use disorder for every U.S. county from 2010 to 2025. The method is meant for policymakers who need local risk maps when direct county data do not exist.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"County-level credible intervals omit Stage I uncertainty because Stage II uses posterior medians as plug-in inputs, so the simulation's 0.89 coverage cannot validate the paper's uncertainty-quantification claim.","rationale":"The reader's weakest assumption is that simulation validation is a self-consistency check because the data-generating process in Section 3.5 uses the same ICAR/RW1 priors and Poisson likelihood as the fitted model. I agree that external validation is missing, but the more decisive technical problem is internal: Stage II is fed posterior medians, not full posterior draws, so the county-level credible intervals cannot include Stage I uncertainty. This directly undermines the abstract's promise to 'quantify associated uncertainties' and the claim that simulation shows 'well-calibrated uncertainty quantification.' The simulation coverage of 0.89 is computed under the same plug-in, so it is not evidence against the concern; if full posterior propagation raises coverage toward nominal and widens intervals, the current intervals are overconfident. The data-dependent prior for Δs,t is a second factor that could compress Stage I uncertainty, but the plug-in issue is sufficient to make the uncertainty claim conditional. The recommended verdict remains CONDITIONAL because the framework is coherent and the fix is straightforward: propagate draws, re-run the simulation, and report the effect on coverage and interval width. The applied estimates may still be reasonable, but their reported uncertainty is not yet credible.","tokens_in":29070,"tokens_out":5203,"duration_ms":55170,"concrete_test":"Modify the Section 3.3 post-processing so that, for each Stage I MCMC iteration m, Stage II is run with ỹ_s,t^(m) = π̂_s,t^(m) · n_s,t and ω̃_s^(m), and pool the resulting county-level posterior samples. On the 100 simulated datasets, compare county-level 95% interval width and coverage with Table 2. If coverage rises from 0.89 toward 0.95 and/or median interval width increases by more than roughly 10%, the plug-in of posterior medians is materially understating uncertainty. As a secondary check, recompute r_s,t from hold-out years before fitting Δs,t to assess the double-use of data in the Eq. 5 prior.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3 states that Stage I produces posterior medians for the state-level risk π̃s,t, the state-level count ỹs,t, and the state-level intercept ω̃s, and Section 3.4 uses these medians as fixed offsets in the county-level model: ỹs,t is treated as the known multinomial total and γs is centered on ω̃s. The reported county-level intervals in Figures 5 and 6 therefore reflect only county-level process uncertainty; they exclude uncertainty in the state-level totals, which are themselves estimated from sparse and definitionally heterogeneous NSDUH counts. This is in direct tension with the paper's stated 'fully Bayesian' propagation of uncertainty and with the central claim that the framework quantifies small-area uncertainty under data sparsity. The simulation evidence is affected in the same way: Table 2's county-level coverage of 0.89 is computed under the same plug-in pipeline, so it cannot establish that the reported credible intervals are well calibrated. Separately, the adjustment-factor prior in Eq. 5 uses r_s,t computed from the same observed counts y_s,t that enter the Stage I likelihood, creating a data-dependent prior that further narrows Stage I intervals. Both issues are fixable, but until full posterior draws are propagated, the uncertainty quantification claim is not supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops B-Step, a two-stage Bayesian spatio-temporal model that downscales state-level NSDUH opioid use disorder (OUD) counts to county-level risk estimates for 3,143 US counties over 2010-2025. Stage I is a state-level Poisson model with a first-order random walk over time and a case-definition adjustment factor; Stage II allocates the resulting state totals to counties using a softmax-normalized Poisson model with BYM spatial effects, RW1 temporal effects, and a population-weighted offset. The authors apply the framework to real surveillance data and evaluate it through 100 simulated datasets, reporting state- and county-level accuracy and coverage metrics.","tokens_in":29355,"tokens_out":3776,"duration_ms":37752,"significance":"If the central claim were fully supported, the paper would make a useful applied contribution: county-level OUD estimates with uncertainty are policy-relevant, and the proposed adjustment for the 2020 NSDUH definitional change addresses a real and under-documented harmonization problem. The paper also has practical strengths: it builds on publicly available NSDUH and Census data, the two-stage modular structure is transparent, and the use of a population-weighted offset to stabilize allocation in sparsely populated counties is a sensible practical fix. However, the validation is internal: the simulation generates data from the same Poisson/ICAR/RW1 structure used by the fitted model, and the reported credible intervals exclude Stage I uncertainty because posterior medians are plugged into Stage II. These limitations directly affect the paper's claims of 'strong accuracy and calibration' and 'fully Bayesian' uncertainty propagation, so the contribution is currently overstated.","major_comments":[{"comment":"The simulation data-generating process in Eq. (10) uses the same Poisson likelihood, the same ICAR/RW1 random-effect structure, and the same covariates as the fitted Stage II model, so the simulation is a self-consistency check rather than an independent validation of the model's assumptions. Consequently, Table 2's favorable state-level metrics and county-level coverage of 0.89 cannot support the Discussion's claim that simulations demonstrated 'strong accuracy and calibration' for real-world OUD estimation. The authors should either reframe the simulation explicitly as an internal-consistency check or supplement it with an external validation, for example by withholding some state-years, comparing with direct county-level estimates where available, or examining sensitivity to alternative data-generating mechanisms.","section":"Section 3.5, Eq. (10)"},{"comment":"Stage II uses posterior medians of the state-level risk, count, and intercept as plug-in inputs: Section 3.3 states that ~ys,t and ~omega_s are posterior medians, and Section 3.4 treats them as fixed offsets in the multinomial/Poisson allocation and the prior for gamma_s. The county-level credible intervals reported in Figures 5 and 6 therefore exclude Stage I uncertainty about the state totals and state intercepts, which is in direct tension with the paper's stated 'fully Bayesian' uncertainty propagation. Table 2's county-level coverage of 0.89 is computed under the same plug-in pipeline, so it cannot establish calibration of the reported county-level intervals. The authors should propagate full posterior draws of ~ys,t and ~omega_s through Stage II, or explicitly describe the reported intervals as conditional on Stage I point estimates.","section":"Sections 3.3 and 3.4, Eqs. (7)-(9)"},{"comment":"The lower bound r_s,t of the Uniform prior for the adjustment factor Delta_s,t is computed directly from the observed counts y_s,2016 and y_s,t that also enter the Stage I likelihood in Eq. (1). This makes the prior data-dependent and uses the post-2020 observations twice, which can artificially narrow Stage I credible intervals and overstate precision for the adjusted post-2020 rates. The simulation study does not address this issue because the data-generating process in Eq. (10) omits the adjustment-factor mechanism entirely. The authors should either specify a prior for Delta_s,t that is conditionally independent of the likelihood (for example, derived from external methodological-validation data) or provide a sensitivity analysis that quantifies the effect of this double use of the data.","section":"Section 3.2, Eqs. (4)-(5)"},{"comment":"The county-level simulation results are not as strong as the text suggests: median relative error is 0.84 and 95% interval coverage is 0.89, which is below nominal and far from 'near-nominal coverage' as stated in Section 5.3. For a county with true risk near the national average, a median relative error of 84% means the posterior median can be off by almost a factor of two. The Discussion's claim of 'strong accuracy and calibration' should be revised, and the authors should report performance stratified by county population size or true risk level, since the current aggregate metrics conceal where the model performs poorly.","section":"Table 2 and Section 5.3"}],"minor_comments":[{"comment":"The manuscript contains several typographical errors: 'mortalty' in Section 3.4, 'calues' in the Figure 6 caption, 'subsetted' in the Figure 5 caption, and 'Y ear' as a partially cut axis label in multiple figures. These should be corrected in a final polish.","section":"Throughout"},{"comment":"The symbol mu_c,t is used both for the latent county-level count and for the Poisson mean, which is confusing; for example, Eq. (8) states 'mu_c,t ~ Poisson(mu_c,t)'. Use distinct notation, such as theta_c,t for the mean, to avoid ambiguity.","section":"Eq. (8) and Table B.3"},{"comment":"The sentence introducing the performance metrics is incomplete: 'For each parameter at both spatial scales (state and county), we computed the following performance metrics across all simulation replicates:' is followed by a blank line and then a bulleted list. The sentence should be finished and the list integrated into the text.","section":"Section 3.5"},{"comment":"The posterior interval for beta_1, the PR misuse coefficient, is very wide (95% CrI [-6.07, 1.45]); the text should describe this as an imprecisely estimated effect consistent with a wide range of values, rather than emphasizing the negative posterior mean without noting the near-zero information in the data.","section":"Table 1"},{"comment":"The manuscript does not mention data or code availability. Since the data are public and the computation uses NIMBLE and R-INLA, providing a repository with the fitted models and processing scripts would strengthen the reproducibility of the applied results.","section":"Reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a genuine applied problem and the two-stage structure is reasonable, but the validation strategy is currently self-consistency rather than independent validation, and the reported uncertainty quantification omits Stage I uncertainty. Both issues are fixable within the scope of the manuscript, so I would be willing to review a revised version that either propagates full posterior draws or carefully re-scopes the claims. The estimate of 0.84 for county-level median relative error should be prominently disclosed in the abstract or conclusions, because it materially qualifies the headline claim of 'strong accuracy and calibration'."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a coherent and genuinely useful two-stage top-down model for getting county-level OUD risk from state-level NSDUH data, and the population-weighted softmax offset plus the definitional adjustment factor are sensible, practical contributions. The state-level Stage I is standard Bayesian time-series disease mapping; Stage II is a reasonable multinomial-Poisson disaggregation. The modular/cut design is a good choice to avoid feedback. The paper deserves serious refereeing, but the central claim that the framework produces well-calibrated county-level uncertainty is not yet supported.\n\nWhat the paper does well: it tackles a real surveillance gap, and the modeling choices are clearly motivated. The population-weighted offset prevents extreme scaling for tiny counties, and the Delta adjustment for the NSDUH definition change in 2020 is a thoughtful practical fix. The authors are honest that county data are unobserved and that they rely on state aggregates. The equations are clear and the two-stage logic is coherent.\n\nThe soft spots are real and, for the most part, fixable. The stress-test note is right: Stage II plugs in posterior medians of pi-tilde, y-tilde, and omega from Stage I (Section 3.3), so the county-level credible intervals in Figures 5-6 reflect only Stage II process uncertainty. The stated 'fully Bayesian' propagation is therefore overclaimed. And because the simulation pipeline uses the same plug-in medians, the 0.89 coverage in Table 2 cannot validate the reported intervals as well-calibrated. That is a load-bearing gap for a methods paper whose selling point is uncertainty quantification.\n\nSecond, the simulation in Eq. 10 draws from the same ICAR/RW1/Poisson structure used in the fitted model, so it is a self-consistency check, not independent validation. County-level MRE of 0.84 and coverage of 0.89 are weak results for a national application with no external benchmark.\n\nThird, the adjustment factor prior in Eq. 5 uses r_s,t computed from the same observed counts that enter the Stage I likelihood. That is a data-dependent prior that narrows intervals in the post-2020 years, and the paper does not acknowledge it as such.\n\nI disagree with the reader on one point: the Stage II constraint and the population-weighted offset are not flaws; they are the paper's genuine contributions. And the absence of code and data is annoying but not disqualifying for a statistical methodology paper.\n\nWho this is for: statisticians working on small-area estimation for substance use or sparse health outcomes. It should go to peer review, but the referees should insist on full posterior propagation, a simulation with a different DGP or external ground truth, and a discussion of the double-use of the adjustment ratio.","headline":"Useful two-stage downscaling framework for sparse-data disease mapping, but the uncertainty quantification claim and internal simulation need more support before I would trust the county-level maps.","tokens_in":29881,"tokens_out":1769,"would_cite":false,"duration_ms":16850,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62M30","62P10"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that county-level opioid use disorder risk can be inferred from state-level surveillance data alone through a two-stage Bayesian top-down model that disaggregates state totals to counties with full uncertainty…","keywords":["Bayesian hierarchical model","small area estimation","opioid use disorder","spatio-temporal model","top-down downscaling","data sparsity","uncertainty quantification","public health surveillance"],"falsifier":"Generate synthetic county counts under a different mechanism, such as extra-Poisson overdispersion or spatial dependence with a different range than the fitted ICAR prior, and check whether the 95% credible intervals still cover the true county risks in about 95% of cases. Alternatively, take a state that does report county-level OUD or a strong proxy such as treatment admissions, fit B-Step using only state totals, and test whether the estimated county risks track the observed county-level values.","tokens_in":28843,"feed_emoji":"🗺️","tokens_out":7920,"duration_ms":68047,"temperature":0.7,"pith_summary":"This paper tries to establish that county-level opioid use disorder (OUD) risk can be reliably inferred from observed state-level case counts even when no county-level counts are directly observed. The method, called B-Step, first estimates state-level risk and then disaggregates the estimated state totals down to counties using covariates, population weighting, and spatio-temporal random effects. Simulation experiments with 100 synthetic datasets show low bias and near-nominal coverage of 95% credible intervals at both spatial scales. If correct, the approach supplies policymakers with uncertainty-aware county-level risk estimates for all 3,143 U.S. counties from 2010 to 2025.","feed_headline":"New model maps opioid use disorder for every U.S. county","feed_subtitle":"A top-down Bayesian approach turns sparse state-level counts into county-level risk estimates with uncertainty for 2010-2025.","key_machinery":"The load-bearing mechanism is the two-stage top-down disaggregation with a population-weighted softmax allocation. In Stage I a Poisson likelihood links observed state counts to latent risk through an adjustment factor $\\Delta_{s,t} \\sim \\mathrm{Uniform}(r_{s,t},1)$ that rescales post-2020 counts; in Stage II the estimated state total is split across counties by Poisson means $\\mu_{c,t} = \\rho_{c,t} \\cdot \\tilde{y}_{s,t}$ with $\\rho_{c,t}$ a softmax over a log-linear predictor that includes a population offset, county covariates, and a BYM spatial plus RW1 temporal random effect. The multinomial-Poisson equivalence makes the county model conditionally equivalent to a multinomial allocation, so county counts reproduce the state total while remaining computationally tractable.","core_discovery":"The paper argues that county-level OUD risk can be estimated entirely from state-level surveillance counts by splitting the inference into two modular stages. Stage I fits a Bayesian Poisson time-series to observed state counts, with a random-walk temporal effect and a uniform-prior adjustment factor that corrects for the post-2020 widening of the case definition. Stage II takes the posterior state totals and disaggregates them to counties through a softmax-normalized Poisson model whose expected county counts sum exactly to the state total, adding a population-weighted offset, county covariates, and a BYM spatial plus first-order random-walk temporal effect. The output is a posterior distribution of risk for each county-year, so uncertainty is explicitly quantified rather than ignored. The claimed result is that this pipeline recovers true simulated risks with low bias and approximately nominal 95% credible-interval coverage at both state and county levels.","pith_inferences":["Because the validation is self-referential (the simulator uses the fitted model's own priors), the accuracy claim should be read as conditional on the model family being correct; a misspecified-data test would be the natural next check.","If the framework generalizes, the same two-stage design could be applied to other health outcomes that are only observed at state or national level, such as hepatitis C or stimulant use disorder, with minimal changes.","The paper's interpolation of covariates for 2024-2025 and its reliance on the random walk for years beyond the data mean the later-year estimates are extrapolations; a back-test that fits up to 2022 and compares 2023 predictions against observed state counts would quantify this.","The wide credible intervals on some state-level covariates suggest those variables carry little information; a sensitivity analysis removing them would reveal how much of the county variation comes from spatial smoothing rather than covariate effects."],"forward_implications":["County-level OUD risk maps for all 3,143 U.S. counties from 2010 to 2025 become available even though county-level case counts are never observed directly.","Each county-year estimate carries a 95% credible interval, so policymakers can distinguish reliably high-burden areas from areas where the data are simply too sparse to know.","County estimates sum exactly to the modelled state totals, preserving internal consistency when state and local numbers are used together.","The population-weighted offset prevents tiny counties from receiving extreme allocated risks solely because of their small size.","The case-definition adjustment factor lets pre- and post-2020 rates be compared on a common scale, so apparent jumps in the epidemic are not artefacts of survey changes."],"supporting_citations":[{"why":"Supplies the simulation-based validation template for disaggregation regression that the paper adapts to the two-stage setting.","marker":"Arambepola et al., 2022"},{"why":"Provides the multinomial-Poisson equivalence used to express county allocation as conditionally independent Poisson likelihoods with softmax probabilities.","marker":"Baker, 1994"},{"why":"Establishes cut-model modular inference, the rationale for keeping Stage II from feeding back into Stage I.","marker":"Plummer, 2015"},{"why":"States the small-area estimation principles of borrowing strength and hierarchical aggregation that the top-down design builds on.","marker":"Rao and Molina, 2015"},{"why":"Documents scaling instability in small-area models with extreme population ratios, motivating the population-weighted offset.","marker":"Gao and Wakefield, 2022"},{"why":"Demonstrates Bayesian downscaling of aggregated counts to finer areas, the class of methods the paper extends to sparse, biased state-level OUD data.","marker":"Python et al., 2022"},{"why":"Prior modular Bayesian implementation for national mortality estimation whose uncertainty-propagation structure is reused for state-to-county disaggregation.","marker":"Peterson et al., 2024"},{"why":"Provides a two-stage Bayesian small-area method for sparse proportions, a direct methodological comparator for the paper's approach.","marker":"Hogg et al., 2023"}],"fun_headline_variants":["Top-down Bayesian model fills in sparse county OUD data","Sparse state counts yield county-level OUD estimates","Bayesian disaggregation maps OUD risk for all counties","From state to county: OUD risk with quantified uncertainty","Data-sparse OUD mapping via top-down Bayesian framework"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the process that generated the real OUD counts matches the Poisson likelihood with ICAR/RW1 random effects used both to fit the model and to generate the simulation data, since no independent county-level ground truth is used for validation.","fun_headline_variants_meta":{"raw":{"variants":["Top-down Bayesian model fills in sparse county OUD data","Sparse state counts yield county-level OUD estimates","Bayesian disaggregation maps OUD risk for all counties","From state to county: OUD risk with quantified uncertainty","Data-sparse OUD mapping via top-down Bayesian framework"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1316,"prompt_tokens":906,"completion_tokens":410,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":522,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":522,"tokens_out":410,"duration_ms":3825,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:27:21.419428+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Generate synthetic county counts under a different mechanism, such as extra-Poisson overdispersion or spatial dependence with a different range than the fitted ICAR prior, and check whether the 95% credible intervals still cover the true county risks in about 95% of cases. Alternatively, take a state that does report county-level OUD or a strong proxy such as treatment admissions, fit B-Step using only state totals, and test whether the estimated county risks track the observed county-level values.","supporting_citations":[{"cited_title":", author Lucas, T.C.D","cited_arxiv_id":null,"evidence_quote":"Supplies the simulation-based validation template for disaggregation regression that the paper adapts to the two-stage setting."},{"cited_title":", author Molina, I","cited_arxiv_id":null,"evidence_quote":"States the small-area estimation principles of borrowing strength and hierarchical aggregation that the top-down design builds on."},{"cited_title":"A spatial variance-smoothing area level model for small area estimation of demographic rates","cited_arxiv_id":"2209.02602","evidence_quote":"Documents scaling instability in small-area models with extreme population ratios, motivating the population-weighted offset."},{"cited_title":", author Bender, A","cited_arxiv_id":null,"evidence_quote":"Demonstrates Bayesian downscaling of aggregated counts to finer areas, the class of methods the paper extends to sparse, biased state-level OUD data."},{"cited_title":", author Guranich, G","cited_arxiv_id":null,"evidence_quote":"Prior modular Bayesian implementation for national mortality estimation whose uncertainty-propagation structure is reused for state-to-county disaggregation."},{"cited_title":"A Two-Stage Bayesian Small Area Estimation Approach for Proportions","cited_arxiv_id":"2306.11302","evidence_quote":"Provides a two-stage Bayesian small-area method for sparse proportions, a direct methodological comparator for the paper's approach."}],"review_version":1}