{"id":"09a65282-74a4-4253-a84a-a9675a72753c","arxiv_id":"2607.07767","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"PSD kernel densities give a convex MCAR imputer with closed-form marginals, KL consistency rates that adapt to smoothness, and competitive energy-distance performance on small real tables.","lead":"The paper turns MCAR missing-data imputation into convex density estimation that matches observed marginals with positive semi-definite kernel models. One fitted density yields single and multiple imputations and shows competitive distributional accuracy on modest tabular datasets.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The stated rate in Theorem 5.4 is slower than the classical minimax rate the paper claims to attain, and the reverse-Pinsker step that converts L2 approximation into KL control of the masked risk is the softest link under the paper's own assumptions.","rationale":"The reader's weakest_assumption correctly flags Assumption (i) and the reverse-Pinsker step as load-bearing for the approximation-error bound. That is the right place to look. The additional, more precise concern is that even under the assumption the paper does not actually deliver the classical minimax rate it advertises; the final exponent is strictly worse, and the degradation is an artifact of the particular analysis rather than an intrinsic limitation of PSD models. Because the paper's strongest claim packages both consistency and the rate, and because the experiments remain exploratory, the verdict stays CONDITIONAL. No derivation error that would force REJECT is visible; the theory still gives a clean consistency result under strong but standard nonparametric assumptions. The concrete test above would settle whether a tighter analysis recovers the classical rate or whether the slower rate is unavoidable with the present proof architecture.","tokens_in":22005,"tokens_out":926,"duration_ms":9355,"concrete_test":"Re-derive the excess-risk bound of Theorem 5.4 while replacing the reverse-Pinsker step of Theorem 5.2 by a direct KL-to-L2 control that uses only the Hölder assumption and the mixture with the uniform (e.g., via the local inequality KL(p∥q) ≤ C∥p-q∥_2^{2} / inf q when q ≥ \nu/V, or via the Pinsker + reverse-Pinsker pair under the paper's own boundedness). If the resulting exponent remains -s/(6s+2d) rather than recovering something closer to -s/(2s+d), the paper's claim of classical minimax rates is incorrect and should be restated as a slower but still consistent rate.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (reader's strongest_claim) rests on Theorem 5.4: with the stated choices of ν, λ, ℓ, η one obtains ∫ KL(p★_S ∥ p̂_S) dμ(S) = O(N^{-s/(6s+2d)}). The paper repeatedly calls this the classical minimax rate (or an adaptive rate that beats the curse of dimensionality for regular densities). That is not accurate. The classical nonparametric KL/minimax rate for Hölder-s densities is O(N^{-s/(2s+d)}); the paper's own approximation theory (Thm 5.1 / B.1) recovers the L2 rate O(ℓ^{-s/d}) that would support the classical rate if the excess-risk analysis closed at that order. Instead the final exponent is degraded by a factor of roughly three because of the particular balancing of learning error (Thm 5.3, which carries an extra 1/ν and log factors) against approximation error (Thm 5.2). The degradation is forced by the reverse-Pinsker step used in Thm 5.2: L((1-ν)p_Q̄ + ν u) ≤ log(C/ν)·∥(1-ν)p_Q̄ + ν u - p★∥_1. That step requires either p★ bounded away from zero or zeros forming a C^{1} manifold (Assumption (i)), and it introduces the log(1/ν) factor that forces the suboptimal balancing ν ~ N^{-s/(6s+2d)}. If the reverse-Pinsker constant is large (or the lower-bound assumption fails even mildly), both the rate and the claim that the estimator is minimax-optimal for the observable marginals become unsupported. The rest of the argument (convexity, closed-form marginals, consistency under the stated assumptions) is sound; the load-bearing soft spot is precisely this rate claim and the reverse-Pinsker bridge that produces it.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper reframes MCAR imputation as estimating a density whose observed marginals match those of the data, using positive semi-definite (PSD) kernel densities. This yields a convex regularized empirical-risk problem (Eq. 4 / Problem P) with closed-form marginals, solved by a damped Newton interior-point method. The same fitted density supports both single (conditional mean) and multiple imputation. Under Hölder smoothness and mesh assumptions, the authors prove consistency of the masked KL risk, culminating in Theorem 5.4’s rate O(N^{-s/(6s+2d)}). Preliminary experiments on synthetic manifolds and eleven real tabular datasets report competitive energy distance and OT scores against mean, IterativeImputer, SoftImpute, and OT-Impute baselines.","tokens_in":22544,"tokens_out":1639,"duration_ms":26868,"significance":"If the technical claims hold, the work offers a rare combination of (i) a convex, distributionally motivated objective for imputation, (ii) closed-form marginals that make the ERM tractable, and (iii) a single coherent density for both single and multiple imputation—properties that joint-model, FCS, low-rank, and deep generative methods rarely share. Full proofs of convexity (Thm 4.1), approximation (Thm 5.1–5.2), generalization (Thm 5.3), and the combined rate (Thm 5.4) appear in Appendices B–C and rest on standard tools (Rademacher complexity, reverse Pinsker, PSD approximation theory). That package is a genuine contribution to principled missing-data methodology. The practical significance is currently limited by the preliminary experimental scale and by overstated rate claims that need correction before the theoretical contribution can be fairly assessed.","major_comments":[{"comment":"Abstract, §1 Contributions (3), and the discussion around Theorem 5.4 repeatedly describe the excess-risk rate as “classical minimax,” “O(1/√N),” or “beating the curse of dimensionality for very regular probabilities.” Theorem 5.4 actually establishes ∫ KL(p⋆_S ∥ p̂_S) dμ(S) = O(N^{-s/(6s+2d)}). The classical nonparametric KL/minimax rate for Hölder-s densities is O(N^{-s/(2s+d)}); the paper’s own L² approximation theory (Thm 5.1 / B.1) would support that order if the excess-risk analysis closed at the same scale. The final exponent is degraded by roughly a factor of three. For large s the stated rate tends to N^{-1/6}, not N^{-1/2}. These statements must be corrected or the stronger rate proved; as written they overstate the result.","section":"Theorem 5.4 / Abstract / §1"},{"comment":"The degradation is forced by the reverse-Pinsker step in the proof of Theorem 5.2: L((1−ν)p_Q̄ + νu) ≤ log(C₂/ν)·∥(1−ν)p_Q̄ + νu − p⋆∥₁. That step requires Assumption (i) (p⋆ bounded away from zero, or zeros forming a C¹ manifold) and introduces the log(1/ν) factor that forces the suboptimal balancing ν ∼ N^{-s/(6s+2d)}. If the lower-bound assumption fails even mildly, both the rate and the claim of minimax optimality for the observable marginals become unsupported. The manuscript should either strengthen the KL control (e.g., via a different inequality or a truncated risk) or clearly label the rate as suboptimal and state the assumption’s necessity more prominently in the main text.","section":"§5.1 Assumption (i), Theorem 5.2, Appendix B.3"},{"comment":"Section A states that the implementation uses α = 0 in Problem P, that anchor budgets are capped at 65–85, that hyper-parameters are tuned by a two-stage CV + alternating minimization, and that results are “preliminary.” The abstract and conclusion nevertheless claim “competitive distributional accuracy” and “strong practical promise.” With α = 0 the log term is unbounded below when Tr(QA_i) = 0, the reported variance across seeds is non-negligible, and the method is restricted to d ≲ 60, n ≲ 1.6k. Either the experimental claims should be tempered to match the exploratory status, or a more complete benchmark (including α > 0, larger d, and runtime/memory profiles consistent with §6.1) should be supplied.","section":"Appendix A / Abstract / §8"},{"comment":"Contribution (3) in §1 lists “an O(1/√N) excess-risk bound” separately from the consistency statement. Theorem 5.3 does give a generalization term of order (ℓ + λ + 1/λ) log^{3/2}(·)/√N, but after balancing with approximation error the final masked-KL rate is the slower quantity of Theorem 5.4. Presenting O(1/√N) as a headline guarantee without the balancing caveat is misleading; the contribution list should be aligned with the theorems that are actually proved.","section":"§1 Contributions / Theorem 5.3"}],"minor_comments":[{"comment":"Notation for the mask distribution switches between μ(S) and m(S); the feature map ϕ and the moment matrix H are redefined with slightly different domains in §4 and §6. A single consistent notation block would help.","section":"§3–§4"},{"comment":"Figure 2 caption claims PSD-Impute “outperforms baselines by ≥15% on distributional metrics”; the plotted bars do not uniformly support a 15% relative improvement across all datasets and missing rates. Soften or quantify per-panel.","section":"Figure 2"},{"comment":"Several references to “classical minimax” and “beating the curse” appear before any rate is stated; a forward pointer to Theorem 5.4 (with the corrected wording) would avoid early over-claim.","section":"Abstract / §1"},{"comment":"Typos: “regulirized” (§4), “dimen-sions” (§5.5), “T r” vs “Tr” inconsistency in Appendix D, and “ice_mi” in Figure 5 caption (presumably IterativeImputer MI).","section":"Throughout / Appendix D"},{"comment":"The complexity derivation in §6.1 quotes memory O(ε^{-24−10d/s}) for the SWM method; a short sanity check against the moderate-d regime used in experiments would make the asymptotic claim more credible.","section":"§6.1"}],"recommendation":"major_revision","confidential_remarks":"The theoretical core (convexity, closed-form marginals, consistency under stated assumptions) is solid and the appendices are thorough; the main obstacles to acceptance are the overstated rate language and the gap between “preliminary” experiments and the abstract’s practical claims. Heavy reliance on the authors’ prior PSD-model papers is natural but should be balanced by clearer comparison to other nonparametric density estimators used for imputation. Scope is appropriate for a stat.ML / missing-data venue once the rate claims are corrected."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful core is a clean recasting of MCAR imputation as convex marginal KL matching with PSD kernel densities. That gives closed-form marginals, a Newton interior-point solver, and one fitted density that produces both single and multiple imputations. The PSD model itself is not new (they cite Marteau-Ferey/Rudi-Ciliberto), but the masked ERM formulation, the structured solver, and the consistency theory for the observable marginals are genuine additions.\n\nWhat works: convexity is real (Thm 4.1), the risk decomposition is standard and carefully done, and Appendices B–C supply full proofs. Experiments are explicitly preliminary (α=0, d≤60, n≤1.6k, heavy hyper-parameter tuning) yet already competitive on energy distance and OT against MICE-style, SoftImpute, and OT baselines, while RMSE stays reasonable. That is enough to show the idea is not empty.\n\nSoft spots, in proportion. The load-bearing Hölder-s > d/2 and “zeros form a C¹ manifold or density bounded away from zero” assumption is strong; without it the reverse-Pinsker step that turns L² approximation into KL control fails. More importantly, Theorem 5.4’s rate O(N^{-s/(6s+2d)}) is not the classical nonparametric minimax rate O(N^{-s/(2s+d)}) the paper repeatedly claims. The degradation comes from balancing the 1/ν and log factors that reverse Pinsker injects; the approximation theory itself would support the better rate if the excess-risk analysis closed tighter. Consistency under the stated assumptions still holds; the optimality language does not. Engineering (code, data, α>0, larger d) is left for later, which the authors admit.\n\nWho it is for: people who care about distributional fidelity under MCAR and who already like kernel densities or convex density estimation. A serious referee should see it. I would bring it to reading group and would cite the method and the convex formulation; I would not cite the rate as minimax without the caveat. Send it to peer review.","headline":"Solid convex PSD-imputation method with real theory and competitive early experiments; the 'minimax' rate claim is overstated by a reverse-Pinsker balancing factor.","tokens_in":23176,"tokens_out":540,"would_cite":true,"duration_ms":6306,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G07","62D10","68T05"],"pacs":[],"model":"grok-4.5","headline":"Imputation under MCAR is convex density estimation that matches observed marginals, solved with PSD kernel models that give both single and multiple fills from one fit.","keywords":["missing data","MCAR imputation","PSD kernel densities","marginal matching","convex optimisation","distributional fidelity","multiple imputation"],"falsifier":"On a compact domain where the true density is only Lipschitz or has a non-manifold zero set, measure whether the observed excess masked KL still decays at the claimed rate as sample size grows; a clear plateau or slower rate would refute the consistency theorem.","tokens_in":22871,"feed_emoji":"📊","tokens_out":603,"duration_ms":6655,"temperature":0.7,"pith_summary":"Missing values break statistical pipelines, and most imputers either chase pointwise errors or lean on restrictive models that never recover a coherent joint. This paper reframes MCAR imputation as estimating a density whose observed marginals match those of the data. Positive semi-definite kernel densities make the resulting empirical risk convex, supply closed-form marginals, and admit a Newton interior-point solver. One fitted density then produces both single imputations (conditional means) and multiple imputations (samples). Theory shows the estimator is consistent for the observable marginals at an adaptive rate that can beat the usual curse of dimensionality when the true density is sufficiently smooth. Early experiments on synthetic manifolds and eleven real tables already match or beat strong baselines on distributional metrics while remaining competitive on RMSE.","feed_headline":"One convex density fills missing values and samples them","feed_subtitle":"PSD kernels turn MCAR imputation into marginal KL minimisation with adaptive rates","key_machinery":"PSD kernel densities: non-negative functions p_Q(x) = φ(x)^T Q φ(x) with Q ≽ 0 and a closed-form normalising constraint, whose every marginal is again an explicit trace formula. This structure turns the masked negative log-likelihood into a convex optimisation problem over the PSD cone.","core_discovery":"Under MCAR, the only recoverable object is the family of observed marginals. Minimising the masked Kullback–Leibler risk over the class of PSD kernel densities yields a convex program whose solution is statistically consistent for those marginals and, from the same density, supplies both single and multiple imputations.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Convex PSD density matches observed margins for MCAR imputation","One PSD kernel density fills and samples missing values under MCAR","Masked KL over PSD kernels yields consistent MCAR imputations","Newton-solvable PSD density recovers joints from MCAR margins","PSD Impute supplies single and multiple fills from same density"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The true density must be Hölder-smooth of order higher than half the dimension and either stay bounded away from zero or have zeros that form a smooth manifold; without that regularity the approximation rates and the conversion from L2 error into KL control fail.","fun_headline_variants_meta":{"raw":{"variants":["Convex PSD density matches observed margins for MCAR imputation","One PSD kernel density fills and samples missing values under MCAR","Masked KL over PSD kernels yields consistent MCAR imputations","Newton-solvable PSD density recovers joints from MCAR margins","PSD Impute supplies single and multiple fills from same density"]},"model":"grok-4.5","effort":"low","cost_usd":0.006234,"raw_usage":{"total_tokens":1543,"prompt_tokens":660,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":62340000,"prompt_tokens_details":{"text_tokens":660,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":799,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":660,"tokens_out":84,"duration_ms":7919,"temperature":1.0,"reasoning_tokens":799,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T18:43:12.364886+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On a compact domain where the true density is only Lipschitz or has a non-manifold zero set, measure whether the observed excess masked KL still decays at the claimed rate as sample size grows; a clear plateau or slower rate would refute the consistency theorem.","supporting_citations":[],"review_version":1}