{"id":"784b5c97-fafe-496d-b6ac-8da52e2777fd","arxiv_id":"2603.23449","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Bayesian posterior contraction and minimax-rate density estimation remain valid under general non-monotone MAR when fully observed rows occur with probability bounded away from zero.","lead":"Under general non-monotone missing-at-random (MAR) mechanisms, the paper proves that Bayesian nonparametric density estimation still recovers the complete-data distribution, at the minimax rate up to log factors. This gives statisticians a principled, theory-backed way to estimate and sample from the unobserved distribution when missingness depends only on observed values.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 5.1 omits the support condition on which its own testing and prior-mass lemmas depend; the stated minimax rate is unproved for Hölder densities with zeros or compact support.","rationale":"The paper's central chain is plausible and largely coherent: fKL is a natural observed-data divergence, positivity of the fully observed pattern gives identifiability, the adapted Hellinger test in Theorem 4.4 is constructed explicitly, and the density-estimation rate follows Ghosal–van der Vaart machinery if the prior-mass condition can be verified. The reader correctly identified positivity (Assumption 2.3) as a substantive limitation; I agree that it is necessary and explicitly assumed. However, I found a distinct gap that the reader listed only as a minor issue: the omitted support condition in Theorem 5.1. This is load-bearing because it is needed in both the testing theorem (Theorem 4.4) and the prior-mass lemma (Lemma C.4). The Dirichlet-mixture prior has full-support densities, so if p_θ* is not strictly positive, the required assumption fails for essentially every prior-supported θ. The theorem statement is therefore broader than what the proof establishes. This does not invalidate the main idea — adding p_θ*>0 or supp(p_θ)⊂supp(p_θ*) would make the proof go through — but it should be stated and acknowledged. Hence the reader's ACCEPT should be conditioned on adding this assumption or supplying a repaired proof.","tokens_in":33226,"tokens_out":21434,"duration_ms":212164,"concrete_test":"Take d=2 and let p_θ* be a smooth compactly supported density satisfying Assumption 5.1, e.g. a narrow Gaussian convolution of the uniform density on [-1,1]^2. Let the missingness mechanism satisfy MAR with P(M=0|X)=1/2. Independently verify the proof's key step: for θ = a Gaussian mixture with full support R^2, check whether q_θ(x,m) = p^{(m)}_θ(x^{(m)}) p_θ*(x^{(-m)}|x^{(m)}) P(M=m|x) defined in Eq. (31) is a probability density (integrate over x and sum over m). If it is not, Lemma C.4 and Theorem C.2 cannot be applied, and Theorem 5.1 as stated lacks the support assumption needed by its own lemmas. This analytical check settles whether Theorem 5.1 requires an added condition such as p_θ*>0 on R^d or a new proof avoiding the support assumption.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Theorem 5.1 claims the rate (5) for every p_θ* satisfying Assumptions 2.1–2.3 and the Hölder class Assumption 5.1. But the proof of Theorem 5.1 invokes Lemma C.4, Theorem C.2, Corollary C.3, and Theorem 4.4, all of which require an extra condition: there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ) ⊂ supp(p_θ⋆) for all θ ∈ Θ. The Dirichlet-mixture prior used in Theorem 5.1 puts positive mass on densities with full support R^d. If p_θ* has any zero set — e.g. a smooth compactly supported density satisfying Assumption 5.1 — then supp(p_θ) ⊄ supp(p_θ*) for every such θ, so Theorem 4.4 cannot deliver the testing condition and Lemma C.4 cannot be used to verify Assumption 4.1(i). Concretely, the construction in Eq. (31) of q_θj(x,m) = p^{(m)}_{θj}(x^{(m)}) p_θ⋆(x^{(-m)}|x^{(m)}) P(M=m|x) is not a probability density when p_θj has mass where p_θ⋆'s marginal is zero; p_θ⋆(x^{(-m)}|x^{(m)}) is undefined or the integral exceeds 1. Thus the fKL/eV prior-mass bound in Eq. (18) does not follow. The statement of Theorem 5.1 does not include or acknowledge this support condition, and the abstract's claim of estimating 'the uncontaminated density' under general MAR is broader than the theorem as proved. This is a real proof gap, not merely a typo: it disables both pillars (testing and prior mass) of the contraction argument for a natural class of densities.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a general Bayesian nonparametric posterior contraction theory for data with general non-monotone missing-at-random (MAR) mechanisms. It introduces a missing-data-adapted Kullback-Leibler divergence fKL and an adapted Hellinger distance eH, proves that the Hellinger testing condition holds under MAR plus a positivity condition on the propensity score, states a general rate theorem, and specializes it to density estimation with a Dirichlet-mixture-of-normals prior. The advertised main result is that the complete-data density can be estimated at the minimax rate up to logarithmic factors under Hölder smoothness. The paper also provides an MCMC algorithm and simulation evidence.","tokens_in":33641,"tokens_out":8206,"duration_ms":75396,"significance":"If the main theorem were correct as stated, this would be the first nonparametric posterior-contraction result for general non-monotone MAR, and the testing construction is a genuinely interesting contribution. The paper is mostly rigorous, with detailed proofs and an implemented algorithm; the simulation study is honest about settings that violate the assumptions. However, as explained in the major comment, Theorem 5.1 is not proved as stated because the proof relies on a support condition that is absent from the theorem statement and is not satisfied by the Dirichlet-mixture prior for a natural class of Hölder densities. The central idea is defensible, but the main density-estimation claim needs to be corrected or qualified.","major_comments":[{"comment":"The proof of Theorem 5.1 rests on Theorem C.2, Corollary C.3, and Lemma C.4, all of which require the additional condition that there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ) ⊂ supp(p_θ⋆) for every θ in the model. In particular, the construction q_θj(x,m) = p^(m)_θj(x^(m)) p_θ⋆(x^(-m)|x^(m)) P(M=m|x) in Eq. (31) is only a probability density when supp(p_θj) ⊆ supp(p_θ⋆). The Dirichlet-mixture-of-normals prior used in Theorem 5.1 assigns positive mass only to densities with full support R^d. Thus if p_θ⋆ is compactly supported — which Assumption 5.1 explicitly permits — the support condition fails for every θ in the prior support. Consequently, Theorem 4.4 cannot supply the testing condition and Lemma C.4 cannot be used to verify the prior-mass condition Assumption 4.1(i). The theorem as stated is therefore unproved. This is load-bearing: it invalidates the claimed minimax-rat","section":"§5.2, Theorem 5.1 and its proof"}],"minor_comments":[{"comment":"In the first inequality of the proof, 'dx^(m)' should be 'dx' (or 'dx^(0)'); the displayed lower bound integrates over the full vector x, not the observed subvector.","section":"Appendix B, Proposition B.1"},{"comment":"Near the end of the proof, the expression '-δε²/6 + ε²δ/24' is written as '= δε²/8'; it should be '-δε²/8'. The intended inequality is clear, but the sign typo is confusing.","section":"Appendix C, Corollary C.3"},{"comment":"The exponent t is given as 't > βd+βd/τ+d+β / 2β+d 2', which is hard to parse. It should be written as t > 2(βd+βd/τ+d+β)/(2β+d); also 'desribed' is a typo.","section":"Theorem 5.1 statement"},{"comment":"In the last display of the proof, the factor e^{-n\\barε_n^2(2+C1)} appears; the preceding steps have e^{+n\\barε_n^2(2+C1)}. The convergence conclusion is unaffected because ε_n ≥ \\barε_n, but the sign should be corrected for clarity.","section":"Theorem 4.3 proof"}],"recommendation":"major_revision","confidential_remarks":"The support-condition gap is the main obstacle. It is a genuine proof gap, but it is local to Theorem 5.1 and can be fixed by restating the theorem with the required support assumption or by changing the prior. The general framework and testing theorem appear sound. I would encourage the editors to request a revision that addresses this issue explicitly rather than rejecting the paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — here's my read on arXiv:2603.23449.\n\nThe paper does something genuinely new: it gives the first nonparametric posterior-contraction theorem for the complete-data distribution under general non-monotone MAR, with a rate that matches the complete-data minimax rate up to logs (modulo a 1/√δ constant). The machinery is a missingness-adapted KL divergence (fKL), a Hellinger-type distance eH with the sandwich bounds δH² ≤ eH² ≤ H², and a constructive testing theorem under MAR plus positivity. The high-level chain is coherent, and the density application in Theorem 5.1 is plausible.\n\nWhere I'd push back: the stress-test note about the missing support condition in Theorem 5.1 is the main thing to check. The proof uses Theorem 4.4 and Lemma C.4, both of which require supp(pθ) ⊂ supp(pθ⋆) for some continuous pθ⋆. The statement of 5.1 doesn't include that. But I think the concern is less damning than it first appears: Assumption 5.1 includes the moment condition P[(L/p)^2 + (|D^k p|/p)^(2β/k)] < ∞, which effectively requires pθ* > 0 a.e. If so, supp(pθ*) = R^d, and the support condition is automatic for any prior density with full support. The authors should still say this out loud, because the paper currently borrows the condition without comment, and a reader shouldn't have to infer it. If there is a way to satisfy Assumption 5.1 with a zero set of positive measure, then the gap is real; I don't think there is, but a referee should confirm formally.\n\nOther soft spots: the proof of Theorem 5.1 is terse — it imports a lot from Ghosal-van der Vaart, and there are garbled exponents in the statement (the formula for t is mangled). None of this looks fatal; it needs cleanup and a careful line-by-line check of the constants. The positivity assumption (δ>0) is strong but explicitly acknowledged, and they show it's necessary for identification, so that's fair. The novelty claim seems justified against Chen-Sadinle and Takai-Kano.\n\nBottom line: this paper deserves a serious referee. If the support condition is resolved and the proof details are tightened, it's a solid contribution. I'd bring it to the reading group and would cite it if I worked on missing data.\n\nRecommendation: send it to peer review, with a request that the authors clarify the support condition in Theorem 5.1 and fix the typos.","headline":"First nonparametric Bayesian contraction under general non-monotone MAR; rate is plausible, but the support condition in Thm 5.1 needs to be made explicit.","tokens_in":34191,"tokens_out":7309,"would_cite":true,"duration_ms":69581,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62G20","62C10","62G07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proves that under general non-monotone missing-at-random missingness, a Bayesian Dirichlet-mixture-of-normals posterior can estimate the complete-data density at the same minimax rate as if the data were fully observed, up to log","keywords":["missing at random","non-monotone missingness","posterior contraction","Bayesian nonparametrics","density estimation","Dirichlet mixture of normals","propensity score","minimax rate"],"falsifier":"Take any MAR mechanism with P(M=0|X=x)=0 on a set A of positive probability, and two densities that agree on every observed marginal but differ on A. Simulate data from such a mechanism and check whether the posterior of the proposed Dirichlet-mixture sampler fails to contract to the true density in Hellinger distance — or, equivalently, whether two different complete-data densities produce the same observable likelihood. If the posterior still concentrates near the true density, the identifiability claim would be contradicted; if it drifts, the positivity condition is confirmed as load-bearin","tokens_in":33081,"feed_emoji":"📊","tokens_out":8284,"duration_ms":75964,"temperature":0.7,"pith_summary":"The paper targets a long-standing gap: under general, non-monotone missing-at-random (MAR) missingness, no nonparametric consistency result for the complete-data distribution existed, and common M-estimators can be inconsistent. The authors establish that Bayesian posterior inference built on the ignorable likelihood remains valid: the posterior contracts around the true complete-data density at the same rate as if the data were fully observed, up to logarithmic factors. The key adaptation is a missingness-aware Kullback-Leibler divergence that replaces the usual KL divergence in the prior-mass condition, and a constructive test showing that Hellinger testing still works automatically under MAR when the probability of a fully observed row is bounded away from zero. In the density-estimation application with a Dirichlet mixture of normals, the contraction rate matches the minimax rate up to logs. A practical consequence is an algorithm that takes incomplete rows and returns samples from a consistent estimate of the unmasked distribution.","feed_headline":"Missing-at-random data no longer blocks density estimation","feed_subtitle":"Bayesian posteriors converge at the fastest possible rate when the chance of a fully observed row stays positive.","key_machinery":"Two objects carry the argument. First, the relative KL divergence fKL(Pθ*∥Pθ) = E_{(X,M)~Pθ*} log[pθ*^{(M)}(X^{(M)})/pθ^{(M)}(X^{(M)})], which is nonnegative under MAR and has a unique zero exactly when the complete-data densities coincide, provided the propensity score is positive. It replaces the ordinary KL in the prior-mass part of the contraction theorem. Second, the pattern-weighted Hellinger distance eH², which sums squared differences of observed marginal densities weighted by P(M=m|x^{(m)}); MAR lets it be sandwiched as δ H² ≤ eH² ≤ H², where δ is the propensity lower bound. This sandwich yields uniform tests of power with type-I error e^{-n δ ε²/8}, recovering the classical automat","core_discovery":"The central claim is that, contrary to what the fragmented non-monotone MAR literature suggested, the missingness mechanism does not change the minimax exponent for density estimation: with a positive propensity score, the posterior based only on the observed fragments concentrates in Hellinger distance at n^{-β/(2β+d)} up to log factors. This is achieved by showing that the two classical ingredients of Bayesian nonparametric contraction — prior mass in a KL-type neighbourhood and existence of uniform tests — have missing-data analogues that are no harder to satisfy than their complete-data versions. The prior-mass condition is expressed through a new relative KL divergence that measures how","pith_inferences":["A sharp experiment suggested by the paper: run the proposed estimator under a MAR mechanism whose propensity score is zero on a region of positive mass; the posterior should drift to an indistinguishable alternative, confirming that the positivity condition is not merely technical.","The paper itself notes it is unclear whether a prior that satisfies the prior-mass condition for complete data automatically satisfies the adapted prior-mass condition for missing data; verifying this on a case-by-case basis is a natural next step.","Because the estimator generates fresh samples from the learned mask-free density rather than imputing missing entries, it reframes imputation as a distributional task — whole-distribution distances, not cell-wise accuracy, become the right benchmark.","A semiparametric Bernstein–von Mises result under MAR, which the paper flags for future work, would add confidence intervals to everything estimated from the recovered density."],"forward_implications":["The complete-data density can be estimated consistently at the minimax rate up to log factors under general non-monotone MAR, so any downstream estimate built from the density inherits this rate.","The ignorable Bayesian likelihood — which ignores the missingness mechanism — is frequentist-valid in a nonparametric sense whenever the propensity score is bounded away from zero.","The proposed algorithm returns posterior samples from a consistent estimate of the unmasked distribution, giving a principled alternative to ad hoc imputation for distributional questions.","The general contraction theorem does not depend on density estimation specifically, so the same framework can be reused for other nonparametric models under MAR."],"fun_headline_variants":["Missing data no match for Bayesian density estimates","Bayes beats missing data at minimax speed","Density estimation stays optimal despite missing-at-random","First proof: MAR data won't slow density estimation","Missing-at-random loses: minimax rates intact"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The complete-data density is recoverable from the observed fragments only because the probability of seeing a fully observed row is assumed to be bounded away from zero at every point; if that probability can be zero on some region, two distributions that differ only there are indistinguishable from the observed data.","fun_headline_variants_meta":{"raw":{"variants":["Missing data no match for Bayesian density estimates","Bayes beats missing data at minimax speed","Density estimation stays optimal despite missing-at-random","First proof: MAR data won't slow density estimation","Missing-at-random loses: minimax rates intact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000581,"raw_usage":{"total_tokens":2545,"prompt_tokens":686,"completion_tokens":1859,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":1788}},"tokens_in":430,"tokens_out":1859,"duration_ms":11851,"temperature":1.0,"reasoning_tokens":1788,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T17:33:33.277346+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take any MAR mechanism with P(M=0|X=x)=0 on a set A of positive probability, and two densities that agree on every observed marginal but differ on A. Simulate data from such a mechanism and check whether the posterior of the proposed Dirichlet-mixture sampler fails to contract to the true density in Hellinger distance — or, equivalently, whether two different complete-data densities produce the same observable likelihood. If the posterior still concentrates near the true density, the identifiability claim would be contradicted; if it drifts, the positivity condition is confirmed as load-bearin","supporting_citations":[],"review_version":1}