REVIEW 1 major objections 4 minor 1 cited by
Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach
T0 review · 1 major / 4 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read This paper proves that under general non-monotone missing-at-random missingness, a Bayesian Dirichlet-mixture-of-normals posterior can estimate the complete-data density at the same minimax rate as if the data were fully observed, up to log
desk verdict First nonparametric Bayesian contraction under general non-monotone MAR; rate is plausible, but the support condition in Thm 5.1 needs to be made explicit. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two objects carry the argument. First, the relative KL divergence fKL(Pθ*∥Pθ) = E_{(X,M)~Pθ*} log[pθ*^{(M)}(X^{(M)})/pθ^{(M)}(X^{(M)})], which is nonnegative under MAR and has a unique zero exactly when the complete-data densities coincide, provided the propensity score is positive. It replaces the ordinary KL in the prior-mass part of the contraction theorem. Second, the pattern-weighted Hellinger distance eH², which sums squared differences of observed marginal densities weighted by P(M=m|x^{(m)}); MAR lets it be sandwiched as δ H² ≤ eH² ≤ H², where δ is the propensity lower bound. This sandwich yields uniform tests of power with type-I error e^{-n δ ε²/8}, recovering the classical automat
What would settle it
Take any MAR mechanism with P(M=0|X=x)=0 on a set A of positive probability, and two densities that agree on every observed marginal but differ on A. Simulate data from such a mechanism and check whether the posterior of the proposed Dirichlet-mixture sampler fails to contract to the true density in Hellinger distance — or, equivalently, whether two different complete-data densities produce the same observable likelihood. If the posterior still concentrates near the true density, the identifiability claim would be contradicted; if it drifts, the positivity condition is confirmed as load-bearin
Extended reading notes
Core claim
The central claim is that, contrary to what the fragmented non-monotone MAR literature suggested, the missingness mechanism does not change the minimax exponent for density estimation: with a positive propensity score, the posterior based only on the observed fragments concentrates in Hellinger distance at n^{-β/(2β+d)} up to log factors. This is achieved by showing that the two classical ingredients of Bayesian nonparametric contraction — prior mass in a KL-type neighbourhood and existence of uniform tests — have missing-data analogues that are no harder to satisfy than their complete-data versions. The prior-mass condition is expressed through a new relative KL divergence that measures how
Load-bearing premise
The complete-data density is recoverable from the observed fragments only because the probability of seeing a fully observed row is assumed to be bounded away from zero at every point; if that probability can be zero on some region, two distributions that differ only there are indistinguishable from the observed data.
Editorial extensions
If this is right
- The complete-data density can be estimated consistently at the minimax rate up to log factors under general non-monotone MAR, so any downstream estimate built from the density inherits this rate.
- The ignorable Bayesian likelihood — which ignores the missingness mechanism — is frequentist-valid in a nonparametric sense whenever the propensity score is bounded away from zero.
- The proposed algorithm returns posterior samples from a consistent estimate of the unmasked distribution, giving a principled alternative to ad hoc imputation for distributional questions.
- The general contraction theorem does not depend on density estimation specifically, so the same framework can be reused for other nonparametric models under MAR.
Reading between the lines
- A sharp experiment suggested by the paper: run the proposed estimator under a MAR mechanism whose propensity score is zero on a region of positive mass; the posterior should drift to an indistinguishable alternative, confirming that the positivity condition is not merely technical.
- The paper itself notes it is unclear whether a prior that satisfies the prior-mass condition for complete data automatically satisfies the adapted prior-mass condition for missing data; verifying this on a case-by-case basis is a natural next step.
- Because the estimator generates fresh samples from the learned mask-free density rather than imputing missing entries, it reframes imputation as a distributional task — whole-distribution distances, not cell-wise accuracy, become the right benchmark.
- A semiparametric Bernstein–von Mises result under MAR, which the paper flags for future work, would add confidence intervals to everything estimated from the recovered density.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a general Bayesian nonparametric posterior contraction theory for data with general non-monotone missing-at-random (MAR) mechanisms. It introduces a missing-data-adapted Kullback-Leibler divergence fKL and an adapted Hellinger distance eH, proves that the Hellinger testing condition holds under MAR plus a positivity condition on the propensity score, states a general rate theorem, and specializes it to density estimation with a Dirichlet-mixture-of-normals prior. The advertised main result is that the complete-data density can be estimated at the minimax rate up to logarithmic factors under Hölder smoothness. The paper also provides an MCMC algorithm and simulation evidence.
Significance. If the main theorem were correct as stated, this would be the first nonparametric posterior-contraction result for general non-monotone MAR, and the testing construction is a genuinely interesting contribution. The paper is mostly rigorous, with detailed proofs and an implemented algorithm; the simulation study is honest about settings that violate the assumptions. However, as explained in the major comment, Theorem 5.1 is not proved as stated because the proof relies on a support condition that is absent from the theorem statement and is not satisfied by the Dirichlet-mixture prior for a natural class of Hölder densities. The central idea is defensible, but the main density-estimation claim needs to be corrected or qualified.
major comments (1)
- [§5.2, Theorem 5.1 and its proof] The proof of Theorem 5.1 rests on Theorem C.2, Corollary C.3, and Lemma C.4, all of which require the additional condition that there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ) ⊂ supp(p_θ⋆) for every θ in the model. In particular, the construction q_θj(x,m) = p^(m)_θj(x^(m)) p_θ⋆(x^(-m)|x^(m)) P(M=m|x) in Eq. (31) is only a probability density when supp(p_θj) ⊆ supp(p_θ⋆). The Dirichlet-mixture-of-normals prior used in Theorem 5.1 assigns positive mass only to densities with full support R^d. Thus if p_θ⋆ is compactly supported — which Assumption 5.1 explicitly permits — the support condition fails for every θ in the prior support. Consequently, Theorem 4.4 cannot supply the testing condition and Lemma C.4 cannot be used to verify the prior-mass condition Assumption 4.1(i). The theorem as stated is therefore unproved. This is load-bearing: it invalidates the claimed minimax-rat
minor comments (4)
- [Appendix B, Proposition B.1] In the first inequality of the proof, 'dx^(m)' should be 'dx' (or 'dx^(0)'); the displayed lower bound integrates over the full vector x, not the observed subvector.
- [Appendix C, Corollary C.3] Near the end of the proof, the expression '-δε²/6 + ε²δ/24' is written as '= δε²/8'; it should be '-δε²/8'. The intended inequality is clear, but the sign typo is confusing.
- [Theorem 5.1 statement] The exponent t is given as 't > βd+βd/τ+d+β / 2β+d 2', which is hard to parse. It should be written as t > 2(βd+βd/τ+d+β)/(2β+d); also 'desribed' is a typo.
- [Theorem 4.3 proof] In the last display of the proof, the factor e^{-n\barε_n^2(2+C1)} appears; the preceding steps have e^{+n\barε_n^2(2+C1)}. The convergence conclusion is unaffected because ε_n ≥ \barε_n, but the sign should be corrected for clarity.
Circularity Check
No significant circularity found; the contraction proof rests on external Ghosal–van der Vaart theory and no fitted quantity is relabeled as a prediction.
full rationale
The central rate theorem (Theorem 5.1) is derived by verifying the three components of the prior-mass-and-testing framework (Assumption 4.1 and the new Hellinger testing result, Theorem 4.4), and the verification uses external, parameter-free bounds from Ghosal and van der Vaart (2017) for Dirichlet-mixture priors and inverse-Wishart priors. The MAR-specific objects fKL, eV, and eH are internally defined distances used to reformulate the prior-mass and testing conditions; they are not fitted outputs that are then fed back into the derivation. Proposition 4.2 only establishes identifiability of the true density from observed fragments under the positivity assumption, not a rate result. Self-citations to Näf et al. (2026) and Chérief-Abdellatif and Näf (2025) are motivational or contextual and do not carry the proof of Theorem 5.1. The reviewer-flagged support-condition gap is real and non-circular: Theorem 5.1 omits the supp(pθ)⊂supp(pθ⋆) continuity condition that Theorem 4.4, Corollary C.3, Lemma C.4, and the construction qθj(x,m)=p^{(m)}_{θj}(x^{(m)})p_{θ⋆}(x^{(-m)}|x^{(m)})P(M=m|x) require, so the stated Hölder-class result is not fully proved for densities with zeros or compact support. This is a correctness/assumption gap, not a circular reduction: the conclusion is not equivalent to an input by construction, and no fitted parameter is renamed as a prediction.
Assumptions & free parameters
assumptions (7)
- domain assumption Rubin MAR: P(M=m|X=x)=P(M=m|X^{(m)}=x^{(m)}) for all m and almost all x (Assumption 2.1).
- domain assumption Empty pattern has zero probability: P(M=1)=0 (Assumption 2.2).
- domain assumption Positivity: P(M=0|X=x)>δ>0 for almost all x (Assumption 2.3).
- domain assumption Well-specified model: P_θ* belongs to the model {P_θ}, and the missingness mechanism is fixed but unknown.
- domain assumption Support and regularity: there exists θ⋆ such that p_θ⋆ is continuous and supp(p_θ)⊂supp(p_θ⋆) for all θ (Theorem 4.4).
- domain assumption Hölder class Assumption 5.1 with mixed derivatives, integrability of L/p, and exponential tail.
- standard math Classical prior-mass and testing framework of Ghosal and van der Vaart (2017): sieve entropy bounds, Dirichlet-process/inverse-Wishart prior-mass bounds, and i.i.d. testing lemmas.
Cite this review
Pith. "Pith review of Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach." pith.science (2026). https://pith.science/paper/IMNVX7K6
@misc{pith2026260323449,
author = {Pith},
title = {Pith review of: Asymptotics of Nonparametric Estimation under general non-monotone MAR missingness: A Bayesian Approach},
year = {2026},
howpublished = {\url{https://pith.science/paper/IMNVX7K6}},
note = {Machine review of arXiv:2603.23449}
}
read the original abstract
Missing values are ubiquitous in statistical practice, with potentially detrimental consequences for any statistical analysis. As such, a wealth of methods and theoretical results have been developed in the last decades. However, many questions remain open, in particular in the case of general non-monotone missing at random (MAR), where nonparametric results are still lacking. In this paper, we extend nonparametric Bayesian theory to this MAR setting. We introduce a general theorem of posterior contraction under MAR and an additional positivity condition and apply this result to density estimation as well as regression problems. In particular, we show that, despite the missing values, the complete-data density can be estimated with the minimax posterior contraction rate up to logarithmic factors. To the best of our knowledge, this is the first nonparametric result showing that the complete-data distribution can be consistently estimated under Rubin's MAR definition. As a consequence, we obtain an algorithm that takes incomplete data and returns a sample from a consistent estimate of the complete-data distribution.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Asymptotics of Nonparametric Estimation under General Non-monotone MAR Missingness: A Nonparametric Maximum Likelihood Approach
A sieve maximum likelihood estimator attains near-minimax Hellinger rates for density estimation under general non-monotone missing at random, with missingness affecting only the constant.
Reference graph
Works this paper leans on
-
[1]
Castillo, I. (2024). Bayesian nonparametric statistics, st-flour lecture notes.arXiv preprint arXiv:2402.16422
arXiv 2024
-
[2]
and Yu, C
Chen, S. and Yu, C. L. (2016). Parameter estimation through semiparametric quantile regression imputation. Electronic Journal of Statistics, 10:3621–3647
2016
-
[3]
Chen, Y.-C. and Sadinle, M. (2019). Nonparametric pattern-mixture models for inference with missing data. arXiv preprint arXiv:1904.11085. Ch´ erief-Abdellatif, B.-E. and N¨ af, J. (2025). Parametric MMD estimation with missing values: Robustness to missingness and data model misspecification.arXiv preprint arXiv:2503.00448. Daniel Malinsky, I. S. and Tch...
arXiv 2019
-
[4]
Frahm, G., Nordhausen, K., and Oja, H. (2020). M-estimation with incomplete and dependent multivariate data.Journal of Multivariate Analysis, 176:104569
2020
-
[5]
and van der Vaart, A
Ghosal, S. and van der Vaart, A. (2017).Fundamentals of Nonparametric Bayesian Inference. Cambridge Series in Statistical and Probabilistic Mathematics. Cambridge University Press
2017
-
[6]
Grzesiak, K., Muller, C., Josse, J., and N¨ af, J. (2025). Do we need dozens of methods for real world missing value imputation?arXiv preprint arXiv:2511.04833
arXiv 2025
-
[7]
and Yang, S
Guan, Q. and Yang, S. (2024). A unified inference framework for multiple imputation using martingales. Statistica Sinica, 34:1649–1673
2024
-
[8]
Little, R. J. A. and Rubin, D. B. (2019).Statistical Analysis with Missing Data., 3rd Edition. Wiley
2019
Show all 28 references
-
[9]
and Fan, Y
Liu, Y. and Fan, Y. (2023). Biased-sample empirical likelihood weighting for missing data problems: an alternative to inverse probability weighting.Journal of the Royal Statistical Society Series B: Statistical Methodology, 85(1):67–83
2023
-
[10]
A., Berrett, T
Ma, T., Verchand, K. A., Berrett, T. B., Wang, T., and Samworth, R. J. (2024). Estimation beyond missing (completely) at random.Preprint arXiv:2410.10704
2024 arXiv
-
[11]
and Frellsen, J
Mattei, P.-A. and Frellsen, J. (2019). MIW AE: Deep generative modelling and imputation of incomplete data sets. InProceedings of the 36th International Conference on Machine Learning, volume 97, pages 4413–4423
2019
-
[12]
and Rubin, D
Mealli, F. and Rubin, D. B. (2015). Clarifying missing at random and related definitions, and implications when coupled with exchangeability.Biometrika, 102(4):995–1000
2015
-
[13]
Murphy, K. P. (2012).Machine Learning: A Probabilistic Perspective. The MIT Press
2012
-
[14]
Neal, R. M. (2000). Markov chain sampling methods for dirichlet process mixture models.Journal of Computational and Graphical Statistics, 9(2):249–265. N¨ af, J. (2026). A practical guide to modern imputation.arXiv preprint arXiv:2601.14796. N¨ af, J., Scornet, E., and Josse, ...
2000
-
[15]
Pinheiro, J. C. and Bates, D. M. (1996). Unconstrained parametrizations for variance-covariance matrices. Statistics and Computing, 6(3):289–296
1996
-
[16]
Qin, J., Zhang, B., and Leung, D. H. Y. (2009). Empirical likelihood in missing data problems.Journal of the American Statistical Association, 104(488):1492–1503. 42
2009
-
[17]
Rubin, D. B. (1976). Inference and missing data.Biometrika, 63(3):581–592
1976
-
[18]
Missing at Random
Seaman, S., Galati, J., Jackson, D., and Carlin, J. (2013). What is meant by “Missing at Random”? Statistical Science, 28(2):257–268
2013
-
[19]
Seaman, S. R. and Vansteelandt, S. (2018). Introduction to double robust methods for incomplete data.Stat Sci, 33(2):184–197
2018
-
[20]
Smith, R. L. (1992). Some interlacing properties of the schur complement of a hermitian matrix.Linear Algebra and its Applications, 177:137–144
1992
-
[21]
Stekhoven, D. J. and B¨ uhlmann, P. (2011). MissForest—non-parametric missing value imputation for mixed- type data.Bioinformatics, 28(1):112–118
2011
-
[22]
and Tchetgen, E
Sun, B. and Tchetgen, E. J. T. (2018). On inverse probability weighting for nonmonotone missing at random data.Journal of the American Statistical Association, 113(521):369–379. Sz´ ekely, G. J. (2003). E-statistics: the energy of statistical samples. Technical Report 05, Bowl...
2018
-
[23]
and Kano, Y
Takai, K. and Kano, Y. (2013). Asymptotic inference with incomplete data.Communications in Statistics - Theory and Methods, 42(17):3174–3190. Van Buuren, S. (2018).Flexible Imputation of Missing Data. Second Edition. Chapman & Hall/CRC Press. Van Buuren, S. and Groothuis-Oudsh...
2013
-
[24]
and Robins, J
Wang, N. and Robins, J. M. (1998). Large-sample theory for parametric multiple imputation procedures. Biometrika, 85(4):935–948
1998
-
[25]
and Rao, J
Wang, Q. and Rao, J. N. K. (2002). Empirical likelihood-based inference under imputation for missing response data.The Annals of Statistics, 30(3):896–924
2002
-
[26]
Yu, J., Ying, Q., Wang, L., Jiang, Z., and Liu, S. (2025). Missing data imputation by reducing mutual information with rectified flows.arXiv preprint arXiv:2505.11749
2025
-
[27]
and Dong, X
Yuan, X. and Dong, X. (2019). Weighted empirical likelihood for quantile regression with non ignorable missing covariates.Communications in Statistics - Theory and Methods, 48(12):3068–3084
2019
-
[28]
and Cand` es, E
Zhao, S. and Cand` es, E. (2025). Imputation-powered inference.arXiv preprint arXiv:2509.13778. 43
2025
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.