REVIEW 2 major objections 4 minor 300 references
KL error of diffusion sampling is governed by one data-geometry curve, the denoising growth complexity.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 00:18 UTC pith:GEY2JBYJ
load-bearing objection Genuinely new DGC-based KL bound with a clean proof and useful multi-block consequences; the 'fully data-certified' claims overreach because the certified estimators require exact denoisers and are not instantiated for learned scores. the 2 major comments →
Denoising growth complexity: Data geometry and certified schedules for diffusion sampling
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central result is that for the SI-Euler scheme—the Euler–Maruyama discretization of the stochastic-innovations SDE associated with the heat path—the KL divergence from the true smoothed law to the sampler output is bounded by a sum of relative stepsizes times DGC increments over the time grid, plus an initialization term. The proof proceeds through a one-step bound: each local KL deficit is at most the relative stepsize times the DGC increment over that interval, and this bound is sharp up to a factor of two as the interval shrinks. The same DGC function is then shown to be estimable from data via denoising increments, with a constant-factor sandwich that yields fully certified s
What carries the argument
Denoising growth complexity, H(a,b) = (1/2)∫_a^b h'(t)/t dt, where h(t) is the minimum mean-squared error of denoising the latent variable from the heat-path observation at time t. It is additive over time intervals and has an equivalent information-theoretic form involving mutual information; in precision coordinates it is controlled by the non-increasing MSE that drives the innovations SDE. Its role is to give a local, interval-wise control of the Euler discretization error and to give a data-estimable target for stepsize selection.
Load-bearing premise
The certificate step requires an i.i.d. sample of the latent variable that is independent of any data used to fit the scores and has a known p-th moment bound; reuse the same sample for both tasks and the Monte Carlo estimate of H is biased and the certified KL guarantee no longer follows.
What would settle it
For a Gaussian prior Z ∼ N(0,1), compute exactly the one-step KL deficit between the innovations transition and its Euler approximation and compare it with the relative stepsize times the DGC increment. If the ratio ever exceeds 1, or fails to approach 1/2 as the interval shrinks, the local bound behind the main theorem is false.
If this is right
- A single geometric schedule can sample to ε accuracy in KL using O(H(δ,T) log(T/δ)/ε plus initialization cost) score evaluations, with linear dimension scaling and no logarithmic overhead in the worst case.
- K-block schedules with optimal geometric multipliers achieve D_KL ≤ 4 C_DGC(P)/N plus initialization, and the optimal K-block partition can be computed by dynamic programming.
- With a hold-out sample of the latent variable satisfying a known p-th moment bound, DGC increments can be estimated so that the final KL guarantee holds with probability at least 1−η, up to a factor-of-two loss plus a confidence correction.
- Analytic upper bounds on H via covariance, rate-distortion, metric entropy, and Poincaré constant recover and sharpen existing diffusion-sampling guarantees, including linear dimension scaling, dimension-free bounds for bounded models, log-K for Gaussian mixtures, and log dependence on the Poincaré constant.
- In log heat-time, single-block cost is governed by ∫q while the fine-partition limit is governed by (∫√q)², so the spread ratio quantifies exactly when adaptive schedules help; for a two-point Gaussian mixture the separation can be from Θ(log(R²/δ)) to Θ(1).
Where Pith is reading between the lines
- If DGC estimation is robust enough, certified schedules could be built directly from raw, unlabelled data by running forward heat paths, without retraining or knowing the denoiser analytically.
- The factor-of-two local sharpness suggests the DGC bound is close to tight for Euler-type samplers, so further speed-ups would need higher-order or randomized-midpoint discretizations of the innovations SDE rather than better Euler stepsize choices.
- The perturbed sandwich for learned denoisers gives a practical training target: reduce the weighted denoiser error below the relevant DGC increment, otherwise certification is impossible; this could be used as a stop-rule during score matching.
- The √q-versus-q comparison predicts that multimodal or hierarchical distributions with well-separated resolution times are exactly where K-block schedules pay off most, a testable prediction on synthetic mixture benchmarks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the denoising growth complexity (DGC), H(a,b) = (1/2)∫_a^b h'(t)/t dt, where h is the MSE of the optimal denoiser along the Gaussian heat flow. The central result, Theorem 1, bounds the KL error of a stochastic-innovations Euler scheme on an arbitrary grid by ∑ (t_j/t_{j+1}-1) H(t_{j+1},t_j) plus the initialization error. The proof is via a one-step defect bound (Lemma 5) obtained from an exact entropy/cross-entropy representation and the conditional I-MMSE identity. From this, the paper derives single-block geometric schedules (Corollary 1), a tail-robust data-dependent estimator and certified single-block procedure (Proposition 1 and Section 3.2.2), multi-block schedules (Theorem 2), certified multi-block schedules (Corollary 2), optimal block-boundary choice by dynamic programming, and a fine-partition limit governed by the log-time DGC density. It also gives information-theoretic upper bounds via covariance, rate-distortion, metric entropy, and the Poincaré constant, recovering and sharpening several existing diffusion-sampling guarantees. The main mathematical inequality is elegant, local, additive, and appears correct.
Significance. If the results hold, Theorem 1 is a significant unification: it provides an explicit, additive, parameter-free KL bound for diffusion sampling, with a short elementary proof, and it recovers or sharpens a range of prior dimension, intrinsic-dimension, mixture, and Poincaré-constant guarantees. The multi-block versus single-block comparison through the DGC spread is conceptually clean and yields concrete logarithmic-to-constant separations. The paper also has the praiseworthy feature of giving explicit constants and identifiable statistical estimators, with no fitted parameters. However, the advertised 'fully data-certified' contribution has a substantial implementability gap in the learned-score setting: the certified estimator requires oracle access to exact conditional mean denoisers. This limits the practical scope of the Q2 contribution until the gap is addressed.
major comments (2)
- [Section 3.2.1 / Proposition 1] Proposition 1 is advertised as a fully data-certified guarantee, but the statistic Q in Eq. (18) is built from the exact denoisers μ_{vℓ}(X_{vℓ}). For an unknown target P_Z these conditional expectations are not available from i.i.d. samples alone. The hold-out discussion in Section 3.3.1 addresses independence between the Monte Carlo sample and the score-training data, but it does not address the more basic fact that exact denoisers are unknown. Consequently, the single-block certified procedure in Section 3.2.2 is an oracle certification; it is not implementable in the primary setting of interest, namely sampling from an unknown distribution with estimated score functions.
- [Section 3.2.3 / Corollary 2] For estimated denoisers, the perturbed sandwich (25) states a population-level bound involving E(a,b), the sum of squared denoiser errors. No finite-sample high-probability upper bound on E(a,b) is derived, and the tail-robust Monte Carlo machinery of Proposition 1 is not re-run for the learned statistic based on eD. Thus the certified multi-block multipliers in Corollary 2, which use the Proposition 1 estimates bH_k, are not certified when scores are learned. The same gap propagates to the data-dependent dynamic program in Section 4.1.3. This is load-bearing for the paper's Q2 claim; a finite-sample control on the denoiser-error term, or a modified estimator that bypasses exact denoiser evaluation, is needed before the certified schedule is implementable with learned scores.
minor comments (4)
- [Eq. (17), Section 2.2.4] The same symbol H is used for the DGC function and for the dyadic approximation H(a,b)=1/2 Σ D(vℓ,vℓ+1)/vℓ, making the sandwich '1/2 H(a,b) ≤ H(a,b) ≤ H(a,b)' confusing. A distinct symbol such as H̄ or H̃ would improve readability.
- [Eq. (32), Proposition 4] The displayed formula has an unbalanced parenthesis/brace in the log term: 'dlog(1+8dκ/ε) + 1' should likely be d log(1+8dκ/ε) + 1 inside the curly braces. Please correct the typesetting.
- [References] Several references have formatting problems: [RBD+22] has garbled author initials, and [L WCC23] contains a stray space. These should be cleaned up.
- [Figure 3] The legend text 'g = 9.275 k SkHk = 12.3' is garbled; presumably it should read ∑ √(S_k H_k) or the equivalent. Please fix the figure caption and labels.
Circularity Check
No significant circularity: Theorem 1 is a proved inequality and the data-certified schemes are confidence-interval constructions; the learned-score limitation is a completeness gap, not a circular reduction.
full rationale
The paper's derivation chain is self-contained. Theorem 1 follows from Lemma 5, an analytic one-step bound Γ_Eul(s,s+h) ≤ (h/s) G(s,s+h) proved in Section 5.1 from the exact identity Γ_Eul(s,s+h)= h/2 g(s) − (1/2)∫_s^{s+h} g(r)dr and monotonicity of the precision-space MSE; no free parameter is fitted to match observed KL error. Corollary 1, Theorem 2, and Corollary 2 are algebraic consequences of Theorem 1 via additivity of H and Cauchy–Schwarz, not restatements of their inputs. The data-dependent certification (Proposition 1, Section 3.2.2, Corollary 2) estimates the fixed population quantity H(a,b) with denoising-increment Monte Carlo and uses explicit tail-robust upper confidence corrections; the final KL guarantee is conditional on those intervals, so the 'prediction' is not forced by construction. The only self-citation, [Wai26], appears as a comparison for Proposition 4 and is not load-bearing. The limitations flagged in the text—Section 3.2.3's perturbed sandwich (25) with no finite-sample upper bound on E(a,b), and Section 3.3.1's note that 'if we also incorporate score-based errors, these samples must not be used to fit the score functions'—are real implementability gaps in the advertised data-certified claims, but they do not make any claimed result equal to its input by definition. Hence no circularity.
Axiom & Free-Parameter Ledger
axioms (7)
- domain assumption The MSE function h is differentiable and its derivative h' is integrable (Section 2.1).
- domain assumption Z has finite second moments (Section 2.1, Theorem 1).
- standard math The stochastic innovations SDE representation dY_λ = m_λ(Y_λ)dλ + dB_λ (Eq. 52a) from nonlinear filtering theory.
- domain assumption For tail-robust estimation, a known p-th moment bound (19) on the denoising function μ_t(X_t) with constant M_p (Section 3.2.1).
- domain assumption For certified guarantees, the samples used to estimate H are independent of the score-fitting data (Section 3.3.1 and Section 3.2.2).
- domain assumption In Proposition 4, the target satisfies the Poincaré inequality (31a) and a one-sided L-smoothness condition (31b).
- standard math I-MMSE identity (Guo-Shamai-Verdú) and the conditional I-MMSE for Gaussian observation processes.
read the original abstract
Two central challenges in diffusion-based sampling are the theoretical one of understanding their remarkable effectiveness even in high-dimensional settings, and the practical one of designing algorithms with certified performance guarantees. We show that these questions are intimately connected via the \emph{denoising growth complexity} ($\mathsf{DGC}$). It is a geometric measure defined by a log-time weighted integral of the derivative of the denoising mean-squared error along the Gaussian heat flow. We show how the $\mathsf{DGC}$ increments lead to a simple and explicit bound on the KL error of an Euler scheme applied to the stochastic innovations representation. The bound is local along the path: each step is controlled by the corresponding $\mathsf{DGC}$ increment and its relative stepsize. This structure allows us to derive KL sampling guarantees for optimized stepsize schedules, both in a simpler single-block setting and in a more refined $K$-block setting. The $\mathsf{DGC}$ function has a natural martingale structure, which we exploit to develop fully data-certified versions of these algorithms. It also admits information-theoretic upper bounds in terms of covariance, rate distortion, metric entropy, and the Poincar'e constant, thereby recovering and sharpening a range of existing diffusion-sampling guarantees, as well as giving new results. In log heat-time, the fine partition limit is governed by an integral involving the square root of the $\mathsf{DGC}$ density, whereas a single-block schedule depends on its ordinary integral. This comparison precisely characterizes when adaptation to data geometry yields substantial computational gains, including logarithmic-to-constant separations for simple Gaussian mixture models.
Figures
Reference graph
Works this paper leans on
-
[1]
Agarwal and S
A. Agarwal and S. Negahban and M. J. Wainwright , journal =. Fast global convergence of gradient methods for high-dimensional statistical recovery , volume =
-
[2]
Agarwal and P
A. Agarwal and P. L. Bartlett and P. Ravikumar and M. J. Wainwright , journal =. Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization , topic =
-
[3]
Agarwal and S
A. Agarwal and S. Negahban and M. J. Wainwright , journal =. Noisy matrix decomposition via convex relaxation: Optimal rates in high dimensions , volume =
-
[4]
Agarwal and S
A. Agarwal and S. Negahban and M. J. Wainwright , booktitle =. Stochastic optimization and sparse statistical recovery: An optimal algorithm for high dimensions , topic =
-
[5]
El Alaoui and X
A. El Alaoui and X. Cheng and A. Ramdas and M. J. Wainwright and M. I. Jordan , booktitle =. Asymptotic behavior of _p -based
-
[6]
Albergo and Nicholas M
Michael S. Albergo and Nicholas M. Boffi and Eric Vanden-Eijnden , journal =
-
[7]
Altschuler, J. M. and Chewi, S. , institution =. Faster high-accuracy log-concave sampling via algorithmic warm starts , url =
-
[8]
A. A. Amini and M. J. Wainwright , journal = annstat, pages =. High-dimensional analysis of semdefinite relaxations for sparse principal component analysis , volume =
-
[9]
A. A. Amini and M. J. Wainwright , journal = annstat, number =. Sampled forms of functional
-
[10]
Anderson, B. D. , journal =. Reverse-time diffusion equation models , volume =
-
[11]
Azangulov and G
I. Azangulov and G. Deligiannidis and J. Rousseau , journal =. Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions , year =
-
[12]
Balakrishnan and M
S. Balakrishnan and M. J. Wainwright and B. Yu , journal =. Statistical guarantees for the
-
[13]
and Deligiannidis, G
Benton, J. and Deligiannidis, G. and Doucet, A. , journal =. Error bounds for flow matching methods , year =
-
[14]
and De Bortoli, V
Benton, J. and De Bortoli, V. and Doucet, A. and Deligiannidis, G. , booktitle =. Nearly d -Linear Convergence Bounds for Diffusion Models via Stochastic Localization , url =
-
[15]
Bhatia and A
K. Bhatia and A. Pananjady and P. L. Bartlett and A. D. Dragan and M. J. Wainwright , booktitle =. Preference learning along multiple criteria: A game-theoretic perspective , year =
-
[16]
arXiv , arxivid =:2208.05314 , journal =
-
[17]
arXiv , arxivid =:2202.02763 , title =
Advances in Neural Information Processing Systems , doi =. arXiv , arxivid =:2202.02763 , title =
-
[18]
De Bortoli and M
V. De Bortoli and M. Hutchinson and P. Wirnsberger and A. Doucet , journal =. Target Score Matching , year =
-
[19]
Brooks and A
S. Brooks and A. Gelman and G. L. Jones and X. L. Meng , publisher =. Handbook of
-
[20]
Cai and R
J. Cai and R. Chen and M. J. Wainwright and L. Zhao , institution =. Doubly high-dimensional contextual bandits:
-
[21]
and Li, G
Cai, C. and Li, G. , journal =. Minimax optimality of the probability flow
-
[22]
Celentano and M
M. Celentano and M. J. Wainwright , institution =. Challenges of the inconsistency regime:
-
[23]
Cetin and L
M. Cetin and L. Chen and J. W. Fisher and A. T. Ihler and R. L. Moses and M. J. Wainwright and A. S. Willsky , journal =. Distributed fusion in sensor networks , volume =
-
[24]
Log-Concave Sampling , url =
Sinho Chewi , note =. Log-Concave Sampling , url =
-
[25]
X. Cheng and P. L. Bartlett , booktitle =. arXiv , arxivid =:1705.09048 , month =
-
[26]
Chen and M
L. Chen and M. J. Wainwright and M. Cetin and A. Willsky , booktitle =. Multitarget-multisensor data association using the tree-reweighted \ max-product algorithm , year =
-
[27]
Chen and M
J. Chen and M. Stern and M. J. Wainwright and M. I. Jordan , booktitle =. Kernel Feature Selection via Conditional Covariance Minimization , year =
-
[28]
Chen and R
Y. Chen and R. Dwivedi and M. J. Wainwright and B. Yu , journal =. Fast
-
[29]
Chen and L
J. Chen and L. Song and M. J. Wainwright and M. I. Jordan , booktitle =. Learning to Explain: An Information-Theoretic Perspective on Model Interpretation , year =
-
[30]
Chen and M
J. Chen and M. I. Jordan and M. J. Wainwright , booktitle =
-
[31]
Chen and L
J. Chen and L. Song and M. J. Wainwright and M. I. Jordan , booktitle =. L-Shapley and C-Shapley: Efficient Model Interpretation for Structured Data , year =
-
[32]
Chen and R
Y. Chen and R. Dwivedi and M. J. Wainwright and B\ . Yu , journal =. Fast mixing of
-
[33]
S. Chewi and M. A. Erdogdu and M. B. Li and R. Shen and M. Zhang , booktitle =. arXiv , arxivid =:2112.12662 , title =
- [34]
-
[35]
Chewi and J
S. Chewi and J. Pont and J. Li and C. Lu and S. Narayanan , month =. Query low bounds for log-concave sampling , year =
-
[36]
S. Chen and S. Chewi and J. Li and Y. Li and A. Salim and A. R. Zhang , booktitle =. arXiv , arxivid =:2209.11215 , title =
-
[37]
M. Chen and K. Huang and T. Zhao and M. Wang , eprint =. arXiv preprint arXiv:2302.07194 , title =
-
[38]
S. Chen and S. Chewi and H. Lee and Y. Li and J. Lu and A. Salim , eprint =. arXiv preprint arXiv:2305.11798 , title =
-
[39]
and Mei, S
Chen, M. and Mei, S. and Fan, J. and Wang, M. , journal =. An Overview of Diffusion Models: Applications, Guided Generation, Statistical Rates and Optimization , url =
-
[40]
Chen and S
F. Chen and S. Chewi and C. Daskalakis and A. Rakhlin , journal =. High-accuracy sampling for diffusion models and log-concave distributions , year =
-
[41]
and Gatmiry, K
Chen, Y. and Gatmiry, K. , institution =. A Simple Proof of the Mixing of
-
[42]
Chen and M
Y. Chen and M. J. Wainwright , institution =. Fast low-rank estimation by projected gradient descent:
-
[43]
Chichignoud and J
M. Chichignoud and J. Lederer and M. J. Wainwright , journal =. A practical scheme and fast algorithm to tune the Lasso with optimality guarantees , volume =
-
[44]
G. Conforti and A. Durmus and M. G. Silveri , eprint =. arXiv preprint arXiv:2308.12240 , month =
-
[45]
Croitoru and V
F.-A. Croitoru and V. Hondru and R. T. Ionescu and M. Shah , doi =. Diffusion Models in Vision: A Survey , volume =. IEEE Transactions on Pattern Analysis and Machine Intelligence , number =
-
[46]
A. S. Dalalyan , booktitle =. Further and stronger analogy between sampling and optimization:
-
[47]
Stochastic Processes and their Applications , number =
User-friendly guarantees for the. Stochastic Processes and their Applications , number =. doi:10.1016/j.spa.2019.02.016 , eprint =
-
[48]
Daskalakis and A
C. Daskalakis and A. G. Dimakis and R. M. Karp and M. J. Wainwright , booktitle =. Probabilistic Analysis of Linear Programming Decoding , year =
-
[49]
Daskalakis and A
C. Daskalakis and A. G. Dimakis and R. M. Karp and M. J. Wainwright , journal =. Probabilistic analysis of linear programming decoding , volume =
-
[50]
A. G. Dimakis and A. Sarwate and M. J. Wainwright , booktitle =. Geographic
-
[51]
A. G. Dimakis and A. Sarwate and M. J. Wainwright , journal =. Geographic gossip:
-
[52]
A. G. Dimakis and A. A. Gohari and M. J. Wainwright , journal =. Guessing facets:
-
[53]
A. G. Dimakis and P. B. Godfrey and Y. Wu and M. J. Wainwright and K. Ramchandran , journal = ieeeit, month =. Network coding for distributed storage systems , topic =
-
[54]
A. G. Dimakis and M. J. Wainwright , booktitle =. Guessing
-
[55]
Dolecek and Z
L. Dolecek and Z. Zhang and V. Anantharam and M. J. Wainwright and B. Nikolic , booktitle =. Analysis of absorbing sets for array-based
-
[56]
Dolecek and Z
L. Dolecek and Z. Zhang and M. J. Wainwright and V. Anantharam and M. J. Wainwright , booktitle =. Evaluation of the low frame error rate performance of
-
[57]
Dolecek and Z
L. Dolecek and Z. Zhang and V. Anantharam and M. J. Wainwright and B. Nikolic , journal = ieeeit, month =. Analysis of Absorbing Sets and Fully Absorbing Sets for Array-Based
-
[58]
Dolecek and P
L. Dolecek and P. Lee and Z. Zhang and V. Anantharam and B. Nikolic and M. J. Wainwright , journal =. Predicting error floors of structured
-
[59]
Doucet and W
A. Doucet and W. Grathwohl and A. D. G. Matthews and H. Strathmann , booktitle =. Score‑Based Diffusion Meets Annealed Importance Sampling , year =
-
[60]
Duan and M
Y. Duan and M. Wang and M. J. Wainwright , journal =. Optimal value estimation using kernel-based temporal difference methods , volume =
-
[61]
Duan and M
Y. Duan and M. J. Wainwright , institution =. Policy evaluation from a single path: Multi-step methods, mixing and mis-specification , year =
-
[62]
Duan and M
Y. Duan and M. J. Wainwright , institution =. Taming data-hungry reinforcement learning? Stability in continuous state-action spaces , url =
-
[63]
Duchi and A
J. Duchi and A. Agarwal and M. J. Wainwright , booktitle =. Distributed dual averaging in networks , year =
-
[64]
J. C. Duchi and A. Agarwal and M. J. Wainwright , journal =. Dual Averaging for Distributed Optimization:
-
[65]
J. C. Duchi and A. Wibisono and M. J. Wainwright and M. I. Jordan , booktitle =. Finite sample convergence rates of zero-order stochastic optimization methods , topic =
-
[66]
J. C. Duchi and M. J. Wainwright and M. I. Jordan , booktitle =. Privacy-aware learning , year =
-
[67]
J. C. Duchi and P. L. Bartlett and M. J. Wainwright , journal =. Randomized smoothing for stochastic optimization , volume =
-
[68]
J. C. Duchi and M. J. Wainwright and M. I. Jordan , booktitle =. Local privacy and statistical minimax rates , year =
-
[69]
J. C. Duchi and M. I. Jordan and M. J. Wainwright and Y. Zhang , institution =. Optimality guarantees for distributed statistical estimation , year =
-
[70]
J. C. Duchi and M. J. Wainwright and M. I. Jordan , journal =. Privacy-aware learning , volume =
-
[71]
J. C. Duchi and M. I. Jordan and M. J. Wainwright and A. Wibisono , journal =. Optimal rates for zero-order optimization: the power of two function evaluations , volume =
-
[72]
J. C. Duchi and M. I. Jordan and M. J. Wainwright , journal =. Minimax optimal procedures for locally private estimation , volume =
-
[73]
J. C. Duchi and M. J. Wainwright , institution =. Distance-based and continuum Fano inequalities with applications to statistical estimation , year =
-
[74]
The Annals of Applied Probability , number =
Nonasymptotic convergence analysis for the unadjusted. The Annals of Applied Probability , number =. doi:10.1214/16-AAP1238 , eprint =
-
[75]
Dwivedi and Y
R. Dwivedi and Y. Chen and M. J. Wainwright and B. Yu , booktitle =. Log-concave sampling:
-
[76]
Dwivedi and N
R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan , booktitle =. Theoretical guarantees for the
-
[77]
Dwivedi and N
R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , institution =. Challenges with
-
[78]
Dwivedi and Y
R. Dwivedi and Y. Chen and M. J. Wainwright and B. Yu , journal =. Log-concave sampling:
-
[79]
Dwivedi and N
R. Dwivedi and N. Ho and K. Khamaru and M. J. Wainwright and M. I. Jordan and B. Yu , journal =. Singularity, Misspecification, and the Convergence Rate of
-
[80]
Dwivedi and K
R. Dwivedi and K. Khamaru and N. Ho and M. J. Wainwright and M. I. Jordan and B. Yu , booktitle =. Sharp analysis of expectation-maximization for weakly identifiable models , year =
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.