{"id":"fa51feb0-6469-4878-96b2-c5efe17c4eaa","arxiv_id":"2501.00988","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"VP and VE schedules capture opposite features of Gaussian mixture data, but specially designed time dilations recover both with a constant number of ODE steps as dimension grows.","lead":"This paper analyzes how the choice of noise schedule in diffusion models separates the recovery of high-level features, such as the balance between groups, from low-level features, such as the width of each group. It constructs time-dilated schedules that recover both with a constant number of discretization steps for Gaussian mixture and Curie-Weiss data in the infinite-dimensional limit.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Weakest point: the Θ_d(1)-step claim is proven only via the iterated limit lim_{Δt→0} lim_{d→∞}; without a uniform error bound, a fixed number of steps retains a d-independent Euler bias and the κ-dependent σ bias.","rationale":"I read the paper in good faith and tried to find a concrete failure of the main construction. The asymptotic dichotomy between VP and VE, and the mechanism by which time dilations stretch the speciation window to a constant time, are internally coherent. I independently re-derived the first-phase transport identity for the dilated VP limit (the ODE dμ/dy = tanh(h + y μ) maps N(0,1) to pN(y,1) + (1-p)N(-y,1)), so the claimed limiting magnetization distribution is correct, even though the appendix lemma cited for it is stated too loosely to be a rigorous justification. The real soft spot is not the deterministic limiting ODEs but the transfer of those ODE limits to finite-step discretizations. The theorems prove statements of the form 'first d→∞, then Δt→0'; the paper then infers that Θ_d(1) steps suffice. That inference requires a uniformity that is neither stated nor proved. In particular, for a fixed number of steps, the Euler discretization of the limiting ODE has an error that does not tend to zero with d, and the κ→∞ limit in the VE schedule leaves a nonvanishing bias for any fixed κ. This is precisely the concern the reader raised, and it justifies a CONDITIONAL verdict rather than unconditional acceptance. It does not, however, amount to a demonstrated counterexample, because the claims can plausibly be read asymptotically with constants depending on target accuracy. The recommended verdict is therefore UNCHANGED: the reader's CONDITIONAL assessment is appropriate.","tokens_in":28440,"tokens_out":39702,"duration_ms":370758,"concrete_test":"Analytical check: write down the N-step Euler map for the dilated VP first-phase ODE dμ/dt = 2κ tanh(h + 2κt μ) with N fixed, and compute the Wasserstein-2 distance between the resulting distribution of μ at t=1/2 and pN(κ,1) + (1-p)N(-κ,1). If this distance tends to 0 only as N→∞, then for every fixed N there is a nonzero discretization bias, confirming that the Θ_d(1)-step statement requires a small-Δt accuracy-dependent caveat. Complementary numerical check: run the full d-dimensional probability-flow ODE for the dilated VE interpolant with N=100 steps fixed and d = 10^3, 10^4, 10^5, 10^6, 10^7; if the estimated p and σ² errors plateau at a d-independent value instead of decaying, the literal claim of vanishing error with Θ_d(1) steps is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central Θ_d(1)-step claim is supported only in the iterated limit lim_{Δt→0} lim_{d→∞} (and, for the dilated VE schedule, also lim_{κ→∞}); no uniform-in-Δt error bound is given. In the proof of Theorem 1, the O(1/√d) terms in equations (38) and (43) are dropped before letting Δt→0, and the proof explicitly retains a discretization error o_Δt(1) that is independent of d. Consequently, for any fixed number N = Θ_d(1) of steps, the d→∞ limit is the N-step Euler solution of the limiting ODE, not the solution of that ODE; the Euler error E(Δt) does not vanish as d→∞. The sentence in Section 3.2, 'Since we take first d→∞ and then Δt→0, this means we can discretize the ODE with Δt ∈ Θ_d(1) and get accurate estimation,' is therefore not a logical consequence of the theorem as stated. The iterated limit only says that for each fixed Δt the d→∞ limit is the Euler map, and then the Euler map converges as Δt→0. To conclude vanishing error with a number of steps bounded independently of d, one needs the d→∞ convergence to be uniform in Δt, or an explicit bound of the form error ≤ C Δt + ε(d) with ε(d)→0 uniformly in Δt. Similarly, the recovery of σ in Theorem 2 holds only after κ→∞; for fixed κ, the second-phase variance is biased by the factor κ/√(κ²+σ²), an additional d-independent error. These gaps do not make the asymptotic VP/VE dichotomy false, but they make the practical Θ_d(1) claim stronger than what is rigorously proved.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies probability flow ODE generative models built from stochastic interpolants, focusing on two high-dimensional test distributions: a two-mode Gaussian mixture and the Curie-Weiss model. It shows that with uniform noise schedules the VP interpolant captures the low-level feature (mode variance σ²) but misses the high-level feature (mode asymmetry p), while the VE interpolant does the opposite. The authors then introduce piecewise-linear time dilations specific to VP and VE and prove, in a combined limit d→∞ followed by Δt→0 (and, for some statements, κ→∞), that the dilated schedules recover both p and σ² and that the resulting probability flow ODE can be discretized with a number of steps that is constant in the dimension, Θ_d(1), rather than Θ_d(√d). The paper also reports numerical experiments on the GM and CW models and on CelebA-HQ with pretrained VP/VE SDEs.","tokens_in":28834,"tokens_out":12689,"duration_ms":128470,"significance":"The paper identifies a clean mechanism for a practically relevant phenomenon: the complementary failure modes of VP and VE schedules and the possibility of curing them by time dilation. The reduction of the high-dimensional probability flow ODE to one-dimensional limiting ODEs in Lemmas 2–5 is elegant, and the explicit limiting equations for the magnetization and the orthogonal variance are concrete and falsifiable. The GM and CW results are stated with full proofs in the appendices, and the numerical experiments support the qualitative phase structure. If the technical gaps in the limit interchanges are closed or properly qualified, the Θ_d(1)-step claim would be a valuable theoretical contribution to the diffusion-model literature.","major_comments":[{"comment":"The Θ_d(1)-step conclusion requires a two-limit argument that is not supplied. The proof derives the limiting ODE by dropping O(1/√d) terms and then lets Δt→0, but to justify the sentence 'we can discretize the ODE with Δt ∈ Θ_d(1) and get accurate estimation' one must show that, for fixed Δt, the d→∞ limit of the Euler discretization of the d-dimensional system equals the Euler discretization of the limiting ODE, with errors that are controlled uniformly in Δt. Equation (44) asserts precisely such a decomposition, µ_{t=1/2} = θ + O(1/√d) + o_Δt(1), but no uniform-in-Δt bound or exchange-of-limits lemma is proved. As written, the theorem establishes the double limit for the continuous-time solution and separately that the Euler method converges for the limiting ODE; it does not by itself establish the claimed O_d(1)-step guarantee for the original discretized dynamics.","section":"§3.2, Theorem 1 proof, Eq. (44)"},{"comment":"The recovery of both features is stated in a combined limit that also sends κ to infinity, but this qualification is not carried through the abstract or the Θ_d(1)-step claim. For fixed κ, Theorem 2 gives lim_{Δt→0} lim_{d→∞} σ_1^{κ,Δt,d} = κ σ / √(κ²+σ²), which is strictly smaller than σ, and Theorem 1 gives M_1 ∼ p_κ δ_1 + (1−p_κ)δ_{−1} with p_κ ≠ p for finite κ. Thus a fixed schedule with a fixed κ does not exactly recover both features in the d→∞ limit; the paper should state explicitly that κ must be sent to infinity (or chosen large) and should quantify the resulting bias, e.g., as O(1/κ). This is a load-bearing qualification for the headline claim that a Θ_d(1)-step discretization captures both p and σ².","section":"§3.2, Theorems 1 and 2"},{"comment":"The time dilations (4) and (5) are only piecewise C¹, with a kink at t = 1/2, while Lemma 2 requires α_τ, β_τ ∈ C²([0,1]). The composed coefficients α_t = 1−τ(t) and β_t = τ(t) are therefore not C², and the velocity field of the interpolant may be discontinuous at the kink. The proofs treat the two phases separately, but the paper does not justify that the probability flow ODE, or its Euler discretization, is well-defined across the kink, nor that the phase-wise limiting ODEs combine to give the stated limiting dynamics. A regularity lemma for absolutely continuous or piecewise smooth time changes, or an explicit smoothing argument, is needed.","section":"§A, Lemma 2 and §3.2, Eqs. (4)–(5)"}],"minor_comments":[{"comment":"The title contains a typo: 'Dimensionss' should be 'Dimensions'.","section":"Title"},{"comment":"The caption reads 'uniformly discretized with step size d = 10^6, Δt = 0.01'; this should be 'dimension d = 10^6, step size Δt = 0.01'.","section":"Figure 2 caption"},{"comment":"The caption begins 'or different number of discretization steps'; it should begin 'For different number of discretization steps'.","section":"Figure 5 caption"},{"comment":"The quantity p_κ is used in the statement M_1 ∼ p_κ δ_1 + (1−p_κ)δ_{−1} before it is defined; the definition 'p_κ is such that lim_{κ→∞} p_κ = p' should appear before this display.","section":"Theorem 1 statement"},{"comment":"In the proof of Proposition 3, the display 'µ_τ/d = α_τ Z + √d β_τ m' is dimensionally inconsistent; from X_τ = α_τ z + β_τ a and µ_τ = r·X_τ/√d one obtains µ_τ = α_τ Z + √d β_τ m. The subsequent comparison α_τ ≈ √d β_τ follows from the corrected equation.","section":"Appendix C, Eq. (25)"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely to be of interest to the journal's readership, and the GM/CW analysis is a strong conceptual contribution. My main reservation is that the abstract and introduction overstate the Θ_d(1)-step result: the rigorous statements require iterated limits and, for some conclusions, an additional κ→∞ limit that are not reflected in the headline claims. I would encourage the authors to add a explicit two-limit lemma or to rephrase the practical claims as asymptotic statements valid for sufficiently large d, small Δt, and large κ."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper has a real and interesting core, and a real gap at the edge of its main claim. The VP/VE dichotomy — VP captures the low-level feature (σ²) but misses the asymmetry p, VE does the opposite — is cleanly demonstrated through exact reductions of the probability flow ODE to one-dimensional ODEs. The time-dilation construction that stretches the vanishing speciation window into a constant-time phase is clever and, as far as I can tell, new relative to Biroli-Mézard and Li-Chen. The paper also re-derives the speciation time rather than importing it, and the related-work section is fair. Credit where due: the Gaussianity preservation in the orthogonal complement, the explicit limiting ODEs, and the Curie-Weiss extension are all solid asymptotic analysis in the stated double limit.\n\nThe soft spot is exactly where the stress-test note lands. Theorems 1 and 2 prove statements of the form lim_{Δt→0} lim_{d→∞} for the discrete trajectory, and the proof drops O(1/√d) terms before letting Δt→0. But the conclusion “we can discretize with Δt ∈ Θ_d(1)” does not follow from that iterated limit. For any fixed number N of steps, the d→∞ limit is the N-step Euler map of the limiting ODE, not the solution of that ODE; the Euler error does not vanish as d→∞. To get the Θ_d(1) statement you would need a uniform-in-Δt error bound, e.g. error ≤ C Δt + ε(d) with ε(d)→0 uniformly in Δt, or an explicit non-asymptotic estimate. This is not a fatal flaw in the dichotomy itself, but it means the abstract oversells what is rigorously shown. A related, smaller issue: Theorem 2 recovers σ only in the additional κ→∞ limit, and for fixed κ the variance is biased by κ/√(κ²+σ²); the paper says this, but the abstract’s phrasing hides it. Minor technical point: the dilation (4) is only piecewise C¹ while the interpolant regularity assumed is C²; likely fixable, but should be acknowledged.\n\nWho gets value: theory-minded diffusion researchers, especially those working on phase transitions and asymptotics. The paper deserves a serious referee — the core idea is worth engaging, and a careful referee can ask for the right weakening or a uniform bound. I would send it to review rather than desk-reject, and I would probably cite the VP/VE dichotomy even with the gap, since the qualitative picture is likely right.","headline":"The VP/VE dichotomy and time-dilation idea are genuinely new and likely correct, but the headline Θ_d(1)-step claim is not actually proved because the iterated limit order does not control the Euler error uniformly in the step size.","tokens_in":29370,"tokens_out":2043,"would_cite":true,"duration_ms":23253,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"With correctly dilated noise schedules, a generative probability-flow ODE for Gaussian mixtures and Curie-Weiss models becomes discretizable with a constant number of steps instead of the $\\Theta(\\sqrt{d})$ steps a uniform grid requires.","keywords":["diffusion models","noise schedules","stochastic interpolants","Gaussian mixtures","Curie-Weiss model","speciation transition","probability flow ODE","high-dimensional asymptotics"],"falsifier":"Fix a small step size such as $\\Delta t=0.01$ and simulate the dilated VE probability-flow ODE for a Gaussian mixture with known $p=0.8$ and $\\sigma^2=0.25$ at dimensions $d=10^4,10^6,10^8$; if, with the step size held fixed, the empirical magnetization at $t=1/2$ does not approach $p\\delta_1+(1-p)\\delta_{-1}$ and the coordinate variance at $t=1$ does not approach $\\sigma^2$ as $d$ grows, then the double limit is not uniform and the $\\Theta_d(1)$-step conclusion fails in the finite-step regime it claims.","tokens_in":28225,"feed_emoji":"⏱️","tokens_out":10009,"duration_ms":86931,"temperature":0.7,"pith_summary":"The paper claims that the well-known trade-off between variance-preserving (VP) and variance-exploding (VE) diffusion schedules is not inherent: a correctly scaled initial noise level together with a time-dilated, non-uniform noise schedule makes one generative probability-flow ODE reproduce both the high-level asymmetry between modes and the low-level within-mode structure of the data. For a two-mode Gaussian mixture and the Curie-Weiss spin model, the paper shows that uniform denoising makes VP capture only the within-mode variance $\\sigma^2$ while losing the mode weight $p$, and makes VE capture only $p$ while losing $\\sigma^2$. With specific dilations, the speciation transition—the moment each sample commits to a mode—occurs at a constant time instead of at a time that vanishes with dimension, and the resulting limiting ODE can be discretized with $\\Theta_d(1)$ steps. If correct, this turns the choice of noise schedule into a principled way to control which features of the data are learned and how many solver steps are needed, rather than an empirical hyperparameter.","feed_headline":"Time-dilated noise schedules cut diffusion steps from √d to O(1)","feed_subtitle":"Stretching the speciation window lets one ODE recover both mode weights and within-mode variance with constant step count.","key_machinery":"The machinery is the stochastic interpolant and its probability-flow ODE, $\\dot X_\\tau = b_\\tau(X_\\tau)$ with $b_\\tau(x)=\\mathbb{E}[\\dot I_\\tau \\mid I_\\tau=x]$, together with a time-dilation $\\tau(t)$ that acts as a non-uniform noise schedule. The paper specializes to $\\alpha_\\tau=1-\\tau$, $\\beta_\\tau=\\tau$ (and the VP variant $\\sqrt{1-\\tau^2}$), with noise scale $c=1$ for VP and $c=\\sqrt{d}$ for VE, and chooses $\\tau(t)$ so that the speciation window has constant length in $t$: for VP, $\\tau$ reaches $\\kappa/\\sqrt{d}$ at $t=1/2$; for VE, the final window $\\tau\\in[1-\\kappa/\\sqrt{d},1]$ is stretched to $t\\in[1/2,1]$. The argument reduces the high-dimensional ODE to two low-dimensional objects: the magnetization $M_t=r\\cdot X_t/d$ (or $\\mu_t=r\\cdot X_t/\\sqrt{d}$ in the first VP phase) that settles the mode, and the orthogonal Gaussian fluctuations that carry $\\sigma^2$. Each phase's limiting ODE is identified with a known one-dimensional interpolant transport, which is what allows explicit formulas for the endpoints.","core_discovery":"The central discovery is that the two phases of generation correspond to two distinct features, and each can be stretched to constant duration by rescaling time near the speciation transition. Starting from the stochastic interpolant $I_\\tau = c\\alpha_\\tau z + \\beta_\\tau a$ with $z$ Gaussian and $a$ drawn from the data, the paper proves for the Gaussian mixture $p\\mathcal{N}(r,\\sigma^2 I_d)+(1-p)\\mathcal{N}(-r,\\sigma^2 I_d)$ with $|r|^2=d$ that a uniform grid has VP's speciation time $\\tau_s=1/\\sqrt{d}$ and therefore cannot resolve $p$ with $\\Theta_d(1)$ steps, while VE captures $p$ but drives the orthogonal fluctuations to zero in the $d\\to\\infty$ limit, losing $\\sigma^2$. The dilated VP schedule $\\tau(t)=2\\kappa t/\\sqrt{d}$ on $[0,1/2]$ followed by a linear ramp, and the dilated VE schedule $\\tau(t)=(1-\\kappa/\\sqrt{d})2t$ followed by a $\\kappa/\\sqrt{d}$ window, both produce limiting two-phase ODEs: a first phase with a $\\tanh$ drift that settles the magnetization $M_t$ to $p\\delta_1+(1-p)\\delta_{-1}$ and a second phase that transports the orthogonal component to variance $\\sigma^2$ (for VP) or to the mode distribution (for VE). For the Curie-Weiss model the same dilated VE schedule resolves both the mode asymmetry and the discrete $\\{\\pm 1\\}$ spin distribution. Because the limiting ODE is independent of $d$, the paper concludes that $\\Theta_d(1)$ discretization points suffice, whereas uniform grids need $\\Theta(\\sqrt{d})$.","pith_inferences":["The time-dilation construction suggests a general recipe for hierarchical data: identify each feature's speciation time and stretch the schedule so every critical window has constant duration; the paper demonstrates this only on two test distributions, but the mechanism is stated in terms of generic two-phase ODEs.","If the order-of-limits issue can be tightened, the $\\Theta_d(1)$ statement becomes a proof that schedule design, not just score estimation or solver order, is the dominant factor in the step-count cost of sampling in high dimension.","A natural test is to apply the dilated VP and VE schedules to pretrained image samplers with fixed step budgets and measure feature-level KL divergences; the paper's CelebA experiment suggests VP and VE would swap which feature improves with steps.","The finite-$\\kappa$ correction is unexplored: the theorems recover $p$ and $\\sigma^2$ only as $\\kappa\\to\\infty$, so the optimal dilation strength for a given feature tolerance and dimension is an open quantitative question."],"forward_implications":["Under a uniform grid, VP and VE solve complementary halves of the Gaussian-mixture problem: VP captures $\\sigma^2$, VE captures $p$, and neither captures both with $o(\\sqrt{d})$ uniform steps.","The dilated schedules given in equations (4) and (5) make both $p$ and $\\sigma^2$ recoverable with $\\Theta_d(1)$ steps for the Gaussian mixture, and the same dilated VE schedule works for the Curie-Weiss model.","The first-phase limiting ODEs for the Gaussian mixture and Curie-Weiss model coincide up to a factor $m$, so the early phase is blind to low-level model details; only the second phase distinguishes the two data distributions.","In practice, a non-uniform grid with half the points spent before $\\tau=\\kappa/\\sqrt{d}$ and half after reproduces the output that a uniform grid needs $\\sim\\sqrt{d}$ times more points to achieve.","Real-image experiments on CelebA-HQ show the same directional pattern: more VP steps improve high-level features without fixing low-level quality, while more VE steps improve low-level features without fixing high-level diversity."],"supporting_citations":[{"why":"Defines the stochastic interpolant construction and the probability-flow ODE whose velocity field is $b_\\tau(x)=\\mathbb{E}[\\dot I_\\tau \\mid I_\\tau=x]$; the velocity-field lemmas used throughout the paper rest on this.","marker":"[Albergo et al., 2023]"},{"why":"Introduced the phase-transition and speciation-time picture for score-based diffusion models that this paper translates to interpolants and extends with time dilations.","marker":"[Biroli et al., 2024]"},{"why":"Introduced the speciation time and the Curie-Weiss analysis, including the conditional-expectation formula for $\\eta_t(x)$ used in the Curie-Weiss theorems.","marker":"[Biroli and M´ezard, 2023]"},{"why":"Supplies the VP and VE SDE terminology, the score-based models used in the connection section, and the pretrained samplers used in the CelebA-HQ comparison.","marker":"[Song et al., 2021]"},{"why":"Provides the DDPM schedule with $\\gamma_{\\min}$ and $\\gamma_{\\max}$ used as the practical VP time dilation in Figure 1 and as the VP sampler in the image experiments.","marker":"[Ho et al., 2020]"},{"why":"Gives a Gaussian-mixture sampling guarantee whose polynomial or $\\Theta(\\sqrt{d})$ step requirements form the baseline that the $\\Theta_d(1)$ result improves on.","marker":"[Gatmiry et al., 2024]"},{"why":"Provides another baseline for learning mixtures of Gaussians via a DDPM objective, used as a comparison point for step-count bounds.","marker":"[Shah et al., 2023]"},{"why":"Supplies the CelebA-HQ dataset used to demonstrate that the VP/VE feature dichotomy persists on real images.","marker":"[Karras et al., 2018]"}],"fun_headline_variants":["Dilated noise schedules cut diffusion steps from √d to O(1)","Time-stretched schedules recover both modes and variance in O(1) steps","Rescaling noise time near speciation makes generation O(1)-step","Constant-step diffusion: stretch the speciation window","Dilated schedules cut ODE steps from √d to O(1)"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the asymptotic limits commute in the order $d$ then step size, so that errors of order $1/\\sqrt{d}$ can be discarded before the discretization error is taken to zero; if those errors are not uniform in the schedule parameters, the claim of $\\Theta_d(1)$ steps may hold only in a narrower regime than stated.","fun_headline_variants_meta":{"raw":{"variants":["Dilated noise schedules cut diffusion steps from √d to O(1)","Time-stretched schedules recover both modes and variance in O(1) steps","Rescaling noise time near speciation makes generation O(1)-step","Constant-step diffusion: stretch the speciation window","Dilated schedules cut ODE steps from √d to O(1)"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00067,"raw_usage":{"total_tokens":3146,"prompt_tokens":1128,"completion_tokens":2018,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":744,"completion_tokens_details":{"reasoning_tokens":1927}},"tokens_in":744,"tokens_out":2018,"duration_ms":13689,"temperature":1.0,"reasoning_tokens":1927,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:41:43.025978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fix a small step size such as $\\Delta t=0.01$ and simulate the dilated VE probability-flow ODE for a Gaussian mixture with known $p=0.8$ and $\\sigma^2=0.25$ at dimensions $d=10^4,10^6,10^8$; if, with the step size held fixed, the empirical magnetization at $t=1/2$ does not approach $p\\delta_1+(1-p)\\delta_{-1}$ and the coordinate variance at $t=1$ does not approach $\\sigma^2$ as $d$ grows, then the double limit is not uniform and the $\\Theta_d(1)$-step conclusion fails in the finite-step regime it claims.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced the phase-transition and speciation-time picture for score-based diffusion models that this paper translates to interpolants and extends with time dilations."},{"cited_title":"and Mézard, M","cited_arxiv_id":null,"evidence_quote":"Introduced the speciation time and the Curie-Weiss analysis, including the conditional-expectation formula for $\\eta_t(x)$ used in the Curie-Weiss theorems."},{"cited_title":"P., Kumar, A., Ermon, S., and Poole, B","cited_arxiv_id":null,"evidence_quote":"Supplies the VP and VE SDE terminology, the score-based models used in the connection section, and the pretrained samplers used in the CelebA-HQ comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives a Gaussian-mixture sampling guarantee whose polynomial or $\\Theta(\\sqrt{d})$ step requirements form the baseline that the $\\Theta_d(1)$ result improves on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides another baseline for learning mixtures of Gaussians via a DDPM objective, used as a comparison point for step-count bounds."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CelebA-HQ dataset used to demonstrate that the VP/VE feature dichotomy persists on real images."}],"review_version":1}