{"id":"e776ae17-d5d5-48c2-8477-9f8091d880bf","arxiv_id":"2608.13123","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A conditional diffusion model is used to sample the full distribution of plausible spectra in analytic continuation, yielding uncertainty estimates and a hardness metric called the uncertainty pseudo-volume.","lead":"This paper trains a diffusion generative model to sample many possible real-frequency spectra that are all consistent with one imaginary-time correlation function, instead of returning a single best guess. The authors use the spread of those samples to put error bars on transport coefficients and to flag which spectral features the data truly constrain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No check that diffusion samples reproduce the input iTCF; without forward-consistency, the ensemble spread and UPV may measure model error, not posterior ambiguity.","rationale":"Good-faith reading: the method is a sensible generative approach, and the synthetic demonstrations are encouraging. The central claim, however, requires the learned conditional distribution to be the posterior over spectra given G under the training prior. The weakest point I find is not the prior family itself, although the Section IVD caveat is real, but the absence of any evidence that sampled spectra map back to the conditioning G. Because the model is trained only with a denoising MSE and sampled by an ODE rollout conditioned on G, there is no hard data-consistency constraint. Without a residual or coverage check, the spread and the UPV cannot be distinguished from model approximation error. This is a necessary condition for the uncertainty-quantification claim; the distribution-shift concern raised by the reader is important but secondary, so I partially agree.","tokens_in":18828,"tokens_out":6351,"duration_ms":61592,"concrete_test":"On the same 1000-sample ensembles used for Figures 3 and 5, compute the forward transform K C_j for each generated spectrum using the same trapezoidal rule as in training, and compare with the conditioning iTCF G. Report the normalized L2 residual ||K C_j - G||/||G|| and the maximum-residual distribution; if the median residual is not at or below the noise floor of G (trapezoidal error for synthetic data, simulation noise for the PIMD case), then the samples are not drawn from p(C|G), and the uncertainty and UPV claims need revision. A complementary calibration check on synthetic data, namely whether the 90% ensemble interval covers the true spectrum in about 90% of held-out cases, would also test the probabilistic claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that the diffusion ensemble samples the posterior p(C|G), so its spread is a theoretically grounded uncertainty quantification and the UPV measures intrinsic inversion hardness. The load-bearing condition is that generated spectra actually lie on the data-consistent manifold, i.e., K C ≈ G for essentially every sample. Nothing in the training or sampling procedure enforces this: the network is trained only to denoise pairs (C, G) from the synthetic generator, and inference is a free ODE rollout from noise without a projection or data-consistency term. If the model has finite capacity, is imperfectly trained, or the conditionally generated C is not mapped back through the forward kernel, samples can deviate from the input G. The paper reports no residual or coverage test showing that the ensemble's forward transforms match G at the noise level of the data. In that case the spread and UPV conflate model misspecification or approximation error with the intrinsic ambiguity of the inverse problem, undermining both the probabilistic interpretation of the error bars and the hardness interpretation of P = 61.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a diffusion-based generative model for analytic continuation, aiming to model the full conditional distribution p(C|G) of real-frequency spectra given an imaginary-time correlation function. The authors train a DiT-based denoiser on procedurally generated spectra built from 1-4 warped Gaussian bumps, condition on a four-channel representation of G(τ), and draw 1000 independent samples per input at inference. They summarize the resulting ensemble with a local PCA and propose a new metric, the uncertainty pseudo-volume (UPV), as a quantitative hardness measure for each inversion. The method is demonstrated on synthetic spectra and on a PIMD simulation of liquid parahydrogen, yielding a self-diffusion coefficient with an error bar and a high-frequency feature flagged as uncertain. The central claims are that the ensemble spread provides theoretically grounded uncertainty quantification and that the UPV diagnoses intrinsic inversion difficulty.","tokens_in":19037,"tokens_out":5525,"duration_ms":50785,"significance":"If validated, this would be a useful contribution to the analytic continuation literature, where joint-posterior uncertainty estimates remain rare. The synthetic experiments are carefully designed, the PCA parsimony analysis is a creative way to summarize correlated posterior structure, and the paper provides full implementation details, training and sampling pseudocode, and an explicit statement of a key limitation regarding training-family representativeness. However, the uncertainty statement is not yet anchored by a forward-consistency or calibration check, and the UPV scale is set relative to the compared spectra. These gaps are directly load-bearing for the paper's central claims, but they are addressable within the manuscript's scope. The method is a clear step beyond pointwise error bars, and the intended significance is evident, but the validation needs to be completed.","major_comments":[{"comment":"The inference procedure never projects the generated spectrum through the forward kernel of Eq. (1), so nothing in the algorithm enforces that the sampled C(ω) reproduces the conditioning G(τ). The paper reports no residual or coverage test showing that the ensemble's forward transforms K C_j match G within the data noise level. Without such a check, the ensemble spread and the UPV in Section IV.C may conflate model approximation error with the intrinsic ambiguity of the inverse problem, which is exactly what the paper claims to measure. I request a forward-consistency experiment (e.g., the distribution of residuals ‖K C_j − G‖ across the ensemble, or the fraction of samples whose forward transform lies within a noise ball around G) and a corresponding calibration/coverage statistic.","section":"III.B, IV.A (Eqs. 4-6)"},{"comment":"The UPV defined in Eq. (11) depends on the free scale parameter λ, which is set after the fact to half the maximal z_i across all compared spectra (Section IV.C). This makes the reported UPV values relative to the particular set of spectra included in the comparison. In particular, the claim in Section IV.D that P = 61 indicates \"moderate\" ambiguity relies on the synthetic examples in Figure 5 being used to fix λ; if the comparison set changes, P is not an absolute measure of inversion hardness. Please either fix λ in advance independent of the dataset, report P on an absolute scale, or demonstrate that the ranking of P across spectra is insensitive to the choice of λ.","section":"IV.C (Eq. 11)"},{"comment":"The training distribution in Section VI.B consists of spectra built from 1 to 4 warped Gaussian bumps whose centers are uniformly sampled in the lower half of the frequency domain (ω ∈ [0, 25]). The authors explicitly state at the end of Section IV.D that the uncertainty estimate is valid only if the model was trained on spectra representative of the system at hand, but no diagnostic is provided to detect when this condition fails for a new iTCF. Given that the abstract and title promise uncertainty quantification for analytic continuation generally, the lack of a distribution-shift test or a characterization of when the method can be trusted is a load-bearing gap. At minimum, the claim of a theoretically grounded confidence estimate should be qualified to the training family, or the paper should add a shift-detection mechanism.","section":"IV.D, VI.B"},{"comment":"No calibration check is reported showing how often the true spectrum falls inside the diffusion credibility band. For example, one could compute the empirical coverage of the 5-95% ensemble interval, or the fraction of ground-truth spectra whose projection on the leading PCA components lies inside the UPV ellipsoid, across the synthetic test set. Without such a check, the probabilistic interpretation of the error bars in Figures 3 and 6 is asserted rather than demonstrated.","section":"IV.A, IV.C"}],"minor_comments":[{"comment":"The axis labels in Figure 2 are garbled (e.g., \"G( )\" and non-rendered glyphs); the four channel names should be typeset properly so the figure is self-contained.","section":"Figure 2"},{"comment":"The sentence \"predicting x0 is mathematically equivalent to predicting the score\" could mislead readers because the score is defined for the marginal q_t(x_t) rather than the joint distribution; please spell out the relationship or cite the specific equivalence conditions from the referenced works.","section":"III.A"},{"comment":"The notation P_d is defined as a family of pseudo-volumes but is used interchangeably with P without the subscript in Figure 5 and Section IV.D; please make the notation consistent.","section":"IV.C, IV.D"},{"comment":"The abstract's phrase \"concrete probabilistic basis\" is stronger than what the training objective in Eq. (3) and the deterministic ODE sampler in Section III.B strictly justify; please soften the wording or state the approximation conditions explicitly.","section":"Abstract, IV.A"},{"comment":"The sampling schedule is linearly spaced in t with t1 = 0.99, while training samples t from Beta(1, 2.5); the mismatch between the training-time distribution and the inference-time schedule is not discussed, and a brief comment on its effect would improve reproducibility.","section":"S2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for physics.comp-ph as a methods paper. The main gap is the lack of a forward-consistency and calibration check, which directly affects the interpretability of the UPV and the error bars. The UPV's dependence on λ is also a scale-anchoring issue that should be addressed in revision. The authors are candid about the training-family limitation; I would encourage the editor to require that this limitation be reflected in the abstract-level claims, and to request the additional validation experiments described in the major comments."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first real diffusion-model treatment of analytic continuation, and it's a sensible fit: instead of a single regression output, you sample a distribution of spectra consistent with the observed iTCF. Second, the paper makes a strong claim that this distribution is 'theoretically grounded' posterior uncertainty, but it never checks that the samples actually reproduce the input imaginary-time data, and there's no calibration or coverage test. The stress-test note is right that the load-bearing assumption is unverified.\n\nWhat's genuinely good: the framing is clear, the synthetic data generation is on-the-fly and reasonable, the four-channel conditioning is thoughtful, and the PCA-based correlated uncertainty plus the UPV metric give practitioners a practical way to see which parts of the spectrum the data actually constrain. The parsimony analysis (k≈3 components capture most variance) is a nice result, and the parahydrogen demo showing an error bar on D is a good illustration. The paper also correctly acknowledges the application is a demonstration, not a benchmark.\n\nWhere it's soft, in proportion: the biggest gap is validation. Nothing in the training or sampling enforces K C ≈ G for the sampled spectra. If the model is imperfect or the input is out-of-distribution, the ensemble spread and UPV could measure model error or prior spread rather than posterior ambiguity. The paper says the uncertainty is only valid for representative training spectra, but doesn't provide a diagnostic to detect when that fails. This is fixable: the authors should report residuals on test samples, measure coverage on synthetic tests where ground truth is known, and show the forward transform of their ensemble matches G. The second soft spot is the UPV's lambda: setting it to half the maximal zi across the compared spectra makes P useful within one figure but not scale-invariant, and a sensitivity analysis is missing. Third, no code or data is released, which makes the specific numbers hard to verify. These are not fatal; the method is plausible and the experiments are honest. But the claims currently outrun the evidence.\n\nWho this is for: anyone working on analytic continuation or inverse problems with generative models; the paper deserves a serious referee. I would ask for the missing consistency and coverage checks, and a definition or sensitivity of UPV that doesn't depend on arbitrary scaling. With that, it could be a genuinely useful contribution.\n\nRecommendation: engage with it, send to peer review, and push on the calibration.","headline":"A genuinely new diffusion-based approach to analytic continuation with a useful UPV diagnostic, but the probabilistic uncertainty claim needs calibration and data-consistency checks before I'd trust it fully.","tokens_in":19535,"tokens_out":2352,"would_cite":true,"duration_ms":22196,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A conditional diffusion model learns the full distribution of real-frequency spectra consistent with an imaginary-time correlation function, giving analytic continuation a principled uncertainty estimate.","keywords":["analytic continuation","diffusion models","uncertainty quantification","imaginary-time correlation functions","inverse problems","uncertainty pseudo-volume","liquid parahydrogen","generative modeling"],"falsifier":"Train the identical framework on a deliberately different spectral family, such as sharp narrow peaks confined to the upper half of the frequency domain, and then feed it iTCFs from that family to check whether the UPV and ensemble spread still track the known non-uniqueness; if the UPV stays small for a clearly ambiguous inversion, the metric is not measuring intrinsic hardness.","tokens_in":18573,"feed_emoji":"📊","tokens_out":8294,"duration_ms":69560,"temperature":0.7,"pith_summary":"This paper argues that the right way to treat the ill-posed inverse problem of analytic continuation is to model the full conditional distribution $p(C|G)$ of real-frequency spectra $C(\\omega)$ given an imaginary-time correlation function $G(\\tau)$, rather than to predict a single spectrum. Regression-based methods, including maximum entropy, collapse the solution space toward a conditional mean and hide the fact that many spectra fit the same data. The authors introduce a diffusion-model framework that samples many plausible spectra for one input, and define a new metric, the uncertainty pseudo-volume, that measures how spread out this solution set is after accounting for correlations between frequencies. Applied to synthetic spectra and to a path-integral simulation of liquid parahydrogen, the framework extracts a self-diffusion coefficient with an error bar and flags a high-frequency peak as unsupported by the data.","feed_headline":"Diffusion model puts error bars on analytic continuation","feed_subtitle":"Sampling many plausible spectra per imaginary-time trace reveals which peaks the data truly support.","key_machinery":"The load-bearing mechanism is a conditional diffusion model with a linear interpolant forward process, $x_t=(1-t)x_0+t\\epsilon$, in which a Diffusion Transformer trained to predict the clean spectrum $x_0$ from $(x_t,t,G)$ approximately minimizes the KL divergence to $p(C|G)$. At inference, an ensemble of reverse ODE trajectories, started from Gaussian noise and advanced with a first-order DDIM step followed by second-order DPM-Solver++ updates, produces samples of the posterior. A spectrum-specific principal component analysis then supplies a locally Gaussian model of the uncertainty, and the UPV metric $P_d=\\prod_i(1+z_i/\\lambda)^{-1}$ formed from the principal-axis percentiles measures the correlated spread that pointwise error bars miss.","core_discovery":"The central claim is that a conditional diffusion model with a linear interpolant forward process can approximate the posterior distribution over power spectra for a given imaginary-time correlation function, and that the spread of samples from this learned distribution is a theoretically grounded uncertainty estimate. Pointwise standard deviations show where the data constrain the spectrum, while principal component analysis of the sampled ensemble reveals that the ambiguity is correlated across frequencies and concentrated in a few modes. The uncertainty pseudo-volume $P_d=\\prod_i (1+z_i/\\lambda)^{-1}$, built from the 5th and 95th percentiles of the principal-component coefficients, quantifies the intrinsic hardness of each inversion without access to the true spectrum. On liquid parahydrogen the model yields a self-diffusion coefficient $D=0.59\\pm0.09$ Å$^2$/ps and a wide uncertainty band around a secondary peak near $\\beta\\hbar\\omega\\approx25$, which the authors flag as a likely spurious artifact.","pith_inferences":["A testable extension is to use the UPV as a design objective, choosing experimental or simulation settings that reduce the pseudo-volume rather than merely reporting it after reconstruction.","Comparing the diffusion ensemble with solution sets produced by constrained stochastic analytic continuation on the same iTCFs would test whether the generative prior covers the full ambiguity or only a data-dependent subset.","Because the training distribution centers all spectral bumps in the lower half of the frequency domain, applying the framework to systems with dominant high-frequency structure would require retraining on a broader procedural family or adding an explicit distribution-shift diagnostic; the paper does not yet provide such a diagnostic.","The same correlated-uncertainty analysis could be applied to other inverse problems with smoothing kernels, such as NMR relaxometry or rheology, where the posterior is likewise low-dimensional despite high-dimensional data."],"forward_implications":["A single trained diffusion model yields distributional output, so uncertainty quantification does not require an ensemble of separately trained regression networks, which the paper finds converge to nearly identical spectra and underestimate ambiguity.","The UPV ranks inversions by difficulty without ground truth, allowing practitioners to flag imaginary-time traces whose reconstruction is dominated by the kernel's information loss.","Because the ambiguity is correlated and low-dimensional, with three to five principal components typically explaining over 90% of the variance, compact summaries of the solution space can replace full per-frequency error bars.","Quantities extracted from the spectrum inherit a principled error bar, as in the parahydrogen self-diffusion coefficient $D=0.59\\pm0.09$ Å$^2$/ps.","The framework is formulated for any strongly smoothing inverse problem, so the same uncertainty-quantification logic transfers to other Laplace-type inversions beyond quantum correlation functions."],"supporting_citations":[{"why":"It supplies the PIMD iTCF of liquid parahydrogen and the maximum-entropy and experimental reference values used in the application.","marker":"[5]"},{"why":"It establishes why spectral reconstruction is fundamentally ill-posed, which motivates the need for distributional modeling.","marker":"[17]"},{"why":"It provides the earlier stochastic analytic continuation approach whose difficulty separating physical features from artifacts the generative method aims to improve on.","marker":"[26]"},{"why":"It represents the regression-based neural network paradigm the paper compares against and whose single-output limitation motivates the generative model.","marker":"[37]"},{"why":"It is another supervised mapping baseline showing the standard approach that collapses the solution space to one spectrum.","marker":"[38]"},{"why":"It supplies the denoising diffusion probabilistic model training objective that the conditional formulation builds on.","marker":"[47]"},{"why":"It justifies the score-based SDE framework and the reverse-time ODE sampling used at inference.","marker":"[64]"},{"why":"It provides the DDIM update used as the first-order reverse step in the sampler.","marker":"[68]"},{"why":"It provides the DPM-Solver++ second-order exponential integrator used to accelerate sampling.","marker":"[71]"},{"why":"It introduces the correlated PCA-based uncertainty view that the paper adapts to spectra and builds the UPV upon.","marker":"[74]"}],"fun_headline_variants":["Diffusion model samples plausible spectra to quantify uncertainty","Error bars for an ill-posed problem? Diffusion model says yes","Beyond one answer: diffusion reveals what spectra are plausible","Spurious peaks flagged by uncertainty from diffusion inversion","Forget one spectrum: generate the whole plausible distribution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The uncertainty estimates are trustworthy only when the training spectra resemble the real system's spectrum, and the model was trained on procedural mixtures of one to four warped Gaussian bumps centered in the lower half of the frequency domain, so a real spectrum with very different structure could yield confident but wrong reconstructions.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion model samples plausible spectra to quantify uncertainty","Error bars for an ill-posed problem? Diffusion model says yes","Beyond one answer: diffusion reveals what spectra are plausible","Spurious peaks flagged by uncertainty from diffusion inversion","Forget one spectrum: generate the whole plausible distribution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000967,"raw_usage":{"total_tokens":4150,"prompt_tokens":1017,"completion_tokens":3133,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":3056}},"tokens_in":633,"tokens_out":3133,"duration_ms":21717,"temperature":1.0,"reasoning_tokens":3056,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:18:44.644113+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the identical framework on a deliberately different spectral family, such as sharp narrow peaks confined to the upper half of the frequency domain, and then feed it iTCFs from that family to check whether the UPV and ensemble spread still track the known non-uniqueness; if the UPV stays small for a clearly ambiguous inversion, the metric is not measuring intrinsic hardness.","supporting_citations":[{"cited_title":"Bingham, T","cited_arxiv_id":null,"evidence_quote":"It supplies the PIMD iTCF of liquid parahydrogen and the maximum-entropy and experimental reference values used in the application."},{"cited_title":"Hamann, T","cited_arxiv_id":null,"evidence_quote":"It establishes why spectral reconstruction is fundamentally ill-posed, which motivates the need for distributional modeling."},{"cited_title":"Gunnarsson, M","cited_arxiv_id":null,"evidence_quote":"It provides the earlier stochastic analytic continuation approach whose difficulty separating physical features from artifacts the generative method aims to improve on."},{"cited_title":"Rothkopf, Bayesian inference of real-time dynamics from lattice qcd, Frontiers in Physics10, 1028995 (2022)","cited_arxiv_id":null,"evidence_quote":"It represents the regression-based neural network paradigm the paper compares against and whose single-output limitation motivates the generative model."},{"cited_title":"Huang and S","cited_arxiv_id":null,"evidence_quote":"It is another supervised mapping baseline showing the standard approach that collapses the solution space to one spectrum."},{"cited_title":"Schweighofer, L","cited_arxiv_id":null,"evidence_quote":"It supplies the denoising diffusion probabilistic model training objective that the conditional formulation builds on."},{"cited_title":"Aarts and A","cited_arxiv_id":null,"evidence_quote":"It justifies the score-based SDE framework and the reverse-time ODE sampling used at inference."},{"cited_title":"Karras, M","cited_arxiv_id":null,"evidence_quote":"It provides the DPM-Solver++ second-order exponential integrator used to accelerate sampling."}],"review_version":1}