{"id":"741e71f1-a9f9-4e23-96ee-8ff388520e99","arxiv_id":"2505.13501","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SPIEDiff uses conditional diffusion models and epinets to robustly learn thermodynamic structure from short-time particle simulations, with quantified epistemic uncertainty.","lead":"SPIEDiff is a machine learning framework that learns long-time macroscopic equations from short-time particle simulations, using diffusion models to add robustness and a lightweight 'epinet' to estimate uncertainty. It is tested on 1D Arrhenius particle systems and matches kinetic Monte Carlo predictions while quantifying its own error.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Epinet uncertainty is trained by distillation to a fixed teacher; the Eq. (10) loss drives phi-dependent variation to zero, so reported CIs are optimization artifacts, not model uncertainty.","rationale":"The reader's physical concern about the fluctuation-dissipation estimator (Eq. 6) is real, but that estimator is inherited from prior Stat-PINNs work and is externally validated in the long-range case, where SPIEDiff reproduces the analytic LRM. The novel contribution of this paper is the diffusion-model-based robustness and, especially, the epinet-based epistemic UQ. That contribution has an internal inconsistency: Eq. (10) trains the epinets by mean-squared distillation to a fixed teacher, so the optimal per-φ prediction is the teacher's conditional mean for every φ. The spread over φ that is reported as epistemic uncertainty is not encouraged by the loss; it is whatever residual variation remains after the network fails to cancel the random-prior term. This is not a calibrated posterior, and a post-hoc calibration would not fix the interpretation because the objective is not a posterior-matching objective. The paper's own Limitations section admits that calibration is needed, but the issue is more fundamental. The mean predictions may still be accurate, and the paper's physical results may survive, so rejection is not warranted; however, the 'quantified epistemic uncertainty' and 'reliable long-time predictions' claims require either a corrected training objective that explicitly targets the predictive distribution or an empirical demonstration that the reported intervals track true model uncertainty, e.g., via ensemble comparison. This keeps the reader's conditional verdict unchanged.","tokens_in":27014,"tokens_out":14440,"duration_ms":172427,"concrete_test":"Train an ensemble of 16 independent SPIEDiff runs (different seeds for base DDPMs and epinets) on the same long-range training data, and at a grid of test points compute the ensemble standard deviation of the mean prediction. Compare the SPIEDiff epinet standard deviation over φ (the reported epistemic uncertainty) with this ensemble spread, e.g., by checking whether the reported 95% CI covers the ensemble mean at at least 80% of test points. If it does not, the reported UQ is not capturing model uncertainty. As a second, cheaper check, train EpinetK1 with the same loss but freeze σL at zero and set the prior scale κ=0; if the reported uncertainty remains essentially unchanged, the uncertainty is not being learned at all.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SPIEDiff provides 'quantified epistemic uncertainty' is not supported by the training procedure. EpinetK1 and Epinetf are trained by knowledge distillation (Eq. 10) against fixed teacher predictions K̂1 and f̂ from a single pre-trained DDPM. For every sample s, all epistemic indices φ are regressed to the same target; the loss is a sum of per-φ squared errors. The minimizer of this loss sets the epinet prediction to the conditional mean of the teacher target for every φ, driving φ-dependent variation—the quantity reported as epistemic uncertainty—to zero. Nothing in Eq. (10) rewards preserving the teacher's output distribution or a posterior over functions. The nonzero standard-deviation maps and 95% CIs in Figs. 2, 4, 6, and 7 are therefore the residual of an underdetermined optimization (cancellation of the random-prior term σP), not a measure of model uncertainty. The paper's own Limitations section concedes that calibration is needed; the deeper issue is that the objective is not aimed at a posterior, so tuning calibration would not make the spread meaningful. Consequently, the claim of reliable uncertainty bounds, and the use of those bounds to support 'trustworthy' long-time predictions, is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SPIEDiff, a machine-learning framework for learning the long-time macroscopic dynamics and thermodynamic potentials of purely dissipative particle systems from short-time kinetic Monte Carlo (KMC) simulations. The method extends the Stat-PINNs approach by replacing deterministic neural networks with conditional denoising diffusion probabilistic models (DDPMs) and by augmenting them with epinets for uncertainty quantification. The dissipative operator is estimated from particle fluctuations through a fluctuation-dissipation relation, and the free energy is learned from the short-time macroscopic evolution via a physics-informed loss. The learned components are then integrated to produce long-time continuum predictions. Experiments on 1D Arrhenius lattice-gas models with long-range, weak short-range, and strong short-range interactions show that the mean predictions agree well with the analytical long-range model or with full KMC simulations, and that SPIEDiff is more robust than Stat-PINNs when the training data are scarce or noisy.","tokens_in":27279,"tokens_out":13578,"duration_ms":136559,"significance":"If the uncertainty-quantification claim were established, this would be a significant contribution to SciML for dissipative systems: the combination of fluctuation-dissipation constraints with generative models addresses the operator/potential non-uniqueness, and the strong short-range interaction result (single-well free energy versus LRM's double well) is a compelling demonstration that the method can discover qualitatively correct thermodynamics when the analytical model fails. The authors share the training datasets publicly, which supports reproducibility, and the reported computational-cost comparison is transparent. However, the central UQ claim is not currently supported by the training procedure or by validation statistics; the mean-model contribution is solid, but the paper's title and abstract overstate the reliability of the uncertainty bounds.","major_comments":[{"comment":"The epinet training objective is a knowledge-distillation regression to a fixed teacher target for each epistemic index φ. For every sample, all φ are regressed to the same target \\hat K1(s) or \\hat f(s); the loss is a sum of per-φ squared errors. The global minimizer sets the epinet prediction equal to the conditional mean of the teacher target for every φ, so the φ-dependent variation—the quantity reported as epistemic uncertainty—is driven to zero (up to model capacity and optimization error). The reported standard-deviation maps and 95% confidence intervals in Figs. 2, 4, 6, and 7 are therefore residuals of an underdetermined optimization, not estimates of model uncertainty. The prior-network term can be canceled by the learnable part without affecting the loss, so the width of the reported intervals is not controlled by a meaningful objective. The paper's limitations statement that calibration is needed does not resolve this: calibrating an artifact does not make it a posterior. To support the title claim, the epinet should be trained with an objective that preserves a principled predictive distribution (e.g., a proper ENN loss, an ensemble teacher with per-φ targets, or explicit posterior-matching), or the output should be reframed as a sensitivity band without the phrase 'epistemic uncertainty'.","section":"Section 4.3, Eq. (10)"},{"comment":"The paper does not provide any quantitative calibration or coverage analysis for the reported 95% confidence intervals. Statements such as 'most of the KMC data points are successfully captured' and 'the uncertainty bounds consistently encompassing the LRM and KMC data points' are not statistical evidence; with sufficiently wide intervals such statements are vacuous, and the figures do not report interval widths or coverage counts. Since 'quantified epistemic uncertainty' is a central claim, the authors should report empirical coverage probabilities over the full space-time validation set for each experiment (long-range, weak and strong short-range, scarcer-data, noisier-data), ideally with a calibration plot, and discuss whether the intervals widen when data are scarcer or noisier.","section":"Section 5, Figs. 2, 4, 6, 7; Appendix F"},{"comment":"The definition of the conditioning and target variables for the base network NNf is ambiguous and appears circular. The auxiliary quantity Υ is defined as the residual Σ_i ⟨γj,γi⟩ Δz_i/Δt, the same quantity that appears in the physics-informed loss, and the network is conditioned on the noisy version Υ_ω of this target. The inference procedure for generating Q (and hence f) from the trained model is not specified precisely: a reader cannot tell what value of Υ is used at inference, how the 2000 realizations described in Appendix B are sampled, and how the reported mean and confidence intervals for f are obtained. Please specify the full forward pass at inference and clarify the role of Υ, or the results cannot be reproduced.","section":"Section 4.3, Eq. (9); Appendix A.1"}],"minor_comments":[{"comment":"The abstract states 'reliable epistemic uncertainty bounds' while the Limitations section concedes that calibration is needed to ensure rigorous statistical coverage; these statements should be reconciled and the abstract's wording softened unless new calibration results are added.","section":"Abstract and Section 6"},{"comment":"The last row is labeled '21–28' but the previous row ends at 21; this should likely be '22–28'.","section":"Table 9"},{"comment":"The KMC 'full simulation' runtimes (3500, 875, and 625 days) are presented without explaining how they are estimated; please add a footnote describing the calculation or extrapolation.","section":"Table 1"},{"comment":"The paper refers to Appendix H of [18] for the exact KMC parameters (h, ∆t, teq, interaction potentials); the journal version should include these parameters or a summary so that the experiments are self-contained.","section":"Appendix E"},{"comment":"The notation \\hat K1(Z1(s)_{is}, K1(s)(ω), ω) uses the noisy variable as the second argument, which is confusing because it is not a conditioning variable; the notation should be aligned with standard DDPM notation.","section":"Eq. (8)"},{"comment":"The line 'with pΩ∼N (0, I)' is missing the argument: it should read p(yΩ) ∼ N (0, I).","section":"Section 3, Eq. (3)"}],"recommendation":"major_revision","confidential_remarks":"The paper is borderline. The mean predictions are convincing and the method is a natural extension of Stat-PINNs, but the UQ claim in the title is not supported by the training objective as written. I recommend major revision rather than rejection because the issue is addressable: either re-train the epinets with a principled objective or substantially weaken the claims and add coverage analysis. The authors' own limitations paragraph already concedes the calibration problem, but the more fundamental issue with Eq. (10) needs to be directly addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful contribution here is the robustness of the learned thermodynamic components. Replacing Stat-PINNs' MLPs with conditional DDPMs gives stable mean predictions for K1 and f when data is scarce or noisy, and the three KMC test cases (long-range, weak/strong short-range) show that the deterministic macroscopic predictions match KMC or the analytic LRM, including the strong short-range case where LRM wrongly predicts a double-well free energy. That part is real and worth building on.\n\nThe epistemic uncertainty, however, does not survive contact with Eq. (10). The epinets are trained by knowledge distillation to a fixed teacher: for every epistemic index phi, the target is the same teacher output. The loss drives each phi-dependent prediction to that common target, so the minimizer has no reason to preserve phi-dependent spread. The reported standard deviation maps and 95% CIs are residuals from an underdetermined optimization, i.e., how much the random prior term didn't cancel, not a measure of model uncertainty. The Limitations section says calibration is needed, but calibration of a spread that was never aimed at a posterior will not make it epistemic. This is a load-bearing flaw for the 'quantified epistemic uncertainty' claim and for the 'trustworthy' language in the abstract.\n\nA secondary caveat: the fluctuation-dissipation estimator (Eq. 6) is inherited from the authors' Stat-PINNs work and presumes local equilibrium and white-in-space-time noise at the measurement scale. That's a reasonable assumption for KMC lattice gases, but it limits the scope.\n\nThat said, the mean-prediction methodology is a genuine step over Stat-PINNs, and the experimental validation is honest and reproducible (the KMC data is from the public Stat-PINNs repo, though no code ships). The computational cost table is useful.\n\nVerdict: I'd send this to peer review because the robustness story is solid and the UQ flaw is fixable in revision by either reframing the epinet output as heuristic sensitivity or training a proper posterior (e.g., ensembling over teacher retrains). But I would not let the current abstract stand.\n\nReading group: maybe. Cite? Only for the diffusion-robustness part, with a caveat.\n\nRecommendation: review, major revision.","headline":"Solid diffusion-based robustness upgrade for Stat-PINNs, but the epinet 'epistemic uncertainty' is a distillation artifact and should be reframed.","tokens_in":27789,"tokens_out":2875,"would_cite":true,"duration_ms":29047,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A statistical-physics diffusion model recovers long-time macroscopic dynamics from short-time particle simulations, with uncertainty bounds.","keywords":["dissipative dynamics","GENERIC formalism","fluctuation-dissipation relation","conditional diffusion models","epistemic uncertainty","epistemic neural networks","coarse-graining","kinetic Monte Carlo"],"falsifier":"Run the fluctuation-dissipation estimator on a particle model with controlled colored noise or with measurements taken before local equilibration, and compare the inferred operator to direct estimates from longer-time statistics; a systematic dependence of the inferred entries on the measurement window $h$ (or a visibly biased free energy against equilibrium sampling) would falsify the premise that short-time fluctuations are white and locally equilibrated.","tokens_in":26826,"feed_emoji":"⚛️","tokens_out":8712,"duration_ms":83163,"temperature":0.7,"pith_summary":"The paper claims that the long-time thermodynamics and kinetics of purely dissipative particle systems can be learned from short-time particle simulations, even when the data are scarce or noisy, by combining fluctuation-dissipation relations with conditional diffusion models and epistemic neural networks. In tests on stochastic Arrhenius particle processes, the learned free energy and dissipative operator match the known analytic long-range model, and in a strong short-range interaction case the learned single-well free energy corrects the analytic model's erroneous double-well prediction while the macroscopic evolution tracks kinetic Monte Carlo simulations. The payoff is practical: once trained, SPIEDiff produces long-time continuum predictions with uncertainty bounds in minutes, compared with days or years for direct particle simulation.","feed_headline":"Minutes of machine learning replace years of particle simulation","feed_subtitle":"A physics-informed diffusion model extracts long-time thermodynamics and kinetics from short, noisy particle runs, with uncertainty.","key_machinery":"The load-bearing mechanism is the infinite-dimensional fluctuation-dissipation relation (Eq. 6): after choosing a finite-element basis, each entry of the discretized dissipative operator is estimated as the covariation of rescaled fluctuations of the coarse-grained particle density, under the assumption that the particle field follows an Itô stochastic PDE with the operator and free energy appearing explicitly. A tridiagonal structure-preserving parameterization then reduces the operator to one off-diagonal function $K_1$, with the diagonal fixed by mass conservation and positive semi-definiteness enforced by $K_1 \\le 0$. The free energy density is learned by fitting the discretized gradient-flow residual over a short macroscopic time step, using a conditional diffusion model whose uncertainty comes from an epinet; a deterministic DDIM sampler with few reverse steps makes prediction with many uncertainty samples cheap.","core_discovery":"The central discovery is that the non-uniqueness problem in learning macroscopic thermodynamics is broken by fluctuation data: the dissipative operator is not inferred from the mean macroscopic evolution but read off from the covariance of rescaled particle-density fluctuations, and the free energy is then learned from short-time mean evolution using the structure-preserving discretized operator. Building the estimator and the free-energy fit on conditional denoising diffusion models, with lightweight epistemic networks trained by knowledge distillation, makes the learned thermodynamic components stable under limited and noisy training data and supplies uncertainty bands. In the strong short-range interaction regime, SPIEDiff recovers a single-well free energy where the analytic long-range model predicts a double well, and the resulting macroscopic dynamics agree with kinetic Monte Carlo simulations that the analytic model misses.","pith_inferences":["Beyond the paper: the same covariance-to-operator route could in principle extend to closed GENERIC systems by learning an entropy functional instead of a free energy, but the paper only treats the purely dissipative isothermal setting.","The paper leaves open whether the method tolerates colored or non-Markovian noise; a stress test varying the measurement window $h$ would reveal whether the fluctuation-dissipation premise breaks.","The authors state in their limitations section that uncertainty calibration needs further work before statistical coverage is guaranteed, so the displayed intervals are most safely read as indicative bounds.","Because DDIM sampling keeps accuracy at two reverse steps, data generation dominates cost, suggesting active selection of initial profiles as the next efficiency lever."],"forward_implications":["Long-time macroscopic evolution of purely dissipative systems becomes predictable from short-time particle data alone, sidestepping the time-scale bottleneck of direct simulation.","The learned thermodynamic potential is identifiable rather than arbitrary: fluctuation statistics select the dissipative operator compatible with the particle process, so the pair $(K_z, F[z])$ is not just one of many fits to the same mean dynamics.","Epistemic uncertainty is quantified for the operator, the free energy, and the propagated dynamics, so continuum predictions carry practical error bars that cover the particle-simulation reference points in the paper's examples.","Where analytic coarse-grained thermodynamics fail qualitatively, as in strong short-range interactions, the data-driven framework can still produce a single-well free energy and correct kinetics.","Computational cost shifts from the particle-simulation timescale (days to years) to minutes of training and prediction, even including the many realizations needed for uncertainty."],"supporting_citations":[{"why":"proves the infinite-dimensional fluctuation-dissipation relation that lets operator entries be read from rescaled fluctuation covariances","marker":"[19]"},{"why":"supplies the Stat-PINNs strategy, the structure-preserving operator parameterization, and the kinetic Monte Carlo training datasets reused here","marker":"[18]"},{"why":"introduces the denoising diffusion probabilistic models used as the conditional base networks","marker":"[20]"},{"why":"introduces epistemic neural networks and the epinet architecture used for uncertainty quantification","marker":"[21]"},{"why":"provides the DDIM sampler that reduces reverse diffusion steps to two for cheap uncertainty sampling","marker":"[44]"},{"why":"derives the analytic long-range mesoscopic model that serves as the ground-truth benchmark for long-range interactions","marker":"[48]"},{"why":"supplies the kinetic Monte Carlo algorithm used to generate particle data and reference long-time evolution","marker":"[54]"},{"why":"knowledge distillation is used to train the epinets against frozen base-network predictions","marker":"[47]"}],"fun_headline_variants":["Short runs, long-time laws: SPIEDiff quantifies uncertainty","Fluctuation data breaks thermodynamic non-uniqueness in SPIEDiff","Physics-informed diffusion: minutes instead of years of particle runs","Epistemic uncertainty captured in learned macroscopic dynamics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The formula that turns fluctuation data into the dissipative operator assumes the particle density obeys an Itô stochastic differential equation driven by white space-time noise and that the system is in local equilibrium when the measurements are taken; if either fails, the estimated operator is biased.","fun_headline_variants_meta":{"raw":{"variants":["Short runs, long-time laws: SPIEDiff quantifies uncertainty","Fluctuation data breaks thermodynamic non-uniqueness in SPIEDiff","Physics-informed diffusion: minutes instead of years of particle runs","Epistemic uncertainty captured in learned macroscopic dynamics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000975,"raw_usage":{"total_tokens":4109,"prompt_tokens":874,"completion_tokens":3235,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":3166}},"tokens_in":490,"tokens_out":3235,"duration_ms":24751,"temperature":1.0,"reasoning_tokens":3166,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:02:48.326549+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the fluctuation-dissipation estimator on a particle model with controlled colored noise or with measurements taken before local equilibration, and compare the inferred operator to direct estimates from longer-time statistics; a systematic dependence of the inferred entries on the measurement window $h$ (or a visibly biased free energy against equilibrium sampling) would falsify the premise that short-time fluctuations are white and locally equilibrated.","supporting_citations":[{"cited_title":"Harnessing fluctuations to discover dissipative evolution equations","cited_arxiv_id":null,"evidence_quote":"proves the infinite-dimensional fluctuation-dissipation relation that lets operator entries be read from rescaled fluctuation covariances"},{"cited_title":"Statistical-Physics-Informed Neural Networks (Stat-PINNs): A machine learning strategy for coarse-graining dissipative dynamics","cited_arxiv_id":null,"evidence_quote":"supplies the Stat-PINNs strategy, the structure-preserving operator parameterization, and the kinetic Monte Carlo training datasets reused here"},{"cited_title":"Epistemic neural networks","cited_arxiv_id":null,"evidence_quote":"introduces epistemic neural networks and the epinet architecture used for uncertainty quantification"},{"cited_title":"Derivation and validation of mesoscopic theories for diffusion of interacting molecules","cited_arxiv_id":null,"evidence_quote":"derives the analytic long-range mesoscopic model that serves as the ground-truth benchmark for long-range interactions"},{"cited_title":"41HFfV0NtIe8IWM5ZdD2pkGyZyk=","cited_arxiv_id":null,"evidence_quote":"supplies the kinetic Monte Carlo algorithm used to generate particle data and reference long-time evolution"}],"review_version":1}