Pith. sign in

REVIEW 2 major objections 5 minor 78 references

Diffusion-path proposals with a Metropolis correction sample multimodal targets without distorting mode weights.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 04:14 UTC pith:HQNTYTYT

load-bearing objection Solid, usable MCMC paper: exact path-space MH on diffusion proposals plus a clean spectral-gap argument that tempering lacks for unequal modes. the 2 major comments →

arxiv 2607.11631 v1 pith:HQNTYTYT submitted 2026-07-13 stat.CO stat.ML

Markov Chain Monte Carlo with Diffusion Paths

classification stat.CO stat.ML MSC 62F1565C0560J22
keywords Markov chain Monte Carlodiffusion pathmultimodal samplingMetropolis–Hastingsscore estimationspectral gaptemperingBayesian posterior
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Classical MCMC struggles with well-separated modes, and the usual fix—tempering—can exponentially warp the relative weights of modes that differ in scale, so the chain mixes torpidly. This paper replaces the tempering bridge with the marginal path of a noising diffusion that carries the target to a Gaussian; that path keeps mixture weights intact and has a spectral gap controlled by inter-mode distance and per-mode Poincaré constants rather than scale disparity. Because the intermediate scores are unknown, they are learned approximately and the continuous path is discretized; the resulting bias is removed by a Metropolis–Hastings step on the whole forward–backward trajectory, which leaves the true target invariant no matter how inaccurate the score or the step size. Acceptance rates are then governed by discretization error and an integrated squared score error, giving concrete tuning rules. On mixtures with unequal variances, skew-normals, Bayesian mixture models, sensor localization, and multimodal profile likelihoods, the corrected sampler recovers mode weights more accurately and crosses modes more often than tempering-based competitors or the unadjusted diffusion sampler.

Core claim

The Metropolis-adjusted diffusion path (MAD-Path) transition is reversible with respect to the target for any approximate intermediate score and any Euler–Maruyama discretization, while the ideal continuous diffusion-path kernel preserves mixture weights and, for unequal-variance Gaussian mixtures with diffusion horizon scaling as log d, has a spectral gap bounded away from zero independently of dimension.

What carries the argument

The MAD-Path proposal: evolve the current state forward under a discretized noising SDE, then reverse under an approximate score, and accept or reject the entire path via a Metropolis–Hastings ratio on the augmented trajectory space (a volume-preserving involution), guaranteeing detailed balance with the target.

Load-bearing premise

Practical acceptance rates depend on first finding the modes and learning intermediate scores via a reverse-KL path-space objective built around a Gaussian-mixture reference; if mode discovery fails or the objective collapses modes, efficiency collapses even though invariance still holds.

What would settle it

On the unequal-variance two-Gaussian mixture in growing dimension, replace tempering with MAD-Path using a correctly learned score and α ∼ 1/d: if the empirical spectral gap (or mode-crossing rate) still decays exponentially rather than staying order-one, the claimed advantage is false.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper proposes MAD-Path, an MCMC method that uses forward–backward diffusion paths as nonlocal proposals and corrects them with a Metropolis–Hastings step on an augmented path space. The ideal continuous diffusion-path kernel is shown to be reversible with respect to the target and to enjoy favorable spectral gaps: under a Poincaré inequality (Theorem 2.3), and for mixtures of Poincaré components with an explicit two-component bound that is insensitive to relative mode weights and scales (Theorem 2.5, Corollary 2.6). In practice, intermediate scores are learned variationally and the SDEs are discretized; Algorithm 1 and Theorem 3.1 establish that the resulting Metropolis-adjusted kernel remains reversible for any approximate score and any Euler–Maruyama step size. Theorems 4.1–4.2 quantify how discretization and score error enter the acceptance probability and yield a high-dimensional tuning guideline. Experiments on unequal-variance Gaussian mixtures, skew-normal mixtures, Bayesian GMMs, sensor localization, and a multimodal SUR profile likelihood show improved mode exploration and weight recovery relative to tempering-based MCMC and unadjusted diffusion samplers.

Significance. The work cleanly separates three contributions that matter for multimodal MCMC: (i) a weight-preserving interpolating path with a spectral-gap analysis that explains why it can succeed where tempering fails on asymmetric modes; (ii) a path-space Metropolis correction that restores exact invariance without requiring tractable endpoint densities; and (iii) quantitative acceptance expansions that give practical step-size and score-accuracy guidance. Theorem 3.1 is particularly useful: correctness is unconditional on score quality, so future improvements in score learning plug in directly. The ideal-kernel analysis (maximal correlation / data-augmentation viewpoint, mixture decomposition, dimension-free gap for α∼d^{-1}) is a genuine addition relative to proximal-sampler literature. Code and detailed experimental protocols are provided. The main limitation—that mixing theory for the adjusted chain and practical efficiency still hinge on mode-aware score learning—is scoped honestly in Sections 5 and 7 and does not undercut the stated claims.

major comments (2)
  1. The spectral-gap results (Theorems 2.3–2.5, Corollary 2.6) apply only to the ideal continuous kernel P, not to the Metropolis-adjusted discretized kernel used in practice. Section 7 correctly flags this as open, but the abstract and introduction sometimes read as if the favorable mixing properties transfer immediately to MAD-Path. A short, explicit caveat in the abstract and at the end of §2 (that the gap bounds motivate the path choice, while acceptance and mixing of the adjusted chain remain separate) would prevent over-reading.
  2. Practical performance is tightly coupled to the two-stage score pipeline in §5 / Appendix C.1 (annealed Langevin mode hunt → GMM reference → reverse-KL control). Theorem 4.2 already shows that acceptance is governed by the symmetrized path KL E_score; the experiments demonstrate MH recovering weights when the unadjusted sampler is biased. Still, the paper would be stronger with one controlled failure-mode experiment (or a clear negative result) when mode discovery is incomplete or the reverse-KL collapses a mode, so readers can see how acceptance and ESS degrade rather than only success cases.
minor comments (5)
  1. Notation for the reverse process switches between Yt, bYt, and the Euler iterates yk; a short notation table early in §3 would help.
  2. Figure 2 caption and Remark 4.1: the coincidence of the 0.234 acceptance optimum with RWM is interesting but easy to misread as implying random-walk behavior; the remark already clarifies this—consider elevating one sentence into the main text near the figure.
  3. In §6.1–6.2, computational budgets are matched by target-score evaluations, which is appropriate for amortized MAD-Path, but wall-clock times are only in the appendix; a one-line summary in the main experimental section would aid comparison.
  4. Typo/consistency: “Poincar´ e” appears with a broken accent in several places (e.g., §2.2); standardize to “Poincaré”.
  5. Corollary 2.6 sets α=κ/d; a brief remark that T=Θ(log d) follows from α=e^{-T} would make the dimension-free claim more immediately usable for practitioners choosing the horizon.

Circularity Check

0 steps flagged

No significant circularity: invariance and spectral-gap claims are derived from detailed balance and maximal correlation, not fitted or self-defined.

full rationale

The paper's central results are self-contained mathematical derivations rather than tautologies. Theorem 3.1 establishes reversibility of MAD-Path for arbitrary approximate scores and Euler–Maruyama discretizations by exhibiting the path-reversal involution R on the augmented measure Π and verifying detailed balance; the argument does not assume score accuracy or recover a training objective. Ideal-kernel spectral gaps (Lemma 2.2, Theorems 2.3–2.5, Corollary 2.6) follow from the classical maximal-correlation characterization of two-block Gibbs samplers plus Poincaré inequalities and Wasserstein overlap bounds; they are not obtained by fitting the target gap. Acceptance analysis (Theorems 4.1–4.2) expands the Metropolis–Hastings ratio under stated regularity and isolates discretization and integrated score error; the resulting O(h^{1/2}) and E_score expressions are asymptotic consequences, not redefinitions of the acceptance probability. Score learning (Section 5) uses a standard reverse-KL path-space control objective with a GMM reference; the Metropolis step remains valid for any learned score, so training does not circularly force the sampler's invariant distribution. Experiments compare external baselines (APT, TT, AIS, SMC, unadjusted diffusion) rather than tautologically recovering training losses. No self-definitional loop, fitted-input-as-prediction, load-bearing self-citation uniqueness claim, or renaming of a known result appears in the derivation chain. Score 0 is therefore appropriate.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 1 invented entities

The central invariance claim needs only standard MH detailed balance on an augmented path measure and volume-preserving path reversal. Mixing claims add Poincaré inequalities on components and Wasserstein separation. Acceptance asymptotics add strong regularity and strong-error assumptions on Euler–Maruyama. Practical performance adds a learned score from path-space VI with a hand-built GMM reference and free algorithmic knobs (T, N, h, network, mode-finding schedule).

free parameters (4)
  • diffusion horizon T
    Chosen per experiment (e.g. T=7 in fetal-deaths study); controls mode merging vs accumulated score error and is not derived from a uniqueness theorem.
  • discretization steps N / step size h
    Tuned for acceptance and cost; high-d product scaling suggests h≍n^{-1} and ~0.234 acceptance, but values remain free design choices.
  • GMM reference components and annealed-Langevin mode-finding schedule
    Number of modes K, temperature ladder, step sizes, merge radius, and minimal cluster size are set by hand to build pref and sref_t.
  • MLP control architecture and training budget
    Two-layer MLP for u, particle counts, and training iterations determine score quality and thus acceptance, without a canonical default.
axioms (5)
  • domain assumption Target density is continuously differentiable; forward/reverse SDEs are well-posed with polynomial-growth coefficients.
    Used for ideal kernel construction (Section 2) and Assumption 4.1 for acceptance analysis.
  • domain assumption Poincaré inequality on p or on each mixture component pc with constants CP,c.
    Load-bearing for Theorems 2.3 and 2.5 spectral-gap lower bounds.
  • domain assumption Euler–Maruyama strong error O(h^{1/2}) and uniform Lq moments for forward and reverse paths.
    Assumption 4.1(iii–iv); standard under global Lipschitz but stronger than many multimodal targets satisfy globally.
  • ad hoc to paper Memoryless condition: noising p and pref yield approximately the same terminal law at finite T.
    Section 5; needed so the variational family can recover the exact reverse process targeting p.
  • standard math Metropolis–Hastings detailed balance on path space with volume-preserving involution R.
    Theorem 3.1 proof; classical MH on augmented space.
invented entities (1)
  • MAD-Path transition (forward EM + approximate reverse EM + path-space MH) independent evidence
    purpose: Define a practical MCMC kernel that uses diffusion-path proposals while remaining exactly invariant to p.
    Algorithmic construction, not a physical entity; independent evidence is the reversibility proof and empirical acceptance/mixing, not an external measurement.

pith-pipeline@v1.1.0-grok45 · 43977 in / 3244 out tokens · 30128 ms · 2026-07-14T04:14:51.401348+00:00 · methodology

0 comments
read the original abstract

Sampling from multimodal distributions is a longstanding challenge for classical local Markov chain Monte Carlo (MCMC) methods. A popular remedy is to introduce a sequence of intermediate distributions that interpolate between the target and a simpler reference. The classical choice, tempering, raises the density to a power, but distorts the relative weights of asymmetric modes and can lead to poor mixing. We instead propose interpolating along the diffusion path, the marginals of a noising diffusion process that carries the target toward a Gaussian. This path preserves the relative weights of the modes and enjoys favorable mixing properties, which we make precise through a spectral-gap analysis of the corresponding ideal transition kernel. Sampling along the path requires its intermediate scores, which can be estimated from the unnormalized target through variational approaches, yielding only an approximate sampler. To remove the resulting bias, we introduce the Metropolis-adjusted diffusion path (MAD-Path) sampler, which corrects the diffusion-path proposal in an augmented path space and leaves the target invariant regardless of the accuracy of the learned score or the discretization error. We further quantify how these two errors affect the acceptance probability, providing guidance for practical tuning. Experiments on a range of Bayesian posteriors show that MAD-Path improves global exploration and mode-weight estimation relative to tempering-based MCMC methods and unadjusted diffusion samplers.

Figures

Figures reproduced from arXiv: 2607.11631 by Han Chen, Jun Yang, Sifan Liu.

Figure 1
Figure 1. Figure 1: Illustration of MAD-Path (Algorithm 1). From the current state x0 ∼ p0, the forward process (top) generates x1, . . . , xN over N discretization steps, transporting the multimodal target p0 (left) toward the approximately Gaussian marginal pT (right). The backward process (bottom) starts from the shared endpoint yN = xN and generates yN−1, . . . , y0, returning y0 as the proposed state. By smoothing the we… view at source ↗
Figure 2
Figure 2. Figure 2: Cost-normalized efficiency versus acceptance rate for the target [PITH_FULL_IMAGE:figures/full_fig_p015_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Results for the two-Gaussian mixture example. Top row: empirical distribution of the [PITH_FULL_IMAGE:figures/full_fig_p020_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Results for the mixture of skew normal example. First row: marginal distribution of the [PITH_FULL_IMAGE:figures/full_fig_p021_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Results for the Bayesian Gaussian mixture model, comparing MAD-Path and APT-DEO [PITH_FULL_IMAGE:figures/full_fig_p023_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: MAD-Path acceptance rate (left) and inter-mode crossing rate (right) with varying horizon [PITH_FULL_IMAGE:figures/full_fig_p024_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Burn-in phase (first 500 iterations) of MAD-Path (green), APT-DEO with initial state at [PITH_FULL_IMAGE:figures/full_fig_p025_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Marginal sample distributions for the SUR example. The top row shows MAD-Path; the [PITH_FULL_IMAGE:figures/full_fig_p026_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Trace plots of the transformed parameters under MAD-Path with varying horizon [PITH_FULL_IMAGE:figures/full_fig_p046_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: The left panel shows the true locations of the eleven sensors, with the three known [PITH_FULL_IMAGE:figures/full_fig_p047_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Trace plots for the first 500 iterations of MAD-Path and APT-DEO with varying [PITH_FULL_IMAGE:figures/full_fig_p047_11.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

78 extracted references · 5 linked inside Pith

  1. [1]

    Ahn, S., Chen, Y., and Welling, M. (2013). Distributed and adaptive darting Monte Carlo through regenerations. InArtificial Intelligence and Statistics, pages 108–116. PMLR

  2. [2]

    Albergo, M. S. and Vanden-Eijnden, E. (2025). NETS: A non-equilibrium transport sampler. In International Conference on Machine Learning, pages 1026–1055. PMLR

  3. [3]

    (2008).Gradient Flows in Metric Spaces and in the Space of Probability Measures

    Ambrosio, L., Gigli, N., and Savar´ e, G. (2008).Gradient Flows in Metric Spaces and in the Space of Probability Measures. Lectures in Mathematics ETH Z¨ urich. Birkh¨ auser, 2 edition

  4. [4]

    Anderson, B. D. (1982). Reverse-time diffusion equation models.Stochastic Processes and their Applications, 12(3):313–326

  5. [5]

    Andrieu, C., Lee, A., Power, S., and Wang, A. Q. (2022). Comparison of Markov chains via weak Poincar´ e inequalities with application to pseudo-marginal MCMC.The Annals of Statistics, 50(6):3592–3618

  6. [6]

    F., Roberts, G

    Atchad´ e, Y. F., Roberts, G. O., and Rosenthal, J. S. (2011). Towards optimal scaling of Metropolis-coupled Markov chain Monte Carlo.Statistics and Computing, 21(4):555–568. 47

  7. [7]

    and Bowman, A

    Azzalini, A. and Bowman, A. W. (1990). A look at some data on the Old Faithful geyser.Journal of the Royal Statistical Society: Series C, 39(3):357–365

  8. [8]

    Berner, J., Richter, L., and Ullrich, K. (2024). An optimal control perspective on diffusion-based generative modeling.Transactions on Machine Learning Research

  9. [9]

    Beskos, A., Pillai, N., Roberts, G., Sanz-Serna, J.-M., and Stuart, A. (2013). Optimal tuning of the hybrid Monte Carlo algorithm.Bernoulli, 19(5A):1501–1534

  10. [10]

    Blessing, D., Qu, X., Neitz, A., and Ziesche, T. (2024). Beyond ELBOs: a large-scale evaluation of variational methods for sampling. InAdvances in Neural Information Processing Systems (NeurIPS)

  11. [11]

    Blessing, D., Richter, L., Berner, J., Malitskiy, E., and Neumann, G. (2026). Bridge matching sampler: Scalable sampling via generalized fixed-point diffusion matching.arXiv preprint arXiv:2603.00530

  12. [12]

    Bradley, R. C. (2005). Basic properties of strong mixing conditions. a survey and some open questions.Probability Surveys, 2:107–144

  13. [13]

    P., Morgan, B

    Brooks, S. P., Morgan, B. J., Ridout, M. S., and Pack, S. (1997). Finite mixture models for proportions.Biometrics, pages 1097–1115

  14. [14]

    and Dembo, A

    Bryc, W. and Dembo, A. (2005). On the maximum correlation coefficient.Theory of Probability & Its Applications, 49(1):132–138

  15. [15]

    Chemseddine, J., Wald, C., Duong, R., and Steidl, G. (2025). Neural sampling from Boltzmann densities: Fisher-Rao curves in the Wasserstein geometry. InThe Thirteenth International Conference on Learning Representations

  16. [16]

    Chen, S., Chewi, S., Li, J., Li, Y., Salim, A., and Zhang, A. R. (2023). Sampling is as easy as learning the score: theory for diffusion models with minimal data assumptions. InInternational Conference on Learning Representations

  17. [17]

    Chen, Y., Chewi, S., Salim, A., and Wibisono, A. (2022). Improved analysis for a proximal algorithm for sampling. InConference on Learning Theory, pages 2984–3014. PMLR

  18. [18]

    Choi, J., Chen, Y., Tao, M., and Liu, G.-H. (2026). Non-equilibrium annealed adjoint sampler. Advances in Neural Information Processing Systems, 38:90956–90988

  19. [19]

    De Bortoli, V., Hutchinson, M., Wirnsberger, P., and Doucet, A. (2024). Target score matching. arXiv preprint arXiv:2402.08667

  20. [20]

    Del Moral, P., Doucet, A., and Jasra, A. (2006). Sequential Monte Carlo samplers.Journal of the Royal Statistical Society Series B: Statistical Methodology, 68(3):411–436

  21. [21]

    Dembo, A., Kagan, A., and Shepp, L. A. (2001). Remarks on the maximum correlation coefficient.Bernoulli, pages 343–350

  22. [22]

    and Richardson, T

    Drton, M. and Richardson, T. S. (2004). Multimodality of the likelihood in the bivariate seemingly unrelated regressions model.Biometrika, 91(2):383–392. 48

  23. [23]

    D., Pendleton, B

    Duane, S., Kennedy, A. D., Pendleton, B. J., and Roweth, D. (1987). Hybrid Monte Carlo. Physics Letters B, 195(2):216–222

  24. [24]

    M., and Vanden-Eijnden, E

    Gabri´ e, M., Rotskoff, G. M., and Vanden-Eijnden, E. (2022). Adaptive Monte Carlo augmented with normalizing flows.Proceedings of the National Academy of Sciences, 119(10):e2109420119

  25. [25]

    Geyer, C. J. (1991). Markov chain Monte Carlo maximum likelihood. InComputing Science and Statistics: Proceedings of the 23rd Symposium on the Interface, pages 156–163

  26. [26]

    Geyer, C. J. and Thompson, E. A. (1995). Annealing Markov chain Monte Carlo with applications to ancestral inference.Journal of the American Statistical Association, 90:909–920

  27. [27]

    J., Salmond, D

    Gordon, N. J., Salmond, D. J., and Smith, A. F. (1993). Novel approach to nonlinear/non- Gaussian Bayesian state estimation. InIEE Proceedings F (Radar and Signal Processing), volume 140, pages 107–113. IET

  28. [28]

    and Noble, M

    Grenioux, L. and Noble, M. (2026). Diffusion-based annealed Boltzmann generators: benefits, pitfalls and hopes.arXiv preprint arXiv:2601.21026

  29. [29]

    Grenioux, L., Noble, M., Gabri´ e, M., and Durmus, A. (2024). Stochastic localization via iterative posterior sampling. InProceedings of the 41st International Conference on Machine Learning, pages 16337–16376

  30. [30]

    J., Miller, B

    Havens, A. J., Miller, B. K., Yan, B., Domingo-Enrich, C., Sriram, A., Levine, D. S., Wood, B. M., Hu, B., Amos, B., Karrer, B., et al. (2025). Adjoint sampling: Highly scalable diffusion samplers via adjoint matching. InInternational Conference on Machine Learning, pages 22204– 22237. PMLR

  31. [31]

    He, J., Du, Y., Vargas, F., Zhang, D., Padhy, S., OuYang, R., Gomes, C., and Hern´ andez-Lobato, J. M. (2025). No trick, no treat: Pursuits and challenges towards simulation-free training of neural samplers.arXiv preprint arXiv:2502.06685

  32. [32]

    Ho, J., Jain, A., and Abbeel, P. (2020). Denoising diffusion probabilistic models.Advances in Neural Information Processing Systems, 33:6840–6851

  33. [33]

    V., Langmore, I., Tran, D., and Vasudevan, S

    Hoffman, M., Sountsov, P., Dillon, J. V., Langmore, I., Tran, D., and Vasudevan, S. (2019). Neutra-lizing bad geometry in Hamiltonian Monte Carlo using neural transport.arXiv preprint arXiv:1903.03704

  34. [34]

    Hoffman, M. D. and Gelman, A. (2014). The no-U-turn sampler: Adaptively setting path lengths in Hamiltonian Monte Carlo.Journal of Machine Learning Research, 15(47):1593–1623

  35. [35]

    Huang, X., Dong, H., Hao, Y., Ma, Y.-A., and Zhang, T. (2024). Reverse diffusion Monte Carlo. InInternational Conference on Learning Representations

  36. [36]

    T., Fisher III, J

    Ihler, A. T., Fisher III, J. W., Moses, R. L., and Willsky, A. S. (2004). Nonparametric belief propagation for self-calibration in sensor networks. InProceedings of the 3rd International Symposium on Information Processing in Sensor Networks, pages 225–233

  37. [37]

    C., and Stephens, D

    Jasra, A., Holmes, C. C., and Stephens, D. A. (2005). Markov chain Monte Carlo methods and the label switching problem in Bayesian mixture modeling.Statistical Science, 20:50–67. 49

  38. [38]

    Kloeden, P. E. and Platen, E. (1992).Numerical Solution of Stochastic Differential Equations. Springer Berlin Heidelberg

  39. [39]

    Lan, S., Streets, J., and Shahbaba, B. (2014). Wormhole Hamiltonian Monte Carlo. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 28

  40. [40]

    T., and Stumpf-F´ etizon, T

    Latuszy´ nski, K., Moores, M. T., and Stumpf-F´ etizon, T. (2026). MCMC methods for multi- modal distributions. InHandbook of Markov Chain Monte Carlo, pages 368–396. Chapman and Hall/CRC

  41. [41]

    T., Shen, R., and Tian, K

    Lee, Y. T., Shen, R., and Tian, K. (2021). Structured logconcave sampling with a restricted Gaussian oracle. InConference on Learning Theory, pages 2993–3050. PMLR

  42. [42]

    S., Wong, W

    Liu, J. S., Wong, W. H., and Kong, A. (1994). Covariance structure of the Gibbs sampler with applications to the comparisons of estimators and augmentation schemes.Biometrika, pages 27–40

  43. [43]

    and Parisi, G

    Marinari, E. and Parisi, G. (1992). Simulated tempering: a new Monte Carlo scheme.EPL (Europhysics Letters), 19(6):451

  44. [44]

    and Fleuret, F

    M´ at´ e, B. and Fleuret, F. (2023). Learning interpolations between Boltzmann densities. Transactions on Machine Learning Research

  45. [45]

    W., Rosenbluth, M

    Metropolis, N., Rosenbluth, A. W., Rosenbluth, M. N., Teller, A. H., and Teller, E. (1953). Equation of state calculations by fast computing machines.The Journal of Chemical Physics, 21(6):1087–1092

  46. [46]

    Miasojedow, B., Moulines, E., and Vihola, M. (2013). An adaptive parallel tempering algorithm. Journal of Computational and Graphical Statistics, 22(3):649–664

  47. [47]

    Neal, R. M. (1996). Sampling from multimodal distributions using tempered transitions. Statistics and Computing, 6(4):353–366

  48. [48]

    Neal, R. M. (2001). Annealed importance sampling.Statistics and Computing, 11(2):125–139

  49. [49]

    Neal, R. M. (2011). MCMC using Hamiltonian dynamics. In Brooks, S., Gelman, A., Jones, G., and Meng, X.-L., editors,Handbook of Markov Chain Monte Carlo. Chapman and Hall/CRC

  50. [50]

    Noble, M., Grenioux, L., Gabri´ e, M., and Oliviero Durmus, A. (2025). Learned reference- based diffusion sampler for multi-modal distributions. InInternational Conference on Learning Representations, volume 2025, pages 47046–47100

  51. [51]

    and Richter, L

    N¨ usken, N. and Richter, L. (2021). Solving high-dimensional Hamilton–Jacobi–Bellman PDEs using neural networks: perspectives from the theory of controlled diffusions and measures on path space.Partial Differential Equations and Applications, 2(4):1–48

  52. [52]

    Okabe, T., Kawata, M., Okamoto, Y., and Mikami, M. (2001). Replica-exchange Monte Carlo method for the isobaric–isothermal ensemble.Chemical Physics Letters, 335(5-6):435–439

  53. [53]

    J., Mohamed, S., and Lakshminarayanan, B

    Papamakarios, G., Nalisnick, E., Rezende, D. J., Mohamed, S., and Lakshminarayanan, B. (2021). Normalizing flows for probabilistic modeling and inference.Journal of Machine Learning Research, 22(57):1–64. 50

  54. [54]

    J., De Bortoli, V., Deligiannidis, G., and Doucet, A

    Phillips, A., Dau, H.-D., Hutchinson, M. J., De Bortoli, V., Deligiannidis, G., and Doucet, A. (2024). Particle denoising diffusion sampler. InInternational Conference on Machine Learning

  55. [55]

    Pompe, E., Holmes, C., and Latuszy´ nski, K. (2020). A framework for adaptive MCMC targeting multimodal distributions.Annals of Statistics, 48(5):2930–2952

  56. [56]

    Qin, Q., Ju, N., and Wang, G. (2025). Spectral gap bounds for reversible hybrid Gibbs chains. The Annals of Statistics, 53(4):1613–1638

  57. [57]

    and Berner, J

    Richter, L. and Berner, J. (2024). Improved sampling via learned diffusions. InThe Twelfth International Conference on Learning Representations

  58. [58]

    O., Gelman, A., and Gilks, W

    Roberts, G. O., Gelman, A., and Gilks, W. R. (1997). Weak convergence and optimal scaling of random walk Metropolis algorithms.The Annals of Applied Probability, 7(1):110–120

  59. [59]

    Roberts, G. O. and Rosenthal, J. S. (2001). Optimal scaling for various Metropolis–Hastings algorithms.Statistical Science, 16(4):351–367

  60. [60]

    Roberts, G. O. and Tweedie, R. L. (1996). Exponential convergence of Langevin distributions and their discrete approximations.Bernoulli, 2(3):341–363

  61. [61]

    and Ermon, S

    Song, Y. and Ermon, S. (2019). Generative modeling by estimating gradients of the data distribution.Advances in Neural Information Processing Systems, 32

  62. [62]

    P., Kumar, A., Ermon, S., and Poole, B

    Song, Y., Sohl-Dickstein, J., Kingma, D. P., Kumar, A., Ermon, S., and Poole, B. (2021). Score- based generative modeling through stochastic differential equations.International Conference on Learning Representations

  63. [63]

    Srivastava, V. K. and Giles, D. E. (1987).Seemingly unrelated regression equations models: Estimation and inference, volume 80. CRC Press

  64. [64]

    Sun, J., Berner, J., Richter, L., Zeinhofer, M., M¨ uller, J., Azizzadenesheli, K., and Anandkumar, A. (2024). Dynamical measure transport and neural PDE solvers for sampling.arXiv preprint arXiv:2407.07873

  65. [65]

    Sundberg, R. (2010). Flat and multimodal likelihoods and model lack of fit in curved exponential families.Scandinavian Journal of Statistics, 37

  66. [66]

    Swendsen, R. H. and Wang, J.-S. (1986). Replica Monte Carlo simulation of spin-glasses. Physical Review Letters, 57(21):2607

  67. [67]

    and Lu, J

    Tan, L. and Lu, J. (2025). Accelerate Langevin sampling with birth-death process and exploration component.SIAM/ASA Journal on Uncertainty Quantification, 13(3):1265–1293

  68. [68]

    G., Roberts, G

    Tawn, N. G., Roberts, G. O., and Rosenthal, J. S. (2018). Weight-preserving simulated tempering.Statistics and Computing, 30:27 – 41

  69. [69]

    G., Roberts, G

    Tawn, N. G., Roberts, G. O., and Rosenthal, J. S. (2020). Annealed leap-point sampler for multimodal target distributions.arXiv preprint arXiv:2012.15830

  70. [70]

    and Hegstad, B

    Tjelmeland, H. and Hegstad, B. K. (2001). Mode jumping proposals in MCMC.Scandinavian Journal of Statistics, 28(1):205–223. 51

  71. [71]

    S., and Doucet, A

    Vargas, F., Grathwohl, W. S., and Doucet, A. (2023). Denoising diffusion samplers. InThe Eleventh International Conference on Learning Representations

  72. [72]

    Vargas, F., Padhy, S., Blessing, D., and N¨ usken, N. (2024). Transport meets variational inference: Controlled Monte Carlo diffusions. InThe Twelfth International Conference on Learning Representations

  73. [73]

    Wibisono, A. (2026). Mixing time of the proximal sampler in relative Fisher information via strong data processing inequality.IEEE Transactions on Information Theory

  74. [74]

    B., Schmidler, S

    Woodard, D. B., Schmidler, S. C., and Huber, M. (2009). Sufficient conditions for torpid mixing of parallel and simulated tempering.Electronic Journal of Probability, 14:780–804

  75. [75]

    O., and Rosenthal, J

    Yang, J., Roberts, G. O., and Rosenthal, J. S. (2020). Optimal scaling of random-walk Metropolis algorithms on general target distributions.Stochastic Processes and their Applications, 130(10):6094–6132

  76. [76]

    Zellner, A. (1962). An efficient method of estimating seemingly unrelated regressions and tests for aggregation bias.Journal of the American statistical Association, 57(298):348–368

  77. [77]

    and Chen, Y

    Zhang, Q. and Chen, Y. (2022). Path integral sampler: A stochastic control approach for sampling. InInternational Conference on Learning Representations (ICLR)

  78. [78]

    M., and Aston, J

    Zhou, Y., Johansen, A. M., and Aston, J. A. (2016). Toward automatic model comparison: an adaptive sequential Monte Carlo approach.Journal of Computational and Graphical Statistics, 25(3):701–726. 52