REVIEW 1 major objections 4 minor 1 cited by
On importance sampling and independent Metropolis-Hastings with an unbounded weight function
T0 review · 1 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read For independent Metropolis–Hastings, sharing one proposal and one uniform variate between two chains gives a coupling that is exactly maximal at every time horizon, and this equality drives the paper's polynomial bounds and unbiased…
desk verdict A strong paper with a genuinely new maximal coupling theorem and clean TV bounds, but one unproved proposition about the symmetrized debiased estimator that needs to be addressed before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the common random numbers coupling of IMH/PIMH, which runs two chains from states x and y using the same proposal x* and the same uniform variable u, so they meet exactly when both accept the proposal; this happens with probability governed by the state with the larger average weight. Theorem 3.1 shows this coupling is maximal for every horizon, giving the exact identity |P^t(x, ·) − P^t(y, ·)|_TV = max(r(x), r(y))^t. The quantitative bounds rest on controlling the expected rejection probability E[r(x)^t] through a Paley–Zygmund anticoncentration inequality, a peeling argument over the range of the average weight, and Markov's inequality on the tails.
What would settle it
Run the common-draws coupling from two states x and y of a PIMH chain with N = 1, estimate r(x) and r(y) directly, and compare the empirical tail P(τ > t) with max(r(x), r(y))^t for t = 1, ..., T; any significant mismatch for a non-atomic proposal would disprove Theorem 3.1. The same check can be repeated for finite N using the average-weight version of r.
Extended reading notes
Core claim
Under the assumption that the target is absolutely continuous with respect to the proposal and the weight ω = π/q integrates to one, the paper proves that the common random numbers coupling of PIMH is maximal for all time: for every t ≥ 1, |P^t(x, ·) − P^t(y, ·)|_TV = P_{x,y}(τ > t) = max(r(x), r(y))^t, where r(x) is the probability that a proposal is rejected from state x. This is, the paper argues, the first known case of an all-time maximal coupling of an MCMC algorithm. From this identity, Proposition 4.3 gives P(τ > t) ≤ C/(√N t^p) under Assumption 2, and Theorem 4.1 bounds the total variation distance of a PIMH chain to its target by C/(√N (1+t)^{p−1}). The paper then uses the lagged coupling to construct unbiased estimators of self-normalized importance sampling targets, analyses their moments, and shows that a symmetrized version has asymptotic inefficiency equal to that of importance sampling as N → ∞.
Load-bearing premise
The load-bearing premise is that the importance weight has a finite p-th moment under the proposal for some p ≥ 2; the paper's polynomial convergence rates and finite-moment unbiasedness all depend on that assumption, and with only a first moment the comparison between IS and IMH biases breaks down.
Editorial extensions
If this is right
- If the weight has a finite p-th moment under the proposal for p ≥ 2, the total variation distance of the PIMH chain to its target is O(N^{-1/2}(1+t)^{-(p-1)}) after t iterations with N particles, with an example showing this t-rate cannot be improved beyond polylogarithmic factors.
- For standard IMH (N = 1), convergence from any starting point x is bounded by (r(x))^t + D t^{-(p-1)} under a p-th moment condition with p > 1.
- For bounded test functions and p > 2, the bias after N iterations of IMH is asymptotically smaller than that of sampling-importance resampling, while IS still wins on asymptotic variance.
- The unbiased estimator built from the lagged coupling has s finite moments when p > s, and under stated moment conditions its mean squared error is asymptotically equivalent to that of self-normalized importance sampling.
- The symmetrized version of the unbiased estimator achieves the same asymptotic inefficiency as importance sampling as N grows, and it can be combined with robust mean estimators to obtain sub-Gaussian concentration even with heavy-tailed weights.
Reading between the lines
- A direct consequence of the exact equality in Theorem 3.1 is a practical diagnostic: from a single run one can estimate r(x) and immediately obtain a computable upper and lower bound on the chain's total variation distance to stationarity at every horizon.
- The sharp transition at p = 2 in the bias comparison between IMH and IS suggests an intrinsic phase boundary: with only a finite second moment the ordering may depend on constants, and it would be worth testing empirically where debiased estimators still help.
- The same common-draws identity could serve as a testbed for other proposal-based algorithms, such as pseudo-marginal or particle Gibbs kernels, to see whether their natural couplings are also exactly maximal or how far they are from maximality.
- The authors compare IMH against self-normalized IS and sampling-importance resampling; an implicit extension is to compare against other resampling schemes, where the r(x) formula could be used to quantify finite-time bias directly.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper studies importance sampling (IS), independent Metropolis-Hastings (IMH), and the particle version PIMH when the weight function is unbounded but has finite moments. It gives an asymptotic bias formula for self-normalized IS (Theorem 2.1), moment bounds for IS (Theorem 2.3), proves that the common-draws coupling of IMH/PIMH is maximal at every time horizon (Theorem 3.1), and derives polynomial total-variation bounds in iteration count t and particle number N (Propositions 4.1-4.4, Theorem 4.1). The paper then constructs debiased estimators UIS and SUIS, states moment and asymptotic-efficiency results for them, and combines SUIS with robust mean estimators. Numerical experiments illustrate the rates and the behavior of the proposed estimators.
Significance. If correct, the maximal-coupling result is a notable contribution: it is one of the few MCMC couplings known to be exactly maximal for all time horizons, and it gives sharp analytical control of the meeting time. The explicit polynomial bounds with dependence on both t and N are useful for understanding PIMH under weak assumptions, and the unbiased IS framework under only finite moments is a promising contribution to bias removal. The paper is largely self-contained: the main theorems are proved in the appendix, and the numerical experiments are consistent with the proved rates. However, the headline efficiency claim for the symmetrized estimator SUIS is not proved, so the full significance of Section 5.4 is not yet established.
major comments (1)
- [Section 5.4, Proposition 5.6] Proposition 5.6 is a headline claim, but it is explicitly 'stated without a proof.' The assertion that SUIS has the same asymptotic inefficiency as IS requires more than a mild variation of Proposition 5.4: the estimator in (40) is an average of A_u(x0,y0,ζ) and A_u(y0,x0,ζ), where one of the two terms is only the cheap IS estimator selected by the relative order of Z(x0) and Z(y0). To conclude that the variance is asymptotically σ_IS^2/(2N) and that the expected cost is 2N, the paper must control the covariance between the full UIS term and the cheap IS term, as well as the effect of the selection indicator; this is not derived. In addition, the numerical experiment in Example 6 uses p=3 with a bounded f, a regime in which the hypotheses of Proposition 5.4 (with r=∞ the condition 2p+4r+4<rp requires p>4) are not satisfied. Please supply a proof of Proposition 5.6, state and prove the required covariance argument, and either provide a theorem covering the p=3 experiment or clearly label that curve as heuristic.
minor comments (4)
- [Section 2, Theorem 2.1] The sentence after (15), claiming that the inverse-moment assumption q(ω^{-η})<∞ 'may be removed at the cost of higher positive moments,' is unsupported. The proof as written relies on Proposition A.1, which uses the inverse-moment condition. Please provide a proof or a reference that states the required positive moment order, or qualify the claim.
- [Equation (24) and Lemma A.4] The definition of β_p is difficult to parse because the exponent in the factor multiplying q(ω^p)^{1/(p-1)} is typeset ambiguously. Write it explicitly, e.g. β_p = 1 - 2^{-(3p-2)/(p-1)} q(ω^p)^{-1/(p-1)}, to avoid ambiguity.
- [Abstract and Section 1.1] There are grammatical errors: 'we find the latter to be have a smaller bias' should be 'we find the latter to have a smaller bias,' and the sentence 'This yields an unbiased estimator that can be implemented whenever self-normalized importance sampling, and thus relates to the second question above' is missing a main verb.
- [Section 5.5.2, Tables 1-2] The numerical comparisons report MSEs without Monte Carlo standard errors or confidence intervals. Some differences, such as between MoM-SNIS and MN-SUIS for p=2.01, are small and may be within simulation noise; reporting standard errors would strengthen the comparisons.
Circularity Check
Score 0: no circularity; the derivation is self-contained, though Proposition 5.6 on SUIS efficiency is stated without proof.
full rationale
No significant circularity. The paper's central chain is self-contained: Assumptions 1-2 fix the model; Proposition 1.1 bounds moments of the average weight; Theorem 3.1 proves equality of the total variation distance with the common-draws coupling tail by explicit lower-bound arguments in Appendix A.3, rather than importing the equality; Proposition 4.2 derives polynomial meeting-time bounds via Paley-Zygmund and a peeling argument; the Section 5 unbiased estimator inherits unbiasedness from telescoping sums, not from an assumed conclusion. Constants such as beta_p, A_p, and C_p are either displayed explicitly or obtained from inequalities with stated dependence on q(omega^p); no constant is fitted to the numerical outputs, and no prediction is used to calibrate the derivation. The self-citations to Deligiannidis and Lee (2018), Middleton et al. (2019), and Wang et al. (2021) supply context or earlier algorithm components; the load-bearing equality of Theorem 3.1 is proved in the appendix rather than reduced to those citations. The one flagged gap is Section 5.4: Proposition 5.6, the asymptotic-efficiency equivalence of SUIS, is explicitly 'stated without a proof.' That is a missing-support or correctness issue, not a circularity, because the claimed limit is not used to define or fit any input, and no equation in the paper reduces to it. The numerical experiments illustrate rather than establish that limit, and some runs use parameter regimes outside Proposition 5.4's sufficient conditions. These concerns lower confidence in the SUIS recommendation but do not make the derivation circular, so the circularity score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption Assumption 1: pi is absolutely continuous with respect to q, the weight omega = dpi/dq is positive q-a.s., and q(omega) = 1.
- domain assumption Assumption 2: q(omega^p) < infinity for some p >= 2.
- domain assumption Theorem 2.1 assumes q(omega^{-eta}) < infinity for some eta > 0, q(|f-pi(f)|omega) < infinity, and q(|f-pi(f)|omega^3) < infinity.
- standard math Standard probability inequalities: Marcinkiewicz-Zygmund, Paley-Zygmund, Minkowski, Holder, Markov, and strong law of large numbers.
- standard math Existing results used as black boxes: Agapiou et al. (2017) IS bias bounds, Deligiannidis & Lee (2018) variance comparison, Roberts & Rosenthal (2004) convergence, Glynn & Rhee (2014) debiasing, Middleton et al. (2019) coupled PIMH, Dau (2022) MoM-SNIS.
Cite this review
Pith. "Pith review of On importance sampling and independent Metropolis-Hastings with an unbounded weight function." pith.science (2026). https://pith.science/paper/BBVEMDZR
@misc{pith2026241109514,
author = {Pith},
title = {Pith review of: On importance sampling and independent Metropolis-Hastings with an unbounded weight function},
year = {2026},
howpublished = {\url{https://pith.science/paper/BBVEMDZR}},
note = {Machine review of arXiv:2411.09514}
}
read the original abstract
Importance sampling and independent Metropolis-Hastings are among the fundamental building blocks of Monte Carlo methods. Both require a proposal distribution that globally approximates the target distribution, and pointwise evaluation of the Radon-Nikodym derivative of the target distribution relative to the proposal, also called the weight function. We study the bias of importance sampling and independent Metropolis-Hastings, without assuming that the weight function is bounded. We show that the common random numbers coupling of independent Metropolis-Hastings is maximal. Using that coupling, we derive polynomial bounds on the total variation distance of the chain to its target distribution. We further consider bias removal techniques using couplings, and provide conditions under which the resulting unbiased estimators have finite moments, and under which their efficiency is comparable to that of importance sampling. Experiments illustrate unbiased estimators of the inverse of a normalizing constant, estimators of nested expectations, and combination of importance sampling with robust mean estimation methods.
Figures
Figures from the paper (1 more)
Forward citations
Cited by 1 Pith paper
-
The $L_p$-error rate for randomized quasi-Monte Carlo self-normalized importance sampling of unbounded integrands
Rigorous Lp error bounds of order O(N^{-β+ε}) for RQMC self-normalized importance sampling with unbounded integrands on unbounded domains, with β near 1 under QMC-friendly growth conditions.
Reference graph
Works this paper leans on
-
[5]
more likely to contain outliers, are down-weighted in the final estimate
Return empmed( ¯X1,..., ¯XK). more likely to contain outliers, are down-weighted in the final estimate. The procedure is described in Algorithm 7. Algorithm 7Minsker–Ndaoud (MN) estimator. 1.Input: i.i.d. samplesX 1,...,X n with meanµand varianceσ 2, confidence parameterδ∈(0,1), power parametera∈N ∗
-
[8]
Compute ˆκ=empmed( ¯X1,..., ¯XK)andd i = (Xi−ˆκ)2 fori∈[n]
-
[9]
Apply Algorithm 6 to(di)n i=1 to obtain˜σ2, a MoM-based variance estimate
-
[10]
(b) Compute the weight: wj = 1 (ˆσ2 j + ˜σ2)a/2
Forj∈[K]: (a) Compute block standard deviation ˆσ2 j = 1 |Bj| ∑ i∈Bj (Xi− ¯Xj)2. (b) Compute the weight: wj = 1 (ˆσ2 j + ˜σ2)a/2
-
[11]
Return ∑K j=1wj ¯Xj ∑K k=1wk . The regularization term˜σ2 stabilizes the weights and prevents small variances from causing numerical instability. The power parameteracontrols the sensitivity of the weights to the block variances. In our experiments, we fixa= 2. We refer to Theorem 3.1 of Minsker & Ndaoud (2021) for theoretical guarantees of this estimator...
work page 2021
-
[12]
Partition[n] ={1,...,n}intoKblocksB 1,...,B K, with eachBk of size⌊n/K⌋≤|B k|≤⌊n/K⌋+ 1
-
[13]
Forj∈[K], compute block empirical mean ¯Xj = 1 |Bj| ∑ i∈Bj Xi
-
[14]
Computeˆκ=empmed( ¯X1,..., ¯XK)
Show all 10 references
-
[15]
Find the solutionαto the equation: n∑ i=1 min ( 1,α(Xi−ˆκ)2) = log(1/δ) 3
-
[16]
Compute the correction term: ˆ∆n = 1 n n∑ i=1 (Xi−ˆκ) ( 1−min ( 1,α(Xi−ˆκ)2))
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.