{"id":"7ec01ec9-57b9-480e-9e98-151a5b8f6298","arxiv_id":"2501.02228","paper_version":2,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using the Moreau-Yosida envelope density as an importance distribution yields an asymptotically normal, finite-variance Markov chain importance sampling estimator that often beats proximal MALA and HMC.","lead":"This paper proposes a new way to estimate expectations from non-smooth Bayesian posteriors: sample from a smoothed version of the target called its Moreau-Yosida envelope, then reweight the samples back to the true target. The method comes with a finite-variance guarantee, a central limit theorem, and large efficiency gains over existing proximal MCMC samplers in tests.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"HMC finite-variance guarantee rests on unverified tail condition (29); Theorem 7's proof establishes (28) only.","rationale":"The reader's weakest assumption and my concern coincide: Theorem 2's finite-covariance CLT is only as strong as the geometric ergodicity of the pi_lambda-invariant chain, and for HMC the paper's sufficient conditions are incomplete because condition (29) is not verified. The conditional theorem itself is not wrong, and the proof of Theorem 2 appears sound once geometric ergodicity is granted. However, the paper goes beyond its own verified assumptions in Section 5.2, particularly in the beta>2 bullet where pi_lambda-HMC is asserted to be geometrically ergodic without demonstrating (29). Since the numerical studies include super-Gaussian targets and report large HMC efficiency gains, the practical central claim depends on this unverified tail condition. Verifying (29) for the envelope densities, or explicitly restricting the HMC guarantee to cases where it is proved, would remove the concern. Until then, a conditional acceptance is the honest recommendation.","tokens_in":30780,"tokens_out":18198,"duration_ms":184380,"concrete_test":"Directly verify condition (29) for pi_lambda-HMC in the Gaussian-envelope case pi_lambda = N(0, Omega + lambda I_d) used in Theorem 3: use the leapfrog proposal expression (26) to write the rejection region R(x), and compute lim_{||x||->infinity} integral_{R(x) cap I(x)} q_H(x,y) dy. If the limit is not zero, the HMC geometric-ergodicity proof fails in the simplest setting. If it is zero, repeat the calculation for the beta=4 super-Gaussian envelope pi_lambda, or numerically estimate the limit in the d=20 setting of Section 6.1 by comparing the batch-means variance estimator of Section C against a 100-chain replication variance for the MY-IS estimator. This would settle whether the unproved condition (29) is innocuous or load-bearing for the HMC-based finite-variance guarantee.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central finite-variance claim (Theorem 2) is conditional on geometric ergodicity of the pi_lambda-invariant MALA/HMC chain. That is where the paper's support is thinnest. For HMC, Theorem 7 inherits two conditions from Livingstone et al. (2019): the drift condition (28) and the tail rejection-region condition (29). The proof of Theorem 7 verifies (28) via (53), (54), and (56), but it never verifies (29). Section 5.2 acknowledges that 'demonstrating (29) can be challenging to establish for general problems,' yet item 3 of that section asserts that pi_lambda-HMC is geometrically ergodic for beta>2 targets because pi_lambda has Gaussian tails, without supplying the missing check. The same pattern appears in MALA: Theorem 5 assumes Roberts--Tweedie's inward-convergence condition (23) without proof for pi_lambda. If (29) or (23) fails for a target covered by the paper's examples, such as the beta=4 super-Gaussian toy in Section 6.1, then the HMC/MALA chain is not known to be geometrically ergodic, Theorem 2 does not provide the promised finite covariance, and the very large HMC efficiency gains in Sections 6.3 and 6.4 could be an artifact of an unchecked tail effect.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an importance sampling framework for targets of the form π(x) ∝ e^{−ψ(x)} with convex, possibly non-smooth ψ, using the Moreau-Yosida envelope density πλ(x) ∝ e^{−ψλ(x)} as the importance distribution. The main theoretical result (Theorem 2) establishes asymptotic normality of the self-normalized estimator θ̂^{MY}_n under geometric ergodicity of the πλ-invariant Markov chain, with an explicit covariance expression. The paper derives sufficient conditions for geometric ergodicity of πλ-MALA (Theorem 5) and πλ-HMC (Theorem 7), provides a closed-form optimal λ for Gaussian targets in the iid case (Theorem 3), and gives a practical tuning rule based on the importance sampling effective sample size. Numerical experiments cover a toy product target, Bayesian trend filtering, nuclear-norm matrix denoising, and a Poisson random effects model, comparing MY-IS with P-MALA, P-HMC, and Barker's algorithm, with 100 replications each and efficiency gains reported.","tokens_in":31110,"tokens_out":7922,"duration_ms":78080,"significance":"The central idea is simple and appealing: because ψλ envelopes ψ from below, the importance weights are uniformly bounded, and the smoothness of πλ makes gradient-based MCMC applicable. If the geometric ergodicity conditions hold, Theorem 2 delivers a genuinely useful finite-covariance CLT for an estimator that can have substantially lower variance than proximal MCMC alternatives. The Gaussian optimal-λ derivation is elegant, the batch-means variance estimator is standard, and the numerical studies are extensive, with code provided. The main weakness is that the advertised 'finite asymptotic variance' guarantee is conditional on geometric ergodicity of the πλ-chain, and the paper's sufficient conditions for that geometric ergodicity are not fully verified for the HMC case (condition (29)) and are assumed for MALA (condition (23)). These gaps are load-bearing for the central claim, especially for the β>2 targets used in the numerical evaluation.","major_comments":[{"comment":"The HMC geometric ergodicity guarantee is not fully established for the classes used in the numerical studies. Theorem 7 is stated conditionally on condition (29), but the proof in Section E.2 verifies only condition (28) (via (53), (54), and (56)); Section 5.2 explicitly acknowledges that demonstrating (29) is challenging. Item 3 of Section 5.2 nevertheless asserts that for β>2 targets πλ-HMC is geometrically ergodic for sufficiently small ε because πλ has Gaussian tails, without verifying (29). Since Theorem 2's finite covariance conclusion requires geometric ergodicity, the HMC efficiency gains reported in Sections 6.3 and 6.4 are not backed by the paper's theorems for the β=4 class of targets; either supply a proof of (29) for the targets considered or explicitly present the HMC results as conditional on (29).","section":"Section 5.2 / Proof E.2"},{"comment":"The MALA geometric ergodicity theorem assumes the inward-convergence condition (23) without verifying it for πλ or for the E(β,γ) class. Item 3 then states that for β>2, πλ-MALA is geometrically ergodic as long as h ≤ 4λ; this conclusion does not follow from Theorem 5 unless (23) holds. If (23) fails, the chain may not be geometrically ergodic and Theorem 2 is not applicable. Please either verify (23) for the examples in Section 6 or label the item as conditional on (23).","section":"Section 5.1, Theorem 5 and item 3"}],"minor_comments":[{"comment":"The displayed acceptance probability α(x,y) is missing the 'min' operator; as written, it is not a proper probability.","section":"Supplement F.1, Algorithm 4"},{"comment":"The indicator notation 1(1 ≤ β < 2) is nonstandard and could be confused with the function value; writing 1_{[1,2)}(β) would be clearer.","section":"Equation (22)"},{"comment":"The phrase 'all choices of λ yield a finite variance estimator' is stronger than what is proved, since Theorem 2 is conditional on geometric ergodicity of the πλ-chain; qualify this statement.","section":"Section 4.1, Remark 1"},{"comment":"The sentence 'For HMC, estimation of components are at least 25 times more efficient' should specify that the comparison is relative to P-HMC, as is done for the MALA case.","section":"Section 6.2"},{"comment":"The notation w^{(i)}_{(l)} is easy to misread; consider defining the ordered-sample weights explicitly for each component before displaying the quantile estimator.","section":"Equation (17)"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid contribution and the core importance-sampling construction is sound. My main concern is the gap between the advertised finite-variance guarantee and the unverified tail conditions behind Theorems 5 and 7. If the authors can prove condition (29) for the target classes considered, or restrict the claims accordingly, I would be willing to support acceptance. The numerical study is thorough, and the code availability is a plus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nTwo things to know about arXiv:2501.02228. The core estimator is genuinely new and useful, and the paper's central finite-variance promise is more conditional than the abstract suggests.\n\nWhat's new: the MY-IS estimator, using the Moreau-Yosida envelope density as an importance distribution for MCMC, with a CLT (Theorem 2) that gives an explicit covariance expression. This is not a trivial repackaging of Pereyra (2016) and Durmus et al. (2022). The proof of Theorem 2 is in the supplement and the algebra checks out. I also worked through the Gaussian optimal-λ derivation in Theorem 3; it is correct. The numerical study is extensive: 100 replications, comparisons against P-MALA, P-HMC, and Barker, with code available. The efficiency gains, especially for HMC, are large and credible.\n\nThe soft spot is the geometric ergodicity support for the proposal chain. Theorem 2's finite variance requires the πλ-invariant MALA or HMC chain to be geometrically ergodic. The sufficient conditions are partial. For HMC, Theorem 7 verifies Livingstone et al.'s drift condition (28) but never checks the tail rejection-region condition (29). Section 5.2 says demonstrating (29) is hard in general, then item 3 for β>2 simply asserts geometric ergodicity because πλ has Gaussian tails. That assertion goes beyond the proof. The MALA result has the same structure: Theorem 5 assumes Roberts-Tweedie's inward convergence condition (23) without verification for πλ. The authors are honest about these gaps in the text, but the abstract's \"guaranteed to have finite asymptotic variance\" overstates what is established. The estimator is still consistent by Theorem 1, and the variance will likely be finite in practice, but the guarantee is conditional.\n\nThe λ selection via ne/n in [0.4, 0.8] is heuristic, though the Gaussian case gives a principled starting point. Minor issue.\n\nWho should read it: methodologists working on MCMC for non-smooth log-concave targets, and practitioners using proximal MCMC. It deserves a serious referee. I would send it to review with a request to either verify (29) for the paper's examples or soften the claims. Accept with minor revisions is reasonable.\n\nRecommendation: engage with it; I'd cite it if I work in the area.","headline":"Solid new IS estimator for non-smooth targets; the CLT is real, but the finite-variance guarantee rests on unverified tail conditions the paper itself flags.","tokens_in":31613,"tokens_out":2647,"would_cite":true,"duration_ms":25269,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","65C05","60J22"],"pacs":[],"model":"deepseek-v4-flash","headline":"Using the Moreau-Yosida envelope density as the importance distribution makes self-normalized importance sampling for non-differentiable posteriors asymptotically normal with finite covariance, and it can substantially beat proximal MCMC…","keywords":["Moreau-Yosida envelope","importance sampling","non-differentiable posterior","geometric ergodicity","MALA","Hamiltonian Monte Carlo","asymptotic variance","proximal MCMC"],"falsifier":"Run MY-IS with a MALA chain on a target where $\\limsup_{\\|x\\|\\to\\infty} \\|\\mathrm{prox}^\\lambda_\\psi(x)\\|/\\|x\\| = 1$ (for instance, a heavy-tailed log-concave $\\psi$ whose proximal map grows along a ridge) and check whether the batch-means estimate of $\\Xi$ in (40) stays finite as $n$ grows; if the drift condition (24) fails, the estimator may remain consistent but its asymptotic variance can diverge, which would show the geometric-ergodicity assumption is doing the work.","tokens_in":30551,"feed_emoji":"📊","tokens_out":9942,"duration_ms":80647,"temperature":0.7,"pith_summary":"This paper proposes a new way to estimate expectations under non-differentiable Bayesian posteriors: draw samples from the Moreau-Yosida envelope density $\\pi_\\lambda(x) \\propto e^{-\\psi_\\lambda(x)}$ instead of the target $\\pi(x) \\propto e^{-\\psi(x)}$, and correct with the self-normalized weights $w_\\lambda(x) = e^{-(\\psi(x)-\\psi_\\lambda(x))}$. The authors prove that the resulting estimator $\\sqrt{n}(\\hat{\\theta}^{\\mathrm{MY}}_n - \\theta)$ is asymptotically normal with finite covariance whenever the $\\pi_\\lambda$-invariant MALA or HMC chain is geometrically ergodic; the envelope inequality $\\psi_\\lambda \\le \\psi$ makes the weights uniformly bounded, so the variance is finite. They give sufficient conditions for that geometric ergodicity and show in several simulation settings that the smoothed importance sampler can be substantially more efficient than proximal MCMC methods that target $\\pi$ directly. If the argument is right, practitioners with $\\ell_1$, fused-lasso, or nuclear-norm priors gain a principled finite-variance alternative to proximal MALA and HMC.","feed_headline":"Smoothing non-differentiable posteriors gives finite-variance sampling","feed_subtitle":"Moreau-Yosida envelopes give smoother targets, beating proximal MCMC variance while keeping asymptotic normality.","key_machinery":"The object carrying the argument is the Moreau-Yosida envelope $\\psi_\\lambda(x) = \\inf_y\\{\\psi(y) + \\tfrac{1}{2\\lambda}\\|y-x\\|^2\\}$ and its density $\\pi_\\lambda(x) \\propto e^{-\\psi_\\lambda(x)}$; the proximal map $\\mathrm{prox}^\\lambda_\\psi(x) = \\arg\\min_y\\{\\psi(y) + \\tfrac{1}{2\\lambda}\\|y-x\\|^2\\}$ supplies the gradient $\\nabla \\log \\pi_\\lambda(x) = (\\mathrm{prox}^\\lambda_\\psi(x) - x)/\\lambda$, which is Lipschitz and computable whenever the proximal map is. The envelope inequality $\\psi_\\lambda \\le \\psi$ gives the uniformly bounded weights $w_\\lambda(x) \\le 1$ that turn the Markov chain central limit theorem into the finite asymptotic covariance $\\Xi$ via the delta method on the augmented statistic $S(x) = (\\xi(x)w_\\lambda(x), w_\\lambda(x))$. Assumption 1 is the tail-growth condition on the proximal map that, through Lemma 1, makes the drift conditions for MALA and HMC tractable: for MALA it becomes the simple step-size rule $h \\le 2\\lambda$, and for HMC it verifies three gradient conditions from Livingstone et al. (2019).","core_discovery":"The central claim is that the Moreau-Yosida envelope density $\\pi_\\lambda(x) \\propto e^{-\\psi_\\lambda(x)}$ is a provably safe importance distribution for a log-concave target $\\pi(x) \\propto e^{-\\psi(x)}$ whose proximal map of $\\psi$ can be computed. Because $\\psi_\\lambda$ lies below $\\psi$, the unnormalized weight $w_\\lambda(x) = e^{-(\\psi(x)-\\psi_\\lambda(x))}$ satisfies $0 < w_\\lambda(x) \\le 1$, and the smoothed density both bounds the importance ratio $\\pi/\\pi_\\lambda$ and gives gradient samplers a well-conditioned target. Theorem 2 shows that the self-normalized estimator (12) is asymptotically normal with the explicit covariance $\\Xi$ in (14), assuming only a second moment under $\\pi$ and a geometrically ergodic $\\pi_\\lambda$-chain; Theorem 3 shows that for a Gaussian target the variance-minimizing $\\lambda$ lies in $[s_1/d, s_d/d]$, so the optimal smoothing decreases with dimension. The paper further proves (Theorems 5 and 7) that under the tail-growth Assumption 1, $\\limsup_{\\|x\\|\\to\\infty} \\|\\mathrm{prox}^\\lambda_\\psi(x)\\|/\\|x\\| < 1$, both MALA and HMC sampling from $\\pi_\\lambda$ are geometrically ergodic for suitable step sizes, so the finite-variance guarantee applies to the algorithms practitioners would run.","pith_inferences":["If the finite-variance guarantee extends to non-log-concave targets via weak convexity, the same estimator could apply to multi-modal posteriors by choosing $\\lambda$ to make $\\pi_\\lambda$ log-concave; the paper leaves this conditional on $\\pi_\\lambda$ being integrable.","The paper's Theorem 3 suggests a default choice $\\lambda \\propto 1/d$ in high dimensions, since the optimal smoothing shrinks with dimension; this could be turned into a pilot-free recipe before any MCMC run.","The weight-boundedness mechanism is general: any approximating density that lies below the target pointwise and is smooth enough for gradient samplers inherits the same finite-variance asymptotic normality, so the envelope trick could be applied to other families like spline or kernel-smoothed potentials.","A testable extension is to estimate the asymptotic covariance $\\Xi$ online and stop when the batch-means estimate stabilizes, giving a principled stopping rule for proximal MCMC that the paper's Theorem 8 would justify."],"forward_implications":["For any non-differentiable log-concave posterior whose proximal map satisfies Assumption 1, practitioners can replace proximal MALA/HMC with $\\pi_\\lambda$-sampling plus weights and obtain Monte Carlo error bars from the explicit covariance estimator in (40); the paper shows this is often many times more efficient.","The same pipeline applies to differentiable targets with heavy or non-Lipschitz tails, where standard MALA and HMC are unreliable; the paper demonstrates this on the Poisson random-effects model, where even the Barker proposal is outperformed.","The tuning guideline $n_e/n \\in [0.4, 0.8]$ gives a simple, self-checking way to balance smoothing against weight degradation, and Theorem 3's inverse-dimension scaling tells users in high dimensions to start with small $\\lambda$.","Marginal quantiles and credible intervals for each coordinate follow directly from the Chen–Shao weighted order-statistic estimator, so the method is usable as an end-to-end Bayesian inference routine, not just for means."],"supporting_citations":[{"why":"Supplies Proposition 1: $\\pi_\\lambda$ is a proper density, converges to $\\pi$ pointwise as $\\lambda \\to 0$, and $\\nabla \\log \\pi_\\lambda = (\\mathrm{prox}^\\lambda_\\psi(x) - x)/\\lambda$.","marker":"Durmus et al. (2022)"},{"why":"Introduced proximal MALA and the envelope-density identity used as the importance distribution; also supplies Result 1 on tail orders.","marker":"Pereyra (2016)"},{"why":"Theorem 4's drift conditions for MALA geometric ergodicity that the paper verifies for $\\pi_\\lambda$.","marker":"Roberts and Tweedie (1996)"},{"why":"Theorem 6's conditions for HMC geometric ergodicity and the gradient conditions (28)-(29) used in Theorem 7.","marker":"Livingstone et al. (2019)"},{"why":"Established MCMC-based importance sampling theory and sufficient variance conditions that motivate the self-normalized estimator.","marker":"Madras and Piccioni (1999)"},{"why":"Provides the weighted order-statistic quantile estimator used for credible intervals.","marker":"Chen and Shao (1999)"},{"why":"Defines the importance-sampling effective sample size $n_e$ used for tuning $\\lambda$.","marker":"Kong (1992)"},{"why":"Gives the asymptotic covariance of self-normalized importance sampling for iid samples, used to derive Theorem 3's optimal $\\lambda$ for Gaussian targets.","marker":"Agarwal et al. (2022)"},{"why":"Proximal HMC algorithm that serves as the comparison baseline in the numerical studies.","marker":"Chaari et al. (2016)"},{"why":"Barker proposal that the paper compares against on the Poisson random-effects model.","marker":"Livingstone and Zanella (2022)"}],"fun_headline_variants":["Smooth priors for MCMC: Moreau-Yosida beats proximal variance","Moreau-Yosida envelopes tame non-smooth MCMC sampling","MY importance sampling gives finite variance and geometric mixing","Smoothing non-differentiable posteriors yields lower MCMC variance","Gradient-friendly MCMC via Moreau-Yosida importance distribution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The finite-variance guarantee holds only when the chain targeting the smoothed density is geometrically ergodic, which the paper verifies under a tail-growth condition on the proximal map, plus an extra tail condition for Hamiltonian Monte Carlo that is not established for all log-concave targets.","fun_headline_variants_meta":{"raw":{"variants":["Smooth priors for MCMC: Moreau-Yosida beats proximal variance","Moreau-Yosida envelopes tame non-smooth MCMC sampling","MY importance sampling gives finite variance and geometric mixing","Smoothing non-differentiable posteriors yields lower MCMC variance","Gradient-friendly MCMC via Moreau-Yosida importance distribution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000644,"raw_usage":{"total_tokens":3005,"prompt_tokens":1033,"completion_tokens":1972,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1881}},"tokens_in":649,"tokens_out":1972,"duration_ms":13585,"temperature":1.0,"reasoning_tokens":1881,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:14:31.567915+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run MY-IS with a MALA chain on a target where $\\limsup_{\\|x\\|\\to\\infty} \\|\\mathrm{prox}^\\lambda_\\psi(x)\\|/\\|x\\| = 1$ (for instance, a heavy-tailed log-concave $\\psi$ whose proximal map grows along a ridge) and check whether the batch-means estimate of $\\Xi$ in (40) stays finite as $n$ grows; if the drift condition (24) fails, the estimator may remain consistent but its asymptotic variance can diverge, which would show the geometric-ergodicity assumption is doing the work.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Proximal HMC algorithm that serves as the comparison baseline in the numerical studies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Proposition 1: $\\pi_\\lambda$ is a proper density, converges to $\\pi$ pointwise as $\\lambda \\to 0$, and $\\nabla \\log \\pi_\\lambda = (\\mathrm{prox}^\\lambda_\\psi(x) - x)/\\lambda$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduced proximal MALA and the envelope-density identity used as the importance distribution; also supplies Result 1 on tail orders."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Theorem 4's drift conditions for MALA geometric ergodicity that the paper verifies for $\\pi_\\lambda$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Theorem 6's conditions for HMC geometric ergodicity and the gradient conditions (28)-(29) used in Theorem 7."},{"cited_title":"and Piccioni, M","cited_arxiv_id":null,"evidence_quote":"Established MCMC-based importance sampling theory and sufficient variance conditions that motivate the self-normalized estimator."},{"cited_title":"and Shao, Q.-M","cited_arxiv_id":null,"evidence_quote":"Provides the weighted order-statistic quantile estimator used for credible intervals."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the importance-sampling effective sample size $n_e$ used for tuning $\\lambda$."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the asymptotic covariance of self-normalized importance sampling for iid samples, used to derive Theorem 3's optimal $\\lambda$ for Gaussian targets."},{"cited_title":"and Zanella, G","cited_arxiv_id":null,"evidence_quote":"Barker proposal that the paper compares against on the Poisson random-effects model."}],"review_version":1}