{"id":"9d04b8f9-7c6d-45e0-bf65-8a5fa94014dd","arxiv_id":"2506.14685","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"For the elliptic data assimilation problem, Gaussian priors defined directly on finite element spaces give discrete posterior means that contract to the ground truth at the optimal continuous-level rates when the mesh size and the number of samples are coupled as h ~ N^{-1/(2α+d)}.","lead":"The paper proves convergence rates for discretized Bayesian data assimilation when the computational mesh and the number of noisy point observations scale together: the discrete posterior mean recovers the true solution at the same rate as the continuous infinite-dimensional theory. The key technical step is a set of tailor-made Gaussian priors defined directly on finite element spaces instead of discretizing a continuous prior.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The polynomial-degree condition in Theorems 4.1–4.2, k≥min{α−1,1}, is too weak: for α>2 it allows k=1, while approximation property (A.4) used in Lemma 4.4 requires k+1≥α; the claimed h^{α−1+β} interpolation-rate bound therefore does not follow for the stated k.","rationale":"I read the paper's core construction as: choose a V_h-adapted Gaussian prior with PDE residuals, bound the deterministic bias via Lemma 4.4, bound the stochastic fluctuation via the empirical-process Lemma 4.10 and the trace bound in Lemma 4.13, then use unique-continuation stability to convert dβ to the desired norm. This structure is coherent and, modulo typographical sign errors, plausible. The reader's weakest assumption about ω being element-aligned and the truth being H^α∩C^0 is a genuine applicability restriction and is explicitly assumed; it is not an internal gap. The most serious internal issue is the polynomial-degree condition: as printed, it is inconsistent with the approximation theory quoted in Appendix A and used in Lemma 4.4. If k is too low, the apparent optimal coupling h∼(σ²/(Nl_N))^{1/(2α+d)} is not supported, and the central rate is not established. Because this is a condition in the theorem statements rather than a flaw in the overall strategy, the correct disposition is still conditional acceptance after correcting min to max (and fixing the sign of the exponents). I agree with the reader's conditional verdict, though I would put the polynomial-degree/interpolation mismatch ahead of the element-aligned-observation restriction as the load-bearing point.","tokens_in":25435,"tokens_out":29542,"duration_ms":286238,"concrete_test":"Take Ω=(0,1), α=2.5, k=1, and a smooth u† with non-zero H^α norm. On a sequence of quasi-uniform meshes h→0, compute ||u†−I_h u†||_{H^1(Ω)} for the quasi-interpolant of Appendix A. If the decay is O(h) rather than O(h^{1.5}), Lemma 4.4's bound fails for k=1, and the theorem condition must be k≥α−1 (i.e., max{α−1,1}). The same check can be done analytically from (A.4): the inequality is only asserted for s≤k+1, so substituting s=2.5>2 is outside its stated range.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Central claim Theorem 4.2(1) asserts an L2(B) rate (σ²/(Nl_N))^{ατ/(2α+d)} with h=(σ²/(Nl_N))^{1/(2α+d)}. The proof routes the deterministic bias through Lemma 4.4, whose key estimate is ||h^β u†_h||²_Vh plus PDE-residual terms ≤ C h^{2(α−1+β)}||u†||²_{H^α}. That estimate is obtained from [6, Prop. 3.3] and the quasi-interpolation bound (A.4), which is stated only for s≤k+1. To get an H1 error of order h^{α−1}, one needs k+1≥α, i.e. k≥α−1. The theorems instead require k≥min{α−1,1}; for α>2 this permits k=1. With k=1 and α=2.5, for example, piecewise-linear interpolation leaves an H1 error of order h, not h^{1.5}, so the h^{2α} bias bound in Lemma 4.4 is false; the trade-off h=(σ²/(Nl_N))^{1/(2α+d)} and the resulting rate are not justified. Separate literal sign errors (h∼(σ²/N)^{−1/(2α+d)} and δ_N∼(σ²/N)^{−α/(2α+d)} in Theorem 4.1, and δ_N with a negative exponent in Theorem 4.2) would make those displayed statements diverge rather than converge, but these are likely typos. The polynomial-degree condition is the load-bearing issue because it concerns an assumption actually used in the proof.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes posterior contraction for Bayesian data assimilation under discretization. Section 3 considers conforming subspaces of the Cameron-Martin space and proves a convergence rate for the discrete posterior mean via Galerkin orthogonality. Section 4 constructs tailor-made Gaussian priors on H1-conforming finite element spaces and states contraction rates for the elliptic data assimilation problem: for the L2(B) semi-norm and beta=1, with h chosen as (sigma^2/(N l_N))^{1/(2alpha+d)}, the posterior mean/MAP estimator contracts at rate (sigma^2/(N l_N))^{alpha tau/(2alpha+d)}; for the H1(B) semi-norm and beta=0, an analogous rate (sigma^2/(N l_N))^{(alpha-1)tau'/(2(alpha-1)+d)} is claimed. The stated rates are intended to match the continuous posterior theory of [23] while avoiding the assumption that the ground truth lies in the Cameron-Martin space.","tokens_in":25682,"tokens_out":9882,"duration_ms":100923,"significance":"If the main theorems are correct, this is a useful contribution: it gives a concrete recipe for finite element priors in Bayesian data assimilation with simultaneous rates in sample size, mesh size, and noise, and it relaxes the requirement that the truth belongs to the Cameron-Martin space. The proof architecture is coherent, separating deterministic interpolation bias, stochastic variance, and an empirical-process lower bound via the matrix Chernoff bound. The paper is also careful to state the coupling between h and N and to compare with the continuous benchmark. However, the central rate is not established as written because the polynomial-degree condition is too weak for the interpolation estimates used in Lemma 4.4, and the theorem displays contain sign errors that make the stated h and delta_N diverge.","major_comments":[{"comment":"The polynomial-degree condition k >= min{alpha-1,1} is insufficient for alpha > 2. The approximation property (A.4) is stated only for s <= k+1, and the proof of Lemma 4.4 requires an H1 interpolation error of order h^{alpha-1} (via [6, Prop. 3.3] and (A.4)) to obtain the bound h^{2(alpha-1+beta)}. For alpha > 2, the condition allows k=1, but then k+1=2 < alpha, so (A.4) does not apply at s=alpha and the H1 error is only of order h, not h^{alpha-1}. For example, with alpha=2.5 and k=1, the bound in Lemma 4.4 would claim h^3 bias while piecewise-linear interpolation only gives h^2 in H1. Consequently the balance h = (sigma^2/(N l_N))^{1/(2alpha+d)} and the resulting rate in Theorem 4.2(1) are not justified. The condition should be k+1 >= alpha (equivalently k >= ceil(alpha-1)) throughout, and the complexity statements should be rechecked with this requirement.","section":"Lemma 4.4; Appendix A, Eq. (A.4); Theorems 4.1 and 4.2"},{"comment":"The displayed formulas for h and delta_N contain negative exponents. Theorem 4.1 states h = h(N) ~ O((sigma^2/N)^{-1/(2alpha+d)}) and delta_N = (sigma^2/N)^{-alpha/(2alpha+d)}, which tend to infinity as N grows; Theorem 4.2(1) similarly states delta_N = (sigma^2/(N l_N))^{-alpha/(2alpha+d)}, and Theorem 4.2(2) states delta'_N = (sigma^2/(N l_N))^{-(alpha-1)/(2(alpha-1)+d)}. These are presumably typographical errors, but as written they contradict the claimed convergence. They must be corrected to positive exponents: h ~ (sigma^2/N)^{1/(2alpha+d)} and delta_N = (sigma^2/N)^{alpha/(2alpha+d)} in Theorem 4.1, and likewise in Theorem 4.2.","section":"Theorems 4.1 and 4.2"},{"comment":"The observation measure is handled inconsistently. In Section 2.1 the sampling distribution lambda is assumed only to be a probability measure equivalent to Lebesgue on omega. In Lemma 4.10, however, the proof states 'Let lambda be the Lebesgue measure in omega' and uses E[A_N] = A, where A is the L2(omega) Gram matrix. If lambda has a density that is not bounded above and below uniformly on omega, the smallest eigenvalue of the expected sampled Gram matrix can be much smaller than the corresponding L2 mass-matrix eigenvalue, and the constant c in (4.39) may not be uniform in h. This is load-bearing for the high-probability variance estimates in Lemmas 4.11 and 4.13. The authors should either assume lambda has density bounded above and below (or is normalized Lebesgue) and state this in the problem setting, or adapt the proof to the genuinely weighted Gram matrix.","section":"Section 2.1 and Lemma 4.10"}],"minor_comments":[{"comment":"In the displayed definition just before (4.42), the set should be {forall v_h in V_h: 1/N sum v_h^2(X_i) >= c ||v_h||^2_{L2(omega)}}, but the text writes '<= c'. This is a typographical inversion of the inequality direction and should be corrected.","section":"Lemma 4.13"},{"comment":"In the final display of the proof, the norm ||u||_{L^infty(Omega)} and ||u||_{H^alpha(Omega)} should refer to the ground truth u^dagger; the variable u is undefined at that point.","section":"Lemma 4.7"},{"comment":"The notation h = h(N) >> O(N^{-1/(2alpha+d)}) is imprecise; it should be written as h(N)/(N^{-1/(2alpha+d)}) -> infinity to make the asymptotic meaning unambiguous.","section":"Notation in Theorem 4.2"},{"comment":"The statement that the dimension of V_h scales as O(N^{d/(2alpha+d)}) ignores the dependence on the polynomial degree k; after correcting k >= ceil(alpha-1), the constant in this dimension bound depends on k (through the number of degrees of freedom per element), which is harmless for the rates but should be acknowledged.","section":"Section 1.2 and Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"I agree with the stress-test assessment: the polynomial-degree condition is the load-bearing gap, and it is fixable by requiring k+1 >= alpha. The sign errors in the theorem displays are easy to correct, and the observation-measure issue needs a stated bounded-density assumption. The manuscript's core architecture is sound enough that I would not reject it, but the main theorems as printed are not supported by the proofs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, Section 4 is a real contribution: the authors build Gaussian priors directly on finite element spaces, with PDE-residual penalties, instead of discretizing a continuous prior. That lets them bypass Cameron-Martin support assumptions and still get contraction rates that match the continuous theory in [23]. Second, the main theorem as written has a load-bearing error: the polynomial-degree condition k≥min{α−1,1} is too weak. The proof of Lemma 4.4 uses h^{α−1} interpolation error in H1, which requires k+1≥α. For α>2 the stated condition allows k=1, and the claimed h^{α−1+β} bias bound fails. So Theorem 4.2(1) is not established for α>2. The fix is likely k≥max{α−1,1}, but this is exactly the kind of assumption that must be checked, not glossed.\n\nWhat the paper does well: the tailor-made prior idea is genuinely new relative to [27], [29], and [23]. The h-N coupling h~(σ²/(N l_N))^{1/(2α+d)} is useful practical guidance. The proof structure—bias-variance decomposition, matrix Chernoff empirical process, unique-continuation stability—is coherent. The self-citations to [6]-[8] are for technical lemmas and are appropriate.\n\nSoft spots, in proportion. The sign errors in Theorem 4.1: h and δ_N are displayed with negative exponents. These are almost certainly typos, but as printed the statements diverge. That has to be fixed, and it is a little embarrassing in the main theorem. The abstract promises posterior contraction, but the theorems state convergence of the MAP or posterior mean. Lemma 4.13 bounds the posterior variance, so a measure-contraction statement is within reach; they should state it. Also, the analysis assumes ω is exactly a union of elements. That's a real restriction, but reasonable if ω is polyhedral and the mesh is fitted.\n\nWho this is for: people working on FEM-based Bayesian inverse problems who want to avoid spectral bases and want to know how to choose mesh size as N grows. It deserves a serious referee: the core idea is sound and the errors are correctable, but the degree condition is load-bearing and must be fixed before the main result can be trusted. My recommendation: send it to review, but with a referee instruction to verify the interpolation condition carefully.","headline":"Genuinely useful FE-prior construction, but the main theorem's polynomial-degree condition is load-bearing wrong and the displayed exponents are typos; fixable, worth refereeing.","tokens_in":26362,"tokens_out":3107,"would_cite":false,"duration_ms":31086,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65N30","62F15","65J22","35R30"],"pacs":[],"model":"deepseek-v4-flash","headline":"Finite-element priors built directly on the mesh recover the optimal posterior contraction rate for elliptic data assimilation without assuming the truth lies in the prior's support.","keywords":["posterior contraction","Bayesian data assimilation","finite element","unique continuation","MAP estimator","discrete priors","ill-posed inverse problems","posterior consistency"],"falsifier":"Run the proposed estimator on the elliptic data assimilation problem with ω chosen to cut diagonally through mesh cells, keeping h at the optimal coupling; if the L2(B) error does not decay like (σ²/(N l_N))^{ατ/(2α+d)} with probability tending to 1, the element-alignment assumption is load-bearing. Alternatively, take u† ∈ H^α with α≤1 and check whether the Chebyshev bound in Lemma 4.7 fails, since it relies on the L∞ control of the interpolant.","tokens_in":25047,"feed_emoji":"📐","tokens_out":7096,"duration_ms":60264,"temperature":0.7,"pith_summary":"The paper proves that a Bayesian data assimilation problem, posed as an ill-posed elliptic inverse problem with unique continuation, can be discretized with finite elements while keeping the optimal posterior contraction rate of the continuous theory. The key is to construct the prior directly on the finite element space, with penalties on regularity and on the PDE residual, rather than discretizing a continuous Gaussian prior. Under a coupling between mesh size h and sample size N, the MAP estimator and posterior mean converge to the ground truth at rate roughly (σ²/(N l_N))^{ατ/(2α+d)} in the L2(B) norm, matching the continuous rate and requiring only H^α∩$C^{0}$ regularity with α>1. This matters because standard finite element spaces are not contained in the Cameron-Martin space of typical priors, so naive discretization can degrade or destroy posterior consistency.","feed_headline":"Discrete priors reach optimal Bayesian contraction rates","feed_subtitle":"For elliptic data assimilation, a mesh-built finite-element prior matches the oracle rate without assuming the truth lies in the prior.","key_machinery":"The central object is a tailor-made discrete Gaussian prior on the finite element space, whose density penalizes three things through the parameter β: a regularity seminorm ||h^β uh||_{V_h} involving jumps of normal gradients across facets; element-wise L2 penalties on $h^{{1+β}}$(-Δuh - fh); and a dual-norm penalty ||h^β(-Δuh - f)||_{W*_h} enforcing the PDE in a weak sense. Working with the associated triple norm |||v|||$_β^{2}$ = (1/N)||v||_{Σ-1}^2 + s_h(h^β v, h^β v) + ||h^β Δv||_{W*_h}^2, the proof splits the error into a deterministic mean-error part (interpolation plus PDE residual, controlled by finite element approximation estimates) and a stochastic variance part (controlled by a matrix Chernoff bound on the empirical Gram matrix restricted to the observation region ω). The unique-continuation stability estimate (Lemma 2.6) then converts L2(ω)-plus-Laplacian control into L2(B) or H1(B) control with Hölder exponents τ, τ'. The assumption that ω is a union of mesh elements makes the basis functions vanishing outside ω, which is what reduces the sampled Gram-matrix control to the subspace active in ω.","core_discovery":"The central claim is that if one defines a Gaussian prior on the finite element space Vh by log π(uh) ∝ -(N/(2σ²))( ||h^β uh||_{V_h}^2 + Σ_K ||$h^{{1+β}}$(-Δuh - fh)||_{L2(K)}^2 + ||h^β(-Δuh - f)||_{W*_h}^2 ) for β=0,1, then the discrete posterior mean (equivalently the MAP estimator for this linear-Gaussian problem) contracts to the true solution u† at the rate predicted by the continuous posterior theory, provided h = O((σ²/(N l_N))^{1/(2α+d)}) for the L2(B)-norm statement and h = O((σ²/(N l_N))^{1/(2(α-1)+d)}) for the H1(B) statement. The proof obtains the convergence with probability tending to 1 over both the random observation points and noise, and it does not require u† to lie in the Cameron-Martin space of the prior, only in H^α(Ω)∩$C^{0}$(Ω) with α>1. The paper also shows that if the mesh is coarser than the optimal coupling, the error is dominated by the finite-element interpolation error $h^{{ατ}}$, which is the optimal rate for this ill-posed problem.","pith_inferences":["If the observation region ω cuts across mesh cells, the identity used to reduce the Gram-matrix control to the active subspace breaks; one testable fix is to enrich the space near ∂ω or to use a quadrature/oversampling approach, at the price of possibly slower rates.","The same recipe—penalize the PDE residual in the natural dual norm and use conditional stability to convert observation error into state error—should transfer to other ill-posed problems such as the Cauchy problem for the Helmholtz equation or the Schrödinger unique continuation problem, where similar stability exponents are available.","Because the rate's exponent contains the unique-continuation stability τ, the paper implies that the achievable contraction rate for this problem is intrinsically limited by stability, not by the Bayesian machinery—a feature shared with deterministic regularization.","One could test the theory numerically by comparing the predicted h ~ N^{-1/(2α+d)} coupling against a fixed, too-coarse mesh: the error should stall at the interpolation floor h^{ατ}, which is directly measurable."],"forward_implications":["For elliptic data assimilation with a mesh-aligned observation region, the discrete posterior mean and MAP estimator achieve the same contraction rate as the continuous posterior, as long as one solves on a mesh with h coupled to N like (σ²/(N l_N))^{1/(2α+d)}.","The finite element dimension only needs to grow as N^{d/(2α+d)}, which is at most N^{1/2} when α>d/2, so computational cost of the linear solve stays comparable to the sampling cost when α>d.","The truth does not need to live in the Cameron-Martin space of the prior: only H^α∩C^0 with α>1 is required, relaxing the usual α>d/2 assumption.","If the mesh is too coarse relative to N, the error becomes purely the finite-element interpolation error h^{ατ}, giving a clean separation between statistical and discretization errors.","The mixed formulation (4.19) gives a practical way to compute the MAP estimator without forming the dual-norm penalty explicitly."],"supporting_citations":[{"why":"Supplies the continuous posterior contraction theory (Theorem 2.5) whose rates the finite-element results are designed to match.","marker":"[23]"},{"why":"Provides the conditional stability estimate for the elliptic Cauchy problem that yields the Hölder exponent τ in Lemma 2.6.","marker":"[4]"},{"why":"Gives the L2(B) unique-continuation estimate (first part of Lemma 2.6) used to lift observation-norm control to the interior norm.","marker":"[22]"},{"why":"Adapts the three-spheres argument to H1(B), giving the second estimate in Lemma 2.6.","marker":"[7]"},{"why":"Provides the finite-element approximation and jump-stabilization estimates used in Lemma 4.4 to bound the PDE-residual penalties at the interpolant.","marker":"[6]"},{"why":"Defines the quasi-interpolation operator with the approximation and stability properties used throughout the proofs.","marker":"[12]"},{"why":"Supplies the matrix Chernoff bound that underpins Lemma 4.10, controlling the empirical Gram matrix in the observation region.","marker":"[32]"},{"why":"Gives the earlier non-asymptotic finite-element posterior convergence analysis that the paper extends to sample-size dependent contraction rates.","marker":"[29]"}],"fun_headline_variants":["Discrete priors achieve optimal Bayesian contraction rates","Mesh-based priors match oracle convergence rates","Bayesian data assimilation: optimal rates via discrete priors","Finite-element priors unlock optimal posterior rates","Optimal contraction from tailor-made discrete priors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proof requires the observation region ω to be exactly a union of mesh elements and the truth to lie in H^α(Ω)∩$C^{0}$(Ω) with α>1; either condition failing—ω cutting through cells, or a truth with α≤1 or without continuity—leaves the stated high-probability rates unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Discrete priors achieve optimal Bayesian contraction rates","Mesh-based priors match oracle convergence rates","Bayesian data assimilation: optimal rates via discrete priors","Finite-element priors unlock optimal posterior rates","Optimal contraction from tailor-made discrete priors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000183,"raw_usage":{"total_tokens":1291,"prompt_tokens":902,"completion_tokens":389,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":518,"completion_tokens_details":{"reasoning_tokens":317}},"tokens_in":518,"tokens_out":389,"duration_ms":3603,"temperature":1.0,"reasoning_tokens":317,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:50:55.251119+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed estimator on the elliptic data assimilation problem with ω chosen to cut diagonally through mesh cells, keeping h at the optimal coupling; if the L2(B) error does not decay like (σ²/(N l_N))^{ατ/(2α+d)} with probability tending to 1, the element-alignment assumption is load-bearing. Alternatively, take u† ∈ H^α with α≤1 and check whether the Chebyshev bound in Lemma 4.7 fails, since it relies on the L∞ control of the interpolant.","supporting_citations":[{"cited_title":"Nickl , Bayesian non-linear statistical inverse problems , EMS press Berlin, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies the continuous posterior contraction theory (Theorem 2.5) whose rates the finite-element results are designed to match."},{"cited_title":"Alessandrini, L","cited_arxiv_id":null,"evidence_quote":"Provides the conditional stability estimate for the elliptic Cauchy problem that yields the Hölder exponent τ in Lemma 2.6."},{"cited_title":"Ultra-weak least squares discretizations for unique continuation and Cauchy problems","cited_arxiv_id":"2407.04571","evidence_quote":"Gives the L2(B) unique-continuation estimate (first part of Lemma 2.6) used to lift observation-norm control to the interior norm."},{"cited_title":"Burman, M","cited_arxiv_id":null,"evidence_quote":"Adapts the three-spheres argument to H1(B), giving the second estimate in Lemma 2.6."},{"cited_title":"Solving the unique continuation problem for Schr\\\"odinger equations with low regularity solutions using a stabilized finite element method","cited_arxiv_id":"2403.16914","evidence_quote":"Provides the finite-element approximation and jump-stabilization estimates used in Lemma 4.4 to bound the PDE-residual penalties at the interpolant."},{"cited_title":"Ern and J.-L","cited_arxiv_id":null,"evidence_quote":"Defines the quasi-interpolation operator with the approximation and stability properties used throughout the proofs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the matrix Chernoff bound that underpins Lemma 4.10, controlling the empirical Gram matrix in the observation region."},{"cited_title":"Sanz-Alonso and N","cited_arxiv_id":null,"evidence_quote":"Gives the earlier non-asymptotic finite-element posterior convergence analysis that the paper extends to sample-size dependent contraction rates."}],"review_version":1}