{"id":"1a38872b-5a5a-4656-99e9-768b279676b8","arxiv_id":"2412.05227","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A fixed-hyperparameter Gaussian-process prior recovers parton distributions from limited Ioffe-time data and yields more realistic small-x uncertainties than simple parametric fits.","lead":"Lattice QCD calculations only provide a handful of Fourier modes of a parton distribution. This paper shows a Gaussian-process prior with hand-picked, physically motivated hyperparameters can reconstruct the full x-shape while keeping the uncertainty under control.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The σ² rule ties posterior uncertainty to a single endpoint via the max(1.3ΔM(ν_max), 0.3|M(ν_max)|) envelope; the §IV closure test's narrow x^α(1-x)^β family cannot validate x-space coverage, and the z=9a case shows the reconstruction shifts sharply when the endpoint changes.","rationale":"The paper does what it claims at the machinery level: eqs. (7)–(9) are the standard analytic Gaussian-process posterior, the σ² scan is numerically trivial, and the closure test shows the posterior mean tracks smooth x^α(1-x)^β inputs with small-x uncertainty growing roughly as 2/ν_max. The real-data section is honest about the parametric Ansatz's failures, including the bimodal z=3a imaginary-part fit. These are real strengths and are acknowledged as such.\n\nThe advertised deliverable, however, is uncertainty control. That claim rests entirely on the §III σ² rule, whose 1.3 and 0.3 factors the authors themselves describe as 'dictated by a perception of what a good extrapolation uncertainty is.' The rule is well-defined — the posterior Fourier variance at fixed ν > ν_max increases with σ², so a unique σ² exists — but it calibrates a ν-space envelope, not x-space coverage. The closure test is the only calibration evidence, and it is weak: the input family is narrow and smooth, so the prior plays a minor role and the small-x uncertainty is largely an artifact of the imposed envelope rather than a reflection of the spread of plausible inputs. The paper's own admission that the forced extrapolation increase makes the reconstruction 'much more faithful' to the input confirms this.\n\nThe z=9a real-data case is the paper's clearest evidence of the rule's sensitivity: using the former-to-last datapoint instead of the last changes the x-space reconstruction markedly. Since the authors call the last-point uncertainty 'likely excessive,' the σ² rule amplifies a suspect statistical fluctuation into a global uncertainty band. The promised comparison with evidence-based hyperparameter sampling is not included, so the hand-picked targets are never benchmarked.\n\nI agree with the reader's weakest-assumption identification; the CONDITIONAL verdict is appropriate. The proposed coverage test is the decisive check: if the posterior 68% bands contain the truth at the expected rate for a broad, realistic input family, the concern is answered; if not, the uncertainty claim should be downgraded.","tokens_in":15760,"tokens_out":20252,"duration_ms":187767,"concrete_test":"Repeat the §IV closure test with a broader, more realistic input family: f_q(x) = N x^α(1-x)^β [1 + c cos(ω ln(1/x))], with α, β, c, ω drawn from uniform distributions α ∈ [-0.5, 0], β ∈ [1.5, 4], c ∈ [0, 0.3], ω ∈ [1, 10]; keep the §III hyperparameter prescription (l = ln 2, g = σ, σ² fixed by the max(1.3ΔM(ν_max), 0.3|M(ν_max)|) rule) unchanged. For ν_max ∈ {4, 10, 25} and 1000 samples each, compute the empirical coverage of the posterior 68% credible band at grid points x ≈ 0.02, 0.05, 0.10, 0.30, 0.51, 0.81. If pooled coverage deviates from 68% by more than about ±8%, the σ² rule does not provide honest x-space uncertainty for realistic inputs; report separately the sub-coverage for x < 0.1 to expose the small-x extrapolation region.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the fixed-hyperparameter GP (§III) yields reconstructions whose uncertainty one can 'understand and keep control easily.' The load-bearing element is the σ² rule: σ² is chosen so that the maximal posterior Fourier-space standard deviation at ν > ν_max equals max(1.3ΔM(ν_max), 0.3|M(ν_max)|). This rule anchors a ν-space envelope to a single endpoint; it says nothing directly about x-space calibration, and the posterior x-space width is an indirect image of that envelope through the log-x kernel and the data covariance.\n\nThe §IV closure test is too easy a target to validate that image. Its inputs are f_q = N x^α(1-x)^β with α ~ N(-0.2, 0.1), β ~ N(2.5, 0.5): a narrow family of smooth functions whose endpoint Fourier uncertainty ΔM(ν_max) is small, so the envelope is set by the two-parameter signal shape and the posterior is close to the data-constrained interpolation. Realistic PDFs have independent small-x and large-x structure, harder small-x growth, and non-smooth features; the mapping from the ν_max endpoint to x-space coverage is untested for such inputs. The authors themselves note it is 'unwarranted to assume that the first principles parton distribution would have the simple features of this two-parameter input distribution,' and that without the forced extrapolation increase the reconstruction would be 'much more faithful' — both admissions that the reported small-x uncertainty is imposed by the rule, not validated against the spread of plausible inputs.\n\nThe z=9a real-data case exposes fragility: ΔM(ν_max) is dominated by the last lattice point, which the authors call 'likely excessive, due to a poor lattice signal-to-noise.' Because σ² inherits that single point's uncertainty, the whole x-space reconstruction widens (Fig. 8 vs Fig. 9).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a non-parametric reconstruction of parton distribution functions f_q(x) from limited Fourier information M(ν), using a Gaussian process prior with hyperparameters fixed by a physical prescription. The prior mean is g(x)=σ, the covariance is a squared-exponential kernel in log-x with correlation length l=ln(2), and σ² is chosen so that the maximum posterior standard deviation in the extrapolation region ν∈[νmax,∞] equals max(1.3ΔM(νmax), 0.3|M(νmax)|). The posterior is a multivariate normal distribution that can be computed analytically, making the method numerically efficient. The authors test the procedure with a closure test using 1000 samples of the two-parameter family f_q=N x^α(1-x)^β for νmax=4, 10, and 25, and apply it to lattice QCD data from [30], comparing with a parametric fit of the form N x^α(1-x)^β. The central claim is that this simple, fixed-hyperparameter prescription gives reconstructions with uncertainties that are easy to understand and control.","tokens_in":16145,"tokens_out":4730,"duration_ms":50496,"significance":"If the central claim holds, the paper offers a computationally inexpensive and transparent alternative to evidence-based hyperparameter sampling for a class of ill-posed inverse problems in lattice QCD. The strengths of the paper are its clean use of the closed-form GP posterior, the explicit algorithm in Section III, the clarity of the closure-test protocol, and the real-data comparison that exposes known pathologies of parametric fits. The paper also gives a falsifiable heuristic, x≈2/νmax in Eq. (12), for the resolution limit. However, the validation is too narrow to fully support the main claim: the closure test uses only one smooth two-parameter input family, and the σ² prescription is calibrated to a single endpoint in Fourier space. The z=9a real-data example in Section V shows that the x-space uncertainty can be strongly altered by the noise of one endpoint. The paper is a useful proposal, but the load-bearing parts of the uncertainty-calibration claim require additional robustness and coverage analysis.","major_comments":[{"comment":"The σ² prescription sets the Fourier-space extrapolation uncertainty using only the endpoint value νmax, through max(1.3ΔM(νmax), 0.3|M(νmax)|). The z=9a real-data case in Section V demonstrates the fragility of this rule: when the last datapoint has an unusually large uncertainty, the reconstruction's x-space uncertainty widens considerably, and replacing the endpoint by the former-to-last datapoint gives a qualitatively different posterior (Fig. 8 vs. Fig. 9). The authors acknowledge this by saying 'we may want to not base our extrapolation on this point.' Since the stated goal is to 'understand and keep control easily of the uncertainty,' the procedure should include a robustness test or a criterion for identifying atypical endpoints, otherwise the uncertainty is controlled only in an average sense and can be dominated by a single noisy measurement.","section":"§III, Procedure item 3; §V, Fig. 8–9"},{"comment":"The closure test is performed with inputs drawn from the single smooth two-parameter family f_q=N x^α(1-x)^β with α~N(-0.2,0.1) and β~N(2.5,0.5). This family does not exercise the small-x/large-x decorrelation that the log-x kernel is designed to encode, because both regimes are tied to the same two parameters. The authors themselves note, at the end of Section IV, that 'it seems unwarranted to assume that the first principles parton distribution would have the simple features of this two-parameter input distribution.' To support the claim of honest uncertainty, the paper should include closure tests with inputs having independent small-x and large-x structure (e.g., sums of terms with different powers or inputs drawn from a broader GP), and it should report coverage probabilities, i.e., the fraction of inputs falling inside the 68% posterior credible intervals as a function of x.","section":"§IV, Closure test"},{"comment":"The paper argues against evidence-based hyperparameter sampling as in [20] on the grounds that it may over-estimate correlation lengths and under-estimate prior variances, but this is supported only by qualitative remarks and by a reference to 'a subsequent publication.' Since the central methodological novelty is the fixed-hyperparameter prescription, the absence of any quantitative comparison with hyperparameter sampling under the same kernel and the same ν range leaves the claimed advantage unsubstantiated. At minimum, the closure test should include a comparison of the fixed prescription with sampling-based hyperparameter inference for at least one νmax, or the paper should state more precisely what evidence would be needed to distinguish the two approaches.","section":"§III, hyperparameter discussion; comparison to [20]"}],"minor_comments":[{"comment":"The discussion of the functional determinant via zeta operators or a discretized path integral is not used in the rest of the paper; it could be shortened or moved to a footnote to avoid distracting the reader.","section":"§II, after Eq. (6)"},{"comment":"The wording 'the maximal standard deviation ... is exactly max(1.3ΔM(νmax), 0.3|M(νmax)|)' is a bit confusing because the right-hand side is itself a maximum; I suggest rephrasing to 'the posterior maximum over ν∈[νmax,∞] equals the larger of 1.3ΔM(νmax) and 0.3|M(νmax)|.'","section":"§III, Procedure item 3"},{"comment":"The caption does not define the normalization used for the y-axis; the text explains it, but the caption should state explicitly that the plotted quantity is (Reconstruction − mean(input))/std(input) so that the figure is self-contained.","section":"Fig. 2 caption"},{"comment":"The sentence 'Both the parametric and non-parametric then coincide well' would benefit from a quantitative measure of agreement (e.g., a χ² or the fraction of overlapping errors) rather than a visual claim.","section":"§V, Fig. 9 discussion"},{"comment":"There is a typo: 'per seas' should be 'per se.'","section":"§VI, GPD paragraph"},{"comment":"Several references are missing years (e.g., [8], [9], [13], [16], [17], [32]); these should be completed in the final version.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is clearly written and the proposed method is a potentially useful contribution to the lattice-QCD inverse-problem literature. My main concern is that the two central validation elements—the σ² rule's dependence on a single endpoint and the narrow closure-test family—are load-bearing for the claim that the uncertainty is controlled and honest. These are fixable within the scope of the paper by adding robustness tests and coverage studies, or by explicitly reframing the claim as a proposal with stated limitations. I would support publication after a revision that addresses these points; the paper need not be delayed for a full comparison with [20], but a quantitative statement about the comparison would strengthen it."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the fixed hyperparameter prescription: log-x kernel with l = ln(2), prior mean g(x) = sigma, and a sigma^2 rule calibrated to the desired Fourier-space extrapolation uncertainty. That is not a new framework — the GP machinery comes from earlier work — but the paper turns it into a turnkey procedure with an analytic Gaussian posterior, which is genuinely useful for lattice practitioners who want a fast, transparent alternative to parametric fits.\n\nWhat the paper does well: the GP algebra is standard and the presentation is clear. The closure tests cover three different nu_max values and show the expected behavior — resolution improving with range — and the real-data application is instructive, especially the comparison with the parametric fit. The z=3a imaginary-part case, where the parametric fit is bimodal and numerically unstable, makes a strong practical point for the GP approach. I also give real credit for the authors' honesty: they explicitly say the sigma^2 choice is \"dictated by a perception of what a good extrapolation uncertainty is,\" and they show the z=9a case where the rule makes the uncertainty jump because of a single noisy endpoint, then show what happens if you use the previous point instead. That is the kind of transparency I want to see.\n\nThe soft spots are real but not fatal. The sigma^2 rule anchors the whole uncertainty to the last data point's error; in the z=9a case that point is noisy, and the reconstruction's width changes visibly when you swap endpoints. That fragility is inherent to a prescription that fixes the envelope at nu_max, and the closure test's input family — smooth, two-parameter x^alpha(1-x)^beta — is too easy a target to validate x-space coverage for realistic PDFs with independent small-x and large-x physics. The paper acknowledges this too. The promised comparison to evidence-based hyperparameter sampling is deferred, so we don't yet know whether this prescription is better or just different. No code is shipped, but the method is simple enough that reproducing it is straightforward.\n\nWho is this for? Lattice QCD people reconstructing PDFs from pseudo-distributions or quasi-PDFs, and anyone facing a similar limited-Fourier inverse problem. It deserves a serious referee. I would ask for a stress test against the evidence-sampling baseline, and more diverse closure-test inputs, but the paper is a solid, useful contribution as it stands.","headline":"A practical, honestly presented GP recipe for PDF reconstruction from limited Ioffe-time data, whose main soft spot is the subjective sigma^2 calibration rule that the authors themselves acknowledge.","tokens_in":16715,"tokens_out":1418,"would_cite":true,"duration_ms":16931,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a Gaussian process prior with fixed, physically motivated hyperparameters turns a limited set of low-frequency Fourier modes of a parton distribution into a complete reconstruction with exactly Gaussian…","keywords":["parton distribution functions","Gaussian process prior","Bayesian inverse problem","Ioffe time","lattice QCD","non-parametric reconstruction","uncertainty quantification","Fourier modes"],"falsifier":"A coverage test would settle whether the uncertainty is controlled: simulate many datasets from known input distributions with realistic small-$x$ behavior, keep only the Fourier modes up to $\\nu_{\\max}=10$, apply the prescription, and count how often the true input lies inside the reported 68% posterior band; coverage clearly below 68% for $x$ below the resolution limit would show the fixed prior is too confident in the extrapolation region.","tokens_in":15576,"feed_emoji":"⚛️","tokens_out":9981,"duration_ms":94908,"temperature":0.7,"pith_summary":"Parton distribution functions (PDFs) encode how quarks and gluons share a hadron's momentum, but first-principles lattice calculations only give access to a limited range of Fourier modes (Ioffe times $\\nu$) of the distribution. This paper tries to turn that incomplete Fourier information into a full PDF without imposing a rigid parametric shape. The proposed route is a Gaussian process prior on $x$-space with a logarithmic correlation kernel, with all hyperparameters fixed by physical rules rather than fitted to the data. Because the hyperparameters are fixed, the posterior is exactly Gaussian and analytic, so the reconstruction and its uncertainty can be computed in about a second and shared exactly. If the prior choices are accepted, a handful of low-$\\nu$ modes plus this prescription yields a PDF reconstruction with controlled, tunable uncertainty down to small $x$, where common parametric fits tend to produce artificially narrow error bands.","feed_headline":"Fixed prior recovers parton distributions from limited Fourier data","feed_subtitle":"Lattice QCD reaches only a few Fourier modes; this prior turns them into a PDF with controlled uncertainty.","key_machinery":"The central object is the Gaussian process posterior evaluated on a grid of $x$ values, with the logarithmic kernel $K(x,x')=\\sigma^2 \\exp\\left(-\\frac{(\\ln x-\\ln x')^2}{2l^2}\\right)$ serving as the prior covariance. The fixed hyperparameter prescription, $l=\\ln 2$, $g(x)=\\sigma$, and the Fourier-extrapolation rule for $\\sigma^2$, is what makes the procedure non-parametric yet fully specified. The posterior mean and covariance are obtained analytically from $\\langle f_q(x_i)\\rangle = g(x_i)+KB^T[C+BKB^T]^{-1}[M-Bg]$ and $H=(K^{-1}+B^T C^{-1}B)^{-1}$, so the entire reconstruction is an exact multivariate normal distribution and no sampling is required. The kernel's logarithmic argument is the physics-carrying ingredient: it decorrelates behavior at small $x$ from behavior near $x=1$ while enforcing smoothness across ratios $e^{-l}<x/x'<e^{l}$.","core_discovery":"The paper's central claim is that the ill-posed inverse problem $M(\\nu)=\\int_{-1}^{1} dx\\, e^{ix\\nu} f_q(x)$ can be regularized non-parametrically by a Gaussian process prior on $x\\in[0,1]$, with covariance $K(x,x')=\\sigma^2 \\exp\\left(-\\frac{(\\ln x-\\ln x')^2}{2l^2}\\right)$, mean $g(x)=\\sigma$, and the fixed choices $l=\\ln 2$, $g(x)=\\sigma$, and $\\sigma^2$ chosen so that the largest standard deviation of the posterior's Fourier transform for $\\nu\\ge\\nu_{\\max}$ equals $\\max(1.3\\Delta M(\\nu_{\\max}), 0.3|M(\\nu_{\\max})|)$. With these choices the posterior is a multivariate normal whose mean and covariance follow from closed-form linear equations, and the paper shows in closure tests and on real lattice data that the reconstruction tracks the input where the data has information and that its uncertainty in $x$ grows roughly as $x\\approx 2/\\nu_{\\max}$ toward small $x$, while remaining under explicit control in the Fourier extrapolation region.","pith_inferences":["Editorial extension: the resolution heuristic $x\\approx 2/\\nu_{\\max}$ gives a concrete planning target—resolving momentum fraction $x_0$ requires Ioffe-time coverage $\\nu_{\\max}\\gtrsim 2/x_0$, which could be used to design future lattice ensembles or to decide when data are sufficient.","Editorial extension: because $\\sigma^2$ is set from the uncertainty of the last data point, the error bars inherit that point's noise; a robustified variant would estimate the endpoint uncertainty from several trailing points or from a smoothed variance, and the paper's $z=9a$ comparison shows how strongly this choice matters.","Editorial extension: the kernel's log-decorrelation assumption is testable—one could introduce a second correlation length or a cross-term linking small and large $x$ and ask whether the data likelihood favors it, which would change the extrapolation uncertainty."],"forward_implications":["If the procedure works as claimed, a lattice calculation reaching only $\\nu_{\\max}\\approx 10$ still yields a useful PDF reconstruction with uncertainty growing near $x\\approx 2/\\nu_{\\max}\\approx 0.2$, and the resolved region extends as $\\nu_{\\max}$ grows.","Because the posterior is a multivariate normal, the full reconstruction can be stored and shared exactly as a mean vector and covariance matrix, without Monte Carlo replicas.","In the small-$x$ region the non-parametric reconstruction avoids the artificially narrow error bands and numerical instabilities of the $N x^\\alpha(1-x)^\\beta$ parametric fit, including the bimodal fit distribution seen for the imaginary part at $z=3a$.","The same fixed prescription transfers to quasi-PDFs by extending the support to $x\\in[0,2]$, where the reconstruction yields tiny contributions outside $[0,1]$ that shrink as the hadron momentum increases."],"supporting_citations":[{"why":"Establishes the non-local matrix element and the Fourier relation $M(\\nu)=\\int dx e^{ix\\nu} f_q(x)$ that define the inverse problem.","marker":"[1-4]"},{"why":"Supplies the Gaussian-process Bayesian prior framework in $x$-space, which this paper adopts but fixes the hyperparameters instead of sampling them.","marker":"[20]"},{"why":"Provides the definition and closed-form posterior machinery for Gaussian processes used in eqs. (7)-(9).","marker":"[26]"},{"why":"Earlier use of Gaussian processes for lattice parton distributions that motivates the present approach.","marker":"[19]"},{"why":"Earlier comparison of Bayesian and neural-network reconstruction methods that frames the non-parametric alternative and motivates avoiding parametric bias.","marker":"[9]"},{"why":"Supplies the real lattice QCD dataset (unpolarized PDF at $a=0.094$ fm, $P=2.47$ GeV) used for the application and comparison with parametric fits.","marker":"[30]"},{"why":"Documents numerical instability in $N x^\\alpha(1-x)^\\beta$ fits, which the proposed method avoids in the comparison.","marker":"[31]"},{"why":"Critical review of the $x^\\alpha(1-x)^\\beta$ parametric model's unrealistic extrapolation uncertainties, justifying the non-parametric prior.","marker":"[28]"}],"fun_headline_variants":["Gaussian prior reconstructs PDFs from sparse Fourier data","Bayesian fix for limited Fourier modes in parton data","Controlled uncertainty in PDF reconstruction via GP prior","From few Fourier modes to full parton distributions","Simple GP regularization for inverse problems in QCD"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's reliability rests on the prior's fixed structure: that parton distributions decorrelate between small and large $x$ on a log scale, and that inflating the last data point's uncertainty by 30% (or 30% of the signal when the signal dominates) is the right way to set extrapolation uncertainty.","fun_headline_variants_meta":{"raw":{"variants":["Gaussian prior reconstructs PDFs from sparse Fourier data","Bayesian fix for limited Fourier modes in parton data","Controlled uncertainty in PDF reconstruction via GP prior","From few Fourier modes to full parton distributions","Simple GP regularization for inverse problems in QCD"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1310,"prompt_tokens":853,"completion_tokens":457,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":382}},"tokens_in":469,"tokens_out":457,"duration_ms":4243,"temperature":1.0,"reasoning_tokens":382,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:49:22.221172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A coverage test would settle whether the uncertainty is controlled: simulate many datasets from known input distributions with realistic small-$x$ behavior, keep only the Fourier modes up to $\\nu_{\\max}=10$, apply the prescription, and count how often the true input lies inside the reported 68% posterior band; coverage clearly below 68% for $x$ below the resolution limit would show the fixed prior is too confident in the extrapolation region.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier use of Gaussian processes for lattice parton distributions that motivates the present approach."}],"review_version":1}