{"id":"e524904c-b965-4478-9583-e6a946faa901","arxiv_id":"2504.21267","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A two-stage likelihood reweighting method recovers the parameters and Bayes factor of an added sinusoid signal in simulated pulsar timing data, matching full Bayesian analyses while claiming about a tenfold speedup.","lead":"Pulsar timing arrays have found a background hum of gravitational waves, and scientists now want to efficiently test whether extra signals hide in that background. This paper shows that a computational trick called likelihood reweighting can screen such extra-signal models about ten times faster than a full analysis, at least in simulated data with a simple periodic signal.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Bayes-factor derivation for the second reweighting (Eqs. 19-22) is internally inconsistent: the factor rho is not uniquely defined by the preceding equations, so the reported evidence ratios are not reproducible from the text as written.","rationale":"The reader's verdict was CONDITIONAL, and the reader's rationale already noted that 'the written derivation of the second-reweighting Bayes factor (Eqs. 20-22) is internally inconsistent.' I agree that this is a genuine defect, and I elevate it to the single most load-bearing concern because the Bayes-factor recovery is a stated headline result and the ambiguity concerns the exact multiplicative factor needed to go from the auxiliary-model evidence to the proposal-model Bayes factor. The reader's explicit weakest_assumption focused instead on target-proposal overlap and weight variance; that is a real limitation but it is acknowledged by the authors and is partly mitigated by the second reweighting. The rho ambiguity is more concrete and more directly tied to the claimed numerical agreement. My proposed test distinguishes a purely typographical/definitional slip from a substantive error in the method. If the correct rho is simply recoverable from context (e.g., the authors implicitly used the normalized pKDE), then the paper needs only a clarification and the CONDITIONAL verdict stands. If the literal equations are what was implemented, the reported Bayes factors may be wrong and the central claim is not supported. I also noted the abstract's 'fiducial SGWB' wording versus the CURN proposal actually used, and the O(10) speedup being an estimate rather than a measurement, but these are secondary and do not change the verdict. Therefore I recommend no change to the reader's CONDITIONAL verdict.","tokens_in":14276,"tokens_out":10996,"duration_ms":120005,"concrete_test":"Recompute the weak-signal Bayes factor from the second reweighting using the two candidate definitions of rho: (A) rho = [integral fKDE(y) pi(y) dy]^-1, which follows from Eq. (20), and (B) rho = [integral pKDE(y) pi(y) dy]^-1 with pKDE(y)=fKDE(y)pi(y), which is the literal reading of Eq. (21). If interpretation A reproduces the hypermodel value B=1.48+-0.04 and interpretation B does not (or vice versa), the text needs a corrected definition of rho. If neither matches, the reported Bayes-factor agreement is suspect and the second-reweighting formula must be revisited. The test should use the same KDE samples and the same proposal chain, changing only the rho factor, and should report the resulting B values alongside the stated uncertainties.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim is that the second reweighting recovers the CURN+sin Bayes factor compatibly with hypermodel MCMC (B = 1.63 vs 1.48, 11.07 vs 10.16, 491 vs 448). That claim rests on Eq. (22), B = rho * mean(w'), where rho is supposed to convert the auxiliary-model evidence ratio Z1/Z0 into the desired Z/Z0. But the derivation of rho is not self-consistent. The authors define pKDE(y) = fKDE(y) * pi(y) in the paragraph before Eq. (12), and later state that fKDE is a probability distribution almost equal to pKDE. Eq. (20) gives Z1/Z0 = integral fKDE(y) pi(y) dy, while Eq. (21) approximates this by integral pKDE(y) pi(y) dy. If pKDE means fKDE*pi, then Eq. (21) contains an extra factor pi(y) inside the integral; if pKDE means the normalized density (which equals fKDE only for uniform priors), then the definition is different and the equality with Eq. (20) is still not generally true. The actual priors on Asin and fsin are log-uniform, not uniform, so the two readings differ by the non-constant factor integral fKDE pi^2 / integral fKDE pi. The symbol rho in Eq. (22) is the inverse of whichever integral is used, and the paper never states which. Without this specification, a reader cannot reproduce the reported Bayes factors, and the apparent ~10% agreement with hypermodel values could be the result of an unstated correct choice of rho. Because the Bayes-factor agreement is one of the paper's headline outcomes, this ambiguity is load-bearing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage likelihood-reweighting method for pulsar timing array (PTA) searches beyond a fiducial stochastic gravitational wave background. Using posterior samples from a common uncorrelated red noise (CURN) model as a proposal, the first reweighting assigns likelihood-ratio weights to estimate the posterior and Bayes factor of a CURN plus global-sinusoid target model. To mitigate the large importance-weight variance that arises when the target is not a small perturbation of the proposal, the authors introduce a second reweighting step: kernel density estimates of the first-stage marginal posteriors of the sinusoid parameters are used to construct an adaptive proposal, and a second set of importance weights is computed. The method is tested on three simulated 20-pulsar datasets with sinusoid amplitudes corresponding to Bayes factors of roughly 1, 10, and 500. The reweighted posteriors are compared visually with full MCMC posteriors, and the reweighted Bayes factors are compared with hypermodel MCMC values. The authors claim that the method gives compatible results and provides at least an order-of-magnitude speedup over full analyses.","tokens_in":14696,"tokens_out":6788,"duration_ms":71805,"significance":"If validated, the method would be practically useful: PTA collaborations could reuse one fiducial CURN chain to screen many beyond-fiducial models, including global sinusoids, modified Hellings-Downs correlations, and continuous-wave-like signals. The second-reweighting innovation, using KDE-based adaptive proposals to temper importance-weight variance, is a reasonable extension of existing likelihood-reweighting work in gravitational-wave astronomy. However, the quantitative Bayes-factor claims currently rest on a derivation that is not self-consistent, and the reported speedup is asserted rather than measured. The central idea is sound and the simulated comparisons are encouraging, but the manuscript needs a corrected derivation and a more careful statement of the speedup claim before the results can be relied upon.","major_comments":[{"comment":"The derivation of the second-stage Bayes factor is not self-consistent and the reported numbers cannot be reproduced as written. With pKDE(y)=fKDE(y)π(y), the product p0(x)pKDE(y) equals L1(x,y)π0(x)π(y)/Z0, so the fraction in Eq. (17) is Z1/Z0, not 1; the displayed equality therefore does not lead to Eq. (18). In addition, Eq. (21) approximates ∫fKDE(y)π(y)dy by ∫pKDE(y)π(y)dy: if pKDE=fKDEπ this introduces an extra factor π(y), while if pKDE is taken to be the normalized density ≈fKDE, the definition in the text is contradicted. Since the priors on Asin and fsin are log-uniform, the two readings differ by the factor ∫fKDEπ^2 / ∫fKDEπ. The symbol ρ in Eq. (22) is the inverse of whichever integral is intended, but the paper never states which. The reported Bayes factors B=1.63, 11.07, and 491 all depend on this unstated choice. Please correct the derivation (e.g., define ρ≡Z0/Z1=[∫fKDE(y)π(y)dy]^{-1} and compute it explicitly) or report evidence estimates obtained by a direct method.","section":"II B, Eqs. (16)-(22)"},{"comment":"The 'at least O(10)-time speedup' claim is asserted rather than demonstrated. The estimate assumes that the number of samples needed for a converged target chain equals that of the proposal chain and that thinning is performed every 10 samples, but no wall-clock measurements, effective-sample-size comparisons, or convergence diagnostics for the target chains are reported. The claim should either be backed by benchmark timings for the actual likelihood evaluations used here or be explicitly qualified as a conditional estimate that applies only when the target likelihood is substantially more expensive than the proposal likelihood.","section":"III B3"},{"comment":"The normalization of pKDE is ambiguous. If pKDE(y)=fKDE(y)π(y) as stated, then pKDE does not integrate to 1 for the non-uniform priors used in this paper, so the instruction 'draw N samples from pKDE(y)' is not well defined without specifying a normalization constant. This ambiguity is directly connected to the ρ inconsistency in Eqs. (19)-(22) and should be resolved in the same revision.","section":"II B, paragraph after Eq. (11)"}],"minor_comments":[{"comment":"The convergence check using one MCMC chain split in half after burn-in and requiring R<1.005 is not the standard Gelman-Rubin diagnostic, which compares multiple independent chains. Please either rename this as a within-chain consistency check or justify why it is adequate for the results presented.","section":"III B"},{"comment":"The standard error ΔB=σw/√N assumes independent samples, but the proposal chain is serially correlated. If the chain is thinned, this should be stated; otherwise an effective-sample-size correction should be used.","section":"II A, Eq. (9)"},{"comment":"The notation 'log-uniform [−18,−11]' is ambiguous. Please state explicitly that log10 of the parameter is drawn from a uniform distribution on the indicated interval, for all such entries.","section":"Table I"},{"comment":"The comparison between reweighted and full-MCMC posteriors is purely visual. Reporting a quantitative measure (e.g., Kullback-Leibler divergence or credible-interval overlap) would make the claimed compatibility more precise.","section":"III B1"}],"recommendation":"major_revision","confidential_remarks":"The paper's novelty is modest relative to the existing likelihood-reweighting literature, but the KDE-based second reweighting is a useful practical extension. The main risk is reproducibility: the Bayes-factor numbers in the headline comparisons depend on an unstated normalization choice. If the authors fix the derivation and either measure or properly qualify the speedup claim, the paper could be suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, the paper does something genuinely useful: it takes the standard likelihood-reweighting idea from Hourihane et al. and adds a second reweighting pass whose proposal is built from KDEs of the first pass's marginal posteriors. That two-pass recipe is new as a packaged tool, and the simulations show it recovers reweighted posteriors and Bayes factors that agree with full hypermodel MCMC to about ten percent for weak and moderate signals. Second, the written derivation of the second-pass Bayes factor is not consistent as it stands, and because that Bayes-factor agreement is one of the paper's two headline outcomes, this matters.\n\nThe importance-sampling identities in Section II are standard and correct. The paper is honest about unstable weights in the strong-signal regime, and it is right that the first reweighting still finds the sinusoid parameters even when the background posterior is badly reconstructed. The simulated datasets are appropriate, and the comparisons against full MCMC are the right check.\n\nThe soft spots are real but addressable. The derivation of rho in Eqs. (19)-(22) is muddled: pKDE is defined as fKDE(y) pi(y), but then the text approximates Z1/Z0 with an integral over pKDE(y) pi(y), which would double-count the prior. Either pKDE is the unnormalized posterior and the pi factor is wrong, or pKDE is meant to be the normalized KDE and the definition is different. The paper never pins this down, so the reported Bayes factors are not reproducible from the text. That is the load-bearing issue. Additionally, no code or simulated data are released, the O(10) speedup is estimated from assumptions rather than measured, and the convergence check splits a single MCMC chain, which is weak. The abstract's 'fiducial SGWB' is actually CURN; the text clarifies this, but it is sloppy.\n\nNone of this makes the core idea wrong. The method is a modest, practical extension of known importance-sampling techniques, and for PTA groups screening many beyond-fiducial models from one chain it could save real compute. With a clean derivation and code release, I'd take the numbers seriously. As written, I'd want a referee to require those fixes. This deserves a serious referee, and I'd bring it to a reading group to work through the evidence-ratio logic.","headline":"A useful two-pass reweighting recipe for PTA model screening, with a clean core idea but a Bayes-factor derivation that needs fixing before the headline numbers can be trusted.","tokens_in":15177,"tokens_out":4496,"would_cite":false,"duration_ms":44690,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage likelihood reweighting recovers the posterior and Bayes factor of a signal-plus-background pulsar timing array model from a fiducial background-only chain, with at least an order-of-magnitude speedup over full MCMC.","keywords":["gravitational wave background","pulsar timing arrays","Bayesian inference","importance sampling","likelihood reweighting","kernel density estimation","Bayesian model selection","Hellings-Downs correlation"],"falsifier":"Run the same two-stage reweighting on a fourth simulated dataset with a sinusoid injected at $\\log_{10} A_{\\mathrm{sin}} = -6.5$, just beyond the paper's strong case, and compare the averaged-weight Bayes factor and the reweighted posterior of $\\log_{10} A_{\\mathrm{SGWB}}$ against a full MCMC run of the same target: if the Bayes factor disagrees by more than the reported differences (e.g., $491\\pm23$ versus $448\\pm42$) or the SGWB posterior separates from the full-search posterior, the claimed screening reliability has a sharp amplitude ceiling.","tokens_in":14016,"feed_emoji":"📡","tokens_out":8658,"duration_ms":85660,"temperature":0.7,"pith_summary":"PTA collaborations have found evidence for a stochastic gravitational wave background, and a natural next question is whether the data contain features beyond the fiducial correlated background: deterministic signals, modified-gravity correlation patterns, or a dominant single supermassive-black-hole binary. This paper argues that such beyond-fiducial models can be screened without running a new Markov chain every time. Starting from posterior samples of a common-uncorrelated-red-noise (CURN) model, it assigns each sample a weight equal to the likelihood ratio between the target model and CURN, yielding reweighted posteriors and a Bayes factor given by the average weight. Because vanilla importance sampling has unstable weights, the authors reweight a second time using a kernel-density estimate of the new parameters, and test the procedure on three simulated datasets with a global sinusoid injected at weak, moderate, and strong amplitudes. They find that the reweighted posteriors and Bayes factors are consistent with full Bayesian runs, that the sinusoid parameters are recovered well even when the signal dominates, and that the procedure offers an at least order-of-magnitude speedup.","feed_headline":"Reweighted PTA samples screen new signals 10x faster","feed_subtitle":"A two-stage likelihood reweighting recovers signal posteriors and Bayes factors from one background-only chain.","key_machinery":"The load-bearing object is the likelihood-ratio weight $w(x,y)=L_{\\mathrm{target}}(x,y)/L_{\\mathrm{proposal}}(x)$, assigned to each proposal MCMC sample so that the weighted samples approximate the target posterior and the sample mean of the weights estimates the Bayes factor $B=Z/Z_0\\approx\\bar{w}$. Because vanilla importance sampling has unstable weights when the models are dissimilar, the second mechanism is a KDE-adapted proposal: one-dimensional kernel density estimates $f_{\\mathrm{KDE}}(y)$ of the new parameters from the first reweighting form a new likelihood $L_1(x,y)=L_0(x)f_{\\mathrm{KDE}}(y)$, and second-round weights $w'=L_{\\mathrm{target}}/L_1$ are computed. The diagnostic $N_{\\mathrm{eff}}=N/(1+(\\sigma_w/\\bar{w})^2)$ tells whether the weight distribution is healthy enough to trust the result. These objects carry the argument: the first weight extracts signal information from the fiducial chain, the second stabilizes the estimate, and $N_{\\mathrm{eff}}$ flags when the method is being pushed past its validity.","core_discovery":"On its own terms, the paper's central claim is that two-stage likelihood reweighting is a reliable screening tool for beyond-fiducial PTA models. For a target model that adds parameters $y$ to a proposal with likelihood $L_0$ and posterior samples $x^{(i)}$, reweighting by $w_i = L_{\\mathrm{target}}(x^{(i)}, y^{(i)})/L_0(x^{(i)})$ produces target posterior samples with the correct weights, and the target/proposal Bayes factor is the mean weight. The second reweighting replaces the uniform prior draws over $y$ with samples from one-dimensional KDEs of the first-round marginal posteriors, forming a new proposal likelihood $L_1 = L_0 \\, f_{\\mathrm{KDE}}(y)$ and weights $w'_i = L_{\\mathrm{target}}/L_1$. In the paper's three sinusoid-injection tests, the Bayes factors from the second reweighting ($1.63\\pm0.03$, $11.07\\pm0.23$, $491\\pm23$) match those from full hypermodel MCMC runs ($1.48\\pm0.04$, $10.16\\pm0.74$, $448\\pm42$), and the reweighted signal posteriors overlap the full searches. The method's stated boundary is model similarity: when the signal dominates, the SGWB posterior recovery degrades, although the signal parameters are still captured.","pith_inferences":["Beyond the paper's tests, a sky-located continuous-wave source is the natural next probe: it adds two sky-position parameters to the sinusoid, and the paper explicitly notes this as a physically interesting extension, so a failure there would mark the practical screening boundary for deterministic-source searches.","The dimension limit the paper states (no more than roughly ten new parameters) implies that modified-gravity correlation models that change many angular correlation coefficients at once are unlikely to be screenable by this route; those models would need a different proposal.","If this workflow is adopted, the fiducial chain becomes a reusable asset: each new target model costs only weight evaluations, so a single background-only analysis can support a fast model survey, and the effective sample size can be monitored in real time as a stopping rule."],"forward_implications":["A single converged CURN chain can be reused to screen many target models, since the weights are computed in parallel and the proposal chain is thinned before reweighting.","Small Bayes factors from reweighting can set upper limits on new-signal parameters, while large Bayes factors flag targets that deserve a full MCMC follow-up.","The second KDE-based reweighting extends the usable range of importance sampling to moderately strong signals, where a single reweighting has visibly unstable weights.","An at least order-of-magnitude speedup over a full target search follows whenever the target likelihood is the expensive part of the calculation.","The signal parameters themselves are recovered even in the strong-signal regime where the SGWB posterior is degraded, so screening can still detect a dominant sinusoid."],"supporting_citations":[{"why":"Supplies the observational context of the detected stochastic background and the ~3.95 nHz sinusoid search that motivates the toy model.","marker":"[1]"},{"why":"Provides the stepping-stone sampling algorithm that the reweighting approach is closely related to.","marker":"[43]"},{"why":"Earlier use of likelihood reweighting for higher-order modes in the LIGO context that this paper extends to PTA data.","marker":"[47]"},{"why":"Earlier use of likelihood reweighting for eccentric binary orbits, another target-model screening application.","marker":"[48]"},{"why":"Shows likelihood reweighting can recover the Hellings-Downs background from a CURN chain, the direct PTA precedent.","marker":"[50]"},{"why":"Review of adaptive importance sampling that motivates the second KDE-based reweighting.","marker":"[51]"},{"why":"Software used to compute the PTA likelihoods in the simulated datasets.","marker":"[58]"},{"why":"The hypermodel method whose Bayes factors serve as the full-analysis reference for comparisons.","marker":"[65]"}],"fun_headline_variants":["Two-stage reweighting finds PTA signals 10x faster","Reweighting screens beyond-fiducial PTA models at 10x speed","Fast PTA model search: reweighting matches full Bayesian","KDE reweighting recovers signal posteriors from one chain","PTA likelihood reweighting: 10x faster, same accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The target model must be a small perturbation of the proposal model so that the likelihood-ratio weights have small variance; the paper states this similarity requirement, and when the signal is strong enough to violate it, the weights and Bayes-factor estimates become unstable.","fun_headline_variants_meta":{"raw":{"variants":["Two-stage reweighting finds PTA signals 10x faster","Reweighting screens beyond-fiducial PTA models at 10x speed","Fast PTA model search: reweighting matches full Bayesian","KDE reweighting recovers signal posteriors from one chain","PTA likelihood reweighting: 10x faster, same accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000275,"raw_usage":{"total_tokens":1693,"prompt_tokens":1042,"completion_tokens":651,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":556}},"tokens_in":658,"tokens_out":651,"duration_ms":6201,"temperature":1.0,"reasoning_tokens":556,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:09:32.624231+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two-stage reweighting on a fourth simulated dataset with a sinusoid injected at $\\log_{10} A_{\\mathrm{sin}} = -6.5$, just beyond the paper's strong case, and compare the averaged-weight Bayes factor and the reweighted posterior of $\\log_{10} A_{\\mathrm{SGWB}}$ against a full MCMC run of the same target: if the Bayes factor disagrees by more than the reported differences (e.g., $491\\pm23$ versus $448\\pm42$) or the SGWB posterior separates from the full-search posterior, the claimed screening reliability has a sharp amplitude ceiling.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the observational context of the detected stochastic background and the ~3.95 nHz sinusoid search that motivates the toy model."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier use of likelihood reweighting for higher-order modes in the LIGO context that this paper extends to PTA data."},{"cited_title":"(13)], where N was mistakenly written as Neff defined below","cited_arxiv_id":null,"evidence_quote":"Shows likelihood reweighting can recover the Hellings-Downs background from a CURN chain, the direct PTA precedent."},{"cited_title":"Vitoratou and I","cited_arxiv_id":null,"evidence_quote":"Review of adaptive importance sampling that motivates the second KDE-based reweighting."}],"review_version":1}