{"id":"0d17005e-5487-4d27-9f70-b5800e9030db","arxiv_id":"2506.19206","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Precomputed inner products make coherent Bayesian gravitational wave searches on astrometric data tractable, yielding Kepler and Roman strain sensitivities around 10^-12.4 and 10^-11.4 from 10 nHz to 1 microHz.","lead":"Using a data-compression trick from pulsar timing searches, the authors show Bayesian coherent gravitational wave searches on astrometric data can be sped up by up to a factor of 100, and they forecast that Kepler and Roman could reach strains near 10^-12.4 and 10^-11.4 in the microhertz band. A generalist might read this because it suggests a way to probe the untapped frequency gap between pulsar timing arrays and LISA using existing or near-future telescopes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 1%-accuracy claim is validated at a single frequency; the same F grid may be insufficient at the high-frequency end of the two-decade band.","rationale":"The paper is transparent and the compression method itself is a credible extension of the PTA inner-product precomputation to astrometric data. The most load-bearing claim, however, is that the reduction remains accurate to within 1% over two orders of magnitude in frequency, and the validation in Section III A is performed at a single frequency. The reader's formal weakest_assumption concerns frequency evolution and chirp mass, which the paper explicitly acknowledges and scopes; that concern is real but already disclosed. The single-frequency validation is more directly tied to the abstract's headline claim and is not adequately addressed by the chirp-mass restriction, since even with constant-frequency injections the interpolation error can grow with f_I T_obs. I therefore identify the untested frequency dependence of the 1% accuracy contour as the primary concern. The proposed concrete test would settle it by recomputing the evidence-error curves at several frequencies. If the error is indeed larger at high f, the appropriate fix is to quote F values per frequency range or restrict the claimed accuracy band; the method's core idea would remain valid. This does not change the reader's CONDITIONAL verdict: the paper should be conditional on demonstrating the accuracy claim across the actual frequency range, either by additional runs or by an analytic bound on the interpolation error.","tokens_in":17130,"tokens_out":8620,"duration_ms":100964,"concrete_test":"Repeat the Section III A procedure at injected frequencies f_I = 10^-8, 10^-7.5, 10^-6.5, and 10^-6 Hz for both surveys, using the same exact nested sampling run to obtain dtheta_E,i and then the same linear-in-log-frequency interpolation at F = 3140 (Kepler) and F = 15500 (Roman). Report |Z(F) - Z_E|/Z_E at each frequency. If the error at 10^-6 Hz exceeds 1%, or if the F required for 1% at 10^-6 Hz is much larger than at 10^-7 Hz, the abstract and Section III A should be revised to state the valid frequency range for the quoted F values and compression factors.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section III A and Figure 2 measure the interpolation error |Z(F) - Z_E|/Z_E for one injected frequency only, f_I = 10^-7 Hz, and the quoted grid sizes (F = 3140 for Kepler, F = 15500 for Roman) are chosen from that single curve. The same F is then used across the log-frequency prior [10^-8, 10^-6] Hz in Section III B, and the abstract claims accuracy 'within 1% over two orders of magnitude in gravitational wave frequency.' That extrapolation is not supported by the presented test. The interpolation error for the Fourier-type inner products is not frequency-independent: the likelihood peak width in log f is approximately 1/(f_I T_obs), so for a fixed log-spaced grid the number of grid points per independent frequency bin decreases linearly with f_I. For Kepler, T_obs is about 1.3e8 s, so f_I T_obs rises from ~1.3 at 10^-8 Hz to ~130 at 10^-6 Hz; at 10^-6 Hz the F = 3140 grid has only ~5 points per independent frequency bin, versus ~50 at 10^-7 Hz. Linear interpolation of an oscillatory inner product with only ~5 points per oscillation is expected to be substantially less accurate, and the paper's own observation that per-sample evidence errors are always negative indicates the coarse grid systematically underestimates narrow likelihood peaks. This is not an internal contradiction of the method, but it means the headline accuracy claim currently rests on a single-frequency validation and has no demonstrated support across the two-decade band.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the pulsar-timing-array inner-product precomputation scheme of Bécsy et al. to coherent gravitational-wave searches in relative astrometry. By rewriting the Gaussian likelihood as a linear combination of precomputed inner products between the astrometric data and sine/cosine filter functions, and by interpolating these inner products on a log-spaced frequency grid, the authors compress the dataset from O(NT) to O(NF). They validate the approximation by comparing the interpolated Bayesian evidence with an exact nested-sampling evidence calculation for an injected f_I = 10^-7 Hz source, reporting <1% evidence error at F = 3140 for Kepler and F = 15500 for Roman. They then use the compressed likelihood with JAXNS to forecast B = 10 strain thresholds across 10^-8 to 10^-6 Hz, obtaining h0 about 10^-12.45 for Kepler and about 10^-11.39 for Roman. The paper also reports poor sky localization and a marginal sensitivity gain from position-targeted searches.","tokens_in":17395,"tokens_out":15557,"duration_ms":157005,"significance":"If the accuracy claim is robust across the full frequency band, the compression step is a practical enabler for coherent microhertz astrometric searches on terabyte-scale datasets, and the Kepler/Roman forecasts provide concrete, falsifiable sensitivity targets. The method is an independent extension of existing PTA techniques, the nested-sampling setup is documented in sufficient detail for reproducibility, and the paper is transparent about the constant-frequency limitation. The main caveats are that the 1% accuracy is demonstrated at only one frequency, the quoted thresholds are for face-on binaries only, and the empirical-Bayes evidence correction is not calibrated; these issues affect the headline quantitative claims.","major_comments":[{"comment":"The headline claim that the compression is accurate 'to within 1% over two orders of magnitude in gravitational wave frequency' is supported by a single injected frequency, f_I = 10^-7 Hz, in Section III.A. For a log-spaced grid, the number of grid points per independent 1/T_obs frequency bin is proportional to (f_I T_obs)^-1. For Kepler (T_obs about 1.3e8 s), the F = 3140 grid has roughly 50 points per bin at f = 10^-7 Hz but only about 5 points per bin at f = 10^-6 Hz; for Roman (F = 15500) the corresponding drop is from about 800 to 90. The systematic negative per-sample evidence error noted by the authors indicates that under-resolved likelihood peaks are underestimated, so the same F cannot automatically be assumed accurate at the high-frequency end. Repeating the evidence-accuracy test at f_I = 10^-8 Hz and f_I = 10^-6 Hz, or deriving a frequency-dependent F requirement, is necessary to support the abstract's two-order-of-magnitude claim.","section":"III.A and Figure 2"},{"comment":"The quoted B = 10 thresholds (h0 about 10^-12.45 for Kepler and about 10^-11.39 for Roman, used in the abstract and conclusion) are averages over sky position only; the full injection grid is run solely for face-on binaries (cos i = 1). The paper's own check at f_I = 10^-7 Hz for Kepler shows that the threshold worsens by Delta log10 h0 about 0.2 at cos i = 0.5 and by a similar amount from 45 degrees to edge-on. For a population with random inclinations, the average threshold strain would therefore be roughly 0.2-0.4 dex higher. The headline numbers should either be labeled as face-on, optimally oriented sensitivities, or the forecasts should marginalize over inclination.","section":"III.B, Figure 3, and Abstract"},{"comment":"In Section II.E, the Bayesian evidence is corrected by multiplying Z_GW by 10^(W-N), where N is the 1-99% width of the h0 posterior obtained from the same data. This is an empirical-Bayes plug-in procedure rather than a Bayes factor under a fixed prior, and the null distribution of the resulting statistic is not characterized. Because the sensitivity thresholds in Section III.B are defined by B = 10, a calibration check with noise-only injections (or an analytic statement of the implied false-alarm rate) is needed to show that the correction does not distort the detection statistic. At minimum, the text should state explicitly that B is a corrected empirical-Bayes quantity rather than a proper Bayes factor.","section":"II.E and III.B"}],"minor_comments":[{"comment":"Each curve in Figure 2 comes from a single nested-sampling run, so the F values quoted for 1% accuracy have no associated uncertainty; a few independent realizations would make the threshold more robust.","section":"III.A and Figure 2"},{"comment":"The statement that the evidence error does not depend on the number of observations per star or on the signal-to-noise ratio is asserted without a supporting test; since this scaling is later used to justify F = 1000 in Section III.B, it should be demonstrated explicitly.","section":"III.A"},{"comment":"The substitution of N_sub = 1000 stars with scaled-down noise for the full catalog is described in one sentence; a single full-star run at one frequency would confirm that the scaling preserves both the sensitivity and the interpolation accuracy.","section":"III.B"},{"comment":"The dashed high-frequency extrapolation to 10^-3.5 Hz is not used for the quoted thresholds; given the text's own caveat that this regime requires frequency-evolution-aware methods, consider removing it or labeling it more prominently as an illustration.","section":"III.B and Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a credible methods contribution and the central compression idea is sound. The main needs are additional validation: a multi-frequency accuracy test, a clearer statement of the face-on assumption in the headline sensitivities, and some calibration of the empirical-Bayes evidence correction. These are addressable within the scope of the manuscript, so I would not recommend rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know up front. The compression method is real: transferring the PTA inner-product precomputation from Becsy et al. to astrometric deflections is a clean, non-obvious move, and the likelihood rewrite is algebraically sound. Second, the headline numbers should be read as order-of-magnitude forecasts, not hard sensitivities. The 1%-accuracy claim is supported at a single frequency, and the forecast pipeline makes optimistic noise/systematics choices that the authors disclose but do not fix.\n\nWhat's genuinely new is applying this compression to relative astrometry, which opens a practical route into the microhertz band with Kepler and Roman data. The paper does a good job of laying out the survey models, the noise scaling, and the nested sampling setup. The finding that small fields of view make sky localization poor, and that targeted galaxy-position searches help only marginally, is a useful and non-obvious result. The authors are also refreshingly honest about the high-frequency end: avoiding frequency evolution forces unrealistically low chirp masses and sub-Mpc distances, and they say so.\n\nThe main soft spot is the accuracy claim. Section III A and Figure 2 measure interpolation error only at f=1e-7 Hz, and the F values that satisfy the 1% criterion there are then used across the entire 10^-8 to 10^-6 Hz prior. The stress-test concern is on target: the number of grid points per independent frequency bin drops by an order of magnitude as f increases, and the paper's own observation that per-sample evidence errors are always negative indicates the coarse grid systematically under-resolves narrow likelihood peaks. This is fixable, but it needs a multi-frequency error scan before the headline claim is credible.\n\nTwo lesser concerns. First, the forecasts assume white noise, full systematics removal, and no mean subtraction for Kepler; those are optimistic, though clearly labeled. Second, the empirical Bayes evidence correction in Section II E uses the posterior width to rescale the evidence; that is an ad hoc step that could bias the B=10 thresholds, though the injection calibration makes it acceptable for a first forecast. No code or data are provided, so nothing can be independently checked.\n\nWho should read this: anyone working on astrometric GW detection or microhertz data analysis. The method is a stepping stone, not the final word. I would send it to peer review with a request for multi-frequency validation and a more careful treatment of the evidence correction. It deserves referee time, but it needs revision before the accuracy and sensitivity claims are taken at face value.","headline":"A credible transfer of PTA inner-product compression to astrometric GW searches, but the headline accuracy claim is validated at only one frequency and the sensitivity forecasts rest on optimistic assumptions.","tokens_in":18010,"tokens_out":3574,"would_cite":true,"duration_ms":37167,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper extends a pulsar-timing-array likelihood compression technique to astrometric deflection data, cutting dataset size by up to ~100 with under 1% accuracy loss, and uses it to forecast Kepler and Roman sensitivity to microhertz…","keywords":["gravitational waves","astrometry","Bayesian inference","nested sampling","microhertz band","inner product precomputation","Kepler space telescope","Nancy Grace Roman Space Telescope"],"falsifier":"Inject into simulated Kepler-like data a binary with initial frequency $10^{-6}$ Hz and a chirp mass large enough that the 0PN phase error over 16 quarters exceeds $\\pi/2$ (the paper's own cutoff), then compare the exact and interpolated-likelihood nested sampling evidences; a discrepancy beyond the claimed 1% at the relevant grid density would show that the constant-frequency assumption is the limiting factor.","tokens_in":16890,"feed_emoji":"🔭","tokens_out":5882,"duration_ms":56380,"temperature":0.7,"pith_summary":"This paper shows that a Bayesian likelihood trick originally developed for pulsar timing arrays can be applied to astrometric deflection data, making coherent gravitational-wave searches with thousands to hundreds of millions of stellar baselines computationally feasible. By precomputing inner products between stellar position measurements and sine/cosine templates at a grid of frequencies, the terabyte-scale dataset is reduced by up to a factor of ~100 with less than 1% loss in accuracy across two orders of magnitude in frequency. The paper then uses this compressed likelihood to forecast the sensitivity of the Kepler and Roman space telescopes to coherent gravitational waves in the microhertz band ($10^{-8}$ to $10^{-6}$ Hz), finding strain thresholds $h_0 \\simeq 10^{-12.45}$ and $10^{-11.39}$ respectively. Because the survey fields cover little sky, sources are poorly localized, and the strain horizon is only a few Mpc for plausible supermassive-black-hole binaries. The result matters because it opens the microhertz gap between pulsar timing arrays and LISA to actual searches with existing and upcoming photometric surveys.","feed_headline":"Astrometric GW searches get a 100x data cut","feed_subtitle":"A pulsar-timing likelihood trick carries over to stellar astrometry, letting Kepler and Roman probe the microhertz band at 1% accuracy.","key_machinery":"The central object is the precomputed inner product $\\langle a|C^{-1}|b\\rangle$ between stellar deflection time series and sinusoidal filter functions, evaluated once on a log-frequency grid. Because a constant-frequency sinusoid of arbitrary phase and amplitude is a linear combination of $\\cos\\omega t$ and $\\sin\\omega t$, the full Gaussian log-likelihood — a sum over $10^5$–$10^8$ stellar baselines and many exposures — collapses to $O(NF)$ precomputed quantities instead of $O(NT)$ raw data points. Linear interpolation across the log-frequency grid lets the frequency remain a continuous search parameter. This compression is what turns a terabyte-to-petabyte dataset into something a nested-sampling search can read from memory.","core_discovery":"The central claim is that the inner-product precomputation scheme for Bayesian likelihoods, previously applied to pulsar timing residuals, carries over to coherent gravitational-wave astrometry essentially unchanged: for a constant-frequency source, the Gaussian log-likelihood of the deflection data is exactly a linear combination of a handful of inner products between the data and $\\cos(\\omega t)$/$\\sin(\\omega t)$ filters. Interpolating those inner products on a log-spaced frequency grid keeps the evidence accurate to within 1% for a barely detectable (Bayes factor $B=10$) source, provided enough grid points ($F=3140$ for Kepler, $F=15500$ for Roman). The paper demonstrates this with injection-based nested sampling runs and produces sensitivity forecasts: averaging over source sky positions, Kepler reaches $h_0 \\geq 10^{-12.45}$ and Roman reaches $h_0 \\geq 10^{-11.39}$ over $10^{-8}$ to $10^{-6}$ Hz, corresponding to detectability of a $10^9\\,M_\\odot$ chirp-mass binary at 3.6 Mpc (Kepler) or 0.31 Mpc (Roman).","pith_inferences":["One could extend the same precomputation to chirping signals by using a bank of frequency-evolution templates, interpolating over chirp mass as well as frequency, which would remove the main restriction at the high-frequency end.","The compression factor depends on the number of exposures per star, so future surveys with longer baselines or higher cadence will benefit even more, while very short surveys may not.","The same likelihood rewrite may apply to other high-baseline, low-SNR astrometric searches, such as exoplanet astrometry or searches for compact dark-matter lenses, wherever the signal is a coherent sinusoid in time.","If real microhertz sources chirp on month timescales, the quoted strain thresholds are optimistic; a direct search over $10^{-6}$–$10^{-4}$ Hz will need the chirping extension before claiming astrophysical constraints."],"forward_implications":["Coherent Bayesian searches of astrometric datasets with $10^5$–$10^8$ stars become computationally tractable, removing the hard-drive bottleneck that previously made MCMC or nested sampling infeasible.","Archival Kepler data and near-future Roman observations can be used to search the $10^{-8}$ to $10^{-6}$ Hz band, partially filling the gap between pulsar timing arrays and space-based interferometers.","The strong dependence of sensitivity on source–telescope separation means searches should marginalize over sky position; position-targeted strategies give only modest gains ($\\Delta\\log_{10} h_0 \\approx -0.2$).","Frequency-targeted searches, using electromagnetic priors, are expected to improve sensitivity substantially, although the paper does not quantify that gain.","The compressed dataset retains per-baseline information (no averaging of stars), so the compression step does not mask spatially correlated systematics."],"supporting_citations":[{"why":"Derives the exact astrometric deflection response of stellar centroids to a coherent gravitational wave, which is the signal model used throughout.","marker":"[14]"},{"why":"Provides the original projection of Roman's astrometric gravitational-wave sensitivity and the simulation code the present forecasts build on.","marker":"[17]"},{"why":"Extends the astrometric sensitivity forecasts and the simulation framework that generates the mock survey data.","marker":"[18]"},{"why":"Develops the inner-product precomputation technique for PTA likelihoods that this paper generalizes to astrometric deflection data.","marker":"[20]"},{"why":"Extends the precomputation to a frequency grid with interpolation, the scheme adapted here to a log-spaced grid.","marker":"[21]"},{"why":"Sets the Roman Galactic Bulge Time Domain Survey parameters (fields, cadence, seasons) used in the mock datasets.","marker":"[24]"},{"why":"Supplies the empirical Kepler astrometric precision curve used to assign per-star noise levels.","marker":"[27]"},{"why":"Provides the 15-year pulsar timing array strain upper limits used as a comparison sensitivity curve.","marker":"[8]"}],"fun_headline_variants":["Fast inner-product trick cuts astrometric GW data 100x","Microhertz GWs: fast Bayesian search cuts astrometric data 100x","Kepler and Roman get 100x data cut for GW astrometry","Astrometric GW search speedup: 100x data cut with 1% accuracy","Pulsar timing likelihood trick works for astrometric GWs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes each candidate gravitational wave keeps a nearly constant frequency over the whole observing window, so that the signal is a pure sinusoid; if a real source's frequency evolves appreciably within that time, the precomputed inner products no longer represent the signal and the compressed likelihood is invalid in that part of the band.","fun_headline_variants_meta":{"raw":{"variants":["Fast inner-product trick cuts astrometric GW data 100x","Microhertz GWs: fast Bayesian search cuts astrometric data 100x","Kepler and Roman get 100x data cut for GW astrometry","Astrometric GW search speedup: 100x data cut with 1% accuracy","Pulsar timing likelihood trick works for astrometric GWs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001672,"raw_usage":{"total_tokens":6712,"prompt_tokens":1108,"completion_tokens":5604,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":724,"completion_tokens_details":{"reasoning_tokens":5504}},"tokens_in":724,"tokens_out":5604,"duration_ms":35651,"temperature":1.0,"reasoning_tokens":5504,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:35:35.957186+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inject into simulated Kepler-like data a binary with initial frequency $10^{-6}$ Hz and a chirp mass large enough that the 0PN phase error over 16 quarters exceeds $\\pi/2$ (the paper's own cutoff), then compare the exact and interpolated-likelihood nested sampling evidences; a discrepancy beyond the claimed 1% at the relevant grid density would show that the constant-frequency assumption is the limiting factor.","supporting_citations":[{"cited_title":"Dom` enech and M","cited_arxiv_id":null,"evidence_quote":"Derives the exact astrometric deflection response of stellar centroids to a coherent gravitational wave, which is the signal model used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends the astrometric sensitivity forecasts and the simulation framework that generates the mock survey data."},{"cited_title":"Aghanim, Y","cited_arxiv_id":null,"evidence_quote":"Sets the Roman Galactic Bulge Time Domain Survey parameters (fields, cadence, seasons) used in the mock datasets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the empirical Kepler astrometric precision curve used to assign per-star noise levels."}],"review_version":2}