{"id":"ef1ac88e-f2de-40fe-a029-9a57c6e894e4","arxiv_id":"2505.11655","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The upgraded RIFT pipeline, built around an adaptive-volume Monte Carlo integrator, accurately and efficiently handles exceptional compact binaries with precession or eccentricity, as shown by PP tests, timing benchmarks, and cross-code comparisons.","lead":"RIFT estimates the properties of colliding black holes and neutron stars from gravitational wave data, and this paper upgrades its sampling methods so it can also handle rare, extreme events like strongly precessing or eccentric mergers. The updated code is checked on synthetic injections and on public LIGO/Virgo events, including direct comparisons with the bilby inference package.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Frame-convention repair for IMRPhenomXPHM is empirically tuned and never tested end-to-end; a silent bias there would invalidate the main precessing comparisons.","rationale":"The reader's CONDITIONAL verdict is well calibrated. The AV integrator is validated on toy problems and zero-spin PP tests, and the timing benchmarks support the headline efficiency claim. However, the paper's own text flags the weakest link: the IMRPhenomXPHM interface was inconsistent, and the repair in Appendix A is explicitly empirical. Since the central precessing demonstration and the bilby comparisons rely on IMRPhenomXPHM, a silent convention error there would bias the main results without being caught by the provided PP tests. The q p-value of 0.017 in Figure 13 is worth noting but is not by itself decisive given the 14 parameters and N=99; the omission of GW200129 from the eccentricity section is a scope limitation rather than a correctness flaw. The frame-convention concern is therefore the load-bearing one, and the proposed independent derivation plus overlap check would settle it. Because the reader already conditioned acceptance on addressing related issues, no verdict change is needed; a positive outcome of the test would remove the main obstacle to full acceptance.","tokens_in":32042,"tokens_out":5477,"duration_ms":58972,"concrete_test":"Independently derive the Appendix A J-to-L transformation for IMRPhenomXPHM from the Euler-angle conventions in Pratten et al. (2021) Appendix C, including the α→α+π correction, without tuning to waveform output; then verify on a grid of precessing parameters (θ_JN, φ_JL, ι, φ_ref, χ_1⊥) that RIFT's mode-resummed h(t) matches lalsuite's direct SimInspiralChooseFDWaveform h(t) with complex overlap >0.9999. If the analytic derivation disagrees with the empirical correction, or overlaps fail, the correction is not trustworthy; a positive result would justify re-running the Figure 13 PP test with IMRPhenomXPHM.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the new RIFT operating point is trustworthy for precessing binaries rests on the Appendix A J-to-L frame correction for IMRPhenomXPHM. Section III A admits that the old interface used an inconsistent convention; Appendix A repairs it with an Euler-angle rotation plus an 'empirically found' α→α+π correction (footnote 1), and validates it only by comparing reconstructed h(t) against the conventional implementation (Figure 17), which shows residual 'small amplitude disagreements' and an 'overall difference in polarization convention.' No PP test or injection-recovery study exercises the repaired XPHM path: the precessing PP plot (Figure 13) uses IMRPhenomPv2, whose mode extraction does not go through the J-frame workaround, and the bilby comparisons in Section VII C are anecdotal. If the empirical correction is wrong for any region of the precessing parameter space, every XPHM-based likelihood, the head-to-head bilby comparisons, and the Figure 1/2 demonstrations are silently biased; internal consistency checks cannot catch this because both code paths share the same convention assumptions after the repair.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This methods paper reports a set of algorithmic and operational upgrades to the RIFT gravitational-wave parameter-estimation code, aimed at making the pipeline efficient and reliable for 'exceptional' sources: loud, precisely measured, precessing, or eccentric binaries. The main technical contributions are a new adaptive-volume (AV) Monte Carlo integrator, a portfolio integrator, normalizing-flow and unreliable-oracle samplers, updates to the external waveform interfaces (including a J-to-L frame correction for IMRPhenomXPHM), a transverse-spin 'puffball' proposal, disabling of automatic mode filtering, and integration with the asimov reproducibility framework. The paper validates these changes with closed-form Gaussian and Rosenbrock toy problems, two end-to-end PP tests (N=99 each, global p=0.26 and p=0.34), timing benchmarks, reanalyses of selected O3 events with eccentric SEOBNRv5EHM, and head-to-head comparisons with bilby on O3 events using IMRPhenomXPHM. The central claim is that the new RIFT operating point is both faster and trustworthy for precessing and eccentric binaries, with roughly an order-of-magnitude improvement in integrator efficiency for typical problems.","tokens_in":1844,"tokens_out":1806,"duration_ms":62167,"significance":"If the claims hold, this is a practically important contribution: it changes RIFT's recommended default operating point and extends its demonstrated range to precessing and eccentric sources with costly time-domain waveforms. The paper's strengths include genuine end-to-end PP testing with production settings, closed-form toy problems with known answers, reproducible asimov-based analyses, and direct cross-code comparisons with bilby. The order-of-magnitude efficiency gain for the AV integrator, if confirmed with the benchmarks, would be valuable to the gravitational-wave inference community. However, the central claim for precessing IMRPhenomXPHM analyses is currently supported only by waveform-level agreement and anecdotal event comparisons, not by an end-to-end test of the repaired frame-convention path; this is a correctness-risk point that needs to be addressed before the paper can serve as the standard reference for the new operating point.","major_comments":[{"comment":"The J-to-L frame correction for IMRPhenomXPHM is load-bearing for the precessing and bilby-comparison claims, but it is validated only at the waveform level. Section III A states that the old interface used a frame convention inconsistent with all other interfaces, and Appendix A repairs it with a J-to-L rotation plus an 'empirically found' alpha-to-alpha-plus-pi correction (footnote 1). The only validation shown is the h(t) comparison in Figure 17, whose caption reports an 'overall difference in polarization convention' and 'small amplitude disagreements' due to data conditioning. The precessing PP test in Figure 13 uses IMRPhenomPv2, whose mode extraction does not pass through the ChooseFDModes J-frame workaround; therefore no PP or injection-recovery test exercises the repaired XPHM path. Because an incorrect empirical correction could silently bias every XPHM likelihood, the Figure 1/2 demonstrations, and the Section VII C bilby comparisons, I ask the authors to add an end-to-end validation (e.g., an XPHM PP test or a set of injection-recovery runs) or otherwise demonstrate that the empirical rotation is exact across the precessing parameter space.","section":"Section III A and Appendix A"},{"comment":"The precessing PP test shows a per-parameter p-value of 0.017 for the mass ratio q, which is well below the 5% level. Since this PP test is the principal end-to-end evidence that the new operating point is unbiased for precessing binaries, the authors need to address this outlier directly. If they interpret it as an expected multiple-comparisons fluctuation, they should show that the number of parameters and their correlations make such a value unremarkable; otherwise the result suggests a small but real miscalibration in the ILE/CIP treatment of q for precessing systems. Reporting only the global p-values (0.26 and 0.34) is insufficient, because they aggregate over parameters and can hide a localized deviation.","section":"Figure 13"},{"comment":"The eccentricity results are presented as a demonstration that RIFT can 'efficiently analyze events with multiple costly models including the effects of precession or eccentricity,' but the eccentric path is not validated end-to-end. Figure 16 verifies only that the modal reconstruction and the direct h(t) interface agree for SEOBNRv5EHM, and Table I reports posteriors for real events without any injection-recovery or PP test. Given that the same appendix-level empirical conventions are used for the gwsignal phase shift, I recommend adding at least one end-to-end validation for the eccentric interface, or explicitly stating in Section VII B that the eccentric reanalyses are preliminary proof-of-concept demonstrations that do not yet establish unbiased inference for eccentric binaries.","section":"Section VII B and Figure 16"}],"minor_comments":[{"comment":"The sentence introducing Eq. (5) contains a duplicated 'where where' that should be corrected.","section":"Eq. (5)"},{"comment":"The citation breaks as '[40?]' in the text; the reference number and the bibliography entry should be reconciled.","section":"Section II F"},{"comment":"The claim of 'roughly an order of magnitude improvement' is based on a comparison with benchmarks reported in a previous paper (Appendix B of [23]), not on a same-hardware, same-input reproduction. A direct side-by-side benchmark on identical data and hardware would make this quantitative claim easier to verify.","section":"Section V A"},{"comment":"The captions describe the precessing PP plot as 'precessing spin extrinsic' and the nonspinning plot as 'zero spin and extrinsic,' but both panels include intrinsic parameters such as chirp mass and mass ratio; please reword the captions to reflect the full parameter set.","section":"Figures 12 and 13"},{"comment":"The column labeled 'BE/QC' is not defined anywhere in the text or caption; please define the abbreviation and explain how the Bayes factor was computed.","section":"Table I"},{"comment":"The caption says the J-to-L corrected IMRPhenomXPHM implementation agrees with the conventional implementation 'Except for an overall difference in polarization convention.' Since the purpose of the correction is to match conventions, please clarify whether this difference is a known residual, an artifact of the plotting convention, or an indication that the correction is incomplete.","section":"Figure 17 caption"},{"comment":"There are numerous typographical errors that should be cleaned up, including 'culiminating' (Introduction), 'techniues' (Section II C), 'ewn contributions' (Section IV), 'ananlysis' (Section V A), 'Becuase' (Section VI B), 'inlcudes' (Table II), 'unphyiscal' (Section VI C), and 'rquire' (Section II J).","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the authors are clearly reporting on their own code, so there is no novelty or attribution concern. My main worry for the editor is that the centrally advertised precessing capability rests on an empirically tuned frame correction that has not been exercised by any end-to-end test; the Section VII C bilby comparisons are anecdotal and would not catch a shared convention error. The q p=0.017 in the precessing PP test is also worth pushing on. Both issues are fixable with additional validation, which is why I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent and candid methods paper. The headline claim—that RIFT can now handle high-SNR, strongly precessing, and eccentric binaries efficiently—is mostly backed up. The AV integrator is the real workhorse: ported from VARAHA, it gives an order-of-magnitude sampling-efficiency gain in the Rosenbrock toy problem, and it's validated end-to-end in two PP tests with N=99 each. The portfolio, unreliable-oracle, and normalizing-flow integrators are less mature, but the paper is appropriately cautious about them.\n\nWhat's new is concrete: the AV port, a recommended operating point, the gwsignal waveform interface, and demonstrations on real O3 events with eccentric and precessing models. The citation pattern looks fine—prior RIFT papers and VARAHA get proper credit, and self-citations are to work that is actually being built on.\n\nWhere I'd push back, in increasing order of importance. First, the q marginal in the precessing PP plot has p=0.017. That's one parameter among fourteen, so it isn't damning, but it deserves a direct explanation or a follow-up test; leaving it unexplained is sloppy. Second, the eccentricity section explicitly omits GW200129, the most interesting eccentric candidate. They say it'll be reported elsewhere, but that weakens the paper's own eccentricity claim. Third and most important, the frame-convention repair for IMRPhenomXPHM is empirically tuned and never tested end-to-end. Section III admits the old interface was inconsistent; Appendix A patches it with an Euler-angle rotation plus an 'empirically found' alpha+pi shift. The validation is a comparison of reconstructed h(t) against the conventional implementation, which shows small amplitude disagreements and a polarization-convention difference. No PP test or injection-recovery uses the repaired XPHM path—the precessing PP plot uses IMRPhenomPv2. If that empirical correction is wrong somewhere, the XPHM-based figures and the bilby comparisons would be silently biased. It's not a fatal flaw; the fix might well be right. But it's a gap in the evidence, and the authors should either add an XPHM end-to-end test or explicitly scope the claim.\n\nThere's also a minor reproducibility niggle: no commit hash or machine-readable settings are pinned, which makes 'reproducible settings consistent with past work' harder to verify.\n\nWho this is for: anyone doing GW parameter inference with expensive waveform models, especially for loud, precessing, or eccentric events. I'd send it to a serious referee and ask for the XPHM test and a comment on the q p-value. With those, it's publishable.","headline":"A solid, honest methods paper: RIFT's new AV-based operating point is likely a real win for loud/precessing/eccentric GW sources, but the empirically patched XPHM frame convention needs an end-to-end test before I'd fully trust the precessing demonstrations.","tokens_in":32833,"tokens_out":2869,"would_cite":true,"duration_ms":27079,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The latest RIFT release claims efficient, accurate inference for loud, precessing, and eccentric gravitational-wave sources by switching to an adaptive-volume Monte Carlo integrator.","keywords":["gravitational-wave parameter estimation","RIFT","simulation-based inference","adaptive volume integration","precessing binaries","eccentric binaries","Monte Carlo integration","Bayesian inference"],"falsifier":"Run the paper's modal-reconstruction comparison across a dense grid of precessing configurations, varying mass ratio, spin magnitude, and orientation, and compare the reconstructed h(t) against the waveform generator's direct output; any orientation with a systematic phase or amplitude offset larger than numerical tolerance would show the empirical rotation correction is incomplete. A complementary test is to run the paper's precessing PP test using a different waveform interface and check whether spin parameters remain unbiased.","tokens_in":31793,"feed_emoji":"🔭","tokens_out":8627,"duration_ms":84249,"temperature":0.7,"pith_summary":"This paper argues that RIFT, a simulation-based gravitational-wave inference code, has been re-tuned so that it can handle the unusual sources that its previous conservative settings struggled with: loud signals, strongly precessing high-mass-ratio binaries, and eccentric mergers analyzed with costly waveform models. The central technical move is a new default Monte Carlo integrator, the adaptive-volume (AV) integrator, which samples the likelihood far more efficiently than the old adaptive-cartesian and Gaussian-mixture samplers. The paper reports that the AV integrator delivers roughly an order-of-magnitude improvement for typical problems, and validates the new operating point with probability-probability tests, synthetic injections, and reanalyses of real O3 events using precessing and eccentric waveform models. If correct, the result means fast, reproducible inference for exceptional sources no longer requires hand-tuned special settings.","feed_headline":"Exotic GW sources now analyzed tenfold faster","feed_subtitle":"RIFT's new adaptive-volume integrator also fixes a waveform frame bug that could bias precessing binaries.","key_machinery":"The adaptive-volume (AV) integrator is the load-bearing object: following an adaptive volume-sampling strategy, it partitions the integration domain into hypercubes, discards cells with negligible probability, estimates the enclosed probability via likelihood thresholds and live points, and sets the next refinement scale from the Monte Carlo uncertainty in the sampled volume. It replaces the earlier adaptive-cartesian and Gaussian-mixture integrators as the default for both the extrinsic marginalization step (ILE) and the posterior-generation step (CIP). Supporting machinery includes the factorized likelihood built from inner products Q, U, V of spherical-harmonic modes with detector data, GPU-accelerated evaluation, dithering plus puffball jitter that now includes transverse spin components, and a corrected L-frame spherical-harmonic convention for the IMRPhenomXPHM interface.","core_discovery":"The paper's central claim is that the latest RIFT release can efficiently and reliably interpret exceptional compact binaries -- sources with very high signal-to-noise, large mass ratio with strongly misaligned spins, or measurable eccentricity -- by replacing the default integrators with an adaptive-volume Monte Carlo integrator and adjusting the exploration defaults. The AV integrator recursively subdivides the extrinsic-parameter integration volume into hypercubes, keeps only cells containing significant probability, and refines the grid based on the Monte Carlo uncertainty in the enclosed volume, giving high sampling efficiency on sharply peaked likelihoods. The paper also fixes a frame-convention error in the IMRPhenomXPHM interface, adds transverse-spin jitter to exploration, disables automatic mode filtering that could bias loud sources, and reports end-to-end PP tests and cross-code comparisons that agree with an independent sampler. The reported practical gain is about an order of magnitude in cost per marginal-likelihood evaluation for typical problems.","pith_inferences":["If the AV integrator's efficiency holds in higher-dimensional settings, RIFT could become a low-latency alternative to neural posterior estimators, since it needs no pretraining and returns calibrated posterior samples.","The empirical J-to-L frame correction should be revalidated against an independent waveform family or numerical-relativity surrogate before it is trusted for discovery-level claims; the paper's check uses a limited set of examples.","The same AV machinery could be applied to generic Bayesian integrals outside gravitational waves, where the paper's generalized inference path already points toward non-GW applications.","Excluding the strongest eccentricity candidate from the eccentric-waveform reanalyses leaves open whether the new settings will support or challenge that claim; running the same pipeline on that event is a direct next test."],"forward_implications":["The default operating point now handles loud sources, including those with signal-to-noise above roughly 30, and edge-on, high-mass-ratio binaries without manual adaptation of extrinsic sampling.","Production analyses can use costly time-domain waveforms for precessing and eccentric binaries at roughly an order of magnitude lower cost per likelihood evaluation.","Probability-probability tests with both zero-spin and precessing injections pass, so the AV integrator is claimed to be safe as the sole integrator for both ILE and CIP.","Previous RIFT analyses that used the IMRPhenomXPHM interface with the old frame convention inherit a bias that the new release corrects.","Automated cross-code comparisons over many events are now feasible out of the box, making disagreements between samplers easier to diagnose.","The eccentricity reanalysis of selected O3 events with a uniform eccentricity prior yields upper limits rather than detections, so the pipeline is ready for the stronger eccentricity candidates that were explicitly left out."],"supporting_citations":[{"why":"Defines the original two-stage RIFT algorithm of marginal-likelihood evaluation followed by interpolation and posterior sampling, which this paper extends.","marker":"[16]"},{"why":"Previous RIFT methods paper; supplies the integrators, architectures, dithering, convergence tests, and benchmarks that this work replaces or improves.","marker":"[23]"},{"why":"Introduces the adaptively-refined volume sampling strategy that the new AV integrator ports and optimizes for GPUs.","marker":"[64]"},{"why":"Establishes the automation interface and prior efficient reanalyses of O3 events that this paper reuses for eccentric and cross-code demonstrations.","marker":"[25]"},{"why":"Provides the new waveform interface that gives RIFT access to the latest precessing and eccentric effective-one-body models.","marker":"[41]"},{"why":"Documents the GPU-accelerated likelihood and inner-product machinery that keeps each marginal-likelihood evaluation cheap enough for the new integrator to exploit.","marker":"[36]"},{"why":"Independent external reanalysis of two O3 events used to check that corrected RIFT settings recover the same results.","marker":"[56]"}],"fun_headline_variants":["RIFT's adaptive-volume integrator speeds up exceptional GW analysis","Faster RIFT for exotic binaries with adaptive sampling and bug fix","RIFT update: 10x faster inference for exceptional gravitational-wave sources","Adaptive-volume integrator speeds RIFT for loud, eccentric, precessing binaries"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole precessing and eccentric inference chain assumes that RIFT's L-frame spherical-harmonic convention matches every external waveform interface; the one known mismatch was repaired with an empirically determined phase correction, and if that correction is wrong for any supported waveform, all likelihoods and posteriors built on it are silently biased.","fun_headline_variants_meta":{"raw":{"variants":["RIFT's adaptive-volume integrator speeds up exceptional GW analysis","Faster RIFT for exotic binaries with adaptive sampling and bug fix","RIFT update: 10x faster inference for exceptional gravitational-wave sources","Adaptive-volume integrator speeds RIFT for loud, eccentric, precessing binaries"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000935,"raw_usage":{"total_tokens":3942,"prompt_tokens":829,"completion_tokens":3113,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":445,"completion_tokens_details":{"reasoning_tokens":3035}},"tokens_in":445,"tokens_out":3113,"duration_ms":22119,"temperature":1.0,"reasoning_tokens":3035,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:50:42.402761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the paper's modal-reconstruction comparison across a dense grid of precessing configurations, varying mass ratio, spin magnitude, and orientation, and compare the reconstructed h(t) against the waveform generator's direct output; any orientation with a systematic phase or amplitude offset larger than numerical tolerance would show the empirical rotation correction is incomplete. A complementary test is to run the paper's precessing PP test using a different waveform interface and check whether spin parameters remain unbiased.","supporting_citations":[{"cited_title":"Wofford, A","cited_arxiv_id":null,"evidence_quote":"Previous RIFT methods paper; supplies the integrators, architectures, dithering, convergence tests, and benchmarks that this work replaces or improves."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the new waveform interface that gives RIFT access to the latest precessing and eccentric effective-one-body models."}],"review_version":1}