{"id":"0fa3f895-0cbc-4ace-aeef-ab569e62d27c","arxiv_id":"2505.02924","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Spatial cross-correlation of LIGO/Virgo/KAGRA events with SDSS AGNs yields a tentative, low-significance excess that implies 30-40 percent of mergers could come from lower-luminosity or lower-Eddington-ratio AGNs.","lead":"The authors compare the sky locations of gravitational wave events from LIGO/Virgo/KAGRA with a catalog of active galactic nuclei and report a tentative excess of dimmer, low-accretion AGNs around the GW sources. This suggests that roughly 30-40 percent of detected black hole mergers could occur in AGN disks, a debated formation channel.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claimed 'evidence' rests on 1.3–1.6σ mock significance, is not trials-corrected, and vanishes when the four best-localized O4 events are removed.","rationale":"The reader correctly identified the null-background specification and the fragility of the result, but the most load-bearing weakness is more direct: the paper's own calibrated null significance is only 1.3–1.6σ, no trials correction is applied across the six sub-catalogs, and Section 4.2 explicitly states that the signal is 'predominated by the four most precisely localized events', with f_agn = 0 after their removal. The f_agn = 0 mock construction using SDSS galaxies largely addresses the large-scale-structure concern raised by the reader, so I would not make that the primary attack. Instead, the central claim of 'evidence' fails the usual threshold for a positive detection. This does not invalidate the method or the data release; it means the current data support an upper limit or a tentative hint, not a robust measurement. The paper is transparent about most caveats and provides reproducible code and data, so a conditional verdict remains appropriate, contingent on a trials-corrected significance and a robustness analysis that does not depend on four events.","tokens_in":20290,"tokens_out":10324,"duration_ms":125710,"concrete_test":"Run 1000 null realizations with f_agn = 0 following Section 4.1, but for each realization compute an extremum statistic over all six sub-catalogs, e.g., the minimum ΔP0.05 across sub-catalogs or the maximum f_peak. Compare the observed minimum ΔP0.05 against this trials-corrected null distribution. Repeat the same procedure after excluding the four best-localized O4 events. If the global p-value exceeds 0.05, or if the restricted sample already gives best-fit f_agn = 0, the claim should be reported as an upper limit rather than evidence for a nonzero AGN-origin fraction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 4.1 state that the lower-Lbol and lower-Eddington-ratio correlations are 'unlikely to arise from random coincidence', but the paper's own null-mock comparison gives only 1.3σ (lower-Lbol) and 1.6σ (lower-Eddington-ratio) credibility against f_agn = 0. Section 4.2 then shows that removing the four best-localized O4 events (ΔVc < 10^6 Mpc^3) makes the best-fit f_agn drop to zero in every sub-catalog; the residual 'narrowing of the upper limit' argument is a post hoc comparison of 29-event versus 90-event samples, not a calibrated significance test. Because six AGN sub-catalogs are searched, and the two positive sub-catalogs overlap by about 40%, the look-elsewhere penalty is non-negligible. A 1.3σ excess that disappears after removing four events is not sufficient to support the headline claim of evidence. The null-background formula B_i = 0.9 f_cover,i (Eqs. 4–5) is a real but secondary concern, because the f_agn = 0 mocks in Figure 4 place mock hosts in SDSS galaxies and therefore partly capture galaxy–AGN large-scale-structure correlation; the more direct failure mode is that the claimed significance is weak and uncalibrated for multiple comparisons.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper constrains the fraction f_agn of LIGO/Virgo/KAGRA events that originate from active galactic nuclei by comparing GW localization volumes (O1-O4a) with SDSS DR16 AGN sub-catalogs split by bolometric luminosity and Eddington ratio. The method extends the Bartos et al. approach by using a 3D Voronoi estimate of the AGN density field, a coverage factor f_cover,i for partial sky coverage, and a modified background probability B_i=0.9 f_cover,i. The authors report non-zero best-fit values f_agn=0.39^{+0.41}_{-0.32} for the lower-Lbol sub-catalog and f_agn=0.29^{+0.40}_{-0.25} for the lower-Eddington-ratio sub-catalog, with zero best fits for the other sub-catalogs and for the full catalog. Mock-injection tests are quoted as rejecting f_agn=0 at 1.3 sigma and 1.6 sigma credibility, and Section 4.2 investigates the origin of the signal by removing the four best-localized O4 events. The paper concludes that a fraction of LVK events come from lower-luminosity or lower-accretion-rate AGNs.","tokens_in":20519,"tokens_out":6151,"duration_ms":68704,"significance":"If the claimed excess is real, the paper would provide the first quantification of the AGN-origin fraction for specific AGN sub-populations, with a clear environmental prediction that lower-luminosity and lower-Eddington-ratio AGNs dominate the AGN channel. The methodological contribution is genuine: the Voronoi-based density estimate, the explicit coverage factor, the mock-injection tests, and the Appendix B unbiasedness argument are useful extensions of earlier work. The authors also make the data publicly available on Zenodo, which supports reproducibility, and their upper limits for the higher-luminosity sub-catalogs are consistent with previous constraints from Veronesi et al. My assessment is that the central evidence claim is not yet supported by the statistics presented; the paper is more defensible as a method demonstration with upper limits and a tentative hint, rather than as evidence of a non-zero AGN fraction.","major_comments":[{"comment":"The quantitative support for the abstract's claim that the correlation is 'unlikely to arise from random coincidence' is a 1.3 sigma (lower-Lbol) and 1.6 sigma (lower-Eddington) rejection based on DeltaP0.05 in the f_agn=0 mocks. These significances are not corrected for the number of AGN sub-catalogs searched, and the two positive sub-catalogs overlap by about 40%, so the effective number of trials is smaller than six but definitely larger than one. A nominal 1.3-1.6 sigma excess, before any trials penalty, is too weak to support the language of evidence; the authors should report trials-corrected p-values or explicitly downgrade the conclusion to a tentative hint.","section":"Section 4.1, Fig. 4"},{"comment":"Removing the four best-localized O4 candidates makes the best-fit f_agn drop to zero in all six sub-catalogs. The non-zero signal is therefore dominated by a very small number of well-localized events, three of which contain AGNs in their error volumes. The subsequent argument that the signal persists because the 90% upper limit narrows when going from the 29-event O1-O3 sample to the 90-event sample is not a calibrated significance test: the comparison is post hoc, and no null distribution for this upper-limit ratio is provided. This does not corroborate the statement in Section 5 that the signal 'remains consistent across different samples.'","section":"Section 4.2, Fig. 5"},{"comment":"The null background B_i=0.9 f_cover,i assumes that under f_agn=0 the expected number of AGNs in a GW error volume is proportional only to the survey-covered fraction of that volume. If the non-AGN BBH population and the AGN population trace the same large-scale structure, the likelihood in Eq. (1) will systematically pull f_agn away from zero. The f_agn=0 mocks that draw hosts from SDSS galaxies may partly capture this correlation, but the paper does not demonstrate that the recovered f_agn is unbiased under the null; a concrete test would be to compare mocks with galaxy-drawn hosts against mocks with uniformly random hosts. Without such a demonstration, the positive best-fit values in Section 4.1 have an unmodeled systematic component in addition to the weak statistical significance.","section":"Section 3, Eqs. (4)-(5)"}],"minor_comments":[{"comment":"The text uses '90% confidence level' for what are Bayesian posterior intervals from the likelihood in Eq. (1); the authors should say 'credible interval' or explicitly state the prior and posterior interpretation.","section":"Abstract and Section 4.1"},{"comment":"The 1.3 sigma and 1.6 sigma statements should specify whether these are one-sided or two-sided Gaussian-equivalent significances; DeltaP0.05 is a lower-tail probability, and the conversion to sigma depends on this choice.","section":"Section 4.1, DeltaP0.05 discussion"},{"comment":"The forecast that f_agn errors scale as 1/sqrt(SNR) is unclear; presumably the intended scaling is with the number of events or the combined signal-to-noise ratio, and the text should be corrected.","section":"Section 5"},{"comment":"The statement that the AGN catalog has not been corrected for Malmquist bias is important because it bears directly on the completeness assumption ci=fcover,i; this limitation should be mentioned in the main-text discussion, not only in the appendix.","section":"Appendix A"},{"comment":"The caption contains a grammatical error ('the cyan solid line correspond') and should define what is meant by 'Mock expectation' in the legend or caption text.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is methodologically interesting and could be publishable after a major revision, but the headline claim currently overstates the statistical evidence. The referee's recommendation is driven by the mismatch between the 1.3-1.6 sigma mock significance and the abstract's 'evidence' language, compounded by the disappearance of the signal when four well-localized events are removed. A revision that repositions the paper as a method demonstration with upper limits and a clearly labeled tentative hint would be appropriate. I see no concerns about novelty or citation practices; the self-citation to Zhu et al. (2024) does not load-bear on the analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper improves the Bartos-style cross-correlation by using a 3D Voronoi tessellation to track a non-uniform AGN density and by splitting the AGN catalog by luminosity and Eddington ratio. That is genuinely new: previous studies assumed uniform density or focused on luminous AGNs. The authors also release their data on Zenodo, and they are transparent about the completeness degeneracy and about what happens when they drop the best-localized events.\n\nThe problem is that the \"evidence\" is much weaker than the abstract suggests. Their own null mocks put the significance at 1.3σ (lower-Lbol) and 1.6σ (lower-Eddington-ratio) — thin under any standard, and they search six overlapping sub-catalogs without a trials correction. The two positive sub-catalogs overlap by ~40%, so the look-elsewhere penalty is real. Most tellingly, Section 4.2 shows that removing four O4 events (really three, since one has no AGN in its error volume) drives the best-fit fagn to zero in every sub-catalog. The \"anomalous variation of the error\" argument they offer in response is a post-hoc comparison of 29-event vs 90-event samples, not a calibrated significance test. The null background Bi = 0.9 fcover,i is a secondary concern: if BBHs and AGNs share large-scale structure, the null is mis-specified, though their mock placement of hosts in SDSS galaxies partly captures that effect.\n\nNone of this makes the paper a waste of time. The method is a real step forward and the direction of the constraint — if anything, lower-luminosity and lower-Eddington-ratio AGNs are the plausible hosts, not the luminous ones — aligns with theoretical expectations and with the upper limits from Veronesi and colleagues. The completeness treatment is conservative (they overestimate ci and so underestimate fagn), though the systematic uncertainty is not folded into the quoted errors.\n\nBottom line: this is a solid methods contribution with a preliminary result that should not be sold as \"evidence.\" It deserves a serious referee, and after the claims are softened and a proper trials-corrected significance is added, it could be a credible paper. I'd want the robustness to the four-event removal addressed head-on, not hand-waved.","headline":"A useful method paper with a headline result that does not survive contact with its own robustness checks: the claimed AGN excess is 1.3–1.6σ, uncorrected for trials, and driven by three well-localized O4 events.","tokens_in":21160,"tokens_out":3861,"would_cite":true,"duration_ms":41181,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that 29-39 percent of 94 LIGO/Virgo/KAGRA gravitational-wave events are spatially correlated with lower-luminosity or lower-Eddington-ratio AGNs, with the signal dominated by four well-localized O4 events.","keywords":["Gravitational wave sources","Black holes","Active galactic nuclei","Sky surveys","Binary black hole mergers","Eddington ratio","Spatial correlation","AGN formation channel"],"falsifier":"The reported excess disappears when the four best-localized O4 events are removed, so the decisive check is external: re-run the same likelihood on the next several hundred well-localized events from ongoing and future runs. If the peak in $f_{\\rm agn}$ does not reappear once localization volumes shrink well below the typical AGN spacing, the correlation is a small-sample artifact. A reader could also build null catalogs that preserve the galaxy density field but randomize which galaxies host AGNs; a persistent excess there would invalidate the independence assumption.","tokens_in":19982,"feed_emoji":"🔭","tokens_out":11298,"duration_ms":105999,"temperature":0.7,"pith_summary":"This paper tries to establish that a sizable minority of the black-hole mergers detected by LIGO/Virgo/KAGRA formed inside active galactic nuclei (AGNs), specifically the dimmer and less actively accreting ones. Using 94 gravitational-wave events from the O1 through early O4 runs and roughly $3\\times10^5$ SDSS AGNs, the authors look for a statistical excess of AGNs inside the three-dimensional localization volume of each merger. They find that a fraction $f_{\\rm agn} = 0.39^{+0.41}_{-0.32}$ (lower-luminosity AGNs) or $0.29^{+0.40}_{-0.25}$ (lower-Eddington-ratio AGNs) is needed to explain the overlap, at 90\\% confidence. Monte Carlo mock catalogs with no AGN-merger link produce smaller overlaps, which the authors take as evidence that the correlation is not random coincidence. If correct, the result identifies a specific astrophysical birth environment for black-hole binaries and predicts that electromagnetic counterparts should be hunted around faint, low-accretion AGNs.","feed_headline":"A third of black-hole mergers may come from faint AGNs","feed_subtitle":"An excess of low-luminosity AGNs around 94 LIGO/Virgo events suggests an active-galaxy birth channel.","key_machinery":"The statistical engine is the event-wise likelihood $\\mathcal{L}(f_{\\rm agn}) = \\prod_i \\left[0.9 c_i f_{\\rm agn} S_i + (1 - 0.9 c_i f_{\\rm agn}) B_i\\right]$, whose maximum over $f_{\\rm agn}$ is the inferred AGN-origin fraction. The signal probability is $S_i = \\sum_j p_i(\\mathbf{x}_j)/n_{\\rm agn}(\\mathbf{x}_j)$, summing over cataloged AGNs inside the 90\\% localization volume, with each AGN weighted by the GW sky-map probability density at its position and downweighted by the local AGN density. A 3D first-order Voronoi tessellation of the SDSS AGN sample supplies $n_{\\rm agn}(\\mathbf{x})$, letting the AGN number density vary over the sky and with redshift instead of assuming it constant as earlier work did. The background probability is $B_i = 0.9 f_{{\\rm cover},i}$, where $f_{{\\rm cover},i}$ is the fraction of the error volume covered by the AGN catalog. This machine converts a stack of imperfect 3D localizations into one number, $f_{\\rm agn}$, plus its uncertainty.","core_discovery":"The paper's central claim is that the spatial distribution of lower-luminosity and lower-Eddington-ratio AGNs is not independent of the localization volumes of the LVK events. Dividing the SDSS DR16 AGN catalog into sub-catalogs by bolometric luminosity and Eddington ratio, the lower-$L_{\\rm bol}$ ($10^{44.5} \\lesssim L_{\\rm bol} \\le 10^{45}$ erg s$^{-1}$) and lower-$\\lambda_{\\rm Edd}$ ($0.01 \\lesssim \\lambda_{\\rm Edd} \\le 0.05$) populations show an excess around GW sources, while the moderate-, higher-, and full-catalog samples peak at $f_{\\rm agn}=0$. The likelihood fit gives $f_{\\rm agn} = 0.39^{+0.41}_{-0.32}$ and $0.29^{+0.40}_{-0.25}$ at 90\\% confidence. The authors argue the excess is unlikely to be random coincidence because mock realizations with injected $f_{\\rm agn}=0$ rarely produce as small a $\\Delta P_{0.05}$ as the real data, and because the inferred distribution narrows rather than broadens when the sample is reduced from the full catalog to O1-O3 alone for the two sub-catalogs in question. The paper also reports that the signal is dominated by the four best-localized O4 events; removing them returns the best-fit $f_{\\rm agn}$ to zero, which the authors interpret as the expected sensitivity of the test to high-quality localizations rather than as evidence against the correlation.","pith_inferences":["The authors do not push the obvious follow-up: a targeted look inside the error volumes of the four best-localized O4 events (S240413p, S250114ax, S250119cv, S230627c) for faint AGNs or AGN flaring would directly test whether these few events really host the claimed population.","Because the null mock catalogs draw hosts from the SDSS galaxy catalog, one natural extension is to scramble AGN positions within that same galaxy density field; if the excess survives such scrambling, the independence assumption in the null, not a physical AGN association, would be the explanation.","If low-Eddington-ratio AGNs are the preferred birthplaces, the controlling variable may be disk gas density rather than total luminosity, which suggests a cross-correlation split by galaxy environment or emission-line properties as a sharper test."],"forward_implications":["If the correlation is real, roughly 30-40 percent of the 94 analyzed events formed in low-luminosity or low-Eddington-ratio AGNs, making the AGN channel a major rather than negligible formation route at these redshifts.","The absence of a signal for luminous AGNs (best-fit zero, with 90% upper limits of 0.16-0.37 depending on the sample) sharpens earlier upper limits and redirects electromagnetic counterpart searches toward fainter nuclei.","The spatial-correlation estimate agrees with an independent hierarchical-Bayesian analysis that finds $f_{\\rm agn} \\approx 0.34$ for O1-O3 events, so two different statistical approaches bracket the same channel.","As the ongoing run completes toward roughly 300 candidates and new AGN catalogs cover more than 70 percent of the sky, the same machinery should shrink the $f_{\\rm agn}$ errors by about half and roughly double the significance of a nonzero fraction if the true value is near 0.2.","The method's sensitivity comes disproportionately from well-localized events, so future runs with improved localization will provide the decisive test of whether the excess persists."],"supporting_citations":[{"why":"Introduces the statistical method of looking for an excess of AGNs inside GW localization volumes that this paper modifies and applies.","marker":"Bartos et al. 2017a"},{"why":"Provides the Monte Carlo mock-event framework used to validate the method and calibrate its null behavior.","marker":"Veronesi et al. 2022"},{"why":"Applies the earlier method to real LVK data and sets the upper limit on the most luminous AGN contribution that this work refines.","marker":"Veronesi et al. 2023"},{"why":"Updates the constraint with a wider-area AGN catalog, giving the comparison upper limits that the paper's results are checked against.","marker":"Veronesi et al. 2025"},{"why":"Supplies the bolometric luminosities and Eddington ratios for the SDSS DR16 AGNs on which the sub-catalog split is based.","marker":"Wu & Shen 2022"},{"why":"Provides the SDSS DR16 quasar/AGN catalog that defines the observed AGN positions and density field.","marker":"Lyke et al. 2020"},{"why":"Supplies the SDSS galaxy catalog used to draw mock host galaxies for the injected-f_agn=0 null simulations.","marker":"Ahumada et al. 2020"},{"why":"Theoretical prediction of an anti-correlation between BBH merger rate and Eddington ratio, which the paper's lower-Eddington-ratio excess is compared with.","marker":"Yang et al. 2019a"},{"why":"Independent hierarchical-Bayesian estimate of f_agn from O1-O3 that agrees with the spatial-correlation value, used as cross-validation.","marker":"Li & Fan 2025"}],"fun_headline_variants":["Spatial excess links black-hole mergers to faint AGNs","Faint AGN excess hints at black-hole merger birthplace","Black-hole mergers cluster around dim active galaxies","Some GW events may originate near low-luminosity AGNs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"When no merger actually occurs in an AGN, the analysis assumes that AGNs and gravitational-wave sources are placed independently, so overlaps are pure chance; if both merely trace the same large-scale cosmic structure, the null expectation is too small and the inferred positive $f_{\\rm agn}$ would be inflated.","fun_headline_variants_meta":{"raw":{"variants":["Spatial excess links black-hole mergers to faint AGNs","Faint AGN excess hints at black-hole merger birthplace","Black-hole mergers cluster around dim active galaxies","Some GW events may originate near low-luminosity AGNs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000534,"raw_usage":{"total_tokens":2694,"prompt_tokens":1194,"completion_tokens":1500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":810,"completion_tokens_details":{"reasoning_tokens":1434}},"tokens_in":810,"tokens_out":1500,"duration_ms":13392,"temperature":1.0,"reasoning_tokens":1434,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:39:09.728566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The reported excess disappears when the four best-localized O4 events are removed, so the decisive check is external: re-run the same likelihood on the next several hundred well-localized events from ongoing and future runs. If the peak in $f_{\\rm agn}$ does not reappear once localization volumes shrink well below the typical AGN spacing, the correlation is a small-sample artifact. A reader could also build null catalogs that preserve the galaxy density field but randomize which galaxies host AGNs; a persistent excess there would invalidate the independence assumption.","supporting_citations":[],"review_version":1}