{"id":"2b1a8e1d-777e-4c91-a715-e770934fe02e","arxiv_id":"2411.19481","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Tidal heating imprints from black hole horizons can be captured by two effective parameters and modeled through merger, offering a future test to distinguish black holes from exotic horizonless objects.","lead":"This thesis studies how the absorption of energy by black hole horizons, called tidal heating, imprints on gravitational-wave signals from merging compact binaries. It finds that next-generation detectors could measure this imprint, and it builds a waveform model that slightly improves agreement with numerical relativity data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Ch. 5 late-inspiral flux model is the load-bearing step for the abstract's horizon-test claim; its acknowledged low-frequency failure and the linear splice Eq. (5.22) leave that demonstration unvalidated.","rationale":"The reader identifies the Ch. 5 late-inspiral flux model as the weakest assumption, and I agree that it is the load-bearing element for the abstract's claim of demonstrating a route to testing horizonless objects from NR data. The model's own acknowledged failure at low frequencies (Sec. 5.3.2) means the claimed 'demonstration' is actually an ad hoc construction unless the splice is validated; the paper provides no such validation. I considered whether the measurability forecasts in Ch. 3 are more load-bearing, since the headline 8% and 2% errors omit spins and tc/phi_c and worsen by about an order of magnitude when those are included (Sec. 3.4.2, Figs. 3.5-3.6). That is a real limitation, but the reader already flags it, and the centrality of the late-inspiral test to the thesis's stated goal makes Ch. 5 the single most consequential weakness. The Ch. 4 model improvement is modest and supported by direct mismatch comparisons against NR data, so it is not the main vulnerability. The concrete tests proposed would settle whether the spliced flux model actually connects to the PN limit and generalizes beyond the six calibration points; until then, CONDITIONAL remains the appropriate verdict, so no change to the reader's verdict is needed.","tokens_in":55980,"tokens_out":3033,"duration_ms":28135,"concrete_test":"Evaluate Eq. (5.9) with the fitted beta_ij (Eq. 5.11) at Mf = 0.022 and compare with the 4PN flux Eq. (5.8). If |(dM/dt)_fit - (dM/dt)_PN| / (dM/dt)_PN exceeds 1 at the splice point, the linear ramp cannot connect smoothly. Then compute the dephasing delta-psi_TH from Eqs. (5.15)-(5.21) using both the ramp and a pure-PN flux below 0.022, and require the two waveforms to agree within the modeled accuracy (e.g., mismatch < 1%) over the band. An out-of-sample check: refit Eq. (5.9) to SXS:BBH:0030 and SXS:BBH:0056, then predict SXS:BBH:0169 (q=2, excluded from calibration) in Mf in [0.03, 0.04]; if the prediction error exceeds the 1-sigma fit errors, the polynomial Eq. (5.10) is overfit.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that tidal heating can distinguish BHs from ECOs using the late inspiral depends on Ch. 5's phenomenological flux model. Eq. (5.9) is fitted to NR horizon data only over Mf in [0.03, 0.04] (Sec. 5.3.1), with four parameters a1..a4 extrapolated in eta via Eq. (5.10). The authors state in Sec. 5.3.2 that at lower frequencies the model deviates by orders of magnitude from the 4PN term; the 'complete' model is then assembled by a linear ramp Eq. (5.22) that forces ai to zero at Mf = 0.022. This ramp is not derived from physics or from NR data, and no validation is shown that the spliced dephasing reproduces either the NR horizon flux or the known PN limit at the interface. Because the late-inspiral dephasing delta-psi_TH(v) in Eqs. (5.15)-(5.21) is built directly on this flux model, any error in the ramp propagates into the claimed horizon test. The 4% median improvement of PhenomD Horizon (Eq. 4.37) is a separate, modest claim and does not rescue the Ch. 5 demonstration.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript, posted as a PhD thesis, develops tidal heating (TH) as a gravitational-wave discriminant for black-hole horizons. Chapter 3 forecasts the measurability of two effective horizon parameters, H_eff5 and H_eff8, using Fisher and Bayesian analyses of a TaylorF2+TH inspiral model in LIGO-Virgo, Einstein Telescope, and Cosmic Explorer, reporting projected 1-sigma constraints of about 8% (H_eff5) and 2% (H_eff8) at 1 Gpc in a fiducial fixed-spin configuration. Chapter 4 builds 'IMRPhenomD Horizon,' a frequency-domain phenomenological inspiral-merger-ringdown model that adds TH phase and amplitude corrections to IMRPhenomD, calibrates it on 20 SEOBNRv2+NR hybrids, validates it on 40 hybrids, and reports a median mismatch improvement over IMRPhenomD of about 4% (aLIGO ZDHP) and 2% (flat noise) against 219 SXS NR waveforms. Chapter 5 models the total horizon flux of nonspinning binaries from SXS apparent-horizon data using a four-parameter ansatz fitted over Mf in [0.03, 0.04], derives the corresponding frequency-domain dephasing, and proposes a linear ramp to splice this to the 4PN low-frequency expression, as a route to testing horizonless exotic compact objects in the late inspiral.","tokens_in":56270,"tokens_out":17080,"duration_ms":132894,"significance":"The Ch. 3 forecasts are useful, internally consistent benchmarks for 3G detectors, and the Bayesian cross-checks (Sec. 3.5) support the Fisher results. Ch. 4 delivers a concrete, self-contained IMR waveform with TH included and benchmarked against external SXS NR data; the full coefficient tables (Appendix C, Eq. 5.11) and the use of public toolkits (GWBENCH, Bilby, the SXS catalog) make the quantitative claims reproducible in principle. The thesis is also commendably transparent: it discloses that the headline errors rise by about an order of magnitude when spins are included (Fig. 3.6), that the amplitude correction is subdominant (Sec. 4.7), and that the Ch. 5 flux model deviates from the PN limit by orders of magnitude outside its fitted band (Sec. 5.3.2). If the Ch. 5 splicing were validated, the framework would be a genuine step toward late-inspiral horizon tests; as it stands, the Ch. 5 demonstration is a proposal with an unquantified error budget, and the abstract's horizon-test claim rests on that unvalidated ground.","major_comments":[{"comment":"The late-inspiral horizon-flux model that underpins the paper's horizon-test proposal is not validated at either end of its frequency range. The four-parameter ansatz (5.9) is calibrated to apparent-horizon data from only six SXS simulations over the narrow band Mf in [0.03, 0.04], and Sec. 5.3.2 states that outside this band the model 'behaves extremely poorly at lower frequencies,' with orders of magnitude of deviation from the 4PN term. The complete dephasing is then assembled by the linear ramp (5.22) on {a_i} between Mf = 0.022 and 0.032. No test is presented that the spliced dM/dt (and hence delta_PSI_TH in Eqs. (5.15)-(5.21)) reproduces either the NR horizon data at Mf approximately 0.032 or the 4PN expression at Mf approximately 0.022, and no held-out SXS simulation is used to check the fit. The concern is compounded by the fact that the low-frequency anchor itself is applied at v approximately 0.41, where truncated 4PN is not expected to be accurate. Because delta_PSI_TH is constructed directly from this flux model, the claim that the late inspiral can be leveraged to test for horizonless compact objects inherits an unquantified error budget. I recommend: (i) validating the spliced model against one or two nonspinning SXS binaries not in Table 5.1; (ii) checking continuity of dM/dt and delta_PSI_TH at both interfaces; and (iii) propagating the 1-sigma fit errors of Fig. 5.1 into the dephasing, together with a report on the conditioning of the four-parameter fit over such a short band.","section":"Sec. 5.3; Eqs. (5.9), (5.15)-(5.22)"},{"comment":"The headline measurability numbers, Delta_H_eff5 approximately 0.05 (8.3%) and Delta_H_eff8 approximately 0.2 (2%) for M = 30 M_sun, q = 1.5, DL = 1 Gpc, are obtained in the reduced five-parameter space {Mc, eta, DL, Heff5, Heff8} with the spins held fixed at chi1 = chi2 = 0.8. The paper's own Sec. 3.4.2 reports that adding the spins together with tc and phi_c (Fig. 3.6) raises the errors by roughly an order of magnitude relative to the seven-parameter Fig. 3.5, and describes Fig. 3.6 as the more realistic estimate. This disclosure is commendable, but Sec. 3.7 and Ch. 6 restate the five-parameter numbers without that qualification, and the abstract's general claim of tight constraints inherits the same framing. The summary should present the spin-marginalized errors as the primary benchmark, or clearly label the 8% and 2% values as fixed-spin, idealized projections.","section":"Secs. 3.4.2, 3.7, and 6"},{"comment":"The central accuracy claim of Ch. 4, a roughly 4% (ZDHP) and 2% (flat noise) improvement in median mismatch against 219 SXS waveforms, is reported without a significance statement, and the comparison set includes the 20 hybrids used to calibrate the model (Table 4.1). The two histograms in Figs. 4.7 and 4.8 overlap substantially, so the shift in medians could be within sampling error or driven by the calibration subset; the single outlier noted in Sec. 4.6 (q = 4, chi1 = 0, chi2 = 0.8; 1.21% at 70 M_sun) shows that the model is not uniformly better. Please report the median improvement computed on the 199 non-calibration waveforms only, add an uncertainty on the median shift (e.g., a bootstrap over the SXS set), and state the fraction of the 219 waveforms for which PhenomD Horizon is the better template.","section":"Sec. 4.7; Eq. (4.37), Figs. 4.7-4.8"}],"minor_comments":[{"comment":"There are several typos and notation slips that should be cleaned up: 'fequencies' (Sec. 3.4.2, first paragraph), 'covarinaces' (Sec. 3.6), 'IMRPheonmD' (Sec. 5.4), and 'post-Newtonain' (Sec. 4.4).","section":"Various"},{"comment":"The instruction that '14.65 is to be added to the tick labels of Mc' is confusing; the chirp-mass axis should simply be relabeled with the correct values.","section":"Fig. 3.7 caption"},{"comment":"For reproducibility, please state the actual values of the Planck-taper window parameters (x1, x2, x3, x4) used in the hybridization, rather than only the functional form.","section":"Sec. 4.4, Eqs. (4.18)-(4.20)"},{"comment":"The restriction to nonspinning binaries and the simplifying assumption H1 = H2 = H should be stated at the beginning of Sec. 5.3, where the modeling strategy is set, rather than appearing in the middle of the paragraph describing the mass-flux data.","section":"Sec. 5.3"},{"comment":"The entries of the beta_ij matrix are quoted to inconsistent precision (four entries to two or three significant figures, the rest to six or more), and no uncertainties are given for the coefficients despite the 1-sigma error bars displayed on the individual a_i fits in Fig. 5.1.","section":"Eq. (5.11)"},{"comment":"The column labeled 'chi_PN' should be defined in the caption as the effective spin parameter of Eq. (4.25), since the table otherwise appears to list only component spins.","section":"Table 4.1"},{"comment":"The manuscript should state explicitly at first use that the TH phase expressions, Eqs. (3.4), (3.7), and (4.10)-(4.12), are taken from the authors' own Ref. [114], so that the reader can distinguish restated published material from new derivations in this thesis.","section":"Secs. 3.2 and 4.3.1"}],"recommendation":"major_revision","confidential_remarks":"The core of the thesis has already been released as peer-reviewed or preprint material: Chapter 3 corresponds to Phys. Rev. D 106, 104032 (2022) and Chapter 4 to arXiv:2311.17554, while Chapter 5 is listed as 'under preparation.' The editor should weigh the venue's prior-publication policy when considering this thesis-format submission. The scientific assessment above is unaffected by this, but it is worth noting that the only load-bearing component that has not passed independent review elsewhere is precisely the Ch. 5 splicing model that needs the most work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The waveform and measurability chapters are solid, and the late-inspiral horizon test in Chapter 5 is a genuinely new but unvalidated model. The thesis is honest about where it fails, but the abstract and conclusion overstate what has been demonstrated.\n\nWhat's new: Chapter 5's horizon-flux ansatz (Eq. 5.9), the fitted parameters, and the dephasing formulas (5.15–5.21) do not appear in the earlier literature. That is a real first step toward NR-informed horizon tests. What the thesis does well: Chapter 4's IMRPhenomD Horizon is a proper frequency-domain model, calibrated against 20 hybrids and validated against 40, with mismatches mostly around 0.1%. The median improvement over PhenomD against 219 SXS NR waveforms is about 4% — modest, but real and honestly reported. Chapter 3's Fisher/Bayesian comparison is careful, and the text is transparent about the parameter-space dependence: the headline 8% and 2% errors come from a five-parameter space with spins fixed; once spins are included, errors rise by an order of magnitude. The paper says this plainly.\n\nThe soft spots are in the gap between the abstract and Chapter 5. The thesis states that its NR-fitted flux model deviates by orders of magnitude from the 4PN term at lower frequencies, and the \"complete\" model is assembled by a linear ramp (Eq. 5.22) that has no derivation and no validation at the interfaces. The dephasing that would carry the horizon test is built on that unvalidated splice. There is also no injection/recovery study showing the spliced model can actually distinguish BBHs from horizonless ECOs; the chapter ends with \"can be used,\" not with a demonstrated test. The conclusion's \"rigorously\" is not supported yet.\n\nThe self-citation in Chapters 3 and 4 is not a problem — the earlier work is published, and the waveform model is benchmarked against external SXS data. The real issue is that Chapter 5 is calibrated, not validated, on the same NR data it is meant to describe.\n\nWho gets value: anyone working on tidal heating, phenomenological BBH waveform models, or 3G detector projections. As a journal submission it would need major revision — either validation of the splice or an explicit statement that the horizon-test claim is deferred. I would still send it to peer review rather than desk reject: the measurability study and the waveform model deserve referee attention, and the new model raises questions worth engaging. Just don't expect the abstract's horizon-test claim to survive unchanged.","headline":"Solid measurability and waveform modeling, but the late-inspiral horizon test is an unvalidated, self-admittedly ad hoc model.","tokens_in":56838,"tokens_out":7743,"would_cite":false,"duration_ms":59764,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Tidal heating could reveal whether black holes really have horizons","keywords":["tidal heating","black hole horizons","exotic compact objects","gravitational waves","inspiral-merger-ringdown waveform","horizon parameters","numerical relativity","post-Newtonian theory"],"falsifier":"Compute the total mass change dM/dt from high-resolution numerical-relativity simulations of nonspinning binaries at frequencies between Mf = 0.01 and merger, and check whether the linear ramp of Eq. (5.22) reproduces the simulated data within errors. If the ramp fails to connect the PN and NR regimes, the late-inspiral dephasing model would be invalidated.","tokens_in":55727,"feed_emoji":"🌊","tokens_out":2682,"duration_ms":25950,"temperature":0.7,"pith_summary":"This thesis argues that tidal heating—the absorption of orbital energy by a black hole's horizon—leaves a measurable imprint on gravitational waves from merging compact binaries. If the compact objects are horizonless exotic alternatives to black holes, that imprint is weaker, so measuring it offers a way to tell black holes apart from their mimickers. The author derives two effective horizon parameters, Heff5 and Heff8, shows they could be tightly constrained with future detectors, and builds a phenomenological inspiral-merger-ringdown waveform model that includes tidal-heating corrections, improving agreement with numerical-relativity data.","feed_headline":"Tidal heating could expose black-hole mimickers","feed_subtitle":"New waveform model measures horizon absorption, separating black holes from horizonless objects.","key_machinery":"The central mechanism is tidal heating: the horizon of a black hole absorbs energy and angular momentum from the companion's tidal field, altering the inspiral rate and thus the gravitational-wave phase and amplitude. The horizon parameters Heff5 and Heff8, which enter the phase at 2.5PN and 4PN order, quantify the fraction of that absorption; they are defined as mass- and spin-weighted combinations of the individual horizon parameters, and their measured values distinguish black holes from horizonless compact objects.","core_discovery":"Tidal heating imprints a phase and amplitude correction on gravitational-wave signals from binary black holes, characterized by two effective horizon parameters Heff5 and Heff8, which take known values for BHs and smaller values for horizonless exotic compact objects. The thesis claims that these parameters are measurable, especially in future third-generation detectors like Einstein Telescope and Cosmic Explorer, with projected 1-sigma errors as small as about 8% for Heff5 and 2% for Heff8 for a fiducial binary. It further constructs a frequency-domain waveform model, IMRPhenomD Horizon, that adds these corrections to the inspiral phase and amplitude, and shows it improves mismatch against numerical-relativity waveforms by about 4%, making tidal heating a practical discriminant for black holes.","pith_inferences":["The two effective parameters may serve as a model-independent diagnostic: measuring both simultaneously could distinguish a single-parameter deviation from general relativity from a genuine two-parameter horizon modification.","Extending the horizon-flux modeling to spinning binaries would likely boost the signal, since higher spins strengthen the phase correction and widen the measurable parameter space.","A practical next step would be to reanalyze existing gravitational-wave events with the new waveform model, searching for statistically significant deviations of Heff5 and Heff8 from the black-hole prediction.","If tidal-heating corrections are as significant as modeled, this approach could complement other black-hole tests such as quasinormal-mode ringdown and tidal Love numbers, covering binaries too heavy for inspiral-only analyses and too light for ringdown-only tests."],"forward_implications":["If the horizon parameters can be measured, gravitational-wave observations could directly test for the presence of horizons in stellar-mass binaries.","In third-generation detectors, Heff5 and Heff8 could be constrained tightly enough to distinguish black holes from exotic compact objects for binaries out to a gigaparsec.","Adding tidal-heating corrections to phenomenological waveforms reduces systematic errors, improving parameter estimation and tests of general relativity in strong-field regimes.","The late-inspiral horizon-flux model could be attached to IMRPhenomD Horizon to search for deviations from the black-hole prediction in heavier binaries.","Current and future detectors could use these methods to probe the nature of compact objects across the entire detectable mass range."],"supporting_citations":[{"why":"Introduces the horizon parameters Heff5 and Heff8 and the phase-shift formula used throughout Chapter 3.","marker":"[114]"},{"why":"Provides the leading post-Newtonian expressions for mass and spin evolution of black holes that underlie the tidal-heating flux model.","marker":"[115]"},{"why":"The SXS catalog supplies the numerical-relativity waveforms and apparent-horizon data used for hybrid construction and horizon-flux fitting.","marker":"[212]"},{"why":"The IMRPhenomD waveform model is the foundation whose inspiral and amplitude are recalibrated to incorporate tidal heating.","marker":"[194, 195]"},{"why":"Provides the SEOBNRv2 effective-one-body approximant used as the point-particle baseline for the hybrids.","marker":"[186]"},{"why":"The Einstein Telescope detector design is used for measurability projections of the horizon parameters.","marker":"[119]"},{"why":"The Cosmic Explorer detector design is used for measurability projections of the horizon parameters.","marker":"[120]"},{"why":"Supplies the most recent analytical horizon-flux expressions that the numerical-relativity model is compared against.","marker":"[234]"}],"fun_headline_variants":["Tidal heating reveals black-hole mimickers in gravitational waves","New model detects black-hole mimickers via tidal heating in GW signals","Tidal heating imprints measurable signatures of black-hole mimickers","Gravitational-wave model separates black holes from exotic mimickers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"That the late-inspiral horizon-flux model, fitted to a narrow frequency band of numerical-relativity data, can be spliced to the low-frequency post-Newtonian result through a linear ramp—the thesis itself states this complete dephasing model is an ad hoc construction that behaves poorly at lower frequencies.","fun_headline_variants_meta":{"raw":{"variants":["Tidal heating reveals black-hole mimickers in gravitational waves","New model detects black-hole mimickers via tidal heating in GW signals","Tidal heating imprints measurable signatures of black-hole mimickers","Gravitational-wave model separates black holes from exotic mimickers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000311,"raw_usage":{"total_tokens":1799,"prompt_tokens":1003,"completion_tokens":796,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":619,"completion_tokens_details":{"reasoning_tokens":725}},"tokens_in":619,"tokens_out":796,"duration_ms":7481,"temperature":1.0,"reasoning_tokens":725,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:08:38.874690+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute the total mass change dM/dt from high-resolution numerical-relativity simulations of nonspinning binaries at frequencies between Mf = 0.01 and merger, and check whether the linear ramp of Eq. (5.22) reproduces the simulated data within errors. If the ramp fails to connect the PN and NR regimes, the late-inspiral dephasing model would be invalidated.","supporting_citations":[],"review_version":1}