{"id":"959f5906-5f8e-4948-a3d2-3333abefffcd","arxiv_id":"2508.05742","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Simulated LSST light curves show that host-galaxy contamination increases false positives and reduces binary SMBH recovery, and that extending survey duration from 5 to 10 years improves detection more than increasing cadence.","lead":"This simulation study tests how host-galaxy light contamination and the Rubin Observatory's observing schedule affect our ability to find pairs of giant black holes. It finds that longer observation windows matter more than observing frequency, and that contamination worsens detection and recovery.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulation's assumed quasar variability model may drive the duration-over-cadence conclusion; no validation against real light curves is visible in the abstract.","rationale":"The reader's weakest_assumption correctly identifies that the simulation's fidelity is the key risk, specifically mentioning the possibility of underestimated red noise. My stress-test sharpens this into a concrete mechanism: the duration-vs-cadence ranking from sinusoidal fits is directly controlled by the assumed noise PSD, because longer baselines help only if the stochastic noise is sufficiently well-behaved. The abstract contains no evidence that the model was validated against real quasar variability, so the central claim remains unverified. The reader's UNVERDICTED verdict is appropriate; my concern reinforces it rather than moving it. I do not change the verdict, because the abstract provides insufficient information to either confirm or refute the concern. A concrete test on real or alternative synthetic data would be needed to move the verdict.","tokens_in":776,"tokens_out":3431,"duration_ms":38890,"concrete_test":"Re-run the simulation pipeline with two alternative quasar variability models: (A) the current empirical model, and (B) a broken power-law PSD with low-frequency slope -1.5 and amplitude calibrated to the ensemble PSD of SDSS Stripe 82 quasars. Use the same sinusoidal detection method and survey designs (cadences 3/6/12 days; durations 5/10/20 years). If the relative ordering of false-positive suppression and period-recovery improvement between duration and cadence changes across models—e.g., cadence matters more under model B—the paper's central claim is not robust. An orthogonal check: inject the same binary signals into real ZTF light curves of non-binary quasars and apply the same detection pipeline; compare the duration/cadence ranking with the simulation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that monitoring duration affects binary SMBH detection more than effective cadence—is derived entirely from simulated LSST light curves using 'simple sinusoidal curve fits' and variability models 'motivated by empirical observations.' The abstract provides no validation of these simulations against real quasar light curves (e.g., SDSS Stripe 82, ZTF). Simple sinusoidal fits are known to be highly susceptible to spurious periodicities from red noise; the false-positive rate as a function of light-curve duration depends critically on the amplitude and spectral index of the assumed noise. If the quasar variability model underestimates low-frequency red noise (e.g., by choosing too short a characteristic timescale in a DRW model), longer durations will appear to suppress false positives more than they would in real LSST data, artificially favoring duration over cadence. Similarly, host-galaxy contamination is modeled with 'empirical distributions,' but its effect on parameter recovery could be mis-estimated if the simulated host-galaxy magnitudes/colors do not match the LSST sample. Because the paper's headline conclusion is a relative ranking of survey parameters, even moderate mismodeling of the noise could flip the ranking. This is the most load-bearing assumption: if the simulation's noise model is not faithful, the duration-over-cadence result may not transfer to real LSST operations.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on simulations of millions of LSST light curves for single and binary supermassive black hole candidates, using empirically motivated distributions of quasar variability and host-galaxy contamination, and then applies simple sinusoidal curve fits as a computationally inexpensive detection method. The main findings are that host-galaxy contamination increases false positives and degrades binary parameter recovery, especially for lower-mass, lower-luminosity systems, and that monitoring duration affects binary detection more than survey effective cadence: extending the light curve duration from 5 to 10 years yields the most dramatic improvement in false-positive rejection and binary-period recovery, with further gains from 20-year durations, particularly for periods longer than a decade.","tokens_in":1098,"tokens_out":3121,"duration_ms":33928,"significance":"If the simulation faithfully represents the LSST data stream, the result has direct practical importance for survey strategy, arguing for long monitoring baselines over dense sampling for sinusoidal binary searches. The work's strengths are its scale (millions of light curves), the forward-modeled ground truth, and the simplicity of the detection method, which can be benchmarked against other approaches. However, the central ranking is only as trustworthy as the simulated quasar variability and host-galaxy contamination; the abstract does not demonstrate validation against observed light curves, and sinusoidal fitting is sensitive to red-noise false positives. The significance is therefore conditional on the full paper's presentation of the noise model and validation.","major_comments":[{"comment":"The headline claim that monitoring duration affects binary detection more than effective cadence rests entirely on the simulated quasar variability. The abstract does not report the noise model (e.g., DRW parameters, characteristic timescale, spectral index) or any validation against real quasar light curves such as SDSS Stripe 82 or ZTF. Since red noise at low frequencies can produce spurious periodic signals that mimic binaries, and the false-positive rate as a function of duration depends strongly on the assumed power spectrum, an unvalidated noise model could flip the duration/cadence ranking. Please provide the model parameters and a direct comparison to observed ensemble variability statistics.","section":"Abstract"},{"comment":"The detection method is described only as 'simple sinusoidal curve fits.' The abstract does not define the detection threshold, the false-positive criterion, or how the period search is conducted. Without this, the claim that false-positive rates are suppressed with duration cannot be assessed; for a fixed period search, longer light curves provide more cycles, but the false-positive rate also depends on the number of trial periods and the noise spectrum. The full paper must specify these choices to establish that the improvement is not an artifact of the search setup.","section":"Abstract"},{"comment":"Host-galaxy contamination is modeled with 'empirical distributions,' but the abstract does not state how host-galaxy light is added, whether the host contribution is constant or variable, or how the detection method treats it. The conclusion that lower-mass, lower-luminosity binaries are most affected is plausible if contamination dilutes the periodic signal, but it depends on the assumed host-galaxy flux fraction relative to the quasar. The authors should quantify the host-galaxy flux distribution and show that it matches the expected LSST sample in terms of redshift and magnitude limits.","section":"Abstract"}],"minor_comments":[{"comment":"The term 'effective cadence' is used without definition; please clarify whether it refers to median observing interval, number of visits per year, or a survey-strategy metric.","section":"Abstract"},{"comment":"The abstract reports 'millions of light curves' but does not give statistical uncertainties on the recovery or false-positive rates; including confidence intervals would strengthen the quantitative claims.","section":"Abstract"},{"comment":"The claim that binary SMBHs 'dominate the low-frequency gravitational wave background' is a strong statement; consider citing the relevant pulsar timing array and population synthesis literature.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"This review is based solely on the abstract because the full text was not available. The paper's central conclusion about duration versus cadence is plausible but is critically dependent on the fidelity of the simulated quasar variability, about which the abstract provides no quantitative detail. The lack of visible validation against real light curves is the main correctness risk. I would not reject the paper on the abstract alone, but I cannot recommend acceptance without examining the full simulation and fitting details."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is clear: for simple sinusoidal detection methods, survey duration matters more than cadence, and host-galaxy contamination hurts low-mass binaries most. That's a genuinely useful, actionable message for LSST planning. The paper earns credit for a systematic parameter study with empirically motivated distributions and millions of simulated light curves; the specific claim that the 5-to-10-year extension gives the biggest gain is the kind of concrete guidance observers want.\n\nThe soft spot is exactly where the stress-test lands: the entire ranking of duration versus cadence depends on the simulated quasar variability, especially the low-frequency red noise. Simple sinusoid fits are notoriously vulnerable to spurious periodicities from red noise, and if the simulated noise underestimates long-timescale power, longer light curves will suppress false positives more than they would in real data. The abstract gives no hint of validation against real quasar light curves (Stripe 82, ZTF, etc.). That's not necessarily a fatal flaw—many simulation papers rely on literature noise models—but it is the load-bearing assumption, and the referee needs to see that the noise parameters are realistic or that the conclusions are tested against alternative noise models.\n\nHost-galaxy contamination is a second concern, though less central. The abstract says the distributions are empirically motivated, but without details we can't judge whether the simulated host galaxies match LSST's actual sample. That could bias the parameter-recovery results, especially for low-mass systems.\n\nI can't verify any of this from the abstract alone, and neither can the reader. The verdict of 'unverified' is fair. But the paper deserves a serious referee: the question is important for LSST cadence decisions, the approach is appropriate if not novel, and the conclusions are specific enough to be checked. I'd want the referee to focus on the noise model, the false-positive definition, and any validation against real light curves. If those hold up, this is a solid contribution; if not, it's still a useful cautionary tale.\n\nFor peer review: send it out. It's not a desk reject.","headline":"Useful, concrete simulation results for LSST binary SMBH searches, but the duration-over-cadence conclusion hinges on the realism of the quasar noise model—needs full-paper scrutiny.","tokens_in":1504,"tokens_out":1482,"would_cite":false,"duration_ms":17409,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that, for simple sinusoidal fits to simulated LSST light curves, survey duration matters more than effective cadence for detecting binary supermassive black holes, with the clearest gain when monitoring extends from 5 to 10","keywords":["binary supermassive black holes","LSST","time-domain surveys","quasar variability","host-galaxy contamination","survey cadence","survey duration","sinusoidal detection"],"falsifier":"A direct test: on archival quasar light curves with two fixed baselines of 5 and 10 years at matched cadence, measure the false-positive rate of sinusoidal candidates; if the 10-year baseline does not suppress false positives relative to the 5-year baseline as strongly as the simulations predict, the central claim that duration dominates cadence would be refuted.","tokens_in":740,"feed_emoji":"🔭","tokens_out":5914,"duration_ms":61745,"temperature":0.7,"pith_summary":"This paper is about how the upcoming LSST survey will find binary supermassive black holes—pairs of black holes orbiting each other after galaxies merge—by spotting periodic brightness changes in quasar light. The authors simulate millions of LSST light curves of single and binary quasars, using empirically motivated distributions of quasar variability and host-galaxy contamination. They compare two survey design choices: how often the sky is observed (cadence) and how long the monitoring runs (duration). The central finding is that duration matters more than cadence: extending the light curves from 5 to 10 years sharply reduces false positives and improves recovery of binary parameters, especially for periods longer than a decade. Host-galaxy contamination is shown to raise false-positive rates and degrade parameter recovery, with the strongest damage for lower-mass, lower-luminosity binaries.","feed_headline":"Ten years of monitoring, not cadence, finds binary black holes","feed_subtitle":"Simulated LSST light curves show longer baselines suppress false positives and recover long-period binaries.","key_machinery":"The load-bearing object is the simple sinusoidal fit applied to simulated light curves: a periodic model of the form $A\\sin(\\omega t + \\phi)$ is used to search for binary SMBH candidates against the background of intrinsic single-quasar variability and host-galaxy starlight. The simulations vary survey cadence (time between observations) and total duration (length of the monitoring baseline) across millions of light curves, separating which design parameter controls detectability. Host-galaxy contamination is modeled as dilution of the quasar's periodic signal, which reduces its effective amplitude and mimics noise.","core_discovery":"For simple sinusoidal fits to simulated LSST light curves, the paper establishes that monitoring duration is the dominant survey design lever for binary supermassive black hole identification. Host-galaxy contamination—starlight from the host galaxy diluting the quasar signal—increases false-positive rates and lowers binary parameter recovery, with the effect strongest for lower-mass, lower-luminosity binaries. Increasing duration from 5 to 10 years yields the largest single improvement in false-positive suppression and period recovery, while extending to 20 years gives additional gains. Periods longer than a decade especially require long baselines to be recovered at all.","pith_inferences":["Editorial inference: if the duration-over-cadence result holds outside the simulation, similar sinusoidal searches for quasi-periodic eruptions or changing-look active galactic nuclei should also favor 10-year baselines over denser sampling.","Editorial inference: the host-galaxy contamination result means simple photometric surveys will preferentially find high-luminosity binaries; demographic inferences from such samples would need to correct for this selection effect.","Editorial inference: the 5-to-10-year gap suggests a practical survey strategy: allocate additional LSST time to maintaining a long timeline for a modest sample rather than increasing cadence, and combine with host-galaxy subtraction to recover lower-mass binaries."],"forward_implications":["For LSST binary black hole searches, a 10-year observing baseline should be prioritized over denser sampling; the biggest gain in false-positive rejection comes at the 5-to-10-year transition.","Binary periods longer than a decade are the main beneficiaries of longer duration; surveys with short baselines will systematically miss the most massive, longest-period binaries.","Host-galaxy light must be modeled or subtracted in detection pipelines; ignoring it will inflate false positives and bias against low-mass, low-luminosity binaries.","Simple sinusoidal detection methods remain a feasible computational approach for the first pass over LSST data, provided the survey duration is sufficient."],"supporting_citations":[],"fun_headline_variants":["Duration beats cadence for binary black hole detection","Host galaxy starlight boosts false binary black hole signals","Decade-long monitoring key to finding binary black holes in LSST","Five-year Rubin runs miss long-period binary black holes","Long baselines suppress fake binary black hole detections"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The relative importance of duration over cadence rests on the assumption that the simulated quasar variability and host-galaxy contamination faithfully represent real LSST light curves; if actual red noise is stronger or binary signals deviate from sinusoids, the ranking may reverse.","fun_headline_variants_meta":{"raw":{"variants":["Duration beats cadence for binary black hole detection","Host galaxy starlight boosts false binary black hole signals","Decade-long monitoring key to finding binary black holes in LSST","Five-year Rubin runs miss long-period binary black holes","Long baselines suppress fake binary black hole detections"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1361,"prompt_tokens":794,"completion_tokens":567,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":538,"completion_tokens_details":{"reasoning_tokens":488}},"tokens_in":538,"tokens_out":567,"duration_ms":6347,"temperature":1.0,"reasoning_tokens":488,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:10:23.910630+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test: on archival quasar light curves with two fixed baselines of 5 and 10 years at matched cadence, measure the false-positive rate of sinusoidal candidates; if the 10-year baseline does not suppress false positives relative to the 5-year baseline as strongly as the simulations predict, the central claim that duration dominates cadence would be refuted.","supporting_citations":[],"review_version":1}