{"id":"a83c8ae7-40df-4875-bac9-33270981e574","arxiv_id":"2506.04986","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A forward model that includes CHIME's beam response shows the FRB population can follow star formation, contrary to several earlier claims, while the implied source rate is too high for core-collapse magnetars alone.","lead":"The authors built a Monte Carlo simulation with CHIME's beam and selection effects and found that the simple idea that fast radio bursts track star formation cannot be ruled out by current data, though a small delay fits slightly better. This matters because earlier studies claimed star-formation tracking was excluded, so this shifts which FRB progenitor scenarios remain on the table.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'SFH cannot be rejected' claim rests on an uncalibrated p>=0.01 acceptance rule, and the paper's own footnote 12 concedes the resulting uncertainty; the non-rejection may reflect low test power rather than population consistency.","rationale":"The reader's weakest assumption (beam model and s(DM)) is well-founded, and the paper explicitly flags the beam model as early. However, the single most load-bearing assumption is methodological: the acceptance criterion p2DKS>=0.01 is used as if it certified consistency, but no calibration is provided. A KS p-value is uniform only under the null; under a false model with tuned parameters, high p-values can be common. The paper's own footnote 12 acknowledges the artificial uncertainty, and the delayed-model 'preference' is even weaker because it compares p-values from two uncalibrated acceptance maps. Correcting the beam model would change the p-values, but unless the acceptance procedure is calibrated, the claim 'SFH cannot be rejected' remains ambiguous. The verdict should therefore remain CONDITIONAL, as the reader set, but for the sharper reason of uncalibrated statistical inference rather than the beam uncertainty alone. The proposed calibration test would settle whether the non-rejection is informative.","tokens_in":18113,"tokens_out":11876,"duration_ms":145286,"concrete_test":"Calibrate the acceptance rule by drawing, say, 100 synthetic observed samples from a deliberately non-SFH redshift model (e.g., a constant comoving-rate distribution with the same energy and DM_host priors), running the paper's exact SFH Monte Carlo procedure on each, and counting the fraction for which at least one parameter draw gives p2DKS>=0.01. If this false-acceptance fraction is comparable to or greater than the 5% implied by the threshold, the paper's non-rejection of SFH is uninformative; the authors should also report the number of trials used in each acceptance map so the passing fraction can be evaluated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 3, parameter combinations are accepted when the bootstrap 2D KS p-value satisfies p2DKS >= 0.01, and the SFH model is declared 'cannot be rejected' because such combinations occupy a 'considerable' region (Fig. 3). This is a calibration-free rule: for a false model, the distribution of p2DKS over random parameter draws is not known, and with enough draws some p>=0.01 hits occur by chance. The paper never states how many parameter combinations were attempted or what fraction of the prior passes the threshold, so the accepted region cannot be distinguished from the tail of a uniform p-value distribution. The 95% ranges quoted for alpha and DMhost are histograms of accepted draws with p>=0.05, not frequentist or Bayesian intervals, and they can be dominated by the prior boundaries, as the unconstrained Ec illustrates. The authors' own footnote 12 concedes the 'artificial uncertainty' from the p-value threshold. Consequently, the headline result that the SFH model is consistent with the first CHIME/FRB catalog does not yet distinguish 'data prefer SFH' from 'this test would fail to reject many false models'. The preliminary CHIME beam model and the James (2023) s(DM) assumed in Eqs. (16) and (15) are separate, potentially large biases, but correcting them would still leave the acceptance-rule calibration problem unresolved.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses Monte Carlo simulations that forward-model FRB populations through CHIME/FRB selection effects (beam response, DM-dependent selection, S/N threshold) to test whether the redshift distribution of the first CHIME/FRB catalog bursts tracks the cosmic star formation history directly or with a delay. The authors find that the SFH model cannot be rejected within a considerable parameter region, that a delay of about 1 Gyr is preferred though not required, and they derive a local volumetric rate of approximately 2.3e5 Gpc^-3 yr^-1 above 1e38 erg, which they argue challenges core-collapse magnetars as the sole FRB progenitor channel.","tokens_in":18369,"tokens_out":4482,"duration_ms":54263,"significance":"The question addressed is central to FRB progenitor physics, and the paper is unusually explicit about selection effects, using the CHIME beam model and a DM-dependent selection function. The non-rejection result, if confirmed, would overturn several earlier claims that the SFH model is excluded, and the local rate estimate provides a concrete falsifiable constraint on progenitor models. The main weakness is that the central non-rejection claim rests on an uncalibrated p-value acceptance rule and on preliminary instrument models; these issues are correctable but load-bearing.","major_comments":[{"comment":"The claim that the SFH model 'cannot be rejected' rests entirely on selecting parameter draws with p_2DKS >= 0.01 and on observing that accepted draws occupy a 'considerable' region. The paper does not report the total number of parameter draws, the fraction of the prior volume accepted, or the distribution of p-values under a known false model. With uncalibrated thresholding, a false model can produce p >= 0.01 in some draws by chance, and footnote 12 itself concedes that the resulting parameter ranges carry 'artificial uncertainty.' The authors should calibrate the acceptance rule, for example by running the same pipeline on simulated populations with a deliberately wrong redshift distribution and reporting the acceptance fraction as a function of sample size, and they should show the actual p-value distribution rather than only accepted contours. Without this, the headline result cannot be distinguished from low test power.","section":"Section 3, Figs. 3 and 5, footnote 12"},{"comment":"The quoted '95% confidence' ranges alpha = -2.1(+0.5,-0.4) and DM_host < 239 pc cm^-3 are histograms of accepted draws with p_2DKS >= 0.05, not confidence intervals in a frequentist or Bayesian sense. They depend on the arbitrary p-value threshold and on the prior boundaries, as the unconstrained log(E_c) demonstrates. These quantities should be re-labeled as summaries of the accepted region or replaced with a properly calibrated inference procedure.","section":"Section 3, Table 2 and text near Fig. 3"},{"comment":"The beam model is described in Section 5 as 'an early version,' yet Eq. (16) uses it to convert every true fluence to a measured lower-limit fluence and Eq. (15) uses the same gain G to set detection. If the gain pattern or the boresight normalization is inaccurate, the simulated F_nu and S/N distributions are systematically biased, and the non-rejection of the SFH model could be an artifact of the beam model rather than a property of the FRB population. A robustness check against the released CHIME beam-corrected fluences or against variations of the beam model is needed to support the central claim.","section":"Section 2.4, Eq. (16), and Section 5"},{"comment":"The DM selection function s(DM) is imported from James (2023) as a fourth-order polynomial and applied through the relation S/N_bias ~ s^{2/3}(DM). The paper does not test the sensitivity of the accepted parameter region to this function, even though low-DM selection strongly shapes the DM_E distribution and hence the inferred redshift distribution. The authors should vary the polynomial parameters or compare with CHIME's injection-based selection function to demonstrate that the non-rejection of the SFH model is robust to this modeling choice.","section":"Section 2.4, Eq. (15) and Table 1"},{"comment":"The spectral index is fixed to gamma = -0.65, and footnote 5 states that tests with gamma = -1.0 and -1.5 show 'only minor differences' without presenting those results. Given the known degeneracy between spectral index and source evolution discussed in Shin et al. (2023) and Hoffmann et al. (2025), the redshift-distribution conclusion requires those tests to be shown, at least as an appendix figure or table.","section":"Section 2.1, Eq. (7), and footnote 5"}],"minor_comments":[{"comment":"The phrase 'a parallel-computable code hat runs on a thread-ripper CPU' should read 'that runs on a thread-ripper CPU.'","section":"Section 2, first paragraph"},{"comment":"The caption says 'Accepted range of free parameters for SFH model,' but the table refers to the delayed SFH model; the caption should be corrected.","section":"Table 3 caption"},{"comment":"The notation 'DM_WM' should be 'DM_MW' for the Milky Way contribution, for consistency with the text.","section":"Section 2.2, Eq. (8)"},{"comment":"The dispersion parameter is denoted sigma_DM in the surrounding text but appears as sigma_IGM in the equation; the notation should be made consistent.","section":"Section 2.2, Eq. (10)"},{"comment":"The explanation of counting only the first burst of repeating sources is repeated nearly verbatim in the Conclusions and in footnote 13; one of these repetitions should be removed or condensed.","section":"Section 5 and footnote 13"}],"recommendation":"major_revision","confidential_remarks":"The paper's honest self-criticism in footnote 12 identifies the central methodological gap. The manuscript is publishable in principle, but the main non-rejection claim needs a calibration/power analysis; otherwise the conclusion is too weak to support the paper's title and abstract. The fit with the journal's scope is appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the most physically complete forward model yet applied to CHIME Catalog 1, and it correctly shows that earlier SFH rejections leaned on simple selection corrections. But the central non-rejection claim is statistically softer than the abstract implies.\n\nWhat is actually new: they explicitly include the CHIME beam gain map, the DM-dependent selection function, and the fact that catalog fluences are lower limits. That is a real step beyond the papers they criticize, and their treatment of the fluence issue is fair. The parallel code and the systematic parameter exploration also go beyond the sparse grids used in several prior Monte Carlo studies. I also credit the honest engagement with the literature, including the independent Wang & van Leeuwen (2024) result and the careful note that the derived local rate challenges the core-collapse magnetar channel.\n\nThe soft spot is the acceptance rule. In Section 3 they accept parameter combinations with p2DKS >= 0.01 and then say the SFH model cannot be rejected because accepted draws occupy a \"considerable\" region. That is not calibrated: for a false model, a fraction of random draws will clear any threshold, and the paper never reports total trials or the fraction accepted. Their own footnote 12 concedes the \"artificial uncertainty\"; that is a real admission, not a minor caveat. The 95% ranges quoted for alpha and DM_host are histograms of accepted draws, not intervals with frequentist or Bayesian meaning, and the unconstrained log Ec shows the prior boundary problem. The beam model is also explicitly preliminary, and the DM selection function and spectral index are fixed inputs from other work. None of these alone sinks the paper, but together they mean the headline result \"SFH not ruled out\" does not yet distinguish a true preference for SFH from a test that would fail to reject many false models. The rate estimate is useful and roughly consistent with prior work, but it inherits the same selection-modeling caveats.\n\nMissing code is a real reproducibility cost. The sample file and the beam model are public, but the simulation pipeline is not, so a referee cannot check the KS procedure or the acceptance calculation without reimplementing it.\n\nWho this is for: FRB population theorists and anyone working with CHIME Catalog 1. It deserves a serious referee, not a desk reject. My recommendation: send it to review, but ask for (1) calibration of the acceptance threshold, ideally reporting total draws and the accepted fraction, (2) sensitivity tests to the s(DM) function and gamma, (3) an updated beam model when available, and (4) release of the simulation code. With those, the non-rejection claim could become solid; without them, it stays conditional.","headline":"A genuinely more careful CHIME forward model, but the headline \"SFH not ruled out\" rests on an uncalibrated p>=0.01 acceptance rule and an early beam model; worth reviewing, but it needs revision.","tokens_in":18981,"tokens_out":1673,"would_cite":true,"duration_ms":24904,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that, once CHIME's beam response and selection effects are modeled in detail, FRB sources can still be born in lockstep with cosmic star formation, while the implied local source rate is too high for core-collapse…","keywords":["fast radio bursts","FRB population synthesis","star formation history","CHIME/FRB catalog","selection effects","magnetar progenitors","volumetric rate","redshift distribution"],"falsifier":"Use CHIME's calibration or injection system to measure the true off-meridian beam gain and check whether recovered fluences obey $F_\\nu^\\mathrm{measured}=G(x,y)/G(0,y)F_\\nu^\\mathrm{true}$, or repeat the KS analysis on the subset of bursts with independently measured beam-corrected fluences; if the no-delay model is rejected with corrected fluences, the acceptance here was a selection-modeling artifact, and if it survives, earlier rejections were artifacts of using lower limits.","tokens_in":17801,"feed_emoji":"📡","tokens_out":12481,"duration_ms":136424,"temperature":0.7,"pith_summary":"The paper asks whether the bursts in the first CHIME/FRB catalog can discriminate between FRB sources born in lockstep with cosmic star formation and sources that only ignite after a delay. Using Monte Carlo mock catalogs that mimic CHIME's beam pattern, its radio-frequency-interference filtering, and the fact that catalog fluences are lower limits, the authors test both scenarios against 223 detected bursts. They find that a population whose birth rate simply follows the star formation history cannot be excluded: a sizeable region of parameter space reproduces the observed fluence and dispersion-measure distributions, even though a delay of about a gigayear fits the data somewhat better. The same simulations imply a local volumetric rate of roughly $2.3^{+2.4}_{-1.2}\\times10^5$ events per $\\mathrm{Gpc}^3$ per year above $10^{38}$ erg, consistent with earlier estimates but high enough to strain the idea that every source is a young magnetar from a core-collapse supernova. If true, this would mean earlier rejections of the no-delay model were driven by selection-effect modeling rather than by the data.","feed_headline":"FRB births may track star formation after all","feed_subtitle":"CHIME's 223 bursts fit a star-tracking FRB population, but the implied source rate is too high for magnetars alone.","key_machinery":"The load-bearing machinery is a Monte Carlo population-synthesis pipeline that turns a hypothetical FRB redshift distribution and energy function into a mock CHIME sample. The key identities are the detection criterion $\\mathrm{S/N} = G F_\\nu / \\sqrt{B N_p (T_\\mathrm{rec}+T_\\mathrm{sky}) \\sqrt{w}} \\times s(\\mathrm{DM})^{2/3}$ and the fluence-to-boresight relation $F_\\nu^\\mathrm{measured} = G(x,y)/G(0,y)\\,F_\\nu^\\mathrm{true}$, where $G(x,y)$ is the beam-corrected gain at the burst's sky position. The measured fluence is always a lower limit because the gain peaks at the meridian, which lets the simulation use the catalog's lower-limit fluences honestly. A DM-dependent selection function $s(\\mathrm{DM})$ encodes the telescope's strong incompleteness at low dispersion measures, and mock bursts are kept only if their S/N exceeds 12 and their fluence is below 100 Jy ms. Simulated and observed samples are then compared with 1D and 2D Kolmogorov-Smirnov tests on $F_\\nu$ and $\\mathrm{DM}_E$, with bootstrap p-values.","core_discovery":"The central claim is that the observed redshift distribution of CHIME FRBs does not force a delay relative to the cosmic star formation history once the telescope's selection effects are modeled in detail. The authors show that the no-delay, SFH-tracking model passes Kolmogorov-Smirnov tests against the catalog's fluence and extragalactic DM distributions over a broad range of energy-function parameters; the delayed model with a typical delay of about 1 Gyr fits a bit better but is not required. They also derive a local event-rate density of $2.3^{+2.4}_{-1.2}\\times10^5\\,\\mathrm{Gpc}^{-3}\\,\\mathrm{yr}^{-1}$ for sources with isotropic energy above $10^{38}$ erg. Because the core-collapse supernova rate is of the same order, the channel that makes young magnetars cannot by itself supply the required number of independent FRB sources; the paper concludes that additional or alternative progenitor channels are needed even though prompt, star-formation-associated progenitors remain viable.","pith_inferences":["Beyond the paper, a clean test is to repeat this analysis on the CHIME/FRB catalog's beam-corrected fluence subset: if the no-delay model survives, earlier rejections were selection artifacts; if it fails, the early beam model used here is the weak link.","The claimed insensitivity of the energy-function slope ($\\alpha\\approx-1.8$) to the redshift model suggests that the energy function and the delay time can be constrained separately, which a future hierarchical fit could exploit to break the degeneracy between $\\alpha$ and $\\log E_c$.","If repeating FRB sources are numerous and each source bursts many times, the distinction between a source-rate density and a burst-rate density becomes essential; a survey counting first bursts only may be measuring the density of active sources times their burst rate rather than a formation rate, and this needs to be checked with repetition statistics.","A natural extension is to allow the spectral index $\\gamma$ to float per source instead of fixing it at $-0.65$; the paper's robustness check with $\\gamma=-1.0$ and $-1.5$ suggests the qualitative conclusion would hold, but a free $\\gamma$ would also test whether the slight preference for a delay is driven by the assumed spectral correction."],"forward_implications":["If the no-delay model is actually viable, FRB sources could be young, short-lived objects formed in star-forming regions, so a star-forming host galaxy does not by itself exclude such progenitors.","A delay of about 1 Gyr being preferred, though not required, leaves room for a mix of prompt and delayed channels, including old stellar populations such as those in globular clusters.","Treating CHIME catalog fluences as true values biases redshift-distribution conclusions, so future analyses should use beam-corrected fluences or marginalize over the beam model.","The local source rate above $10^{38}$ erg is too high for core-collapse magnetars to be the sole source population, so theoretical effort should go into additional channels or into mechanisms that boost the apparent source rate.","The inferred all-sky rate above 5 Jy ms, $216^{+21}_{-19}\\,\\mathrm{sky}^{-1}\\,\\mathrm{day}^{-1}$, is within a factor of two of the CHIME/FRB team's estimate, meaning the sample normalization is roughly consistent."],"supporting_citations":[{"why":"Supplies the 474 non-repeating and 62 repeating bursts in the first catalog, the lower-limit fluences and DMs, and the field of view used for rate estimates.","marker":"CHIME/FRB Collaboration et al. 2021"},{"why":"Provides the CHIME beam model from which the beam-corrected gain G(x,y) at 600 MHz and the boresight fluence ratio are taken.","marker":"Merryfield et al. 2023"},{"why":"Gives the DM-dependent selection function s(DM) and the s^{2/3} S/N scaling used to model RFI-related incompleteness.","marker":"James 2023"},{"why":"Defines the refined 225-burst sample and supplies an earlier S/N-based population synthesis whose conclusions about the SFH model are compared and broadly reproduced.","marker":"Shin et al. 2023"},{"why":"Supports the adopted spectral-index correction gamma=-0.65 under the rate interpretation of FRB spectra.","marker":"James et al. 2022b"},{"why":"Provides the analytic fit to the cosmic star formation history used as the no-delay redshift distribution.","marker":"Yüksel et al. 2008"},{"why":"Supplies the lognormal delay-time distribution used in the delayed-SFH model.","marker":"Wanderman & Piran 2015"},{"why":"Gives the statistical distribution of intergalactic dispersion measure for a given redshift that converts redshifts to DM_E.","marker":"Macquart et al. 2020"},{"why":"Provides the S/N formula and the relation between measured and boresight fluences used in the detection criterion.","marker":"Andersen et al. 2023"},{"why":"A prior Monte Carlo study that ruled out the SFH model with simpler selection handling; the paper argues the discrepancy arises from those simplifications.","marker":"Zhang & Zhang 2022"}],"fun_headline_variants":["FRB rates may match star formation, but magnetars fall short","CHIME data: FRBs track star formation, but need extra sources","Star-forming FRBs fit CHIME, yet magnetars can't explain rate","No delay needed for FRBs, but magnetar-only fails","FRB births track stars, but magnetar rate mismatch persists"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on assuming that the early CHIME beam model and the adopted DM-dependent selection function correctly convert true fluences and burst positions into catalog fluences and detection probabilities; if the true beam response or incompleteness differs, the non-rejection of the star-formation-history model could be an artifact of the selection modeling.","fun_headline_variants_meta":{"raw":{"variants":["FRB rates may match star formation, but magnetars fall short","CHIME data: FRBs track star formation, but need extra sources","Star-forming FRBs fit CHIME, yet magnetars can't explain rate","No delay needed for FRBs, but magnetar-only fails","FRB births track stars, but magnetar rate mismatch persists"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1419,"prompt_tokens":963,"completion_tokens":456,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":579,"completion_tokens_details":{"reasoning_tokens":364}},"tokens_in":579,"tokens_out":456,"duration_ms":5133,"temperature":1.0,"reasoning_tokens":364,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:29:17.456454+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use CHIME's calibration or injection system to measure the true off-meridian beam gain and check whether recovered fluences obey $F_\\nu^\\mathrm{measured}=G(x,y)/G(0,y)F_\\nu^\\mathrm{true}$, or repeat the KS analysis on the subset of bursts with independently measured beam-corrected fluences; if the no-delay model is rejected with corrected fluences, the acceptance here was a selection-modeling artifact, and if it survives, earlier rejections were artifacts of using lower limits.","supporting_citations":[{"cited_title":"P., Shin, K., et al","cited_arxiv_id":null,"evidence_quote":"Provides the CHIME beam model from which the beam-corrected gain G(x,y) at 600 MHz and the boresight fluence ratio are taken."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the DM-dependent selection function s(DM) and the s^{2/3} S/N scaling used to model RFI-related incompleteness."},{"cited_title":"W., Bhardwaj, M., et al","cited_arxiv_id":null,"evidence_quote":"Defines the refined 225-burst sample and supplies an earlier S/N-based population synthesis whose conclusions about the SFH model are compared and broadly reproduced."},{"cited_title":"& Piran, T","cited_arxiv_id":null,"evidence_quote":"Supplies the lognormal delay-time distribution used in the delayed-SFH model."},{"cited_title":"C., Patel, C., Brar, C., et al","cited_arxiv_id":null,"evidence_quote":"Provides the S/N formula and the relation between measured and boresight fluences used in the detection criterion."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A prior Monte Carlo study that ruled out the SFH model with simpler selection handling; the paper argues the discrepancy arises from those simplifications."}],"review_version":1}