{"id":"f66d00e4-f8d9-4e91-a5c2-239a814d49a7","arxiv_id":"2608.08733","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Using neural-posterior-estimation simulation-based inference on the MOJAVE FSRQ sample, the authors infer a Lorentz factor distribution slope b = -1.32 (+0.20, -0.19), consistent with X-ray binary jets at 2 sigma.","lead":"This paper applies simulation-based inference, a machine-learning statistical method, to model the population of radio-bright black hole jets in the MOJAVE survey. It finds the jets' speed distribution has a power-law slope consistent with stellar-mass black hole jets at the 2 sigma level, supporting similar jet physics across mass scales.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The b posterior is calibrated to the simulator, not to the true parent population; the MMD test (p=0.994) is too weak to rule out model-form misspecification, and a known mismatch exists between simulated and observed apparent speeds.","rationale":"The reader's weakest-assumption analysis already identifies the central issue: the inference assumes the Lister et al. (2019) generative model is the true parent population, and the calibration tests only validate the flow conditional on that simulator. My stress-test confirms this and sharpens it with a specific, acknowledged mismatch between the simulated observable and the observed data: the use of maximal apparent speeds from multiple jet components versus a single simulated β_app. This mismatch directly affects the observable most sensitive to the Lorentz-factor exponent b, and it is acknowledged in the paper's own discussion (Section 4.2). The MMD test is not strong enough to retire the concern, because it checks membership of the observed data in the simulated distribution, not the correctness of the generating assumptions. The paper has real strengths: the calibration methodology is unusually thorough, multiple seeds are used, the architecture is documented in detail, and the posterior is not running against prior boundaries. These support the internal consistency of the SBI pipeline, but they do not address the external validity of the simulator. The reader's conditional verdict already requires a robustness analysis of model-form assumptions, which is exactly what would settle this concern. I therefore see no reason to change the verdict: it should remain conditional on that robustness check and on the code-release condition. The agreement is marked 'agree' because the reader's weakest assumption names the same generative-model dependence as the primary source of potential bias, and my concrete test is a direct operationalization of that concern.","tokens_in":11309,"tokens_out":6539,"duration_ms":82926,"concrete_test":"Re-run the SBI pipeline on the same 174 FSRQs while changing one generative assumption at a time: (i) extend Γ_min to 1.05 and Γ_max to 100; (ii) replace the pure power-law N(Γ) with a broken power law or log-normal; (iii) perturb the inclination distribution to p(θ) ∝ sinθ(1 + ε cosθ) for small ε; (iv) simulate the observed maximal β_app by drawing multiple components per source from the same Γ and θ and taking the maximum, rather than drawing a single β_app. For each alternative, report the b posterior and a posterior predictive check on the observed β_app and redshift distributions. If any alternative with a comparable MMD p-value shifts b by more than the quoted ±0.2, the central slope claim and the XRB consistency are not robust to model-form uncertainty.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, b = -1.32 (+0.20, -0.19) and 2-sigma agreement with XRB jets, rests on the Lister et al. (2019) generative model being the true parent population: N(Γ) ∝ Γ^b on [1.25, 50], isotropic inclinations with p(θ)dθ = sinθ, pure luminosity evolution, and selection solely by the 1.5 Jy flux limit (Table 1, Section 2.1). The extensive calibration tests (SBC, TARP, coverage, L-C2ST; Section 2.2.2, Table A2) verify only that the flow recovers the posterior of the simulator; they cannot detect whether the simulator itself is misspecified. The MMD test (p = 0.994) is offered as evidence against misspecification, but this is a goodness-of-fit test with limited power: failing to reject the null only means the 174-object observed sample is not obviously atypical of the simulated ensemble, and a p-value this close to 1 is not itself evidence for correctness. The paper explicitly acknowledges one known mismatch in Section 4.2: the observed apparent speeds are maximal values over multiple jet components, while the simulator draws a single β_app per source from one Γ and θ. This can bias b toward flatter power laws, which is precisely the direction that would ease the comparison with the XRB posterior. Consequently, the reported posterior width and asymmetry may be well-calibrated conditional on the model, but the point estimate can be precise-and-biased if the functional form or the maximal-speed statistic is wrong. This is the load-bearing weakness because the paper's main numerical result and the cross-mass comparison both inherit every untested assumption of the generative model.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies simulation-based inference (SBI) with neural posterior estimation to model the parent population of AGN jets using the MOJAVE 1.5 Jy FSRQ sample. The authors adopt the generative model of Lister et al. (2019): a power-law Lorentz factor distribution N(Γ) ∝ Γ^b on [1.25, 50], isotropic inclinations, a power-law luminosity function with pure luminosity evolution, and selection by a 1.5 Jy flux limit. They train a neural spline flow with per-object and permutation-invariant embeddings, and validate it with SBC, TARP, empirical coverage, an MMD misspecification test, and L-C2ST. The main quantitative claim is b = -1.32 (+0.20, -0.19), with posteriors that are broader and more asymmetric than previous grid-based estimates. They further claim that this AGN Lorentz factor distribution is consistent at the 2σ level with the XRB jet population posterior taken from a companion paper (Lilje et al. 2025). The paper concludes that SBI is well suited to population studies of jetted AGN with multiple observables.","tokens_in":11706,"tokens_out":3014,"duration_ms":34430,"significance":"If the central claim holds, the paper makes a useful methodological contribution by showing that SBI can produce calibrated, non-Gaussian posteriors for AGN jet parent-population parameters where grid-based Anderson-Darling fits cannot. The extensive calibration suite (SBC, TARP, coverage, MMD, L-C2ST) is a genuine strength, and the authors are appropriately cautious in several places about flow pathologies. The claim that AGN and XRB jet speed distributions are compatible at 2σ, if robust, would strengthen the case for common jet physics across black hole mass scales. However, the headline value of b is a fitted quantity whose accuracy depends on the simulator being an adequate model of the true parent population, and the paper's own discussion acknowledges a known mismatch in how apparent speeds enter the analysis. Because the 2σ agreement is the paper's main astrophysical conclusion, the robustness of b to this and other model-form choices is load-bearing.","major_comments":[{"comment":"The MMD test is reported as yielding p = 0.994 and is described as 'strongly disfavours misspecification'. This interpretation is not supported: a high p-value only means that the 174-object observed sample is not unusual under the simulated distribution; it does not establish that the simulator is a good proxy for the true parent population, and the test has limited power against many structured misspecifications. Please replace this wording with a more neutral statement, and ideally provide a power analysis or a demonstration that the MMD test can reject alternative simulators (e.g., with different Γmin, broken power laws, or different missingness mechanisms).","section":"Section 2.2.2 and Section 3"},{"comment":"The paper acknowledges that observed apparent speeds are maximal values over multiple jet components, while the simulator draws a single β_app per source from one Γ and θ, and states that this may bias b toward flatter power laws. Since the central scientific result is the value of b and its 2σ agreement with the XRB posterior, this is a load-bearing caveat, not a peripheral one. The direction of the suspected bias is exactly the direction that would ease agreement with the XRB sample. Please quantify the impact: for example, modify the simulator to draw multiple components per source and take the maximum, or explicitly demonstrate that the inferred b shifts by less than the reported uncertainty under such a change. Without this, the central claim can be precise but biased.","section":"Section 4.2"},{"comment":"The posterior for b is only meaningful if the assumed generative model is approximately correct: N(Γ) ∝ Γ^b on [1.25, 50], isotropic inclinations, pure luminosity evolution, and selection solely by the 1.5 Jy flux limit. The calibration tests validate the flow against its own simulator and cannot detect misspecification of this model class. The paper's MMD test is too weak to resolve this. I recommend adding sensitivity analyses that vary the functional form (e.g., a broken power law or a high-Γ cutoff) and the fixed boundaries (Γmin, Γmax), and showing how the inferred b changes. The current manuscript gives no quantitative indication of the robustness of b to these choices.","section":"Table 1 and Section 2.1"}],"minor_comments":[{"comment":"The name 'Kalmogorov-Smirnov' should be 'Kolmogorov-Smirnov'.","section":"Section 2.2.2"},{"comment":"The statement that a L-C2ST p-value of 0.82 allows one to 'say with 95% confidence that the estimator is valid' is statistically imprecise; a p-value above 0.05 only means the test does not reject validity, not that validity has 95% confidence.","section":"Section 3"},{"comment":"The missingness mechanism is described as randomly choosing 10%–50% of apparent speed values to be missing, but the actual missingness in the MOJAVE sample may not be missing at random. A brief justification or a sensitivity test of the missingness fraction and pattern would strengthen the preprocessing description.","section":"Section 2.2.1"},{"comment":"The comparison to Lister et al. (2019) uses the spacing between adjacent grid points as the literature uncertainty; this is not a statistical uncertainty and should be labelled as such, since the SBI posteriors are on a different footing.","section":"Section 3"}],"recommendation":"major_revision","confidential_remarks":"The XRB comparison in Figure 2 relies on a companion paper by the same authors (Lilje et al. 2025, arXiv:2511.19362) that is not yet peer-reviewed. The 2σ agreement claim inherits any errors in that work, and the manuscript should be explicit about this dependency. The paper is otherwise within the scope of the journal and the methodology is well presented, but the load-bearing sensitivity of b to the maximal-apparent-speed mismatch needs to be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this is a methodologically serious reanalysis of the MOJAVE FSRQ parent population, and the headline b = -1.32 (+0.20, -0.19) is probably the most carefully calibrated estimate of that slope to date. The XRB comparison is suggestive, not established, because the simulator-level validation cannot fix a known model-data mismatch that pushes b in the direction of agreement.\n\nWhat is genuinely new: previous grid-based errors from Lister et al. (2019) were essentially grid spacings; the SBI posteriors here are broader, asymmetric, and explicitly calibrated. The validation is unusually thorough for this literature—SBC, TARP, empirical coverage, L-C2ST, and an MMD misspecification test are all reported, and the authors check that fixing p = 2.0 leaves the results stable. Those tests do what they claim: they show the flow recovers the posterior of the simulator. They do not, and cannot, show the simulator is the right description of the sky.\n\nThat limitation matters more than the paper lets on. The MMD p-value of 0.994 is read as \"strongly disfavours misspecification,\" but a non-rejection on 174 objects is weak evidence for the model class. More telling is the acknowledged mismatch in Section 4.2: observed apparent speeds are maximal values over multiple jet components, while the simulator draws one beta_app per source. The paper concedes this can bias b toward flatter power laws—the same direction that brings AGN closer to the XRB posterior. So the central numerical result is well-calibrated conditional on Lister 2019's generative model, but the point estimate could be precise-and-biased if that model or the maximal-speed treatment is wrong.\n\nThe XRB comparison is a same-group continuation and also depends on the posterior from Lilje et al. (2025). That is not circular, but it is another reason to want independent confirmation before taking the 2-sigma consistency as evidence for scale-invariant jet physics.\n\nI would send this to peer review and require, as a condition, release of the simulator, training scripts, seeds, or pretrained flow artifacts, plus a robustness appendix varying Gamma_min, Gamma_max, and the single-component versus maximal-beta_app choice. The methodology and the correction to previous error bars are valuable; the cross-mass consistency claim needs one more layer of scrutiny.","headline":"A careful, well-calibrated SBI reanalysis of MOJAVE FSRQs; the b posterior is credible conditionally, but the known maximal-speed bias runs in the direction that makes the 2-sigma XRB agreement look better than it may be.","tokens_in":12236,"tokens_out":2852,"would_cite":true,"duration_ms":35552,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Simulation-based inference puts AGN jet-speed slope at b = -1.32, matching XRBs.","keywords":["active galactic nuclei","jets","Lorentz factor distribution","simulation-based inference","normalising flows","neural posterior estimation","MOJAVE","X-ray binaries"],"falsifier":"Rerun the same SBI pipeline on the same MOJAVE data with the simulator's $\\Gamma_{\\min}$ changed from 1.25 to 2 and with a broken power-law; if the posterior for $b$ shifts by more than the quoted uncertainty, the central claim depends on an unverified cutoff.","tokens_in":11119,"feed_emoji":"🕳️","tokens_out":5697,"duration_ms":55322,"temperature":0.7,"pith_summary":"This paper tries to establish the parent distribution of jet speeds for radio-loud active galactic nuclei, after correcting for the severe observational bias imposed by a flux limit and relativistic beaming. Using likelihood-free simulation-based inference—neural posterior estimation with a normalising flow—on the 174 FSRQs of the MOJAVE 1.5 Jy sample, it claims the intrinsic Lorentz factor distribution is $N(\\Gamma)\\propto\\Gamma^b$ with $b=-1.32^{+0.20}_{-0.19}$. The authors argue this calibrated posterior is substantially broader and more asymmetric than earlier grid-based estimates, so previous error bars understated uncertainty. If true, the AGN and X-ray binary jet-speed distributions are consistent at the $2\\sigma$ level, hinting at shared jet physics across black hole masses. The wider point is that simulation-based inference is a workable route to population inference when multiple observables and selection effects make the likelihood intractable.","feed_headline":"AGN jet-speed slope lands at b = -1.32, matching XRBs","feed_subtitle":"Full posteriors from simulation-based inference make supermassive and stellar-mass black hole jets statistically consistent.","key_machinery":"The load-bearing machinery is the combination of a Monte Carlo parent-population simulator and a conditional neural spline flow trained by neural posterior estimation. The simulator generates synthetic MOJAVE-like samples of 174 FSRQs from priors on the Lorentz-factor slope $b$, luminosity-function slope $\\gamma$, beaming exponent $p$, and redshift-evolution parameters $k$ and $\\eta$, applies the $1.5\\,$Jy flux cut, and produces per-object observables (flux, redshift, apparent speed, beamed luminosity, plus a missing-data flag). A permutation-invariant encoder aggregates the per-object features into a context vector, and the normalising flow learns $p(\\theta | x)$ directly; calibration is checked with simulation-based calibration, coverage, TARP, an MMD misspecification test, and a local classifier two-sample test. The flow's job is to turn simulator outputs into a full posterior that captures degeneracies and asymmetry without an analytic likelihood.","core_discovery":"The central discovery is a new posterior for the underlying jet-speed (Lorentz factor) distribution of the MOJAVE FSRQ population, obtained without writing down a likelihood. The paper's headline result is $b = -1.32^{+0.20}_{-0.19}$ for $N(\\Gamma) \\propto \\Gamma^b$, with $\\Gamma$ between 1.25 and 50, which is consistent at $2\\sigma$ with the Lorentz-factor exponent inferred for X-ray binary jets. Equally central is the claim that the posterior is non-Gaussian and wider than the grid-based parameter estimates of Lister et al. (2019); the flow-based inference also exposes a strong degeneracy between the two redshift-evolution parameters $k$ and $\\eta$, which the Anderson-Darling grid method could not characterise. The authors interpret this as showing that previous comparisons underestimated the statistical uncertainty and that, once properly accounted for, AGN and XRB jet speed populations are statistically compatible.","pith_inferences":["If the posterior is as broad as reported, then earlier claims of a significant AGN–XRB speed difference may have been driven by underestimated uncertainties; re-testing with the full posterior would quantify this.","The low value of $\\Gamma_{\\min}=1.25$ is degenerate with the slope $b$; applying the same flow to simulated samples with $\\Gamma_{\\min}=2$ or 5 would test how much of the result depends on this choice.","The consistency of $b$ across radio-, X-ray- and $\\gamma$-ray-selected samples hints that the Lorentz-factor distribution is set by jet launching physics rather than by the band used to select sources, but confirming this requires samples selected at other wavelengths with the same missing-data treatment.","The missing-data handling via a $-1$ flag and indicator channel could be reused for other surveys with partial measurements, and the permutation-invariant encoder is a template for population-level SBI beyond jets."],"forward_implications":["Previous grid-based error bars for AGN parent-population parameters are too small; reanalyses using full posteriors will widen confidence intervals in population comparisons.","The AGN and XRB jet Lorentz-factor distributions are consistent at $2\\sigma$, so differences between them cannot be claimed as significant without stronger data.","The method can be applied to other flux-limited, multi-observable astrophysical populations where the likelihood is intractable but forward simulation is possible.","Calibration diagnostics (SBC, TARP, MMD, L-C2ST) should accompany neural posterior estimates to guard against over-confidence.","The $k$–$\\eta$ degeneracy shows that redshift evolution is not independently constrained by this sample; external redshift-evolution measurements would be needed to break it."],"supporting_citations":[{"why":"Supplies the generative parent-population model and the MOJAVE 1.5 Jy FSRQ sample used for inference.","marker":"Lister et al. 2019"},{"why":"Provides the XRB Lorentz-factor distribution posterior that the AGN slope is compared against.","marker":"Lilje et al. 2025"},{"why":"Frames simulation-based inference and justifies the likelihood-free approach for intractable likelihoods.","marker":"Cranmer et al. 2020"},{"why":"Introduces neural spline flows, the density estimator used as the inference backend.","marker":"Durkan et al. 2019"},{"why":"Defines the automatic posterior transformation objective used to train the flow.","marker":"Greenberg et al. 2019"},{"why":"Supplies the MMD misspecification test whose p-value is used to argue the simulator covers the observed data.","marker":"Schmitt et al. 2021"},{"why":"Supplies the local classifier two-sample test used to validate the estimated posterior.","marker":"Linhart et al. 2023"},{"why":"Provides machine-readable MOJAVE data with apparent speeds used as observables.","marker":"Homan et al. 2021"}],"fun_headline_variants":["Jet-speed slope b = -1.32 unites AGN and XRB outflows","Simulation-based inference tightens black hole jet comparisons","AGN jet speeds: b = -1.32, now consistent with XRBs","Likelihood-free modeling reveals unbiased jet-speed distribution","Non-Gaussian posteriors resolve AGN vs XRB jet tension"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the Lister et al. (2019) simulator—power-law Lorentz factors on $[1.25,50]$, isotropic inclination, pure luminosity evolution $(1+z)^k$, and selection purely by the $1.5\\,$Jy flux cut—is the true parent population; if any of those forms or cut-offs are wrong, the calibrated posterior can be precise but biased.","fun_headline_variants_meta":{"raw":{"variants":["Jet-speed slope b = -1.32 unites AGN and XRB outflows","Simulation-based inference tightens black hole jet comparisons","AGN jet speeds: b = -1.32, now consistent with XRBs","Likelihood-free modeling reveals unbiased jet-speed distribution","Non-Gaussian posteriors resolve AGN vs XRB jet tension"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000626,"raw_usage":{"total_tokens":2935,"prompt_tokens":1023,"completion_tokens":1912,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":639,"completion_tokens_details":{"reasoning_tokens":1818}},"tokens_in":639,"tokens_out":1912,"duration_ms":12475,"temperature":1.0,"reasoning_tokens":1818,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:26:06.606666+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the same SBI pipeline on the same MOJAVE data with the simulator's $\\Gamma_{\\min}$ changed from 1.25 to 2 and with a broken power-law; if the posterior for $b$ shifts by more than the quoted uncertainty, the central claim depends on an unverified cutoff.","supporting_citations":[],"review_version":1}