{"id":"5e32bfce-4bc4-4cdd-b219-7e5290d6e8d6","arxiv_id":"2602.07114","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":18,"one_line_summary":"The splashback radius of AMICO KiDS-1000 clusters, measured with weak lensing and cluster-galaxy clustering, agrees with ΛCDM predictions, with clustering achieving 10% precision.","lead":"Galaxy clusters are surrounded by a boundary called the splashback radius, where infalling matter piles up after its first orbit. This paper measures that edge for about 9,000 clusters in the KiDS survey using two independent probes—the distortion of background galaxy shapes and the clustering of galaxies around clusters—and finds both agree with cosmological simulations, with the clustering probe reaching 10% precision.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"r_sp is derived from a DK14 profile whose steepening parameters are prior-dominated; the paper itself defers simulation validation, so the claimed 10–14% constraints are not yet established.","rationale":"The reader's weakest assumption—that the DK14 profile accurately describes the true stacked mass distribution and extrapolates to the scales needed for r_sp—is exactly where the paper is most vulnerable. The manuscript contains no simulation-based validation of the inference chain, and its own discussion explicitly defers such tests to future work. This is the single most load-bearing concern because every headline result (r_sp, Γ, R_sp–ν relation) is a derived quantity from the fitted profile; if the model shape or extrapolation is wrong, the inferred splashback parameters are biased regardless of statistical precision. The prior-dominated nature of the transition parameters (F_t, β, γ0) amplifies this risk: the data do not independently determine the feature they are meant to measure. I do not see this as a fatal flaw—the analysis is careful and the authors are appropriately cautious—but it justifies the CONDITIONAL verdict. A mock-recovery test would either confirm the pipeline or expose the bias, so the verdict need not change now; it should remain conditional pending that test.","tokens_in":32543,"tokens_out":6954,"duration_ms":80566,"concrete_test":"Run an end-to-end mock recovery: select N-body haloes with known r_sp, assign AMICO-like richness with the assumed log-normal scatter, apply KiDS-1000 footprint, photo-z errors, and the same λ*/z binning; generate stacked g_t and w_cg with the same radial bins and covariance construction; fit with the identical pipeline (Eqs. 19 and 28, same priors). Compare recovered r_sp per stack and the A_sp–B_sp relation to the true values. If the recovery is biased by more than ~1σ of the quoted 10–14% error in any stack, the central claim fails; if unbiased, the concern is retired.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim equates the fitted DK14 profile's minimum-log-slope radius (Eq. 36) with the true splashback radius. The data only cover R∈[0.4,5] h−1 Mpc, and the steepening that sets r_sp is governed by transition parameters (F_t, β, γ0) whose posteriors remain close to the simulation-calibrated Gaussian priors (Table 2, Fig. C.5), with γ0 effectively unconstrained. No end-to-end mock validation is presented. The manuscript itself concedes (Sect. 6) that the DK14 model is extrapolated to very large scales for both probes and that the impact of this extrapolation 'shall be assessed through simulations'; it also concedes that w_cg may be biased by anisotropic projection/selection effects. Thus the quoted 10–14% precision is largely inherited from the model and priors, not independently demonstrated by the data. A systematic failure of DK14 for the AMICO-selected population would shift both probes' r_sp coherently and could mimic or destroy the ΛCDM agreement. The g_t–w_cg consistency is also weakened by the shared g_t mass-richness priors used in the w_cg fit (Sect. 4.6).","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes stacked weak-lensing reduced shear g_t and cluster-galaxy correlation function w_cg for 8730 AMICO clusters in KiDS-1000, binned in richness and redshift. Using a Diemer & Kravtsov (2014) profile with a transition factor and an outer infall term, the authors model the stacked observables and define the splashback radius as the minimum-log-slope radius of the 3D density profile. They report constraints on r_sp, the mass accretion rate Gamma, and the relation between R_sp = r_sp/r_200m and peak height nu_200m, claiming that the two probes are mutually consistent and agree with LambdaCDM simulation predictions, with per-stack precision of 14% (g_t) and 10% (w_cg).","tokens_in":33044,"tokens_out":3551,"duration_ms":37751,"significance":"If the model-derived r_sp indeed recovers the true splashback radius of the underlying halo population, this is a competitive and observationally valuable measurement, leveraging a large optical cluster sample and combining two complementary probes. The analysis is thorough in its treatment of known systematics: bootstrap and jackknife covariance matrices, propagation of shear-calibration, SOM redshift-distribution, miscentering, and mass-richness relation uncertainties, plus robustness tests against cosmology and model choices in Appendix C. The explicit comparison with the L25 mass calibration is a strength. However, the central claim depends on the DK14 profile shape and on priors calibrated with the same simulations used for the theoretical comparison; the paper's own appendices and discussion concede that several key parameters are prior-dominated and that end-to-end simulation validation is deferred. The significance of the quoted precision therefore remains conditional.","major_comments":[{"comment":"r_sp is not directly observed; it is the minimum-log-slope radius of the DK14 profile. The steepening that sets this radius is controlled by F_t, beta, and gamma_0, which are assigned Gaussian priors from DK14 (Sect. 4.6). Table 2 and Fig. C.5 show that the posteriors of these parameters remain close to the priors, with gamma_0 effectively unconstrained by either probe. Since the data only cover R in [0.4,5] h^-1 Mpc and the model is integrated to 40 h^-1 Mpc in Eq. (13), the claimed 10–14% precision per stack is substantially inherited from the simulation-calibrated profile shape rather than demonstrated by the data. The manuscript itself states in Sect. 6 that the impact of the DK14 extrapolation to very large scales 'shall be assessed through simulations.' An end-to-end mock validation—injecting clusters with known r_sp and verifying posterior recovery—is required to support the preci","section":"§4.1–4.5, Eq. (36), Table 2, Fig. C.5"},{"comment":"The mass accretion rate Gamma is obtained by inserting the model-derived R_sp into the More et al. (2015) fitting formula, and the comparison model of Diemer (2020) is calibrated on the same simulation suite. The text explicitly says the agreement 'is expected' for this reason. Therefore the Gamma constraints reported in Table 1 and Fig. 3 are not an independent test of LambdaCDM; they are a consistency check that is partly circular. The abstract and results should state this limitation, or the analysis should derive Gamma through an independent route before claiming a constraint.","section":"§5, Eq. (45)"},{"comment":"The w_cg analysis uses the g_t posteriors on the log lambda*–log M_200m relation (A, B, C, sigma_intr) as priors, and the text assumes M_200m = M_gt = M_wcg. The two probes are therefore not independent: a systematic error in the mass-richness calibration, or in the lensing masses, would shift both r_sp estimates coherently. The 'consistent results' claim in the abstract is weakened by this shared calibration. The authors should either run w_cg with uninformative mass-richness priors as a robustness check, or quantify the correlation between the two probe results.","section":"§4.6"},{"comment":"The paper acknowledges that anisotropic projection and selection effects may bias the w_cg measurements and that the impact on r_sp 'will be tested' with future dedicated mocks. Since w_cg provides the tighter constraint (10%), the central w_cg-based result—including the R_sp–nu relation and the possible dynamical-friction offset—rests on an unquantified systematic. A first-order assessment using the existing L25 anisotropic-boost model, or a simple test with mock galaxy catalogues, should be included before asserting that the w_cg constraints are unbiased.","section":"§6"}],"minor_comments":[{"comment":"Typo: 'statical part of the covariance' should read 'statistical part.'","section":"§4.6, Eq. (40)"},{"comment":"In the Planck18 robustness paragraph, there is a duplicated phrase: 'than the one assumed in our baseline analysis, analysis, namely Omega_m = 0.22.'","section":"Appendix C"},{"comment":"The choice R_max = 40 h^-1 Mpc for the surface-density integration is not justified in the text. Given that the fitting range is [0.4,5] h^-1 Mpc and that Section 6 flags the large-scale extrapolation as a concern, a sentence explaining why 40 h^-1 Mpc is sufficient (or a convergence test) would help.","section":"§4.2, Eq. (13)"},{"comment":"The caption states that error bars include 'residual uncertainties coming from systematic errors,' but the text (Sect. 4.6) models these as an additive covariance term rather than as error bars. Consider aligning the caption phrasing with the covariance treatment.","section":"Fig. 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The work is technically careful and the dataset is well suited to this measurement, but the central claim of 10–14% precision on r_sp is not yet established because the estimator is model-derived and its key shape parameters are prior-dominated, with no end-to-end mock validation. The g_t–w_cg consistency is also weakened by shared mass-richness priors. These are fixable with additional tests, but they are load-bearing for the abstract and conclusions; hence major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a solid, detailed measurement paper on the AMICO KiDS-1000 sample. The new bits are the larger sample, the wcg analysis, and the two-probe comparison, and the headline claim — that both probes give splashback radii consistent with each other and with LCDM — mostly holds up. But the 10–14% precision is model-limited, not purely data-limited, and the paper itself defers the simulation tests that would establish that. Treat the precision as conditional.\n\nWhat's done well: the analysis is methodical. Covariances from bootstrap and jackknife, systematic terms for shear bias, photo-z, and projection effects; careful treatment of miscentering; a consistent mass-richness model calibrated on the same data. The wcg measurement is genuinely new for this sample and does constrain the outer profile parameters (be, se) that lensing cannot. The mass constraints agree with Lesci et al. 2025, and the rsp values broadly line up with earlier work.\n\nWhere I'd be cautious: rsp is a derived quantity. The data used are stacked profiles over 0.4–5 Mpc/h, and the transition parameters (Ft, beta, gamma0) that set the location of the minimum log-slope are prior-dominated — Table 2 and Fig. C.5 show the posteriors sit close to the simulation-calibrated priors, with gamma0 effectively unconstrained. The quoted 14% / 10% per-stack precision therefore contains a substantial contribution from the priors. The paper acknowledges the model is extrapolated to large radii and says the impact 'shall be assessed through simulations'; that is future work, not part of this result. A systematic failure of the DK14 shape would shift both probes coherently, so the gt–wcg consistency is not a strong validation. On top of that, wcg may be affected by anisotropic selection/projection effects, and the authors assume a free galaxy bias absorbs it — an assumption they say they will test later. The wcg fit also borrows the mass-richness posteriors from gt, so the two-probe agreement is not fully independent. The Gamma constraints are obtained by plugging Rsp into Eq. (45), a relation calibrated on the same simulations used for the comparison; the agreement there is unsurprising.\n\nNone of this makes the paper wrong. It is an honest, careful analysis within a standard framework. The issue is the framing: the abstract presents 10–14% as the measurement precision, but it is the precision of a model that has not yet been tested on mocks. A serious referee should ask for end-to-end mock validation, or at least a prior-dependence test, before these numbers are used as external constraints.\n\nThis paper is worth a serious referee and will likely be useful for cluster cosmology and Euclid planning. I would cite it as the current KiDS-1000 AMICO splashback measurement. For a reading group it's a decent case study in model-dependent inference.","headline":"Solid two-probe splashback measurement on AMICO/KiDS-1000, but the quoted 10–14% precision is partly prior-driven and the model is not yet mock-tested.","tokens_in":33555,"tokens_out":4587,"would_cite":true,"duration_ms":43518,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The splashback radius of galaxy clusters, measured two independent ways, agrees with cold-dark-matter predictions at 10–14% precision.","keywords":["galaxy clusters","splashback radius","weak gravitational lensing","cluster-galaxy correlation","mass accretion rate","dark matter haloes","large-scale structure","cosmology: observations"],"falsifier":"Measure the stacked density profile of the same clusters non-parametrically (e.g., by Abel-inverting the lensing signal with no assumed shape) and compare the radius of steepest slope with the profile-fitted rsp; if they disagree by more than the quoted 10–14% uncertainty, the result is model-dependent. Alternatively, apply the same pipeline to mock observations with known true splashback radii and check recovery within 1σ.","tokens_in":32490,"feed_emoji":"🔭","tokens_out":7386,"duration_ms":68093,"temperature":0.7,"pith_summary":"This paper tries to establish that the splashback radius — the boundary where matter has just completed its first orbit around a galaxy cluster — can be recovered from stacked observations of thousands of clusters, and that two independent probes give the same answer. The authors model the weak-lensing shear signal and the cluster-galaxy correlation function with a common density profile, and show both return the splashback radius, the mass accretion rate, and the relation between the normalised splashback radius and cluster peak height. If correct, the result means the dynamics of the infall region in real clusters, traced by both matter and galaxies, follows the standard cosmological prediction, with per-stack precision of 14% from lensing and 10% from clustering.","feed_headline":"Cluster splashback radius pinned to 10% precision","feed_subtitle":"Weak lensing and galaxy clustering agree on the cluster edge, matching dark-matter-only simulations.","key_machinery":"A single-cusp density profile with a transition term and an outer power law (Eq. 8), evaluated as an ensemble average over richness and redshift bins. The splashback radius is not measured directly but defined as the minimum of the logarithmic slope of the total density profile, and the ensemble-averaging step (Eq. 34) converts individual halo profiles into predicted observables free of radial binning effects.","core_discovery":"For 8,730 rich clusters in the redshift range 0.1–0.8, the paper models the stacked reduced shear and the projected cluster-galaxy correlation function with the same truncated Einasto-plus-outer-power-law density profile, marginalising over selection effects, photometric-redshift scatter, miscentring, and the richness–mass relation. The fitted profiles place the average splashback radius at rsp/r200m at values consistent with the ΛCDM theoretical relation as a function of peak height, and give a mass accretion rate consistent with simulation predictions; the two probes agree within 1σ, with the clustering probe yielding a slightly smaller rsp, interpreted as dynamical friction on satellite g","pith_inferences":["If the fitted-profile rsp depends on cosmology as the Planck18 test suggests (lower rsp for higher Ωm), then with the statistical power of upcoming surveys, rsp–ν200m constraints could become a competitive standalone cosmological probe.","The lensing-vs-clustering offset can be tested directly in simulations with galaxy formation: if dynamical friction is the cause, the offset should grow with satellite galaxy mass and with the magnitude gap between the brightest galaxy and its satellites.","Replacing the analytic outer power law with a matter-power-spectrum two-halo term, or measuring the profile non-parametrically, would test whether the extrapolation beyond 5 h−1 Mpc biases rsp; this is the most natural next step to check the model dependence.","A joint fit of lensing and clustering with an explicit dynamical-friction parameter, rather than separate fits, could break the degeneracy between infall profile shape and galaxy bias and yield a single, robust rsp measurement."],"forward_implications":["If true, optically selected clusters have a universal splashback boundary that tracks the ΛCDM Rsp–ν200m relation, with no residual trend in redshift or richness beyond mass.","Galaxy clustering is a sharper splashback probe than lensing (10% vs 14% precision) and also constrains the amplitude and slope of the infalling-matter profile, which lensing alone cannot.","The small but systematic offset between lensing and galaxy-traced rsp implies that galaxies trace a splashback boundary biased slightly inward by dynamical friction — a bias that will matter for any cluster-based cosmology using galaxy positions.","The inferred mass accretion rates agree with simulation expectations, supporting the use of rsp as a mass-accretion probe with cosmological sensitivity to Ωm and σ8.","The results confirm earlier X-ray, SZ, and optical measurements within 1–2σ, so the splashback radius is now measured consistently across cluster selection methods."],"fun_headline_variants":["Splashback radius pinned to 10% with lensing+clustering","Two probes agree on cluster splashback at 10% precision","Splashback radius from 8,730 clusters: 10% error","Lensing and clustering pin splashback radius to 10%","Splashback radius measured to 10% from 8,730 clusters"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The analytic density profile (Eq. 8) describes the true stacked mass distribution from 0.4 to 5 h−1 Mpc and stays valid when extrapolated outward, so that the radius of minimum logarithmic slope of the fitted profile equals the true splashback radius of the halo population.","fun_headline_variants_meta":{"raw":{"variants":["Splashback radius pinned to 10% with lensing+clustering","Two probes agree on cluster splashback at 10% precision","Splashback radius from 8,730 clusters: 10% error","Lensing and clustering pin splashback radius to 10%","Splashback radius measured to 10% from 8,730 clusters"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000451,"raw_usage":{"total_tokens":2197,"prompt_tokens":925,"completion_tokens":1272,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1177}},"tokens_in":669,"tokens_out":1272,"duration_ms":9214,"temperature":1.0,"reasoning_tokens":1177,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T03:40:20.002811+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the stacked density profile of the same clusters non-parametrically (e.g., by Abel-inverting the lensing signal with no assumed shape) and compare the radius of steepest slope with the profile-fitted rsp; if they disagree by more than the quoted 10–14% uncertainty, the result is model-dependent. Alternatively, apply the same pipeline to mock observations with known true splashback radii and check recovery within 1σ.","supporting_citations":[],"review_version":1}