{"id":"e853d136-c190-41be-b049-f5413e8102da","arxiv_id":"1908.10114","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper provides weak lensing masses for 39 APEX-SZ clusters and a three-colour background selection method that reduces cosmic-variance scatter in the mean lensing depth by 20-40%.","lead":"This paper measures weak lensing masses for 39 galaxy clusters in the APEX-SZ survey using three-colour imaging to select background galaxies. It quantifies how much cosmic variance in source redshift estimates affects the masses and reports a 20-40% reduction in that scatter.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection identified; independent test of the 20-40% scatter-reduction claim would still strengthen the paper.","rationale":"The reader's verdict is CONDITIONAL with MODERATE confidence, and the identified weakest assumption is the representativeness of the COSMOS photo-z catalogue for the redshift distribution in every cluster field. My stress-test converges on a related but narrower point: the quantitative 20-40% reduction claim is measured with the same catalogue used for calibration and validation. This is a genuine soft spot, but it is (a) explicitly acknowledged in the text, (b) limited to the headline methodological claim, and (c) not damaging to the mass measurements, which are cross-checked against three independent surveys and show reasonable agreement. The paper includes multiple internal consistency checks: stacked shear profiles of low-β vs high-β subsamples show no residual contamination; density profiles show depletion rather than excess toward the centre; concentration ratios are consistent with the adopted c-M relation; masses with and without the concentration prior agree. These checks do not directly validate the scatter reduction, but they do support the robustness of the mass scale that is the main product of the paper. The honest assessment is that the central methodological claim would benefit from an independent cross-validation, but the paper's own caveats and the favourable comparison with literature masses mean the overall verdict of CONDITIONAL need not change. I therefore recommend UNCHANGED, with the concrete leave-one-field-out test as the natural way to settle the residual concern.","tokens_in":42599,"tokens_out":1629,"duration_ms":15605,"concrete_test":"Recompute smeas/scos for the nine COSMOS subfields using a reference catalogue built from the other eight subfields only (leave-one-field-out), so that the βg cylinder weights for each test field never use that field's own galaxies. If the 20-40% reduction persists under this cross-validation, the circularity concern is settled. As a second check, repeat the measurement on CFHTLS Deep fields by constructing the βg reference from COSMOS alone and applying it to CFHTLS fields, comparing ⟨βmeas⟩/⟨βtrue⟩ field-to-field scatter against the COSMOS-internal value.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central methodological claim is that the multi-colour βg estimator reduces cosmic-variance-induced scatter in the mean lensing depth by 20-40%. The estimate is based on nine 30′×30′ subfields of COSMOS, where the same COSMOS photo-z catalogue supplies both the reference redshifts used in the βg cylinder weighting (Eq. 10) and the 'true' redshifts defining ⟨βtrue⟩. Any residual correlation between the weighting reference and the test fields, or an under-estimation of subfield covariance (which the authors themselves note, Sec. 5.2: 'given the size of the COSMOS field, we expect some correlation between subfields potentially causing an underestimation of the scatter'), could bias the measured 20-40% reduction. The CFHTLS check (Sec. 5.2.1) is a separate, four-field consistency test with only a 25-30% quoted reduction scaling, not a direct independent validation of the same quantity. The mass measurements themselves are benchmarked against CCCP, LoCuSS, and WtG and appear robust; the load-bearing concern is isolated to the headline scatter-reduction claim, whose validation is circular in the sense that the reference catalogue and the test catalogue are the same data product. The paper's own statements flag the correlation caveat, so this is a limitation acknowledged in the text rather than an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a weak lensing analysis of 39 galaxy clusters from the APEX-SZ survey, using deep three-band optical imaging from WFI and Suprime-Cam. The central methodological novelty is a background-selection scheme in which each source galaxy receives an individual estimate of the angular diameter distance ratio βg, computed by averaging COSMOS photo-z galaxies in a colour-colour-magnitude cylinder around the source (Eq. 10). The authors quantify the variance in the mean lensing depth caused by cosmic variance using nine 30x30 arcmin COSMOS subfields, and report that their selection reduces the field-to-field scatter in ⟨β⟩ by 20-40%, depending on cluster redshift and imaging depth. They derive NFW-based cluster masses with and without a mass-concentration prior, check for residual contamination with shear and density profiles, and compare their masses against LoCuSS, CCCP and WtG. The paper also investigates photometric calibration effects, shear calibration bias, selection-induced bias, and profile-fitting systematics, producing a systematic error budget for the mass scale.","tokens_in":42825,"tokens_out":2947,"duration_ms":32934,"significance":"If the central claim holds, the multi-colour βg estimator suppresses one of the principal systematics in cluster weak lensing mass calibration, and the 39 cluster masses in Table 2 provide a useful anchor for SZ and X-ray scaling relations. The paper's strengths include a careful contamination analysis (individual and stacked shear profiles, galaxy density profiles), consistency checks of recovered concentrations against the Bhattacharya et al. (2013) c-M relation, and comparisons with three independent weak lensing studies showing agreement at the 10-20% level. The paper is also transparent about many systematic uncertainties, presenting a detailed error budget in Sec. 5.7. However, the headline scatter-reduction claim is validated only with a test that uses the COSMOS catalogue both as the reference for the βg estimator and as the 'truth' for the subfields, so the published 20-40% reduction requires stronger independent support before it can be considered established.","major_comments":[{"comment":"The central claim that the background selection reduces cosmic-variance scatter by 20-40% is based on a test that is partly circular: the βg estimator in Eq. (10) weights COSMOS photo-z galaxies in a colour-colour cylinder around each source, and the same COSMOS catalogue is then used to define the 'true' mean lensing depth ⟨βtrue⟩i in each of the nine subfields. Because the reference catalogue contains the subfields themselves, any shared large-scale structure or common calibration error will reduce the measured scatter s_meas relative to the cosmic-variance scatter s_cos, biasing the ratio s_meas/s_cos low. The authors acknowledge that the subfields are correlated (Sec. 5.2: 'we expect some correlation between subfields potentially causing an underestimation of the scatter'), but the shared-reference issue is more fundamental and is not quantified. To support the abstract's quantitative claim, the paper should either provide a test where the reference catalogue and the 'truth' are independent, or estimate the magnitude of the circularity bias.","section":"Sec. 5.2, Eq. (10)"},{"comment":"The CFHTLS deep-field check does not provide an independent measurement of the 20-40% reduction. It measures the scatter in mean β among four 1 deg^2 fields using a simple photo-z cut, and then obtains a '25-30% reduction' by combining a naive √2 area scaling with the assumption that the COSMOS-based reduction applies to CFHTLS. This is an extrapolation, not a direct validation of the βg estimator's performance on independent data. Either the estimator should be run on CFHTLS subfields with an independent reference (e.g., using CFHTLS as both reference and truth in a way that avoids the circularity), or the text should clearly label Sec. 5.2.1 as an order-of-magnitude consistency check rather than a validation.","section":"Sec. 5.2.1"},{"comment":"The random realization test described in Sec. 5.2 ('we mimic repeated observations of the same subfields... randomly varied the colours of the individual sources in the photo-z catalogues within their photometric errors') addresses photometric noise but not cosmic variance or the shared-reference bias. It shows that the remaining 60-80% scatter is not driven by photometric scatter, but it does not test whether the subfields are representative of independent cluster fields. The conclusion that 'the remaining 60-80% scatter is driven by variance of the redshift distributions' conflates true cosmic variance with the correlated variation captured by the same catalogue. The authors should clarify this distinction or add an external test.","section":"Sec. 5.2, random realization test"}],"minor_comments":[{"comment":"The abstract contains the typo 'manitude limits' and 'Sunyaev-Zel\\textquotesingle dovich'; these should be corrected to 'magnitude' and 'Sunyaev-Zel\\'dovich'.","section":"Abstract and Sec. 1"},{"comment":"There are several typos in this section, including 'Repetititon test' and 'relativie zero points'; these should be fixed.","section":"Sec. 2.1"},{"comment":"The text uses 'lesning depth' (likely 'lensing depth') and 'string function' (likely 'strong function'); the axes in Fig. 10 are labelled with LaTeX macros such as '/u1D703' and '/u1D737' that are not rendered in the compiled PDF, making the figure difficult to interpret.","section":"Sec. 5.2"},{"comment":"The phrase 'A907 strikes out from the distribution' should probably read 'A907 stands out from the distribution' or 'is an outlier'.","section":"Sec. 6.1.1"},{"comment":"The captions for Fig. 7 and Fig. 8 contain garbled text with repeated cluster names and inline math that is not typeset; they need to be rewritten clearly.","section":"Fig. 7 and Fig. 8 captions"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically solid in its mass measurements and systematic-error treatment, but the novelty claim in the abstract depends on a test that is partly circular. The authors should be asked to either add an independent validation of the scatter-reduction factor or significantly temper the claim. This is not a rejection-level flaw, as the masses themselves appear robust and useful, but the current wording overstates the strength of the evidence for the central methodological result."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a careful, useful mass-calibration paper. The actual new thing is not the three-colour beta estimator per se—Gruen et al. and Rehmann et al. did similar things—but applying it to 39 APEX-SZ clusters and showing, in internal tests, that it lowers cosmic-variance-induced scatter in the mean lensing depth by 20–40%. The cluster masses themselves look solid: they agree with LoCuSS, CCCP, and WtG at roughly the 10–20% level, the default masses with a c–M prior are consistent with free-concentration fits, and the systematic budget (shear calibration, selection bias, profile fitting, zero-point scatter) is unusually complete. That is real value.\n\nThe soft spot is exactly the claim in the abstract. The scatter-reduction test uses nine 30x30 arcmin subfields of COSMOS, with the same COSMOS photo-z catalogue supplying both the reference redshifts in the beta weighting and the \"true\" redshifts for the comparison. That is partly circular, and the authors explicitly note that subfields may be correlated, so their measured scatter is probably a lower limit. The CFHTLS check is a useful sanity check but does not validate the 20–40% number directly. None of this undermines the mass scale—the cross-checks with independent surveys carry that—but the headline methodological claim is not yet nailed down.\n\nI also note that raw catalogues and code were not released, so the numbers were not independently reproduced. That is a minor but real limitation for a methods-heavy paper.\n\nBottom line: this paper deserves a serious referee and will be cited for the masses. I would ask the authors to either get an independent field or simulation for the scatter-reduction claim, or to rewrite the abstract to make the internal-test status clear. It is not a desk reject.","headline":"A careful, useful cluster weak lensing mass catalogue; the headline 20–40% scatter-reduction claim is plausible but only partially validated.","tokens_in":43459,"tokens_out":2113,"would_cite":true,"duration_ms":23218,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A per-galaxy colour-based background selection suppresses cosmic-variance scatter in weak-lensing cluster masses by 20–40%.","keywords":["weak gravitational lensing","galaxy clusters","cosmic variance","photometric redshifts","background selection","Sunyaev-Zel'dovich effect","cluster mass calibration","APEX-SZ"],"falsifier":"Replace COSMOS with an independent, deeper photo-z reference catalogue (e.g., CFHTLS or a dedicated survey of the same fields) and recompute the $\\beta_g$ background selection for the same clusters; if the field-to-field scatter in $\\langle\\beta\\rangle$ is not reduced by 20–40%, or if the mean lensing depth shifts by more than the claimed ~0.5% (about 1.4% in mass), the suppression claim is falsified.","tokens_in":42340,"feed_emoji":"🔭","tokens_out":10243,"duration_ms":85882,"temperature":0.7,"pith_summary":"This paper presents weak-lensing masses for 39 galaxy clusters from the APEX-SZ survey and introduces a background-selection method that assigns each source galaxy its own angular diameter distance ratio $\\beta_g = D_{ds}/D_s$ instead of a single average for the whole field. The central claim is that this per-galaxy colour-based distance assignment reduces the field-to-field scatter in the mean distance ratio caused by cosmic variance by 20–40%, compared with the standard source-sheet approximation. The scatter is about 6% for a cluster at $z = 0.45$ with shallow imaging ($R \\approx 23$), falling to about 1% for deep imaging ($R = 26$), which translates to 8.4% and 1.4% scatter in $M_{200}$. The resulting masses, including an X-ray-selected subsample of 27 clusters, are intended for calibrating SZ and X-ray mass scaling relations at fixed cosmology.","feed_headline":"Colour cuts cosmic-variance scatter on cluster masses by up to 40%","feed_subtitle":"One distance ratio per galaxy, matched to COSMOS colours, shrinks a top systematic in weak-lensing masses.","key_machinery":"The central object is the per-galaxy angular diameter distance ratio $\\beta_g = D_{ds}/D_s$, computed by Eq. (10) as a weighted mean over COSMOS photo-z galaxies inside an elliptical cylinder in colour-colour-magnitude space: $\\beta_g = \\left(\\sum_k w_k \\beta(z_d, z_k)\\right)/\\left(\\sum_k w_k\\right)$, with a two-dimensional Gaussian weight in the colour plane centred on each source. The cylinder replaces the source-sheet approximation in which one average $\\langle\\beta\\rangle$ is assigned to all galaxies, allowing correlated redshift-distribution changes across a cluster field—cosmic variance, magnification bias, and member contamination—to partly average out. The stellar-locus regression that calibrates the observed colours to the COSMOS photometric system is the enabling step that makes the cylinder match meaningful.","core_discovery":"The paper establishes that a three-band background selection based on an elliptical colour-colour-magnitude cylinder around each galaxy (Eq. 10) estimates $\\beta_g$ as a weighted mean over COSMOS photo-z galaxies, recovering the true mean lensing depth $\\langle\\beta\\rangle$ of a cluster field within 0.5% on average. Using nine $30' \\times 30'$ COSMOS subfields for clusters at $z = 0.18$, $z = 0.275$, and $z = 0.45$, the cosmic-variance-induced scatter in $\\langle\\beta\\rangle$ between fields is measured to be about 6% for a cluster at $z = 0.45$ with shallow $R \\approx 23$ imaging, falling to about 1% at $R = 26$, corresponding to 8.4% and 1.4% scatter in $M_{200}$. The per-galaxy $\\beta_g$ selection reduces this scatter by 20–40% relative to using the reference-field mean, and randomized photometry realizations show the remaining scatter is driven by real redshift-distribution variance rather than photometric noise. Combining the selection with a mass-concentration prior from the adopted $c$–$M$ relation yields $M_{200}$ estimates for the 39 APEX-SZ clusters that agree with CCCP and LoCuSS measurements, are lower than Weighing-the-Giants masses, and give concentrations consistent with the adopted $c$–$M$ relation.","pith_inferences":["I infer the method should transfer to other three-band surveys (e.g., wide-field ground-based imaging) as long as a deep, complete photo-z reference catalogue is available; the COSMOS-specific calibration is not essential.","A testable corollary: using an independent reference catalogue (e.g., CFHTLS or a future deep survey) to compute $\\beta_g$ for the same cluster fields should reproduce the 20–40% reduction in field-to-field scatter; if it does not, the suppression is a COSMOS artefact.","Because all masses scale roughly linearly with $\\langle\\beta\\rangle$, the measured ~0.013 mag colour-calibration scatter translates to ~1.3% changes in $\\beta$ and ~2% in mass, so surveys adopting the method should budget colour zero-point calibration as a dominant systematic."],"forward_implications":["The 39 masses in Table 2, including the X-ray-selected subsample of 27 clusters, can be used to calibrate SZ and X-ray mass scaling relations at fixed cosmology.","For deep imaging ($R \\approx 26$), cosmic-variance noise in the mean distance ratio drops to about 1%, so per-cluster mass errors become dominated by shape noise rather than redshift-distribution uncertainty.","Masses derived with and without the mass–concentration prior agree, and recovered concentrations are consistent with the adopted $c$–$M$ relation, supporting the use of such a prior in scaling-relation work.","The 20–40% scatter reduction is a systematic-floor improvement: applying the same $\\beta_g$ selection to larger cluster samples should shrink the cosmic-variance contribution to the mass-scale calibration error."],"supporting_citations":[{"why":"Supplies the COSMOS photo-z catalogue used as the reference redshift distribution for every $\\beta_g$ estimate.","marker":"Ilbert et al. 2009"},{"why":"Provides the CFHTLS deep-field photo-z catalogues used to estimate cosmic-variance scatter in mean lensing depth on reference-field scales.","marker":"Coupon et al. 2009"},{"why":"Gives the earlier ~3% estimate of cosmic variance in a COSMOS-sized field that the paper validates and refines.","marker":"van Waerbeke et al. 2006"},{"why":"Mass-concentration relation used as the prior in the default mass estimates.","marker":"Bhattacharya et al. 2013"},{"why":"Defines the shear measurement pipeline, PSF correction, and shear calibration factor $f_0 = 1.08$ adopted here.","marker":"Israel et al. 2010"},{"why":"Provides the template for systematic-uncertainty accounting and the 0.75–3 Mpc fitting-range test the paper repeats.","marker":"Applegate et al. 2014"},{"why":"Stellar locus regression method used to calibrate the observed colours to the COSMOS photometric system.","marker":"High et al. 2009"}],"fun_headline_variants":["Colour selection cuts weak-lensing mass scatter by up to 40%","Cosmic-variance scatter in cluster masses cut 40% by colour selection","Multi-colour background selection trims lensing mass scatter by 40%","COSMOS colour calibration reduces cluster mass scatter by 40%","Three-band selection shrinks m200 scatter up to 40% in lensing analysis"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the COSMOS photo-z catalogue, after stellar-locus colour calibration, gives the correct redshift distribution for every region of colour-colour-magnitude space in every observed cluster field; if it does not, every per-galaxy distance ratio $\\beta_g$ and hence every cluster mass is biased in the same direction, and the claimed 20–40% scatter reduction would not hold.","fun_headline_variants_meta":{"raw":{"variants":["Colour selection cuts weak-lensing mass scatter by up to 40%","Cosmic-variance scatter in cluster masses cut 40% by colour selection","Multi-colour background selection trims lensing mass scatter by 40%","COSMOS colour calibration reduces cluster mass scatter by 40%","Three-band selection shrinks m200 scatter up to 40% in lensing analysis"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00144,"raw_usage":{"total_tokens":5920,"prompt_tokens":1178,"completion_tokens":4742,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":794,"completion_tokens_details":{"reasoning_tokens":4642}},"tokens_in":794,"tokens_out":4742,"duration_ms":37151,"temperature":1.0,"reasoning_tokens":4642,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:52:53.520000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Replace COSMOS with an independent, deeper photo-z reference catalogue (e.g., CFHTLS or a dedicated survey of the same fields) and recompute the $\\beta_g$ background selection for the same clusters; if the field-to-field scatter in $\\langle\\beta\\rangle$ is not reduced by 20–40%, or if the mean lensing depth shifts by more than the claimed ~0.5% (about 1.4% in mass), the suppression claim is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the earlier ~3% estimate of cosmic variance in a COSMOS-sized field that the paper validates and refines."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the shear measurement pipeline, PSF correction, and shear calibration factor $f_0 = 1.08$ adopted here."}],"review_version":1}