{"id":"338d558d-a529-414d-824c-9e26c2d76ec7","arxiv_id":"2411.12120","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"About 25% of HSC-ACT matched clusters are miscentered by more than 330 kpc; after removing clusters with clear non-astrophysical causes, the miscentered fraction falls to about 10%.","lead":"Using 186 galaxy clusters detected by both the HSC optical survey and the ACT microwave survey, the authors measure how often the optical center sits far from the gas-based center, finding about 25% are miscentered by more than 330 kpc. After removing clusters whose miscentering comes from data problems, the fraction drops to about 10%, and the paper argues that microwave (SZ) centers trace the true cluster potential better than optical centers.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline ~10% cleaned miscentered fraction depends on unblinded, non-quantitative visual labels for 22 of 46 clusters; without an inter-rater or quantitative reclassification, the reduction from 25% is not reproducible.","rationale":"The reader's weakest-assumption analysis identifies the subjective visual classification of the 46 miscentered clusters as the key fragility, and I agree. The paper's most novel and most cited claim is the cleaned miscentered fraction of roughly 10%, and that number is produced by removing clusters whose labels come from non-quantitative visual inspection. The raw 25% measurement is on much firmer ground: it comes from a standard two-component model, uses a plausible fixed sigma1 based on SZ positional uncertainty, and is consistent with a broad prior literature. The lensing signal comparison in Fig. 13 provides genuine independent support that the offset-based split separates populations with different centering quality, though it does not validate the specific astrophysical versus non-astrophysical cause labels. I do not see a more load-bearing concern: the sample-size limitations, missing parameter uncertainties, and arbitrary cutoff are real but secondary; the SZ-center claim is supported by the lensing comparison, albeit with modest significance for the miscentered population (p=0.0276). The cleanest path to resolution is a blinded reclassification exercise. If that exercise reproduces the labels, the conditional concerns are largely resolved; if not, the headline cleaned fraction is unsupported. Since the reader has already conditioned acceptance on exactly this issue, my read does not change the existing verdict.","tokens_in":31838,"tokens_out":4686,"duration_ms":55121,"concrete_test":"Have at least two independent classifiers, blinded to the offset values and to the paper's labels, reclassify the 46 miscentered clusters using a pre-registered decision tree with quantitative inputs: Gaia star-mask radius and alternative-galaxy offset for star-mask cases, HSC deblending flags and object counts for deblending cases, SZ SNR and matched-filter contours for false-signal cases, and richness/photo-z consistency for false-match cases. Compute inter-rater agreement (e.g., Cohen's kappa) and the resulting cleaned-sample fcen by refitting Eq. (1) with sigma1 = 0.15 Mpc. If the consensus labels reproduce fcen within, say, ±0.05 of 0.91, the visual classification is not load-bearing; if the cleaned fraction moves substantially away from ~10%, the headline claim requires revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The raw ~25% miscentered fraction is reasonably supported by the two-component Rayleigh fit in §3 and by consistency with prior miscentering studies, and the lensing comparison in Fig. 13 independently validates the well-centered/miscentered split. The load-bearing weakness is the cleaned-sample claim of §4.4: the reduction from ~25% to ~10% rests entirely on removing 22 of 46 miscentered clusters as having 'clear, non-astrophysical causes' (§4.2, §4.3.1, §4.3.2). These labels are assigned by visually inspecting HSC images after the offset-based miscentered sample is already defined; no pre-registered quantitative criteria, blinded review, or inter-rater agreement is reported. For example, §4.2.1 identifies alternative central galaxies by eye, §4.2.3 judges deblending failures from images, and §4.3.1/§4.3.2 classify false matches and false ACT signals from image inspection and SZ contours. If even a few of the 22 excluded clusters are actually mergers, or if some of the 14 'merger' labels are systematics, the cleaned fraction fcen = 0.91 and the derived 370 kpc cutoff would shift. Since the cleaned ~10% is the paper's main new claim, this subjectivity is the most consequential unresolved assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper cross-matches the HSC CAMIRA optical cluster catalog (S19A) with the ACT DR5 SZ cluster catalog, producing a fiducial sample of 186 clusters in the redshift range 0.1–1.4. The authors fit a two-component Rayleigh model (Eq. 1) to the distribution of optical-SZ centering offsets, obtaining a well-centered fraction fcen = 0.75 and a miscentered scale sigma2 = 0.39 Mpc, corresponding to a miscentered fraction of ~25% beyond a 330 kpc cutoff. They then visually inspect all 46 miscentered clusters and classify the causes into mergers (14), HSC systematics (17, including star masks, artifacts, deblending, and central-galaxy misidentification), false matches/false ACT signals (5), multiple causes (6), and no apparent cause (4). Removing the 22 clusters with 'clear, non-astrophysical causes' yields a cleaned sample of 164 clusters with fcen = 0.91 and a 370 kpc cutoff, i.e., a miscentered fraction of ~10%. Weak-lensing measurements of the well-centered and miscentered samples show suppressed small-scale signal for the latter, and using SZ centers rather than the CAMIRA center for the miscentered sample partially recovers the signal, leading the authors to suggest that SZ centers better trace the cluster potential centroid.","tokens_in":32133,"tokens_out":6034,"duration_ms":58186,"significance":"If the results hold, the paper provides the largest HSC-ACT cross-matched sample for miscentering studies and offers a useful decomposition of apparent miscentering into astrophysical (merger) and systematic (data/algorithm) causes. The raw ~25% miscentered fraction is consistent with earlier studies, and the lensing comparison is a valuable independent check that the well-centered/miscentered split carries physical meaning. The paper also makes its cross-match table available in full (Table A1). The central new claim, however, is the reduction to ~10% miscentered fraction after removing non-astrophysical causes; that claim rests on subjective visual classification with no quantitative criteria, blinded review, or inter-rater check, and the fitted parameters are quoted without uncertainties. These issues do not undermine the raw offset measurement, but they weaken the paper's headline conclusion as currently presented.","major_comments":[{"comment":"The cleaned-sample result (fcen = 0.91, ~10% miscentered fraction; §4.4 and abstract) is obtained by removing 22 of 46 miscentered clusters classified by eye as having 'clear, non-astrophysical causes' (§4.2, §4.3.1, §4.3.2). The classification involves selecting alternative central galaxies from images (§4.2.1), judging deblending failures (§4.2.3), and identifying false matches and false ACT signals from image inspection and SZ contours (§4.3.1, §4.3.2). No quantitative classification criteria, blinded review, or inter-rater agreement are reported, and the text itself acknowledges ambiguity in the 'multiple possible causes' and 'no apparent cause' categories (§4.3.3, §4.3.4). Furthermore, the statement in §4.4 that the cleaned model 'accurately separates' the populations is partly circular, because the same visual labels define which clusters enter the cleaned sample. I request robustness tests that reclassify the borderline clusters (e.g., moving the six 'multiple possible causes' and four 'no apparent cause' clusters into or out of the cleaned sample) and ideally a blinded or criterion-based re-classification; alternatively, the ~10% claim should be presented as conditional on the visual taxonomy rather than as a definitive physical result.","section":"§4.2–§4.4, Table 1"},{"comment":"The maximum-likelihood fit reports fcen = 0.75 and sigma2 = 0.39 Mpc with no uncertainties, and sigma1 is fixed at 0.15 Mpc. The paper's quantitative claims — the 330 kpc well-centered cutoff, the ~25% miscentered fraction in the fiducial sample, and the change to ~10% in the cleaned sample — all derive from these fitted values. Confidence intervals from the likelihood surface or bootstrap, and a sensitivity test of fcen to the assumed sigma1, are needed; without them the reader cannot judge whether the fiducial and cleaned fcen values (0.75 vs 0.91) are significantly different, nor can the consistency with previous studies be properly assessed.","section":"§3.1, Eq. (1)"},{"comment":"The lensing comparison is a valuable independent check, but the concluding claim that 'the ACT SZ centers are a better estimate of the true cluster potential centroid' rests on a chi-square difference with p = 0.0276 for the miscentered population (Fig. 14, right), which is marginal evidence. This test uses only 24 miscentered clusters in the redshift range 0.3 < z < 0.7, and the negative lowest-radius point for the miscentered population is excluded from the plotted and analyzed signal — exactly the radial range where miscentering effects are strongest. The paper should either include that bin through a re-binned or stacked analysis, present the covariance and the lowest-bin data point explicitly, or temper the conclusion to state that the SZ center is 'suggestively' better rather than definitively better.","section":"§5, Figs. 13 and 14"}],"minor_comments":[{"comment":"The 'well-centered cutoff' of 330 kpc is derived from the fitted fcen rather than from an independent observable; the text acknowledges this is 'somewhat arbitrary.' Reporting the miscentered fraction directly with its uncertainty, rather than through a cutoff-dependent definition, would make the headline number more robust.","section":"§3.2"},{"comment":"The star-mask radius formula is presented without a reference at the equation itself; consider citing Coupon et al. (2018) directly at Eq. (2) to make the source of the functional form clear.","section":"§4.2.1, Eq. (2)"},{"comment":"The histogram binning of the offset distributions is not specified; giving the bin width and the number of clusters per bin would improve reproducibility.","section":"Figs. 4 and 12"},{"comment":"The physical offset is computed using the CAMIRA photometric redshift, but the impact of photometric redshift errors on the offset distribution is not discussed; a sentence quantifying or at least acknowledging this uncertainty would be helpful.","section":"§2.2"},{"comment":"The abstract defines the miscentered fraction as 'clusters offset by more than 330 kpc,' but the cleaned sample uses a 370 kpc cutoff; the abstract should note that the threshold changes in the cleaned analysis.","section":"Abstract"},{"comment":"The estimate of ~70 false ACT signals in the HSC footprint and ~13 cross-matched false signals is useful, but the calculation is only partially specified; a brief derivation of the matching-circle coverage fraction would improve transparency.","section":"§4.3.2"}],"recommendation":"major_revision","confidential_remarks":"The raw offset fit and the lensing comparison are solid contributions and consistent with the literature; the paper is a good fit for MNRAS. The main risk is the cleaned-sample claim, which depends on a subjective visual classification that is not currently reproducible. The revision should focus on adding robustness tests for the classification, reporting uncertainties on the fitted parameters, and either strengthening or softening the SZ-center conclusion. I do not see a need to reject the paper, provided these load-bearing points are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the raw ~25% miscentered fraction and the lensing validation are on solid ground, and the per-cluster taxonomy is genuinely useful. But the headline ~10% cleaned fraction depends on the authors' visual classification of 46 images, with no blinding or inter-rater checks, so it should not be treated as established until the labels are made reproducible.\n\nThe new things here: the largest HSC-ACT cross-match for this purpose (186 clusters), a full catalog table with individual miscentering causes, and the cleaned-sample exercise that reduces the inferred miscentered fraction to ~10% by removing 22 clusters judged to have non-astrophysical causes (star masks, artifacts, deblending, false matches, false ACT signals). The two-component Rayleigh fit is standard (Oguri et al. 2018) and gives fcen = 0.75, sigma2 = 0.39 Mpc, consistent with earlier work. The lensing comparison is the best part: miscentered clusters show suppressed DeltaSigma at small radii relative to well-centered ones (p ~ 6.6e-6), and recentering on the SZ position recovers signal for the miscentered sample. That is a real independent check that the offset model split is not pure numerology.\n\nThe soft spots, in order of severity. First, the cleaned ~10% claim: the reduction from 25% to 10% rests entirely on removing 22 of 46 miscentered clusters based on eyeballing HSC images. The paper gives plausible examples, but the criteria for 'deblending failure' or 'false ACT signal' are qualitative, there is no blinded review, and no inter-rater agreement is reported. Some categories (multiple possible causes, no apparent cause) are explicitly ambiguous. If even a handful of those 22 are actually mergers or genuine large offsets, the 10% and the 370 kpc cutoff shift. This is a fixable problem—pre-registered quantitative criteria, or validation against X-ray centroids—but as it stands the headline number is not reproducible. Second, the best-fit fcen and sigma2 are quoted without uncertainties in Section 3.1; for a fitting paper that is a surprising omission. Third, the cutoff at 330 kpc is called arbitrary, and the cleaned-sample statement that the model 'accurately separates' populations is partly circular, since the same model defines the populations. The lensing check mitigates this, so it's a minor point. Fourth, the 'SZ better than BCG' conclusion rests on a single p-value of 0.028 from 24 miscentered clusters; it's suggestive, not decisive, and the paper mostly says 'suggest.'\n\nOverall: this is a careful, useful paper for cluster cosmology and BCG studies. The raw miscentering measurement and the taxonomy deserve citation; the cleaned 10% needs more work before being used as a calibration input. It should go to peer review—with the request that the authors provide uncertainties, quantitative classification rules, and at least a second labeler for the 46 clusters.","headline":"Solid raw miscentering measurement and a useful taxonomy, but the headline ~10% cleaned fraction rests on unblinded visual labels and needs quantitative criteria or external validation before it can be used.","tokens_in":32708,"tokens_out":2578,"would_cite":true,"duration_ms":26442,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Most apparent miscentering of optical galaxy cluster centers is a correctable data artifact, not a sign of cluster mergers, leaving a true miscentered fraction of about 10 percent.","keywords":["galaxy clusters","miscentering","Sunyaev-Zeldovich effect","weak gravitational lensing","HSC survey","ACT survey","CAMIRA","cluster centers"],"falsifier":"A concrete test would be to have independent reviewers re-classify the 46 miscentered clusters from the same HSC/ACT images using the paper's eight categories but with quantitative definitions (e.g., offset to centroid of the nearest Gaia star mask, deblending flag, ACT S/N), and then refit the cleaned sample: if the resulting miscentered fraction does not remain near 10%, the central claim fails. A second, independent falsifier is measuring X-ray centroids for the miscentered clusters: if the SZ center is not closer to the X-ray (potential) center than the optical center is, the claim that SZ centers estimate the true potential centroid is weakened.","tokens_in":31644,"feed_emoji":"🔭","tokens_out":13010,"duration_ms":104601,"temperature":0.7,"pith_summary":"Galaxy clusters are usually centered optically on their brightest central galaxy, but that center can lie hundreds of kiloparsecs from the cluster's true gravitational center—an effect called miscentering that biases cluster lensing and cosmology. Cross-matching 186 clusters in common between the HSC optical catalog and the ACT Sunyaev-Zeldovich catalog, this paper measures a raw miscentered fraction of about 25 percent beyond 330 kpc, consistent with earlier work. Examining each miscentered cluster by eye, the authors find that most large offsets are not caused by cluster mergers but by correctable data systematics: bright-star masks, image artifacts, deblending failures, false matches, and false detections. Removing those cases lowers the miscentered fraction to about 10 percent. The upshot for cluster cosmology is that a large part of the miscentering systematic is algorithmic rather than astrophysical, and that SZ-based centers locate the cluster potential better than optical central galaxies.","feed_headline":"Data artifacts, not mergers, drive most galaxy cluster miscentering","feed_subtitle":"A 186-cluster cross-match finds that removing correctable systematics cuts the miscentered fraction from ~25% to ~10%.","key_machinery":"The load-bearing object is the two-component offset model of Oguri et al. (2018), a mixture of two Rayleigh distributions that separates a well-centered population with characteristic offset $\\sigma_1=0.15$ Mpc (fixed by the SZ positional uncertainty) from a miscentered population with fitted $\\sigma_2=0.39$ Mpc; the fitted fraction $f_\\mathrm{cen}=0.75$ yields the miscentered fraction and defines the 330 kpc well-centered cutoff. Around this model, the paper builds a visual classification scheme that assigns each miscentered cluster to one of eight causes, and a weak lensing comparison of $\\Delta\\Sigma(R)$ measured with optical versus SZ centers that validates that the miscentered clusters are genuinely offset.","core_discovery":"On the paper's own terms, the central discovery is that the offset distribution between CAMIRA optical centers and ACT SZ centers is bimodal—about 75 percent of clusters are well-centered (offsets below 330 kpc) and about 25 percent are miscentered—but that the miscentered population is largely a product of the data and the cluster finder, not of cluster physics. After visually classifying all 46 miscentered clusters, the authors attribute 17 to systematic HSC effects (star masks, observational artifacts, deblending failures, central galaxy misidentification) and 5 to false matches or false ACT signals; only 14 are attributed to ongoing mergers, with 6 having multiple possible causes and 4 showing no apparent cause. Removing the 22 clusters with clear non-astrophysical causes raises the fitted well-centered fraction from 0.75 to 0.91, equivalent to a miscentered fraction of about 10 percent. The weak lensing comparison supports this classification: miscentered clusters show a suppressed signal within ~1 Mpc when centered on the optical galaxy, and re-centering on the SZ position recovers the small-scale signal, indicating that the SZ centroid sits closer to the true potential well.","pith_inferences":["If the cleaned ~10% fraction is reproduced in larger samples, future wide-field optical surveys could reduce miscentering corrections by flagging clusters near star masks and artifacts, but a residual astrophysical floor near 10% would remain.","A blinded, quantitative reclassification of the 46 miscentered clusters would test the paper's central step; the paper gives no such criteria, so the 25% to 10% reduction is not yet independently verified.","The paper's suggestion that SZ centers are better potential centroids could be tested against X-ray centers for the same clusters, especially for the four 'no apparent cause' cases where the offset is astrophysical but not merger-related.","The offset model fixes $\\sigma_1$ from SZ positional uncertainty; a model that also fits $\\sigma_1$ or allows a non-bimodal offset distribution could change the inferred fractions, a point the authors acknowledge near the end."],"forward_implications":["Cluster lensing and richness-mass calibration analyses that assume a 20–40 percent miscentered fraction may be overcorrecting; the true astrophysical fraction may be near 10 percent once data systematics are removed.","Optical cluster finders can be improved by flagging clusters near bright-star masks, artifacts, and deblending failures, and by assigning miscentering probabilities based on these flags rather than treating miscentering as purely astrophysical.","SZ centers, or gas-traced centers generally, are preferable for measuring small-scale cluster lensing signals and for defining cluster centroids in merger systems, where the optical center is not yet relaxed.","The residual ~10 percent miscentered fraction, including clusters with no apparent cause, likely traces genuine astrophysical processes and sets a floor on the systematic that better optical data cannot remove.","Mergers are not strongly correlated with miscentering: the merger fraction is similar for well-centered and miscentered clusters, so merger catalogs alone cannot predict which clusters are miscentered."],"supporting_citations":[{"why":"Supplies the two-component Rayleigh offset model (Equation 1), the CAMIRA cluster catalog and central-galaxy definition, and the comparison value $\\sigma_2\\simeq0.37$ Mpc.","marker":"Oguri et al. (2018)"},{"why":"Provides the ACT DR5 SZ cluster catalog and the nemo algorithm whose positional uncertainty model sets $\\sigma_1=0.15$ Mpc.","marker":"Hilton et al. (2021)"},{"why":"Provides the optical merger catalog used to compare merger fractions between well-centered and miscentered clusters.","marker":"Okabe et al. (2019)"},{"why":"Supplies the lensing measurement methodology followed for the $\\Delta\\Sigma(R)$ comparison.","marker":"Sunayama et al. (2023)"},{"why":"Describes the HSC star mask construction (Equation 2) used to identify star-mask miscentering.","marker":"Coupon et al. (2018)"},{"why":"Provides the HSC Year 3 shape catalog used for the weak lensing measurements.","marker":"Li et al. (2022)"}],"fun_headline_variants":["Systematic errors, not mergers, cause most cluster miscentering","Most galaxy cluster offsets are data artifacts, not mergers","Only 14 of 46 miscentered clusters are actually merging","Cutting systematics lowers cluster miscentering from 25% to 10%","Miscentered galaxy clusters: mostly data artifacts, not physics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline reduction from ~25% to ~10% rests on the authors' visual, unblinded classification of the 46 miscentered clusters into astrophysical versus non-astrophysical causes; if those labels are wrong, the cleaned miscentered fraction is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Systematic errors, not mergers, cause most cluster miscentering","Most galaxy cluster offsets are data artifacts, not mergers","Only 14 of 46 miscentered clusters are actually merging","Cutting systematics lowers cluster miscentering from 25% to 10%","Miscentered galaxy clusters: mostly data artifacts, not physics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000614,"raw_usage":{"total_tokens":2943,"prompt_tokens":1124,"completion_tokens":1819,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":740,"completion_tokens_details":{"reasoning_tokens":1728}},"tokens_in":740,"tokens_out":1819,"duration_ms":14546,"temperature":1.0,"reasoning_tokens":1728,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:53:11.592390+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test would be to have independent reviewers re-classify the 46 miscentered clusters from the same HSC/ACT images using the paper's eight categories but with quantitative definitions (e.g., offset to centroid of the nearest Gaia star mask, deblending flag, ACT S/N), and then refit the cleaned sample: if the resulting miscentered fraction does not remain near 10%, the central claim fails. A second, independent falsifier is measuring X-ray centroids for the miscentered clusters: if the SZ center is not closer to the X-ray (potential) center than the optical center is, the claim that SZ centers estimate the true potential centroid is weakened.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the optical merger catalog used to compare merger fractions between well-centered and miscentered clusters."}],"review_version":1}