{"id":"cb27b89d-fa28-4c50-a0d3-d62f211dddd9","arxiv_id":"2411.11956","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A point-cloud neural network trained on simulated galaxies identifies z~4 protocluster member candidates from broadband photometry with higher purity than density-based methods and yields 121 candidates in HSC-SSP.","lead":"Astronomers trained a point-cloud neural network called PCFNet on simulated galaxies to find protoclusters, the ancestors of galaxy clusters, at redshift about 4 using only photometric images. The method returns higher-purity protocluster member lists than older density-based methods and yields 121 protocluster candidates in the HSC-SSP survey.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation uses light cones that share 13.7% of galaxies with training data; the claimed 5x gain over 2DBM may be leakage, not generalization.","rationale":"Reader identified simulation fidelity as the weakest assumption; I agree that is fundamental, but the more immediate and internally checkable flaw is the training/evaluation overlap disclosed in Section 2.1. The manuscript's dismissal of leakage is unconvincing because the overlap is in galaxy identity and large-scale structure, not merely in photometry. The paper itself provides the numbers needed to quantify the leakage and to design a clean retraining test. This concern is load-bearing: the headline 'five times more' and the low-mass protocluster yield (Fig. 11) rest entirely on the simulation evaluation, and if those numbers shrink when leakage is removed, the applied 121-candidate catalog loses its calibrated basis. I am not claiming the method is wrong; I am claiming the evidence as presented does not yet establish the central claim. The reader's CONDITIONAL verdict is appropriate, and the requested retraining check should be part of the conditions. I also note the abstract/body precision-recall swap and the absence of spectroscopic confirmation, but those are secondary to the leakage issue.","tokens_in":29474,"tokens_out":4921,"duration_ms":43967,"concrete_test":"Retrain PCFNet from scratch after deleting from the training light cones every galaxy whose Millennium Simulation ID appears in any evaluation light cone (at minimum the 11,264 Deep and 10,345 UltraDeep duplicates), and also drop any training light cone that overlaps the evaluation subvolume. Then recompute PR AUC, precision/recall at σth=2.5, and the descendant halo-mass histogram of Fig. 11. If the gap between PCFNet and 2DBM falls below the quoted uncertainties, the headline 'five times more candidates' claim is not supported. If feasible, repeat using a fully independent light-cone mock (e.g., from a different simulation) as an out-of-distribution check.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 discloses that the four evaluation light cones are drawn from the same Millennium Simulation box as the 15 training light cones: 'the same galaxies may appear in training and validation/evaluation data,' with 11,264 of 82,245 Deep (10,345 of 77,986 UltraDeep) evaluation galaxies originating from training galaxies. The authors argue the risk is 'slight' because coordinates and resampled magnitudes differ. This is not sufficient: the underlying dark-matter density field, the catalog of galaxy IDs, and the spatial arrangement of protocluster members are shared. Even if individual magnitudes are re-randomized, the network can memorize specific overdensities and their member configurations, so the PR curve in Fig. 6, the headline precision/recall (Sec. 4.1), and the descendant-mass comparison in Fig. 11 are all computed on non-independent data. Since the claimed 5x improvement over 2DBM and the 7±2x gain for M_halo<1e15 are based on this evaluation, a leakage-driven inflation would directly invalidate the central claim. The 121 observed candidates inherit the same risk because the network has no demonstrated generalization beyond its training distribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents PCFNet, a point-cloud deep-learning classifier that assigns each g-dropout galaxy a probability of being a member of a z~4 protocluster, using sky coordinates, i-band magnitude, (g-i) color, and an MDN-derived redshift PDF for neighbors within 5'. PCFNet is trained on PCcone, a Millennium Simulation plus L-GALAXIES light cone, with labels defined from merger-tree-derived protocluster members (halo mass > 10^14 Msun at z=0). The authors report per-galaxy precision/recall values of 7.5% and 44% (with the labels reversed in the abstract relative to Section 4.1), a PR AUC of about 0.28, protocluster-level completeness/purity of 10.9%/69%, and a factor 7 +/- 2 increase in detected lower-mass descendant halos compared with the 2DBM baseline. Applied to about 17.6 deg^2 of HSC-SSP Deep/UltraDeep, the method yields 121 unconfirmed protocluster candidates, 63% of which match Toshikawa et al. (2024) within 8'. The paper also discusses rest-UV brightness and quenching trends in both simulation and observation.","tokens_in":29719,"tokens_out":8372,"duration_ms":80937,"significance":"If the headline numbers survive independent evaluation, this would be a useful advance: PCFNet is a flexible, publicly coded method that treats protoclusters as point clouds and uses full photo-z PDFs rather than a single redshift, and it is plausibly applicable to other dropout samples and to Euclid, LSST, and Roman. The baselines (2DBM, 3DBM) are reasonable, the labels are defined externally from merger trees rather than by the network, and the paper is transparent about the semi-analytic model's limitations and the unconfirmed status of its observed candidates. The 121-candidate catalog and the low-mass comparison (23 versus 1 at M_halo^z=0 = 10^14-10^14.5 Msun) would be of broad interest to protocluster searches. However, the evaluation's non-independence from the training set and the internal inconsistency in the headline precision/recall numbers currently prevent the central claim from being accepted at face value.","major_comments":[{"comment":"The headline numbers are internally inconsistent. The abstract states 'recall = 7.5 +/- 0.2%, precision = 44 +/- 1%', while Section 4.1 states 'the precision and recall of PCFNet are 7.5 +/- 0.2% and 44 +/- 1%, respectively'; one of these has the labels swapped. The same section reports 2DBM precision/recall of 1.5 +/- 0.1% and 38 +/- 2%, then says PCFNet's recall at the equivalent precision is 'approximately 11 times the recall of the 2DBM'; numerically 16/1.5 ~ 10.7, so the comparison appears to be with 2DBM's precision, not its recall. The abstract's 'five times more protocluster member candidates' is also not supported by the quoted recall values (44% versus 38% is a factor of 1.16); it may refer to precision (7.5/1.5 = 5) or to the halo-mass comparison in Section 6.1, but the sentence as written is misleading. These are the central quantitative claims and must be corrected and made mutually consistent.","section":"Abstract and Section 4.1"},{"comment":"The evaluation is not independent of the training data. Section 2.1 discloses that 11,264 of 82,245 Deep-layer and 10,345 of 77,986 UltraDeep-layer evaluation galaxies originate from the same galaxies used in training, and all light cones are drawn from the same Millennium Simulation box. The authors argue that the leakage risk is 'slight' because coordinates and resampled magnitudes differ, but the underlying dark-matter density field, the galaxy IDs, the halo memberships, and therefore the spatial configurations of protocluster members are shared. The PR curve in Figure 6, the precision/recall values in Section 4.1, and the descendant-mass comparison in Figure 11 are all computed on this partly in-sample evaluation, so the claimed gains over 2DBM could be inflated by memorization of specific overdensities. Please provide a clean evaluation, for example by training on light cones from one simulation box and testing on light cones from a different box or from a disjoint spatial region with no shared galaxy IDs, or by reporting all metrics restricted to the non-shared evaluation galaxies (70,981 Deep and 67,641 UltraDeep).","section":"Section 2.1, Section 4.1, Figure 6, Figure 11"},{"comment":"The transfer of PCFNet to real data is not yet demonstrated. The network is trained purely on PCcone, and Section 2 states that the semi-analytic model 'does not perfectly reproduce the actual universe'; Figure 1 also shows a bright-end excess in the HSC-SSP magnitude distributions relative to PCcone. The only external checks in Section 5.2 are spatial matching to Toshikawa et al. (2024) (63% within 8') and the authors' own statement that the candidates are unconfirmed. To support the abstract's claim that PCFNet can be applied to future surveys, the paper should validate the model on spectroscopically confirmed z~4 protoclusters or on an independent overdensity-selected sample, or otherwise clearly frame the 121 candidates as a predicted catalog requiring follow-up rather than as a validation of the method.","section":"Section 5.2 and Section 2.2"}],"minor_comments":[{"comment":"The footnote says 'Precision represents how completely selected, and Recall does how purely selected', which is the reverse of the standard definitions; please correct it, as it contributes to the headline-number confusion.","section":"Section 4.1 footnote"},{"comment":"There are typos: 'Simualtion Data' should be 'Simulation Data' and 'Outliners' should be 'Outliers'.","section":"Section 2.1 and Section 2.2"},{"comment":"The Deep and UltraDeep fields overlap on the sky (for example, COSMOS is listed as both Deep and UltraDeep), and the table entries sum to 121; please state explicitly whether the 121 candidates are merged across layers and whether 'unique' excludes objects detected in both layers.","section":"Section 5.2 and Table 5"},{"comment":"The normalization of sigma_prob uses the mean and standard deviation of the evaluation data, which are then applied to HSC observations; please justify this choice given the known depth and number-density differences between Deep and UltraDeep, or explain how the domain shift is handled.","section":"Equation (13) and Section 5.2"},{"comment":"The sentence introducing the grouping threshold as 'the minimum number of protocluster members, i.e., gamma N_th in case of lowest completeness' is unclear and should be rewritten for precision.","section":"Section 4.2"}],"recommendation":"major_revision","confidential_remarks":"The central methodological risk is the shared training/evaluation data; if the authors can re-evaluate on a clean split (or honestly report the non-shared subset), the method may be publishable. The internal precision/recall mismatch is easy to fix but must be fixed before publication. The manuscript fits the scope of ApJ, and the public code and transparent disclosure of the shared-galaxy issue are commendable. I would not recommend acceptance before the leakage issue is addressed and the headline numbers are corrected."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely new application of point-cloud deep learning to photometric protocluster detection, with code and a public candidate list. The central performance claim, however, is not credible as written: the abstract and body disagree on which of precision/recall is 7.5% vs 44%, and the 'five times more' and '11 times' statements don't match the numbers in Section 4.1. And because the evaluation shares galaxies and the same simulation box with training, the head-to-head with 2DBM is not an independent test.\n\nWhat's good: PCFNet is a sensible architecture that combines PointNet/DG-CNN with an MDN redshift PDF and persistent-homology peak finding. Training on mock dropout galaxies and using the full line-of-sight PDF rather than a point estimate is a reasonable response to the photometric redshift degeneracy. The paper discloses the train/eval overlap and the semi-analytic model's limitations. The discussion of the bright-end detection bias and the attempt to correct it is honest. The candidate list and code are released, which makes the work useful regardless of the headline.\n\nSoft spots: First, the metric inconsistencies. If the intent is precision 7.5% and recall 44%, then the 'five times more' claim is hard to see: 2DBM at 4σ has recall 38%, so PCFNet's recall is ~1.2x, not 5x. Maybe they mean something like 'five times more candidates per square degree' (which appears as 3.8x in Sec 5.2), but the abstract says 'member candidates.' The '11 times' line is also unmatchable with the numbers given. These need to be sorted out in any revision.\n\nSecond, leakage. 11,264 of 82,245 evaluation galaxies come from the same galaxies in training, and all light cones are carved from the same Millennium Simulation. That means the large-scale structure and the particular protocluster configurations are shared. Re-randomized magnitudes don't erase the overdensity signal; a network can memorize the spatial arrangement of a specific simulated structure. To make the 5x claim stick, they need to evaluate on a truly independent simulation, or at least remove overlapping galaxies and re-run. Until then, the advantage over 2DBM is plausibly inflated.\n\nThird, the 121 observed candidates are unconfirmed. The matching to Toshikawa et al. helps, but real confirmation needs spectroscopy or at least independent multi-wavelength evidence.\n\nNet: the paper deserves a serious referee, but it needs major revision before the performance claims are publishable. The method is promising, and the community-released code/catalog is valuable. I'd want to see the corrected metrics and a clean evaluation before trusting the numbers.","headline":"Promising new ML method for z~4 protocluster search, but the headline performance claims are internally inconsistent and rest on a non-independent simulation evaluation.","tokens_in":30323,"tokens_out":4047,"would_cite":false,"duration_ms":35931,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A point-cloud network trained on mock galaxies finds 121 protocluster candidates at $z\\approx4$ using photometry alone.","keywords":["protoclusters","z≈4","deep learning","point clouds","photometric redshifts","Lyman-break galaxies","HSC-SSP","galaxy quenching"],"falsifier":"A spectroscopic redshift survey of the 121 candidates' members is the direct falsifier: if the fraction of galaxies sharing the core's redshift is statistically indistinguishable from the field, PCFNet's membership probability is not tracing real protocluster structure. A cheaper check is comparing the angular correlation function of $g$-dropouts in PCcone and HSC-SSP on $5'$ scales, since a mismatch would invalidate the training distribution.","tokens_in":29280,"feed_emoji":"🔭","tokens_out":9857,"duration_ms":87711,"temperature":0.7,"pith_summary":"Protoclusters at $z\\approx4$—the overdense regions that evolve into today's galaxy clusters—are rare and hard to identify because photometric redshifts smear member galaxies along the line of sight. This paper claims that a deep-learning model called PCFNet can spot them using only optical broad-band photometry by treating each galaxy as a point in a cloud and feeding the network the sky distribution, $i$-band magnitude, $(g-i)$ color, and a full redshift probability distribution. On mock data from the PCcone light cone, PCFNet recovers about five times more protocluster member candidates than conventional surface-density methods at comparable purity (recall $=7.5\\pm0.2\\%$, precision $=44\\pm1\\%$), and it preferentially finds the less-massive progenitors ($M_\\mathrm{halo}^{z=0}=10^{14}$–$10^{14.5}\\,M_\\odot$) that older searches miss. Applied to $\\sim17\\,\\mathrm{deg}^2$ of HSC-SSP Deep/UltraDeep imaging, it returns 121 protocluster candidates at $z\\approx4$ whose members are brighter in rest-UV than field galaxies. If the claim holds, statistically meaningful protocluster samples at $z\\approx4$ become accessible from imaging alone, without spectroscopy.","feed_headline":"Point-cloud AI finds 121 protocluster candidates at z≈4","feed_subtitle":"Trained on mock photometry, PCFNet finds five times more protocluster members than standard searches, at equal purity.","key_machinery":"PCFNet is a point-cloud classifier: it takes the set of dropout-selected galaxies within a $5'$ radius of a target galaxy as an unordered point cloud, expands each point into 16 features, and processes local neighborhoods with EdgeConv 'skipDG' blocks before a global max-pooling layer and a classifier output a membership probability. The line-of-sight information comes from a Mixture Density Network (MDN) that turns $g,r,i$ magnitudes into a three-Gaussian redshift probability density, so the network can exploit the full multi-peaked redshift uncertainty rather than a single photometric redshift. The grouping stage uses persistent homology peak detection on the significance map to assemble galaxies into protocluster candidates. The training and evaluation data are dropout-selected galaxies from PCcone, a Millennium Simulation plus L-GALAXIES light cone with HSC-SSP-matched depths; the protocluster labels come from merger-tree-defined core galaxies and members within $5.5\\,\\mathrm{cMpc}$.","core_discovery":"The central claim is that protocluster membership at $z\\approx4$ can be predicted per galaxy from photometric data alone, and that a model trained on a realistic mock light cone transfers to real survey data. PCFNet assigns each galaxy a membership probability; thresholding at $2.5\\sigma$ and grouping the candidates with persistent homology yields per-protocluster completeness $10.9\\pm0.8\\%$ and purity $69\\pm4\\%$ in simulation, while a conventional surface-density search reaches comparable purity only at much lower completeness. The same pipeline applied to HSC-SSP Deep/UltraDeep data finds 121 protocluster candidates over $\\sim17.6\\,\\mathrm{deg}^2$, and reaches protoclusters that will become only $10^{14}$–$10^{14.5}\\,M_\\odot$ halos by $z=0$, not just Coma-like superclusters. On the observed candidates, member galaxies show a rest-UV bright-end excess after correcting for the model's brightness-dependent selection bias, which the paper interprets as early, enhanced star formation in protocluster environments.","pith_inferences":["If the mock-to-real transfer holds, running PCFNet on the wider HSC-SSP Wide layer or LSST should yield thousands of $z\\approx4$ protocluster candidates, turning the current small-sample studies into population statistics.","The six-fold recall gap between bright ($i<24.5$) and faint members implies a strong selection function on any luminosity or mass measurement from a learned photometric finder; the simulation-based correction used in the paper is a template future searches will need.","A sharper physical test would be to train the same network on a hydrodynamic light-cone simulation and compare the recovered candidates: overlap of the 121 candidates would show the detections trace real overdensity rather than the semi-analytic model's particular galaxy-halo assignment."],"forward_implications":["The method removes spectroscopy as a prerequisite for building a $z\\approx4$ protocluster sample; inference runs on a single GPU in minutes to hours, so the same network can be applied to LSST, Euclid, and Roman data.","The 121 HSC-SSP candidates give a detection density of $6.9\\,\\mathrm{deg}^{-2}$, about 3.8 times higher than the earlier surface-density search, yielding a larger and less massive-biased sample for follow-up.","The selection reaches protoclusters destined to become $M_\\mathrm{halo}^{z=0}\\sim10^{14}\\,M_\\odot$ groups, where supernova feedback and galactic winds may dominate, so environmental-effect studies can extend below the Coma-like mass scale.","In the simulation, the fraction of protocluster cores associated with quiescent satellites rises with both the $z\\approx4$ halo mass and the accumulated $z=0$ halo mass, giving a quantitative prediction for future observations.","PCFNet is not tied to one dropout color: retraining on other dropout selections should extend the search to $z\\approx2$–$8$, as the paper notes."],"supporting_citations":[{"why":"Builds and validates PCcone, the light-cone mock used for training and evaluation.","marker":"Araya-Araya et al. 2021"},{"why":"Supplies the Millennium Simulation dark-matter backbone and merger trees from which protocluster definitions and PCcone are derived.","marker":"Springel et al. 2006"},{"why":"Provides the L-GALAXIES semi-analytic galaxy formation model that assigns photometric properties in PCcone.","marker":"Henriques et al. 2015"},{"why":"Defines the g-dropout selection criteria and HSC-SSP dropout samples used for both mock and real data.","marker":"Ono et al. 2018"},{"why":"Introduces PointNet, the point-cloud architecture PCFNet is based on.","marker":"Qi et al. 2016"},{"why":"Introduces DG-CNN and EdgeConv, the local neighborhood graph operation PCFNet uses.","marker":"Wang et al. 2018b"},{"why":"Introduces Mixture Density Networks, used to turn g,r,i magnitudes into the redshift PDFs fed to PCFNet.","marker":"Bishop 1994"},{"why":"Supplies the conventional surface-density protocluster search (2DBM) that PCFNet is compared against.","marker":"Toshikawa et al. 2018"},{"why":"Releases the HSC-SSP PDR3 S20A Deep/UltraDeep photometric catalog to which PCFNet is applied.","marker":"Aihara et al. 2022"},{"why":"Defines the protocluster effective radius and halo-mass scale that set the member-galaxy radius and the low-mass target.","marker":"Chiang et al. 2013"}],"fun_headline_variants":["AI point cloud finds 121 protocluster candidates at z≈4","Deep learning uncovers 121 protoclusters from photometry alone","PCFNet: AI triples protocluster member yield at z~4","Machine learning reveals 121 protoclusters in HSC Deep survey"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"PCcone, the semi-analytic Millennium-Simulation light cone, reproduces the real $z\\approx4$ dropout galaxy population closely enough that a network trained on it can recognize protoclusters in HSC-SSP photometry; the paper acknowledges the semi-analytic model does not perfectly match the actual universe.","fun_headline_variants_meta":{"raw":{"variants":["AI point cloud finds 121 protocluster candidates at z≈4","Deep learning uncovers 121 protoclusters from photometry alone","PCFNet: AI triples protocluster member yield at z~4","Machine learning reveals 121 protoclusters in HSC Deep survey"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000225,"raw_usage":{"total_tokens":1583,"prompt_tokens":1180,"completion_tokens":403,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":796,"completion_tokens_details":{"reasoning_tokens":324}},"tokens_in":796,"tokens_out":403,"duration_ms":4319,"temperature":1.0,"reasoning_tokens":324,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T18:04:54.328733+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A spectroscopic redshift survey of the 121 candidates' members is the direct falsifier: if the fraction of galaxies sharing the core's redshift is statistically indistinguishable from the field, PCFNet's membership probability is not tracing real protocluster structure. A cheaper check is comparing the angular correlation function of $g$-dropouts in PCcone and HSC-SSP on $5'$ scales, since a mismatch would invalidate the training distribution.","supporting_citations":[],"review_version":1}