{"id":"35576422-7ebe-46d5-94ef-d4585a19153f","arxiv_id":"2607.08820","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Observational JWST/NIRCam selections recover almost none of TNG300's M⋆≥10^11 M⊙ galaxies at z~5 and their descendants rarely become the most massive systems at z=0 unless they experience late merger growth.","lead":"JWST color cuts miss most of the truly massive z~5 galaxies in the TNG300 simulation and the ones they do catch rarely grow into today's most massive systems. The work shows late mergers, not early mass rank, decide who ends up biggest, and offers cleaner NIRCam cuts for observers.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the dust-model caveat already flagged by the reader.","rationale":"The reader's weakest_assumption correctly isolates the dust-model limitation that the paper itself flags in §5.3. That limitation affects only the absolute recovery numbers, not the rank-order or evolutionary conclusions that form the paper's strongest claims. Because the authors already frame the recovery fraction as a lower limit and because the late-time merger result is dust-independent, no further load-bearing concern is required. The CONDITIONAL verdict with high confidence is therefore appropriate and needs no adjustment.","tokens_in":32292,"tokens_out":501,"duration_ms":4812,"concrete_test":"Re-run the S1 and S6 selections after uniformly increasing the dust-to-metal ratio (or birth-cloud optical depth) for galaxies with M☉>10^10.5 M⊙ until the simulated β–M_UV relation matches observations at z=5; if the fraction of M☉≥10^11 M⊙ galaxies recovered by S1 rises above ~50% while S6 purity falls below 0.5, the incompleteness claim would need quantitative revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's two central claims rest on (1) incompleteness of published NIRCam cuts for the most massive TNG300 galaxies at z~5 and (2) the poor predictive power of early mass rank for z=0 rank. Both are demonstrated cleanly inside the simulation. The only material soft spot is the acknowledged underestimation of UV attenuation for bright/massive systems (steeper simulated β vs. M_UV at z=4–6; §5.3), which the authors themselves treat as making the recovery fraction a lower limit. That caveat does not invert the qualitative conclusions: even if more of the M☉≥10^11 M⊙ galaxies reddened into the S1 wedge, the improved magnitude-dependent cuts (S6/S7) still exclude the low-mass dusty contaminants, and the late-time merger requirement for becoming a z=0 top-mass galaxy is independent of the dust model. No hidden inconsistency or unstated assumption undermines the argument as written.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper applies five published JWST/NIRCam color–magnitude selections to TNG300 galaxies with synthetic dust-attenuated photometry (Model C of Vogelsberger et al. 2020 / Shen et al. 2020) and finds that the Pérez-González et al. (2023) cut (S1) is the most inclusive among those that recover any objects. Even so, only 1 of 18 galaxies with M⋆ ≥ 10^11 M⊙ at z ∼ 5 satisfies S1; the rest are ∼0.5 mag bluer under the adopted dust model. The authors introduce magnitude-dependent S6/S7 wedges in (m200W − m277W) vs m150W that improve purity/completeness for M⋆ > 10^10.5 M⊙ at z = 5 while excluding low-mass dusty contaminants. Using SUBLINK merger trees they then show that the most massive NIRCam-selected (and purely mass-selected) galaxies at z = 7, 5, 4 and 2 rarely remain among the most massive systems at z = 0; only those that experience substantial late-time (z ≲ 0.2) merger-driven growth do so. The work therefore cautions both that current observational NIRCam cuts miss the most massive high-z galaxies and that high-z mass rank is a poor predictor of z = 0 mass rank.","tokens_in":32561,"tokens_out":1244,"duration_ms":10162,"significance":"If the incompleteness and late-time-merger results hold, the paper supplies a concrete, simulation-grounded caution for interpreting JWST photometric massive-galaxy candidates as either a complete census of the high-mass end or as direct progenitors of present-day BCGs/cluster centrals. Strengths include the transparent purity/completeness optimization that yields the new S6/S7 wedges, the explicit 1/18 recovery statistic, the multi-redshift growth-history comparison (Figs. 7–9), and the authors’ own quantification of the dust-model bias as a lower-limit statement (§5.3). The analysis is cleanly executed inside a single, well-documented simulation and is immediately useful to observers designing follow-up selections.","major_comments":[{"comment":"§5.3 and top-right panel of Fig. 2: the central incompleteness claim (1 of 18 galaxies with M⋆ ≥ 10^11 M⊙ recovered by S1) is load-bearing, yet the paper itself states that Model C underestimates UV attenuation for bright/massive systems (steeper simulated β–MUV at z = 4–6). While the authors correctly label the recovery fraction a lower limit, the manuscript never quantifies how large a reddening shift would be required to move the remaining 17 objects into the S1 wedge, nor whether that shift would also pull low-mass dusty satellites into the same region. A short sensitivity test (e.g., uniform or mass-dependent ΔE(B−V) or Δβ applied to the high-mass tail) is needed before the absolute statement “current selections are not identifying the most massive high-redshift galaxies” can be taken at face value outside TNG300+Model C.","section":null},{"comment":"§3.3 and Fig. 4: S6/S7 are optimized on the same TNG300 photometry to which they are then applied; purity (0.70) and completeness (0.39/0.88) are therefore in-sample figures of merit. The paper presents them as “improved” observational selections, but without a hold-out redshift, an independent simulation, or an explicit statement that they remain untested against low-z interlopers (already noted in the text), their claimed superiority for real JWST catalogs is not yet demonstrated. Either a cross-validation step or a clearer framing as “simulation-motivated proposals requiring observational validation” is required.","section":null}],"minor_comments":[{"comment":"Abstract and §3.1: the abstract says “1 of the 18 galaxies” while the body text sometimes says “1 of the 17”; reconcile the exact count of M⋆ ≥ 10^11 M⊙ systems at z = 5.","section":null},{"comment":"Table 1 / Fig. 1: S4 and S5 return zero objects even at z > 5; a one-sentence quantitative statement of how far the simulated colors fall short of the published cuts would help readers judge whether the non-detection is decisive or merely a dust-model artifact.","section":null},{"comment":"Fig. 6 caption: the dynamic-range statement is useful, but the physical scale bars (0.5 cMpc) are hard to read in the rendered panels; enlarge or move them.","section":null},{"comment":"§2.2: the redshift-dependent dust-to-metal ratio 0.9 × (z/2)^−1.92 is given without an immediate citation to the calibration paper; add the reference for reproducibility.","section":null},{"comment":"Throughout: “NIRCam-selected” is used both for the observational S1–S5 cuts and for the new S6/S7 wedges; a brief terminological distinction (e.g., “literature-selected” vs “simulation-optimized”) would reduce ambiguity.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The dust-model caveat is already flagged by the authors and does not invert the qualitative conclusions, so I do not regard it as grounds for major revision or rejection. The paper is a solid, timely contribution that fits the journal; the two major points above are fixable with modest additional analysis or clearer framing and should not delay publication once addressed."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is straightforward. Published JWST/NIRCam color cuts, when applied to dust-attenuated TNG300 photometry, recover only 1 of 18 galaxies with M⋆ ≥ 10^11 M⊙ at z~5 under the best of the five literature criteria (Pérez-González et al.). The rest sit ~0.5 mag too blue. The authors then build magnitude-dependent wedges (S6/S7) that raise purity or completeness for M⋆ > 10^10.5 M⊙ while dropping the low-mass dusty contaminants that the literature cuts pick up. They also track descendants and show that early mass rank rarely maps to the top of the z=0 mass function unless substantial late-time (z ≲ 0.2) merger growth occurs.\n\nWhat is new is the systematic side-by-side application of the five published cuts inside one simulation with consistent synthetic photometry, the explicit recovery fractions, the new wedges optimized on purity/completeness, and the clean demonstration that late mergers, not high-z mass rank, decide who ends up most massive today. The merger-tree work and the color-magnitude optimization are cleanly done; the figures make the incompleteness and the growth histories easy to see. The citation pattern is appropriate and the free parameters of the dust model are stated.\n\nThe soft spot is real but already quantified by the authors: the adopted dust model underestimates UV attenuation for bright/massive systems (steeper simulated β vs M_UV), so the absolute recovery fraction is a lower limit. That does not invert the qualitative claims. Even if more of the ultra-massive galaxies reddened into the S1 wedge, the improved magnitude-dependent cuts still exclude the low-mass contaminants, and the late-merger requirement for becoming a z=0 top-mass galaxy is independent of dust. No other load-bearing flaw appears.\n\nThis is for people who select high-z massive candidates or who map them onto local BCGs/protoclusters. It is immediately usable and deserves a serious referee. I would cite the recovery fractions and the descendant result, and I would bring the new wedges to reading group.","headline":"Solid simulation test of published NIRCam cuts: they miss most of the true high-mass end at z~5 in TNG300, and early mass rank is a poor predictor of z=0 rank unless late mergers intervene.","tokens_in":33171,"tokens_out":566,"would_cite":true,"duration_ms":6554,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"JWST color cuts miss most of the truly massive high-redshift galaxies, and those galaxies rarely become today's giants without late mergers.","keywords":["high-redshift galaxies","JWST/NIRCam","photometric selection","galaxy formation","IllustrisTNG","dust attenuation","stellar mass assembly","progenitors of massive galaxies"],"falsifier":"If deeper multi-band imaging or spectroscopy of a complete mass-selected sample at z ~ 5 shows that the majority of galaxies above 10^11 solar masses already satisfy the published red color cuts (or that their descendants systematically rank among the top ten most massive systems at z = 0 without late mergers), the central claims would be contradicted.","tokens_in":33225,"feed_emoji":"🌌","tokens_out":966,"duration_ms":80696,"temperature":0.7,"pith_summary":"Early JWST/NIRCam surveys found many high-redshift massive-galaxy candidates that earlier UV surveys missed, raising questions about whether galaxy formation models are incomplete and whether those objects are the direct ancestors of today's most massive galaxies. This paper tests five published photometric color selections by applying them to galaxies in the large-volume TNG300 simulation that have been given synthetic dust-attenuated NIRCam photometry. The most inclusive published cut still recovers only one of the eighteen simulated galaxies that already exceed 10^11 solar masses at z ~ 5; the rest are systematically bluer under the adopted dust model. The authors therefore introduce magnitude-dependent color cuts that better isolate massive systems while rejecting dusty low-mass interlopers. Tracking the selected galaxies forward in the merger trees shows that they almost never finish as the most massive galaxies at z = 0. Only those that experience substantial merger-driven growth at very late times (z less than or equal to 0.2) join the ranks of the present-day giants. The work therefore both improves how observers should select high-redshift massive galaxies and warns against reading today's most massive systems as simple descendants of the JWST candidates.","feed_headline":"JWST color cuts miss most truly massive early galaxies","feed_subtitle":"Simulation shows they rarely become today's giants without late mergers","key_machinery":"Synthetic dust-attenuated JWST/NIRCam photometry for TNG300 galaxies (full Monte-Carlo radiative transfer) combined with SUBLINK merger trees that link each high-redshift selected galaxy to its z = 0 descendant.","core_discovery":"Published JWST/NIRCam color selections do not recover the most massive high-redshift galaxies in the TNG300 simulation (only 1 of 18 systems with stellar mass above 10^11 solar masses at z ~ 5 satisfies the best cut), and the galaxies that are selected rarely evolve into the most massive galaxies by z = 0 unless they undergo substantial late-time merger growth.","pith_inferences":["If the dust underestimation is as large as the UV-slope discrepancy suggests, many of the observed 'little red dots' and extremely red candidates may still be lower-mass or AGN-contaminated systems once better dust models are applied.","Wide-area cosmic-web maps at z > 2 (e.g., from future infrared surveys) could supply the environmental context needed to flag which massive high-redshift galaxies are most likely to experience the late mergers that produce today's giants.","The same simulation-based selection-refinement method can be repeated at z = 7 and z = 4 to produce redshift-specific purity cuts for ongoing deep fields."],"forward_implications":["Observers should replace single-color cuts with magnitude-dependent Balmer-break selections if they want higher purity and completeness for massive z ~ 5 galaxies.","Number-density comparisons between JWST candidates and theoretical models must treat the published selections as incomplete at the high-mass end.","Searches for progenitors of brightest cluster galaxies cannot rely solely on the most massive high-redshift galaxies inside overdensities.","Claims that a given high-redshift massive galaxy is a direct progenitor of a present-day ultra-massive galaxy require independent evidence of late-time merger activity."],"fun_headline_variants":["JWST color cuts recover only 1 of 18 massive z~5 galaxies","NIRCam selections miss most truly massive high-z galaxies","Selected early massive galaxies rarely become z=0 giants","Best JWST cut finds just one TNG300 galaxy above 10^11 Msun","High-z massive galaxies need late mergers to top ranks today"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The dust model underestimates how much ultraviolet light is blocked in the most massive high-redshift galaxies, so the true number of systems that would pass the color cuts is higher than the simulation reports.","fun_headline_variants_meta":{"raw":{"variants":["JWST color cuts recover only 1 of 18 massive z~5 galaxies","NIRCam selections miss most truly massive high-z galaxies","Selected early massive galaxies rarely become z=0 giants","Best JWST cut finds just one TNG300 galaxy above 10^11 Msun","High-z massive galaxies need late mergers to top ranks today"]},"model":"grok-4.5","effort":"low","cost_usd":0.004656,"raw_usage":{"total_tokens":1434,"prompt_tokens":891,"num_sources_used":0,"completion_tokens":95,"cost_in_usd_ticks":46560000,"prompt_tokens_details":{"text_tokens":891,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":448,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":891,"tokens_out":95,"duration_ms":3680,"temperature":1.0,"reasoning_tokens":448,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T06:26:11.078675+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If deeper multi-band imaging or spectroscopy of a complete mass-selected sample at z ~ 5 shows that the majority of galaxies above 10^11 solar masses already satisfy the published red color cuts (or that their descendants systematically rank among the top ten most massive systems at z = 0 without late mergers), the central claims would be contradicted.","supporting_citations":[],"review_version":1}