{"id":"b34bd85e-a861-4a33-b0a9-4196bc99c0c5","arxiv_id":"2608.02898","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GARDENS-Wide, a 100 square degree mock redshift survey of dusty star-forming galaxies, predicts proto-cluster contraction, a z about 2 peak in star-forming fraction, and TolTEC LSS detection yields.","lead":"A new 100 square degree mock survey of dusty galaxies, built from a dark matter simulation, tracks how clusters and their bright infrared galaxies grow across cosmic time. It gives concrete forecasts for what the TolTEC telescope should see, which observers can test.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed decline of the star-forming fraction to ~20-30% at z>4 in proto-clusters extrapolates the quenching model beyond its z<=4 calibration and conflicts with the paper's own example proto-clusters being ~70-80% quiescent at z~5.4-5.6.","rationale":"The paper is a solid, reproducible mock catalogue with a clear methodology and public data. The cluster assembly analysis is a valuable use of the MDPL2 simulation. I focused on the headline claim of star-forming fraction evolution. The strongest central claim is the redshift evolution of the SF fraction in proto-clusters, which is directly used in the abstract and conclusions. The paper explicitly states (Section 2.1) that the quiescent fraction is calibrated only out to z~4, yet Figure 7 and the abstract make statements out to z~5.5-6. The example proto-clusters shown in Figures 4 and 5 are ~68-80% quiescent at z~5.4-5.6, which conflicts with the general expectation and with observations that high-redshift proto-clusters are star-forming dominated. This suggests the quenching model is over-extrapolated. The reader's weakest_assumption about the model in rare environments is on point, but I sharpen it to a specific, testable failure mode: the uncalibrated high-z quenching. I did not make the beta circularity the primary concern because the intrinsic L_IR classification of LIRGs/ULIRGs does not depend on beta; beta only affects observed fluxes and TolTEC predictions. The size evolution claim is a robust N-body result, though the exact 22 Mpc value may be resolution-sensitive; that is secondary. The proposed concrete check uses existing public data and observations, so it can settle the concern without waiting for TolTEC. If the check shows the model's z>4 quiescent fraction is far too high, the abstract and summary statements about the high-z decline should be revised or heavily caveated, but the lower-redshift peak at z~2 would likely survive. This is consistent with the reader's CONDITIONAL verdict: accept with major revisions.","tokens_in":27324,"tokens_out":20661,"duration_ms":182480,"concrete_test":"Using the public GARDENS-Wide catalogue, compute the median quiescent fraction (sSFR below the Pacifici et al. 2016 threshold) for M* > 10^9.5 galaxies in the assembly histories of the 32 rich clusters at z = 4.5, 5.0, and 5.5. Compare these values with (a) the observationally measured quiescent fraction of massive galaxies at z~4-5 from deep near-IR surveys (e.g., Schreiber et al. 2018; Tanaka et al. 2019) and (b) the star-forming fraction in spectroscopically confirmed z>4 proto-clusters (e.g., SPT2349-56 at z=4.3). If the mock quiescent fraction exceeds observed upper limits by more than a factor of ~2, the high-z SF fraction decline is an extrapolation artifact; if it matches available constraints, the claim is supported. In the absence of z>4 data, the paper should explicitly state that the z>4 fraction is uncalibrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 states that the galaxy formation model is calibrated to reproduce the quiescent fraction only for 0<=z<=4. Nevertheless, Section 4.1 and Figure 7 report star-forming fractions out to z~5.5-6, and the abstract claims a decline to ~20% by z~5. The example assembly histories in Figures 4 and 5 show the richest and a typical rich proto-cluster at z=5.4-5.6 with only ~20-32% star-forming members, i.e. ~68-80% quiescent at epochs where massive quiescent galaxies are observed to be extremely rare and known z>4 proto-clusters are dominated by starbursts. Because the quenching prescription is unconstrained at these redshifts, the predicted high-z decline is an extrapolation artifact candidate, not a calibrated prediction. This directly affects a headline result: the redshift evolution of the star-forming galaxy fraction in cluster progenitors. The same SFRs drive the LIRG/ULIRG classifications, so the infrared population fractions in proto-clusters inherit this uncertainty. The paper does not compare its high-z proto-cluster quiescent fraction with any z>4 observationally confirmed proto-cluster or with the deep field quiescent fraction at z>4.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents GARDENS-Wide, a 100 square degree mock redshift survey of dusty star-forming galaxies (DSFGs) built from the MDPL2 N-body simulation, with galaxy populations assigned via subhalo abundance matching and a semi-empirical model for star formation and quenching. The mock includes gravitational lensing and is publicly released. The authors validate the mock against observed number counts at 500 microns, 1.1, 1.4, and 2.0 mm, then use the large area to identify gravitationally bound systems (pairs, groups, poor and rich clusters) and trace their assembly histories through merger trees. They report that star-forming galaxies contribute about 35 per cent of cluster members at low redshift, rising to 60-65 per cent at z~2 and then declining to about 20-30 per cent by z~5; that LIRGs peak at about 20-40 per cent near z~2; that ULIRGs remain subdominant; and that proto-clusters of rich clusters contract from about 22 comoving Mpc at z~5.5 to about 5.3 Mpc at z~0. They also find that ULIRGs become increasingly centrally concentrated at z>1.5, and they provide predictions for the TolTEC Large-Scale Structure survey, including source counts, redshift distributions, detectable galaxy pairs, and proto-cluster member separations.","tokens_in":27588,"tokens_out":6552,"duration_ms":54411,"significance":"If the mock is reliable, it is a valuable community resource: the public catalogue, the wide-area lightcone, and the TolTEC predictions provide concrete, falsifiable expectations for an upcoming survey. The use of a 100 deg^2 volume to identify 32 rich clusters and trace their assembly histories is a genuine advance over smaller-area mocks, and the comparison of proto-cluster SFRD contributions with Chiang et al. (2017) is informative. However, the central astrophysical claims about the redshift evolution of star-forming and IR-luminous populations in cluster progenitors rest on two load-bearing assumptions that are not fully supported: (i) the quenching model is calibrated only at z<=4, yet the headline decline at z>4 is extrapolated; and (ii) the number-count validation at 1.1-2.0 mm is partly circular because the dust emissivity index beta=2.2 was chosen to match those counts. The TolTEC predictions inherit these uncertainties, although the survey-specific predictions (e.g., median redshifts, pair resolution fractions) are less sensitive to the high-z extrapolation.","major_comments":[{"comment":"The predicted decline of the star-forming galaxy fraction to about 20-30 per cent at z>4 is an extrapolation of the galaxy formation model beyond its calibration range. Section 2.1 states that the quiescent fraction is reproduced only for 0<=z<=4, yet Section 4.1 and Figure 7 present star-forming fractions at z up to 5.5-6, and the abstract and conclusions advertise the decline to about 20 per cent by z~5 as a result. The example assembly histories in Figures 4 and 5 show quiescent fractions of about 68 per cent at z=5.60 and 80 per cent at z=5.42, which are in tension with the observed rarity of massive quiescent galaxies at z>4 and with the starburst-dominated nature of known z>4 proto-clusters. Because this trend drives the LIRG/ULIRG fractions at high redshift, it is load-bearing for the central claim. Please either restrict the high-redshift claims to z<4, validate the high-z quiescent fraction against independent data (e.g., deep-field quiescent fractions or confirmed z>4 proto-clusters), or present the high-z decline as an unconstrained extrapolation and quantify its systematic uncertainty.","section":"Section 2.1, Section 4.1, Figures 4-5, Figure 7, Abstract"},{"comment":"The validation against the 1.1, 1.4, and 2.0 mm number counts is partly circular. The text states that the dust emissivity index beta=2.2 was adopted specifically because it 'provides the best agreement with the observed number counts at 1.1, 1.4, and 2.0 mm.' Therefore the good agreement at those wavelengths in Figure 1 is a fit to the same data, not an independent prediction. This weakens the claim that the mock 'reproduces' the observed counts. Please recast the validation to identify which statistics are genuinely predicted (for example, the 500 micron counts, the redshift distributions, and the TolTEC source-count predictions), and add a sensitivity test showing how the predicted counts vary within the observationally allowed range of beta.","section":"Section 2.2, Figure 1"},{"comment":"The high-redshift value of the star-forming fraction is reported inconsistently. The abstract and the conclusions state that the fraction declines to about 20 per cent by z~5, while Section 4.1 says it 'gradually declines to about 30 per cent, forming a tail that extends to z~5.5.' This is a load-bearing number in the summary of the paper, so the discrepancy should be corrected and the final value should be reconciled between the abstract and the body.","section":"Abstract, Section 4.1, Section 6"},{"comment":"The central spatial claims (contraction of proto-cluster radii and central concentration of LIRGs/ULIRGs) rely on a radius metric defined as the maximum 3D comoving distance from the central halo to the most distant member halo. The robustness test using the mean distance to the three or four most distant members shows differences of up to 40 per cent for the LIRG and ULIRG subpopulations, and the maximum-distance metric is sensitive to outliers and infalling halos. Given that these trends are headline results, the paper should report how the conclusions in Section 4.2 and the TolTEC angular-size predictions change under the alternative radius definitions, and discuss whether the contraction and central-concentration findings are robust.","section":"Section 4.2, Figure 8"}],"minor_comments":[{"comment":"In the upper-left panel, the legend labels a dataset as 'Ward+22 (317deg2, lensed candidates)', but the reference list contains Ward et al. (2021) and Ward et al. (2024); please correct the year or add the missing reference.","section":"Figure 1"},{"comment":"Table 1 lists 'Magnification (mu)' with a minimum of 1.2 for both GARDENS-Wide and GARDENS-Deep, but the text in Section 2.2 explains that 1.2 is the minimum amplification for lensed sources only. Please clarify in the table caption that the catalogues include both unlensed (mu=1) and lensed (mu>1.2) populations.","section":"Table 1"},{"comment":"The paper reports median fractions with 16th-84th percentiles but does not state how many independent rich clusters contribute to each redshift bin in Figure 7. With only 32 rich clusters in total, the scatter at high redshift may be driven by a handful of systems. Please state the number of independent systems per bin, or add a note about the statistical weight of the median.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript would benefit from a direct comparison of its predicted z>4 proto-cluster quiescent fractions with observational samples of confirmed z>4 proto-clusters (e.g., starburst-dominated overdensities) and with deep-field quiescent galaxy fractions. As written, the headline decline to ~20 per cent star-forming at z~5 is an extrapolation from a model calibrated at z<=4, and the internal inconsistency between the abstract (20 per cent) and Section 4.1 (30 per cent) should be fixed before publication. The circularity in the number-count validation is also worth addressing explicitly, as it affects how the community will use the catalogue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers something useful: a 100 square degree mock redshift survey of dusty star-forming galaxies, public, built on MDPL2 with SHAM galaxy assignment and calibrated dust emission. What is new is the combination of wide area, proto-cluster assembly histories, and specific TolTEC LSS detection forecasts. The validation against number counts at 500 micron, 1.1, 1.4 and 2.0 mm is honest work, including a sensible blending test at SPT resolution and a resolution check against the deeper GARDENS-Deep catalogue. The authors get credit for releasing the catalogue and for flagging the excess of high-magnification lensed sources, which they attribute to the point-mass lens model.\n\nThe soft spots are real but not fatal to the catalogue. First, the number count validation is partly circular: beta=2.2 was chosen specifically to match the 1.1, 1.4 and 2.0 mm counts, so those matches are not independent predictions. The paper acknowledges the exploration but does not quantify how much freedom the beta distribution leaves. Second, and more important, the high-redshift star-forming fraction trend is an extrapolation. Section 2.1 states the quiescent fraction is calibrated only for 0<=z<=4, yet Figure 7 and the abstract report star-forming fractions out to z~5.5-6, including a decline to ~20% by z~5. The example proto-clusters at z~5.4-5.6 already show ~68-80% quiescent, which is exactly what the uncalibrated quenching model produces. No comparison is made to observed z>4 proto-clusters, which are generally starburst-dominated. That does not invalidate the mock, but the paper should present this as model extrapolation, not a calibrated prediction. Third, the cluster sample is small: 32 rich clusters, and the richest cluster cannot be followed to z=0 because of the lightcone geometry. The size evolution medians have large scatter, and the small-N caveat should be stated more prominently.\n\nWho gets value: anyone working on submillimetre galaxy evolution, proto-cluster identification, or TolTEC survey strategy. The public catalogue alone justifies attention.\n\nRecommendation: send it to peer review with major revision. Require the authors to separate fitted from predicted quantities, to label the z>4 star-forming fraction as extrapolation, to compare against any available z>4 observational proto-cluster data, and to add systematic error estimates for the IR parameters. The methodology is transparent and the authors are competent; this is a solid paper that needs honest framing.","headline":"A genuinely useful public mock catalogue and TolTEC forecasts, but the headline z>4 star-forming fraction trend rests on a quenching model calibrated only to z<=4 and needs to be flagged as extrapolation.","tokens_in":28180,"tokens_out":1930,"would_cite":true,"duration_ms":20092,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 100-square-degree mock survey of dusty star-forming galaxies reproduces observed millimeter counts and predicts that proto-clusters of rich clusters contract from roughly 22 comoving Mpc at $z\\sim5.5$ to about 5.3 comoving Mpc at…","keywords":["dusty star-forming galaxies","proto-clusters","galaxy cluster assembly","submillimetre number counts","semi-empirical galaxy-halo model","mock redshift survey","TolTEC predictions","ULIRGs"],"falsifier":"Compare the mock's TolTEC predictions to the real survey: if the median number of 1.1 mm detectable members in $z\\sim3.2$ rich-cluster proto-clusters is not near 8, or if the median angular separation of those members is far from the predicted 5\\u20136.6 arcmin at $z\\sim2\\text{--}3.5$, the spatial and luminosity mapping fails. A direct spectroscopic check at $z\\sim2$ should find star-forming fractions near 60\\u201365 percent, not the 35 percent of low-redshift clusters.","tokens_in":27126,"feed_emoji":"🔭","tokens_out":9160,"duration_ms":76992,"temperature":0.7,"pith_summary":"This paper presents GARDENS-Wide, a 100 square degree mock redshift survey of dusty star-forming galaxies built by assigning galaxies to dark-matter halos through abundance matching and calibrating their infrared emission. The paper claims that this mock reproduces the observed number counts at 500 $\\mu$m, 1.1, 1.4 and 2.0 mm well enough to trust its spatial predictions. Using the mock's merger histories, it claims that proto-clusters of rich clusters contract from roughly 22 comoving Mpc at $z\\sim5.5$ to about 5.3 comoving Mpc at $z\\sim0$, while star-forming galaxies rise from about 35 percent of members at low redshift to 60\\u201365 percent at $z\\sim2$ before declining. The payoff would be a concrete, testable map of what wide-area submillimetre surveys should see in and around forming clusters.","feed_headline":"Proto-clusters shrink from 22 to 5 Mpc in a new mock sky","feed_subtitle":"A 100-square-degree mock of dusty galaxies predicts what TolTEC will see as clusters assemble.","key_machinery":"The load-bearing machinery is a semi-empirical galaxy-halo connection: subhalo abundance matching links halo circular-velocity history to stellar mass, growth histories set star formation rates, an obscured fraction converts part of that into infrared luminosity, and dust temperatures and a grey-body SED with emissivity index $\\beta=2.2\\pm0.34$ turn luminosities into millimetre fluxes. On top of this, the new analysis relies on merger-tree walking (via the halo finder's UPID\\u2013ID hierarchy and consistent tree associations) to collect every progenitor halo of a present-day cluster into a proto-cluster at each epoch, and defines proto-cluster size as the maximum 3D comoving distance from the central halo to its most distant member halo. A point-mass gravitational lensing model with minimum amplification $\\mu=1.2$ modifies the bright-end number counts. Together these pieces convert a dark-matter simulation into a population of luminous infrared galaxies whose environments and histories can be counted.","core_discovery":"On the paper's own terms, the discovery is a set of evolutionary trends extracted from the wide-area mock: within the assembly histories of systems that become rich clusters, the median proto-cluster radius, measured as the maximum three-dimensional comoving distance between the central halo and the most distant member, decreases from $\\gtrsim20$ comoving Mpc at $z\\gtrsim5$ to approximately 5.3 comoving Mpc at $z\\sim0$ (the physical radius instead grows to a maximum of about 6.6 Mpc at $z\\sim2$ and then contracts). Simultaneously, the median star-forming fraction among members peaks at 60\\u201365 percent at $z\\sim2$, LIRGs contribute 20\\u201340 percent with a similar peak, ULIRGs stay below roughly 10 percent, and HyLIRGs are rarer than 0.5 percent. At $z>1.5$ ULIRGs in rich proto-clusters are centrally concentrated, typically lying within 10\\u201365 percent of the proto-cluster radius, and the paper translates these trends into survey predictions: over 60 square degrees the TolTEC Large-Scale Structure survey should detect about 104,000 sources at 1.1 mm, 50,000 at 1.4 mm, and 11,000 at 2.0 mm, with median redshifts of 2.9, 3.1 and 3.3, and with typical separations between detectable proto-cluster members of 5\\u20136.6 arcmin at $z\\sim2\\text{--}3.5$.","pith_inferences":["If the mock is representative, the predicted angular scales imply that a single-dish camera with arcminute mapping capability can outline proto-cluster structure without interferometric follow-up; the resolved members are sparse but well separated.","The central concentration of ULIRGs at $z>1.5$ suggests that deep pointed observations of the inner 10\\u201365 percent of a proto-cluster radius will recover a disproportionate share of the intense obscured star formation, a strategy the paper does not itself propose.","The $z\\sim2$ peak in star-forming fraction coincides with the cosmic star-formation peak, hinting that the most massive clusters assemble their star-forming populations just before the main epoch of environmental quenching; this causal reading goes beyond the paper's correlations.","The paper's own point-mass lensing model overproduces highly magnified sources, so the predicted bright-end counts and the roughly 300 strongly lensed sources per 60 square degrees are likely upper limits; a more realistic lens population would be a direct test."],"forward_implications":["Wide-area millimetre surveys should search for proto-clusters as extended structures of a few arcminutes to about 16 arcmin, with search apertures matched to the predicted angular radii.","At $z\\sim2$, dusty star-forming galaxies are the dominant tracer of cluster assembly, so submillimetre selection is the most efficient route to proto-cluster discovery in that epoch.","Most detectable star-forming galaxy pairs at 1.1 mm will be blended (about 64 percent), so flux densities of bright compact pairs will be systematically overestimated unless higher-resolution follow-up is used.","TolTEC's sensitivity, not its angular resolution, will limit proto-cluster identification, since typical separations of detectable members are about 5\\u20136.6 arcmin, far above the 5-arcsec beam.","About 72 proto-clusters of future rich clusters (48 at $z\\geq2$, 19 at $z\\geq4$) should be identifiable in a 60 square degree survey before any flux cut, providing a sample for studying the assembly of the most massive systems."],"supporting_citations":[{"why":"Supplies the prior 5.3-square-degree realization whose methodology GARDENS-Wide extends and updates.","marker":"NM24"},{"why":"Supplies the large dark-matter simulation from which the lightcone is built.","marker":"Klypin et al. 2016"},{"why":"The semi-empirical SHAM galaxy-halo connection that assigns stellar masses and SFRs.","marker":"Rodríguez-Puebla et al. 2016a, 2017"},{"why":"Provides the empirical obscured-fraction relation used to set SFRIR from stellar mass and redshift.","marker":"Whitaker et al. 2017"},{"why":"Provides the LIR-to-peak-wavelength scaling that drives assigned dust temperatures.","marker":"Casey et al. 2018"},{"why":"Provides the SFR-LIR conversion that turns obscured star formation into infrared luminosity.","marker":"Kennicutt 1998"},{"why":"Serves as the comparison simulation whose proto-cluster SFRD contribution is benchmarked against.","marker":"Chiang et al. 2017"},{"why":"Provide the halo finder and merger-tree algorithm that define system membership and assembly histories.","marker":"Behroozi et al. 2013a, 2013b"},{"why":"Describes the TolTEC instrument whose survey depths and beams ground the observational predictions.","marker":"Wilson et al. 2020"},{"why":"Defines the TolTEC Large-Scale Structure survey area and depths used for predictions.","marker":"Montaña et al. 2019"}],"fun_headline_variants":["Proto-clusters shrink from 22 to 5 Mpc in mock sky","Star-forming fraction peaks at z=2 in cluster mock","TolTEC to see 104k sources at 1.1 mm per mock","ULIRGs concentrate in rich proto-clusters at z>1.5"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the semi-empirical model, calibrated on average galaxy populations (stellar mass functions, quiescent fractions, the star-forming main sequence, and luminosity functions), remains valid in rare, massive proto-cluster environments at high redshift where those calibration data have little leverage, with the specific SFRIR cap of 6000 $M_\\odot$ yr$^{-1}$ and the chosen dust parameters as part of that same assumption.","fun_headline_variants_meta":{"raw":{"variants":["Proto-clusters shrink from 22 to 5 Mpc in mock sky","Star-forming fraction peaks at z=2 in cluster mock","TolTEC to see 104k sources at 1.1 mm per mock","ULIRGs concentrate in rich proto-clusters at z>1.5"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000354,"raw_usage":{"total_tokens":2024,"prompt_tokens":1145,"completion_tokens":879,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":761,"completion_tokens_details":{"reasoning_tokens":797}},"tokens_in":761,"tokens_out":879,"duration_ms":7770,"temperature":1.0,"reasoning_tokens":797,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:54:55.329343+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the mock's TolTEC predictions to the real survey: if the median number of 1.1 mm detectable members in $z\\sim3.2$ rich-cluster proto-clusters is not near 8, or if the median angular separation of those members is far from the predicted 5\\u20136.6 arcmin at $z\\sim2\\text{--}3.5$, the spatial and luminosity mapping fails. A direct spectroscopic check at $z\\sim2$ should find star-forming fractions near 60\\u201365 percent, not the 35 percent of low-redshift clusters.","supporting_citations":[],"review_version":2}