{"id":"486e5bef-1a95-4c2e-936b-aa6c487c9537","arxiv_id":"2507.04927","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new large cosmological simulation with weak stellar feedback reproduces the stellar masses, sizes, star formation rates, and halo relations of galaxies at z roughly 8 to 12 seen by JWST.","lead":"Phoebos is a new 100 Mpc cosmological simulation of galaxy formation that uses intentionally weak stellar feedback and detailed non-equilibrium gas cooling. It reproduces several JWST-era observations of galaxies at z > 8, supporting the idea that early galaxies grew fast through highly efficient star formation with only mild regulation from feedback.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The size–mass slope claim assumes a mass-independent offset between simulated R_half_mass and observed R_eff, but Phoebos' fixed 0.3 kpc softening makes resolution-driven inflation mass-dependent; no Phoebos resolution series for stellar sizes is shown.","rationale":"Phoebos' central claim is that a 100 Mpc simulation with weak stellar feedback and non-equilibrium cooling reproduces several JWST-era observables, and that this supports efficient early star formation. The most load-bearing evidence for the 'reproduces' statement is the set of agreements in Figures 2, 6, 8, and 13. Among these, the size–mass comparison is the one that explicitly requires an a priori correction of the data before agreement can be claimed. The paper itself states that observed galaxies appear more compact than the simulation and that this is 'likely due' to definitional and resolution effects; the slope claim then depends on the correction being a pure intercept shift. That correction is not derived, and the available high-resolution checks are too sparse to constrain its mass dependence. The reader's weakest assumption identifies exactly this step, and I agree with that identification. I do not see a reason to move the verdict: the paper is already CONDITIONAL, and the concern strengthens the condition rather than overturning the overall assessment. The SMF/SHMR tension with Stefanon et al. (2021) noted in §3.2 is a second unresolved point, but the size-slope issue is the more direct threat to the abstract's headline list of reproduced observables.","tokens_in":30880,"tokens_out":10899,"duration_ms":133078,"concrete_test":"Compute the stellar size–mass relation at z=8 in the PhoebosLR run (with the same star-formation threshold and cooling setup as Phoebos, or in a new LR run with that threshold) and fit log10 R_star,1/2 = a + b log10 M_star; compare b and its scatter with the Phoebos fit. If |b_LR - b_Phoebos| exceeds about 0.1, the size–mass slope is not converged and the mass-independent offset assumption is untenable. As an analytic cross-check, estimate the Bate & Burkert (1997) inflation factor as a function of R_star,1/2/epsilon for epsilon=0.3 kpc over the mass range in Figure 6; if the fractional inflation varies substantially across that range, a constant y-intercept shift cannot recover the intrinsic slope. The test requires making the z=8 lower-resolution snapshots and galaxy catalogs available, consistent with the stated data-availability policy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Figure 6 the paper claims that Phoebos 'recovers the observed slope' of the stellar size–mass relation (abstract; §3.1) by comparing simulated stellar half-mass radii with observed effective radii and showing the Yang et al. (2025) relation with 'a modified y-intercept to better align with the simulations.' The justification is that the offset is a mass-independent 25 per cent half-light-to-half-mass difference plus a resolution-driven puffing of the same kind as Bate & Burkert (1997). This is the load-bearing step. The softening is fixed at 0.300 kpc (physical, z<9; Table 1) while the simulated radii in Figure 6 span roughly 0.1–1 kpc, so a fixed softening should inflate the smallest, lowest-mass galaxies proportionally more than the largest ones. The offset is therefore expected to be mass-dependent, and a constant y-intercept shift cannot establish the slope unless that mass dependence is shown to be negligible. The two higher-resolution simulations cited (GigaEris, MassiveBlackPS) provide only a few points and do not constitute a Phoebos resolution sequence for sizes. The lower-resolution runs in Appendix A are used only for halo mass functions, not for stellar sizes or stellar mass functions. If the offset is mass-dependent, the claimed slope recovery is not established, and the abstract's list of reproduced observables loses one member. The concern is about an unverified assumption, not about the simulation's other merits: the lack of redshift-dependent tuning and the multiple independent comparisons are genuine supporting evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Phoebos, a [100 cMpc]^3 cosmological hydrodynamical simulation with particle masses mDM=1.360e6 Msun and mgas=8.473e5 Msun, designed to study galaxy formation at z~8-12. It uses ChaNGa with multi-phase, non-equilibrium cooling, stochastic star formation with epsilon_SF=0.1, and blastwave supernova feedback, deliberately avoiding an effective equation of state and AGN feedback. The authors compare the simulation with JWST-era observations of the stellar mass function, stellar-to-halo mass relation, cosmic star formation rate density, star-forming main sequence, specific star formation rate, stellar size-mass relation, and HI fractions. They report good agreement for the stellar mass function, SHMR, and sSFR, claim recovery of the observed size-mass slope, and interpret the results as evidence for highly efficient star formation with only weak stellar feedback at z>8. Lower-resolution runs are used for halo-mass-function convergence, and the paper discusses limitations including SFRD overprediction and the need for stronger feedback at lower redshifts.","tokens_in":31227,"tokens_out":6829,"duration_ms":78471,"significance":"If the main claims hold, Phoebos would be a valuable resource: a large-volume simulation that reproduces several independent high-redshift observables without redshift-dependent recalibration, thereby supporting the emerging picture of a weak-feedback, efficient star-formation regime at cosmic dawn. The paper is honest about several discrepancies (e.g., SFRD relative to Bouwens et al. 2023, SHMR relative to Paquereau et al. 2025, the HI fraction drop) and provides convergence checks for halo mass functions. The central interpretive claim, however, partly rests on the size-mass slope comparison, whose validity currently depends on an unverified assumption about a mass-independent offset between simulated half-mass radii and observed effective radii. The no-tuning claim is also complicated by a redshift-dependent cooling floor unique to the Phoebos run. These issues are addressable but must be resolved before the paper's headline conclusions can be fully accepted.","major_comments":[{"comment":"The claim that Phoebos 'recovers the observed slope' of the stellar size-mass relation is established by shifting the y-intercept of the Yang et al. (2025) relation (the 'loosely dotted line' in Fig. 6) to align with the simulated half-mass radii. The shift is justified by a mass-independent ~25 per cent half-light-to-half-mass offset plus resolution-driven puffing, but the Phoebos softening is fixed at 0.300 kpc physical for z<9 (Table 1), while the simulated radii in Fig. 6 span roughly 0.1-1 kpc. A fixed softening should inflate the smallest, lowest-mass galaxies proportionally more than the largest ones, making the net offset mass-dependent unless shown otherwise. The GigaEris and MassiveBlackPS points have different softening and resolution and do not constitute a Phoebos resolution sequence for stellar sizes; the PhoebosLR/ULR runs in Appendix A are used only for halo mass functions. Because this is one of the headline successes listed in the abstract, I ask the authors to quantify the mass dependence of the resolution offset (e.g., from a size resolution sequence or from high-resolution zooms of representative Phoebos systems) or to compare mock-observed effective radii before claiming slope recovery.","section":"§3.1, Fig. 6; also Abstract and §4"},{"comment":"Section 4 states that Phoebos succeeds 'without any redshift-dependent calibration or tuning to high-redshift data.' However, Section 2 reports that the Phoebos run uses a cooling temperature floor of 300 K up to z~10, whereas PhoebosLR and PhoebosULR use a floor of 10 K. This is a redshift-dependent modification specific to the main run, and its effect on the z>8 star formation histories, stellar masses, and sizes is not discussed. Please either demonstrate robustness of the headline results to this floor or qualify the no-tuning claim accordingly.","section":"§2, footnote 6; §4"}],"minor_comments":[{"comment":"There is a missing space in 'smoothed-particlehydrodynamics code ChaNGa', and the axis labels use 'M' where 'M☉' would be clearer.","section":"§2, Table 1 and Fig. 2"},{"comment":"Footnote 8 is grammatically incomplete: 'For a more detailed discussion see e.g. Kravtsov (2013) and Szomoru et al. (2013), the half-light radius...' should be rewritten as one or more complete sentences.","section":"§3.1, footnote 8"},{"comment":"There are typos: 'the radii within Phoebos should be seen considered as an upper limit' and 'Similary, the haloes at z = 12'; please correct both.","section":"§3.1"},{"comment":"In the sentence 'especially at = 12, Thesan-Zoom and Phoebos...' the variable z is missing; it should read 'especially at z = 12'.","section":"§3.3.1"},{"comment":"The phrase 'treatment of of a relatively weak stellar feedback' contains a duplicated 'of'.","section":"§4"},{"comment":"The caption spells 'MilleniumTNG'; it should be 'MillenniumTNG'.","section":"Figure 1 caption"},{"comment":"The text says the origin of the HI drop 'will explore potential explanations in subsequent sections,' but the subsequent discussion is brief and confined to Appendix A2; please either expand the discussion or adjust the forward reference.","section":"§3.4 and Appendix A2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript fits the scope of MNRAS and the simulation is likely to be useful to the community. The main risk is the size-mass slope claim, which currently rests on an unverified mass-independent offset; I would not block publication once the authors provide a quantitative resolution study or explicitly reframe the claim. The special 300 K cooling floor in the Phoebos run should also be clarified, as it bears on the 'no tuning' narrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, Phoebos is a genuinely useful new resource: a 100 cMpc volume at ~8.5e5 M_sun gas resolution, run with non-equilibrium primordial cooling, weak blastwave feedback, and no AGN. That combination is new at this volume, and it makes the paper a natural reference point for the JWST overmassive-galaxy debate. Second, the paper's headline about reproducing the observed size–mass slope is the one claim I would not take at face value yet. The stress-test note is right: with a fixed 0.3 kpc softening, resolution-driven size inflation should be stronger in smaller, lower-mass galaxies, so a constant y-intercept shift cannot establish the slope unless the mass dependence is shown to be negligible. The paper points to GigaEris and MassiveBlackPS as supporting evidence, but those are different simulations, not a Phoebos resolution series. That is a genuine gap, though it is a gap in one figure's interpretation, not in the whole simulation.\n\nWhat the paper does well: the multiple independent comparisons are the strongest part. The SMF, SHMR, sSFR, and SFR–stellar mass slope all land within scatter of recent JWST-era observations, and the subgrid parameters were inherited from lower-redshift calibrations rather than tuned to z > 8. That makes the broad agreement a real prediction, not a fit. The authors also disclose the known tensions: SFRD overpredicts Bouwens, SHMR overpredicts Paquereau, and the sharp HI drop around 1e8 M_sun is left unexplained. That is honest and useful.\n\nThe soft spots beyond the size slope: the comparisons are mostly qualitative, with no quantitative goodness-of-fit; the adjusted y-intercept for Yang's relation is a visible fudge; and the data statement (“on reasonable request”) is worse than the field standard. None of these undercut the central point that Phoebos is a credible weak-feedback baseline for early galaxy formation. The size claim just needs more work.\n\nWho this is for: anyone working on high-redshift galaxy formation, JWST number counts, or simulation comparison papers. It will get cited as the Phoebos reference. It deserves a serious referee: the simulation is real, the claims are mostly testable, and the weaknesses are fixable. I would send it to review and ask the authors to add a resolution test for stellar sizes or at least clearly parametrize the mass dependence of the offset before the slope claim goes into the abstract.","headline":"Phoebos is a valuable new simulation resource and the weak-feedback baseline is worth taking seriously, but the size–mass slope claim rests on an unverified mass-independent offset assumption that needs addressing before that result is secure.","tokens_in":31792,"tokens_out":1289,"would_cite":true,"duration_ms":17135,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 100 Mpc simulation with weak stellar feedback reproduces JWST-era galaxy counts, sizes, and star formation at z>8.","keywords":["Phoebos simulation","high-redshift galaxies","galaxy formation","stellar feedback","galaxy sizes","star formation efficiency","cosmic dawn","JWST"],"falsifier":"Forward-model the Phoebos galaxies at z~8 into mock JWST images, fit effective radii from the light, and compare the resulting size–mass relation with the observed one without any free shift of the y-intercept; if the slope or normalization still disagrees, the central claim of reproducing the observed size–mass relation is not established.","tokens_in":30601,"feed_emoji":"🔭","tokens_out":16014,"duration_ms":135679,"temperature":0.7,"pith_summary":"The paper introduces Phoebos, a 100 Mpc cosmological hydrodynamical simulation built to test whether weak stellar feedback can explain the abundance and properties of massive galaxies seen by JWST at $z\\gtrsim8$. At variance with most large-volume simulations, Phoebos uses non-equilibrium radiative cooling without an effective equation of state, and its stellar feedback is intentionally mild. The central claim is that this setup reproduces the observed stellar mass function, stellar-to-halo mass relation, size–mass slope, and specific star formation rate without any redshift-dependent calibration, supporting a picture of rapid, highly efficient star formation during cosmic dawn. The authors also report tentative signs that at lower redshift Phoebos overproduces cosmic star formation, suggesting feedback must strengthen with time.","feed_headline":"Simulation reproduces early-galaxy data with weak stellar feedback","feed_subtitle":"Phoebos, a 100-Mpc simulation run without calibration, matches JWST stellar masses, sizes, and star formation at z>8.","key_machinery":"The central object is the Phoebos simulation itself: a [100 cMpc]^3 box with $2904^{3}$ dark-matter and $1944^{3}$ gas particles, with dark-matter and gas particle masses of $1.360\\times10^6\\,M_\\odot$ and $8.473\\times10^5\\,M_\\odot$, run with a smoothed-particle hydrodynamics code using a Wendland C4 kernel, non-equilibrium primordial cooling with self-shielding, tabulated metal-line cooling, no pressure floor, and stochastic star formation at a density threshold of $1.0\\,m_p\\,\\mathrm{cm^{-3}}$ and efficiency 0.1. The feedback recipe is a blastwave supernova model whose energy injection is weak in the dense, rapidly cooling gas of $z\\gtrsim8$ galaxies, because the cooling and dynamical times are shorter than the typical timescale between supernova explosions. That short-timescale argument, rather than a tuning to high-redshift data, is the load-bearing mechanism that lets Phoebos form galaxies fast enough to match JWST observations.","core_discovery":"Phoebos is claimed to reproduce, at $z\\gtrsim8$ and without any redshift-dependent calibration, the observed stellar mass function, the stellar-to-halo mass relation, the slope of the stellar size–mass relation, and the specific star formation rate distribution. At $z\\sim12$ it matches the highest-redshift stellar-to-halo mass constraints better than other current simulations, while at $z\\sim8$ its stellar mass function and star formation rates sit within the observational scatter. After accounting for a ~25 per cent offset between half-light and half-mass radii and for resolution-driven size inflation, the slope of the simulated size–mass relation tracks the observed one up to stellar masses around $10^{9.5}\\,M_\\odot$, in contrast to a strong-feedback comparison simulation. The paper interprets this as evidence that early galaxy growth happens in a weak-feedback regime where gas cools and collapses faster than supernova feedback can respond, making star formation highly efficient.","pith_inferences":["A direct test of the size claim would be to forward-model Phoebos galaxies into mock JWST images and measure effective radii from the light, rather than assuming a mass-independent 25 per cent offset; if the slope changes when the shift is fitted freely, the match would not survive.","The redshift-dependent feedback transition implied by the star-formation excess could be checked by running Phoebos to $z=0$ and comparing its stellar mass function with the local one; a large overprediction would confirm that a stronger low-redshift feedback mode is needed.","The weak-feedback, fast-assembly scenario predicts low outflow mass loading and possibly high escape fractions of ionizing radiation from $z\\gtrsim8$ galaxies, which future 21-cm reionization observations and JWST/NIRSpec metallicity measurements could test.","If feedback effectiveness is as environment- and redshift-dependent as this simulation suggests, simulations calibrated mainly to the local universe may systematically mispredict the earliest galaxies, and the slope of the stellar mass function at $z\\gtrsim12$ becomes a sharp discriminator."],"forward_implications":["The abundant massive galaxies JWST finds at $z\\gtrsim10$ can be understood within standard structure formation as the product of highly efficient star formation, without requiring modified cosmology or exotic accretion physics.","Once galaxy sizes are defined consistently, early galaxies grow in radius with stellar mass along the observed slope, with the overall normalization set by the half-light/half-mass offset and resolution inflation.","Galaxies at $z\\gtrsim8$ assemble their stellar mass on timescales shorter than the Hubble time, so rapid assembly is the rule rather than the exception during cosmic dawn.","The predicted excess of cosmic star formation at lower redshift implies that the weak-feedback regime must be temporary, with feedback becoming stronger at later epochs.","Resolving the multi-phase interstellar medium with non-equilibrium cooling appears to be at least as important as the strength of stellar feedback for producing the high-redshift galaxy population."],"supporting_citations":[{"why":"Provides the observed size–mass relation at z~7.5 whose slope Phoebos claims to reproduce after shifting the y-intercept.","marker":"Yang et al. (2025)"},{"why":"Supplies the observed SFR–stellar mass, sSFR, and size data used for comparisons at z~8–12.","marker":"Morishita et al. (2024)"},{"why":"Provides the z~8 stellar mass function and stellar-to-halo mass relation that Phoebos claims to match.","marker":"Stefanon et al. (2021)"},{"why":"Gives the z~8–13 observed stellar mass function used to check the simulation's number densities.","marker":"Harvey et al. (2025)"},{"why":"Underlies the argument that resolution inflates simulated galaxy sizes when smoothing is small relative to softening.","marker":"Bate & Burkert 1997"},{"why":"Supports the ~25 per cent offset between half-light and half-mass radii used to align observed and simulated sizes.","marker":"Kravtsov (2013)"},{"why":"Supplies the blastwave supernova feedback recipe whose weakness in dense gas is the paper's physical mechanism.","marker":"Stinson et al. (2006)"},{"why":"Provides the feedback-free galaxy formation scenario invoked to interpret the simulation's efficient star formation.","marker":"Dekel et al. (2023)"},{"why":"Gives the Thesan-based size–mass and SFR relations used as the main simulation comparison that Phoebos outperforms.","marker":"Shen et al. (2024)"},{"why":"Provides the high-redshift stellar-to-halo mass relation constraint used to validate Phoebos at the high-mass end.","marker":"Shuntov et al. (2025)"}],"fun_headline_variants":["Phoebos simulation reproduces early galaxies with weak feedback","Weak feedback in Phoebos matches JWST galaxy findings","No-calibration Phoebos matches early galaxy data","Phoebos: weak feedback drives early galaxy formation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the difference between observed effective radii and simulated half-mass radii is dominated by a mass-independent offset of about 25 per cent plus resolution-driven size inflation, so that after shifting the observed relation's zero-point the slopes can be compared; if the offset varies with mass or is partly physical, the claimed size–mass slope match does not follow.","fun_headline_variants_meta":{"raw":{"variants":["Phoebos simulation reproduces early galaxies with weak feedback","Weak feedback in Phoebos matches JWST galaxy findings","No-calibration Phoebos matches early galaxy data","Phoebos: weak feedback drives early galaxy formation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000441,"raw_usage":{"total_tokens":2300,"prompt_tokens":1074,"completion_tokens":1226,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":690,"completion_tokens_details":{"reasoning_tokens":1161}},"tokens_in":690,"tokens_out":1226,"duration_ms":9906,"temperature":1.0,"reasoning_tokens":1161,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:36:55.529145+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Forward-model the Phoebos galaxies at z~8 into mock JWST images, fit effective radii from the light, and compare the resulting size–mass relation with the observed one without any free shift of the y-intercept; if the slope or normalization still disagrees, the central claim of reproducing the observed size–mass relation is not established.","supporting_citations":[{"cited_title":"The THESAN project: galaxy sizes during the epoch of reionization","cited_arxiv_id":"2402.08717","evidence_quote":"Gives the Thesan-based size–mass and SFR relations used as the main simulation comparison that Phoebos outperforms."}],"review_version":1}