{"id":"0fbd2401-77ba-4ad9-a403-7b7da4287d35","arxiv_id":"2411.14683","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"High-resolution simulations of six ultra-faint dwarf galaxies reproduce observed metallicities better when supernova progenitors are sampled individually, but they still cannot form enough stars with [Fe/H] above -2.","lead":"This paper runs high-resolution simulations of six ultra-faint dwarf galaxies and finds that individually sampling supernova progenitors raises simulated stellar metallicities closer to observed values. It also shows that applying the same profile-fitting methods as observers shrinks simulated galaxy sizes, narrowing a long-standing mismatch.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The size-match claim compares the inner component's scale radius (1.68 r_e) to observed single-component half-light radii, but the two-component model's actual half-light radius is much larger; this is not a like-for-like comparison.","rationale":"The reader's weakest assumption concerns reionization and instantaneous feedback, which the authors themselves flag as making their metallicities an upper limit. That concern is real but already incorporated into the conditional verdict. The size claim, in contrast, is presented as a direct methodological improvement without a comparable caveat, yet it appears to compare an inner-component scale radius to observed total half-light radii. This is an internal-consistency issue rather than a disagreement with consensus, and it can be settled by a simple numerical check. If the check confirms that the quoted r_h is not the two-component model's half-light radius, the paper's second main pillar is substantially weakened, requiring the authors to reframe the size comparison or re-fit the observed data with the same two-component method. The paper still contains valuable results on metal enrichment and on how the choice of fitting method changes measured sizes, so rejection is not warranted; a conditional acceptance with mandatory revision is the appropriate outcome, matching the reader's verdict.","tokens_in":38839,"tokens_out":16028,"duration_ms":162416,"concrete_test":"For each UFD analog, numerically solve 2π∫_0^R [exp(-r/r_e) + B exp(-r/r_s)] r dr = 0.5 * 2π(r_e^2 + B r_s^2) using the best-fit r_e, r_s, and B from Table 3, and compare the resulting model half-light radius R_half,model with the quoted r_h = 1.68 r_e and with r_h,dir from Table 2. If R_half,model exceeds the quoted r_h by more than a factor of two for Halo4 or Halo5, then the Fig. 17 comparison is not measuring the same quantity as the observed single-component half-light radii, and the size-match claim requires re-evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's size conclusion (Section 3.3.1, Fig. 17, Table 3) relies on reporting r_h = 1.68 r_e for the inner component of the two-component profile in Eq. (6), Σ(r) ∝ exp(-r/r_e) + B exp(-r/r_s). For a two-component model, the radius enclosing half the total model light is not 1.68 r_e unless B = 0. The Table 3 parameters imply the outer component contains a comparable or dominant share of the total light: for Halo4, B(r_s/r_e)^2 ≈ 0.009*(850/70)^2 ≈ 1.3, and for Halo5 it is roughly 16. Numerically integrating the Halo4 profile gives a true model half-light radius near 500 pc, not the quoted 117 pc. Thus the 'closer match' to observed UFD sizes is obtained by comparing a component-scale radius—not the galaxy's half-light radius—to observed half-light radii that are derived from single-component fits. The one-component fitting with background, which is the actual observational method used for the comparison catalogs, gives r_h,fit(w/) ≈ 180 pc for Halo4 (Table 2), still larger than typical observed UFDs. The central size claim is therefore vulnerable to being a definitional artifact.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents high-resolution cosmological zoom-in simulations of six ultra-faint dwarf galaxy (UFD) analogs with gas mass resolution ~60 M_sun, focusing on stellar metallicities and sizes. The central claim is that individually sampling supernova progenitors from the IMF raises the average stellar metallicity of the simulated UFDs by 1-1.5 dex compared to earlier simulations, bringing them closer to (though still below) the observed mass-metallicity relation. The MDFs are compared to observations after excluding [Fe/H] < -4 stars, which the authors argue align well with observed UFDs. For sizes, the authors show that applying observational profile-fitting methods, especially a two-component exponential profile, yields half-light radii substantially smaller than direct mass-based estimates, improving agreement with observations, though the most compact observed systems (r_h < 50 pc) are not reproduced. The paper also discusses the role of multiple progenitor halos and dry mergers in shaping stellar properties and extended structures.","tokens_in":39217,"tokens_out":3012,"duration_ms":31908,"significance":"If the central claims hold, the paper makes a useful contribution to the ongoing problem of reconciling simulations of UFDs with observations. The individual IMF sampling for SNe is a promising methodological improvement and is honestly tested against parameter variations. The emphasis on applying consistent observational fitting methods to simulation output is timely and important, as is the comparison to the two-component profile of Jensen et al. (2024). The simulations are forward models with subgrid parameters taken from local calibrations or literature rather than fitted to the UFD data, which strengthens the credibility of the qualitative trends. However, the quantitative claims rest on a small sample of six halos, and the size comparison contains a definitional issue that undermines the strongest size-related statement.","major_comments":[{"comment":"The quoted r_h values from the two-component fit are the effective radii of the inner component only (r_h = 1.68 r_e), not the half-light radius of the total model. In a two-component profile of the form Sigma(r) = exp(-r/r_e) + B exp(-r/r_s), the radius enclosing half the total model light is not 1.68 r_e when B > 0. Using the Table 3 parameters for Halo4 (e.g., r_e ~ 70 pc, r_s = 850 pc, B = 0.009), the outer component contains a comparable or dominant share of the light, and a direct numerical integration gives a true model half-light radius of several hundred parsecs rather than 117 pc. Comparing the inner component's scale radius to observed half-light radii (which are derived from single-component fits) is therefore not a like-for-like comparison, and the visual impression in Fig. 17 that the two-component method 'closely matches' observed sizes is largely an artifact of this definition. The authors should either compute and report the actual half-light radius of the two-component model, or explicitly frame the quoted r_h as the inner component's scale radius and restrict the comparison to observations analyzed with a two-component profile.","section":"Section 3.3.1, Table 3, Eq. (6), Fig. 17"},{"comment":"Halo6 is excluded from the size-luminosity analysis with the stated justification 'to avoid redundancy' because its halo mass is similar to Halo5. This is not a physical or statistical selection criterion, and dropping a data point for this reason can bias the comparison, especially when the sample contains only five halos. The authors should either include Halo6 in Fig. 17 and Table 2 (showing its size alongside the others), or provide a quantitative reason (e.g., b/b_max similarity or a pre-defined selection rule) for the exclusion.","section":"Section 3.3.1, paragraph on Halo6"},{"comment":"The claim of an 'excellent match' with observed MDF parameters is conditional on excluding stars with [Fe/H] < -4, which is an observational completeness limit. While the authors are transparent about this, the central MZR improvement claim (Fig. 6) is based on only six halos, and the simulated averages remain 0.5-1 dex below observed values for all but Halo1. The paper should more prominently separate the two statements: (i) the IMF-sampling mechanism raises metallicities relative to previous simulations, which is supported by the internal comparison, and (ii) the simulated UFDs actually match the observed MZR, which is not supported by the current data because the offsets and small sample size prevent a statistically meaningful match. The discussion in Section 4 partially addresses this, but the abstract and conclusions should be calibrated to the evidence.","section":"Section 3.2.1 and Fig. 8"},{"comment":"The modeling assumptions of uniform reionization at z=6 and instantaneous SN feedback (no delay time, no radiative transfer) are acknowledged by the authors, who state that including radiative transfer would likely lower the average metallicity and that their values may be an upper limit. This is an honest and important caveat, but it also means that the claimed improvement over previous simulations could be reduced or reversed under a more complete feedback treatment. The authors should consider adding a direct quantitative estimate of the sensitivity to the delay-time treatment (e.g., a test run with delayed SNe but no RT) or at least clearly mark the metallicity predictions as upper limits in the abstract and in Fig. 6, since the current abstract presents the higher metallicities as the main result without this qualification.","section":"Sections 2.1, 2.3, and 4"}],"minor_comments":[{"comment":"The two bullets beginning 'We find extended structures of varying degrees in all halos of our UFD analogs' are duplicated verbatim; one of them should be removed.","section":"Section 4, bullet list"},{"comment":"The artificial background star density Sigma_b,0 = 0.1 arcmin^-2 is described as 'arbitrary but representative.' Since the resulting r_h,fit(w/) depends on this choice and the paper also shows that the fitted size varies with background density, the authors should provide a short sensitivity test or a literature-based justification for the adopted value beyond the Draco reference.","section":"Section 3.3.1, Eq. (3) and Table 2"},{"comment":"The comparison of sigma_[Fe/H] for observed UFDs is made against values derived from a two-dimensional Gaussian likelihood on the simulation MDFs, but the observed sample has small number statistics and detection limits; the authors could state whether the observational uncertainties are comparable to the reported differences of ~0.2 dex.","section":"Section 3.2.2, Fig. 8"},{"comment":"The uniform V-band mass-to-light ratio of 2 applied to simulated galaxies is a simplification; a brief comment on the uncertainty this introduces in the luminosity values used in Fig. 17 would be helpful.","section":"Section 3.3.3"},{"comment":"The paper would benefit from a table explicitly listing all subgrid parameters (epsilon_ff, n_H,th, Z_crit, IMF slopes, mass ranges, SN energy, and the adopted background density) so that the reader can quickly assess the free parameters. Many are already given in the text, but a consolidated summary table would improve readability.","section":"Throughout"},{"comment":"In Eq. (1), n_H is used without explicitly stating its units; the text later refers to n_H,th = 100 cm^-3, but the equation should include the normalization unit for clarity.","section":"Section 2.2, Eq. (1)"}],"recommendation":"major_revision","confidential_remarks":"The definitional issue with the two-component half-light radius in Section 3.3.1 is the most serious technical concern and should be resolved before publication; the authors need to either recompute the true half-light radius or reframe their comparison. The paper otherwise presents an honest and useful exploration of UFD formation with a novel IMF-sampling method, but the small sample (six halos, with one excluded from the size analysis) and the acknowledged modeling limitations (instantaneous feedback, uniform reionization) mean that the strongest quantitative claims should be tempered. The paper fits well within the scope of ApJ and the topic is of broad interest to the dwarf galaxy community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth reading for the metallicity section; the size section is not ready as written. The individual IMF sampling for SN feedback is a genuinely new and clearly explained step, and the resulting 1–1.5 dex rise in average metallicities, with the physical story about discrete SNe allowing stars to form from recently enriched gas, is plausible and an advance over earlier simulations. The multi-progenitor assembly picture is well documented, and the authors are honest about their approximations: instantaneous SN feedback, uniform reionization, no radiative transfer, and they note their metallicities are likely upper limits.\n\nThe soft spots are real. The size claim, central to the abstract and Fig. 17, does not survive scrutiny. They fit the two-component profile of Eq. (6) and report r_h = 1.68 r_e for the inner component. But for Halo4 and Halo5 the outer component carries a comparable or dominant share of the total light (B(r_s/r_e)^2 is roughly 1.3 and 16, respectively). The actual half-light radius of that two-component model is much larger—I get ~500 pc for Halo4, not the quoted 117 pc. Comparing an inner-component scale radius to observed single-component half-light radii is not like-for-like, so the apparent resolution of the size discrepancy is mostly a definitional effect. The one-component fits with background (Table 2) give r_h,fit ~180 pc for Halo4, still larger than typical observed UFDs; the conclusion that the two-component method brings simulated sizes into agreement is therefore overstated.\n\nOther limitations are proportional: only six halos, one dropped from the size analysis; the MDF match relies on excluding [Fe/H] < -4 stars; the background density in the fitting is hand-chosen and affects results; no public data or code. These are not fatal for the metallicity part, but they cap the significance.\n\nRecommendation: send it to peer review. The metallicity result is a solid incremental contribution that deserves referee time. The size analysis needs major revision (report the actual model half-light radius, or explicitly frame the inner component as a separate quantity and compare accordingly) before publication. I would not desk reject, but I would not cite the size claim in its current form.","headline":"The metallicity story is credible and worth engaging, but the size claim is a definitional artifact: the quoted 'half-light radii' are inner-component scale radii, not the model's actual half-light radii.","tokens_in":39766,"tokens_out":2244,"would_cite":false,"duration_ms":25910,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"High-resolution cosmological simulations of ultra-faint dwarf galaxies show that individually sampling supernova progenitors from the initial mass function raises average stellar metallicities by 1–1.5 dex, and that using observational…","keywords":["ultra-faint dwarf galaxies","cosmological simulations","mass-metallicity relation","metallicity distribution function","half-light radius","supernova feedback","reionization","galaxy formation"],"falsifier":"Rerun the Halo4 zoom-in with delayed supernova feedback and radiative transfer while keeping all other settings fixed. The paper predicts the average stellar metallicity will fall below its fiducial $[\\mathrm{Fe/H}] \\approx -2.65$; if it instead stays level or rises, the claim that individual IMF sampling is what lifts simulated UFDs toward the observed mass-metallicity relation would be falsified.","tokens_in":38659,"feed_emoji":"🌌","tokens_out":13900,"duration_ms":124566,"temperature":0.7,"pith_summary":"This paper tries to establish that two long-standing mismatches between simulated and observed ultra-faint dwarf galaxies (UFDs)—too little iron and too large sizes—can be substantially reduced by changing how supernovae and galaxy sizes are treated. In six high-resolution cosmological zoom simulations with gas-particle masses near 60 solar masses, the authors replace the usual practice of releasing all supernova energy from a single stellar-population particle with a scheme that samples individual massive stars from the initial mass function and lets each one explode separately. That change raises the average stellar metallicity of the simulated galaxies by 1–1.5 dex relative to earlier simulations, bringing them closer to the observed mass-metallicity relation, though still 0.5–1 dex metal-poor. On the size side, the paper shows that fitting the same exponential density profiles observers use, especially a two-component profile, gives half-light radii of roughly 85–145 pc, considerably smaller than direct mass-based estimates and close to the upper range of observed UFDs. The paper concludes that part of the gap is real and part is a methodological artifact, while the most compact observed systems ($r_h \\lesssim 50$ pc) remain unexplained.","feed_headline":"One simulation tweak lifts faint dwarf metallicities 1.5 dex","feed_subtitle":"Individual SN sampling and observational profile fitting narrow both gaps; the smallest dwarfs stay out of reach.","key_machinery":"The argument runs on two pieces of machinery. The first is individual IMF sampling: rather than treating a 60-solar-mass star particle as a single stellar population that releases all supernova energy at once, the code draws individual stars from a Salpeter IMF (a standard stellar mass distribution) and lets each star in the 8–40 solar-mass range explode as a separate core-collapse supernova, while pair-instability supernovae cover 140–260 solar-mass Population III stars. This discrete injection of energy and metals is what raises the average metallicity, because new stars form from gas recently enriched by one or two explosions instead of being ejected by a combined superbubble. The second is a two-component exponential surface-density profile, $\\Sigma(r) \\propto e^{-r/r_e} + B e^{-r/r_s}$, fitted by maximum likelihood; the inner scale radius gives a half-light radius $r_h = 1.68 r_e$, while the outer scale radius $r_s \\sim 0.6$–$1.7$ kpc captures the extended stellar halo produced by dry mergers of progenitor halos. This profile converts simulated star particles into the same observable used for real UFDs and is what shrinks the derived sizes.","core_discovery":"The paper's central claim is that the stellar metallicities and sizes of ultra-faint dwarf galaxies in cosmological simulations are governed by two things previous simulations did not handle correctly: the discrete, star-by-star nature of supernova enrichment and the method used to measure a galaxy's size. With individual IMF sampling, stars form from gas that has been enriched by one or a few nearby supernovae, rather than being overwhelmed by the combined feedback of an entire stellar population, and the simulated galaxies reach average metallicities of $\\langle [\\mathrm{Fe/H}]\\rangle \\approx -2.2$ to $-3.0$, about 1–1.5 dex higher than earlier simulation suites. The same simulations still lack the observed population of relatively metal-rich stars with $[\\mathrm{Fe/H}] \\geq -2$, because cumulative supernova feedback together with reionization quenches star formation before enough metals accumulate; the maximum values reached are $[\\mathrm{Fe/H}]_{\\max} \\approx -1.5$ to $-1.6$ in the most favorable starbursts. For sizes, the paper argues that the discrepancy is largely a measurement artifact: direct half-mass radii are inflated when stars are spread over several merged progenitor halos, whereas the observational maximum-likelihood fit to a single exponential profile, and even better a two-component exponential profile, yields half-light radii of $r_h \\approx 85$–$145$ pc that sit at the upper end of observed UFD sizes. The most compact observed UFDs, with $r_h \\lesssim 50$ pc, are still not reproduced.","pith_inferences":["If the paper's radiative-transfer caveat is right, then matching observed UFD metallicities may require patchy reionization or a top-heavy IMF; this predicts that UFDs with similar stellar masses but different reionization histories should differ systematically in average [Fe/H].","The two-component fitting result implies that single-component fits in observational catalogs may systematically miss a diffuse outer component; re-fitting existing UFD photometry with two exponentials could reveal extended structures in systems currently classified as compact.","The strong correlation between the number of supernovae during a starburst and the maximum metallicity reached suggests a stochastic, environment-driven ceiling on enrichment, so the scatter in the UFD mass-metallicity relation carries information about local gas density rather than only halo mass.","The simulated absence of a metallicity gradient—because high-density gas blobs, not stars, migrate outward—can be tested by measuring abundances of stars beyond roughly three half-light radii in galaxies like Tucana II; a confirmed gradient would require a formation channel this simulation lacks."],"forward_implications":["Individually sampling supernova progenitors raises average UFD metallicities by 1–1.5 dex relative to earlier simulations, so the mass-metallicity gap is roughly halved rather than closed.","Comparing simulations and observations with identical profile-fitting methods is necessary; single-exponential fits without background stars overestimate half-light radii by up to a factor of six.","A two-component exponential profile is the more faithful description of simulated UFDs, capturing both a compact inner galaxy and an extended outer halo; applying it to observed galaxies could reveal hidden outer components.","Dry mergers between progenitor halos, not tidal stripping, produce extended stellar structures and erase metallicity gradients in isolated UFD analogs.","Stars with $[\\mathrm{Fe/H}] \\geq -2$ remain hard to form because supernova feedback and reionization quench star formation quickly; reducing supernova energy can produce them but overproduces stellar mass."],"supporting_citations":[{"why":"Supplies the ultraviolet background that turns on at z=7 and quenches star formation by z=6.","marker":"Haardt & Madau (2012)"},{"why":"Provides the observed UFD stellar metallicities and the maximum-likelihood MDF fitting method used for comparison.","marker":"Fu et al. (2023)"},{"why":"Introduces the two-component stellar density profile and the evidence for extended outer components in observed dwarfs.","marker":"Jensen et al. (2024)"},{"why":"Supplies the maximum-likelihood exponential profile fitting method used to derive half-light radii from star counts.","marker":"Martin et al. (2008)"},{"why":"Provides the thermal supernova energy injection scheme that avoids over-cooling in the interstellar medium.","marker":"Dalla Vecchia & Schaye (2012)"},{"why":"Gives the stellar mass distribution from which individual supernova progenitors are sampled.","marker":"Salpeter (1955)"},{"why":"Supplies the metal yields of Population III core-collapse supernovae.","marker":"Heger & Woosley (2010)"},{"why":"Supplies the metal yields of Population II core-collapse supernovae.","marker":"Portinari et al. (1998)"}],"fun_headline_variants":["Star-by-star supernovae lift dwarf galaxy metallicities 1.5 dex","Single-supernova sampling fixes UFD metallicity mismatch","Simulated UFDs close metallicity gap with individual SN","Size mismatch in UFDs traced to measurement method","New simulation narrows UFD metallicity and size gaps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results depend on the assumption that the early Universe was completely reheated and reionized all at once by redshift 6, and that supernovae dump their energy instantly with no radiation transport; if real reionization was patchy or supernova feedback was delayed and radiative, the simulated metallicities could shift enough to erase the claimed agreement with observations.","fun_headline_variants_meta":{"raw":{"variants":["Star-by-star supernovae lift dwarf galaxy metallicities 1.5 dex","Single-supernova sampling fixes UFD metallicity mismatch","Simulated UFDs close metallicity gap with individual SN","Size mismatch in UFDs traced to measurement method","New simulation narrows UFD metallicity and size gaps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001259,"raw_usage":{"total_tokens":5326,"prompt_tokens":1280,"completion_tokens":4046,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":896,"completion_tokens_details":{"reasoning_tokens":3960}},"tokens_in":896,"tokens_out":4046,"duration_ms":26212,"temperature":1.0,"reasoning_tokens":3960,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T15:01:42.960503+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Rerun the Halo4 zoom-in with delayed supernova feedback and radiative transfer while keeping all other settings fixed. The paper predicts the average stellar metallicity will fall below its fiducial $[\\mathrm{Fe/H}] \\approx -2.65$; if it instead stays level or rises, the claim that individual IMF sampling is what lifts simulated UFDs toward the observed mass-metallicity relation would be falsified.","supporting_citations":[],"review_version":1}