{"id":"7cccc69f-3b20-466c-aa96-0c1e9049060f","arxiv_id":"2506.15785","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Applying a realistic surface brightness cutoff to simulated ultra-faint galaxies shrinks their inferred sizes and masses, improves agreement with observed dwarfs, but biases dynamical mass estimates so that dark matter densities are overestimated.","lead":"Simulated ultra-faint dwarf galaxies look more like real ones when a realistic surface brightness detection limit is applied, but the same cut systematically inflates estimates of their dark matter content. The result matters because dark matter constraints from faint dwarf galaxies may be biased by hidden stellar halos outside the detection limit.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The internal mechanism is clear, but the R_SB isophotal cut is only a speculative proxy for how real UFDs are detected; the order-of-magnitude bias claim depends on that mapping.","rationale":"The stress-test should target the condition most needed for the central claim to be true. The paper's internal result is solid: imposing an isophotal cut shrinks R_1/2, leaves sigma_los roughly flat, and makes the Wolf estimator decrease by less than the true enclosed mass, so M_est/M_true rises (Figs. 6-7). This mechanism would survive even in a Milky Way-host environment, because tidal stripping would also produce smaller tracer radii; it would change the magnitude but not the direction of the estimator bias. The reader's MW-host concern is thus real but secondary. The weaker link is the paper's own admission in Section 4.1 that the R_SB cut is a 'speculation' as a proxy for star-count detection of UFDs. Real UFD surveys count resolved stars, and the detectability edge is a statistical function of the stellar luminosity function and background, not a smooth 32.5 mag/arcsec^2 isophote. At Mstar ~ 300 Msun, where the sample has only ~10 star particles, the isophote is particularly noisy. Since the abstract and Section 4.2 quote order-of-magnitude biases in halo mass estimates, the quantitative claim depends on this mapping. A forward-model star-count test is the natural check: it uses the same simulations and asks whether the observer-defined edge produces the same bias. If it does, the paper's caution is strengthened; if it does not, the paper should be framed as a controlled experiment on a specific cut rather than a model of real UFD detectability. The conditional verdict is appropriate either way, so no change to the reader's verdict is recommended.","tokens_in":21028,"tokens_out":11807,"duration_ms":143856,"concrete_test":"Forward-model resolved-star observations for the 19 simulated galaxies: assign luminosities to star particles using a Kroupa IMF and an old, metal-poor isochrone, place each galaxy at a typical UFD distance, add a realistic foreground/background stellar density and photometric errors, and run a matched-filter CMD detection algorithm to define the observed stellar edge. Then recompute R_1/2, sigma_los, M_est_1/2 (Eq. 1), and M_true_1/2 using only detected stars, and compare the resulting M_est/M_true distribution and inferred NFW halo masses with the R_SB-based values in Figs. 7 and 8. If the resolved-star edges reproduce the median ratio and outliers, the R_SB mapping is validated; if the edges are larger or smaller, the numerical bias estimate needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 4.2) is that excluding the stellar halo via the R_SB cut moves simulated UFDs closer to observations while degrading the Wolf et al. (2010) mass estimator, by up to an order of magnitude for low-mass systems. For this claim to apply to real UFDs, R_SB must reproduce the actual observational edge. Observed UFDs are found as overdensities of resolved stars, not by integrated surface brightness, and the paper concedes in Section 4.1 that the comparison is 'a reasonably good approximation' and 'speculate[s]' that it is valid. The particle-based mu_V = 32.5 mag/arcsec^2 isophote with M/L=1 does not capture the statistics of star-count detection: the real limiting radius is set by Poisson fluctuations of a small number of bright RGB stars against foreground/background, and depends on distance, depth, and the IMF. For the lowest-mass simulated galaxies (Mstar ~ 250-400 Msun, ~10-30 star particles), the isophotal crossing radius is dominated by counting noise and may not correspond to any measurable radius. The reported M_est/M_true ratios (median 1.48 to 2.79, Fig. 7) and the halo-mass inflation in Fig. 8 are therefore quantitative predictions of a particular cut, not yet validated as quantitative predictions about observed UFDs. This does not break the internal logic, but it is the least secure link in the chain from simulation to the abstract's caution about real mass estimates.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses 19 ultra-faint galaxies (M_star ~ 250-40,000 Msun) drawn from three FIRE-2 cosmological zoom-in simulations with baryonic mass resolution of 30 Msun to compare two ways of defining a simulated galaxy's edge: a fixed cut at 15% of the halo virial radius (R_15%) and an isophotal cut at mu_V = 32.5 mag arcsec^-2 (R_SB) assuming M/L = 1. It finds that applying the surface brightness cut reduces stellar masses, half-mass radii, and line-of-sight velocity dispersions, moving the simulated galaxies closer to observed UFDs in the mass-size plane, the sigma_los-M_star plane, and the circular-velocity-versus-half-light-radius plane. It also finds that the Wolf et al. (2010) mass estimator becomes less accurate after the cut, with the median ratio M_est_1/2 / M_true_1/2 increasing from 1.48 to 2.79, and that for the lowest-mass galaxies the inferred dark matter halo masses can be overestimated by an order of magnitude. The paper concludes that surface brightness limits must be taken into account when using UFD mass estimates to constrain dark matter physics or the low-mass threshold of galaxy formation.","tokens_in":21357,"tokens_out":5284,"duration_ms":64121,"significance":"If the result holds, the paper provides a concrete, physically motivated mechanism that can simultaneously reconcile simulated UFDs with observed compact sizes and warn against interpreting UFD dynamical masses at face value. The numerical experiment is clean and internally consistent: the surface brightness threshold is externally motivated, the Wolf estimator is tested against the true simulated enclosed mass rather than being used to define the conclusion, and no fitting to the target observed result is involved. The simulations and initial conditions are publicly available, which is a strength for reproducibility. The main limitations are the small sample (19 galaxies from three volumes), the lack of a Milky Way-mass host for most of the galaxies, and the reliance on a particular isophotal cut as a proxy for how real surveys detect UFDs; these limitations primarily affect the strength of the quantitative claims rather than the internal logic of the experiment.","major_comments":[{"comment":"The central claim that observed UFD mass estimates are biased high by missing stellar halos depends on R_SB faithfully reproducing the edge set by real surveys. Observed UFDs are discovered as overdensities of resolved stars, so the limiting radius is set by Poisson fluctuations in small numbers of bright RGB stars against foreground and background contamination, not by an integrated isophote at mu_V = 32.5 mag arcsec^-2 with M/L = 1. The paper itself states in Sec. 4.1 that the comparison is 'a reasonably good approximation' and that the authors 'speculate' it is valid. For the lowest-mass simulated galaxies, which contain roughly 10-30 star particles, the isophotal crossing radius may be dominated by counting noise and may not correspond to any measurable observed radius. Because the quantitative ratios in Figs. 7 and 8 are predictions of this particular cut, I ask for a direct validation using mock star-count surveys or mock images with resolved stellar populations, or alternatively a clear reframing of the abstract's caution as conditional on isophotal detection.","section":"Sec. 2.1 and Sec. 4.1"},{"comment":"The simulated sample consists of 19 galaxies drawn from three isolated low-mass halos, and Sec. 4.1 explicitly notes that the simulations do not include a Milky Way-mass host galaxy. Observed UFDs are predominantly satellites of the Milky Way or of other Local Group galaxies, where tidal forces and the host potential can strip or alter outer stellar populations. If real UFD outer halos are truncated by tides, the magnitude and possibly the direction of the reported bias could differ from the isolated case. I recommend testing the R_SB analysis in simulations that include a Milky Way-mass host, or at least comparing the radial stellar profiles of these isolated UFDs with those of observed satellites to justify the extrapolation to the observed population.","section":"Sec. 2 and Sec. 4.1"},{"comment":"The order-of-magnitude halo-mass inflation in Fig. 8 is derived by comparing M_est_1/2 to NFW enclosed-mass curves assuming a fixed concentration c = 24 from the extrapolated Neto et al. (2007) relation. The concentration-mass relation at UFD masses is not well constrained, and changing c shifts the inferred M_halo for a given (M_est_1/2, R_1/2) point. Since the abstract's strongest caution concerns real mass estimates, the sensitivity of the Fig. 8 inference to c and to the assumption of an NFW profile should be quantified, for example by repeating the inference over the plausible range of concentration at these masses.","section":"Sec. 3.4, Fig. 8"}],"minor_comments":[{"comment":"The phrase 'R 156% cut' appears to be a typo; it should read 'R_SB cut'.","section":"Sec. 3.1"},{"comment":"The color coding in Fig. 7 is described in the text as the change in the log mass ratio between the two cuts, but the figure caption and text do not specify the color scale or state explicitly that red points correspond to larger ratios for R_SB and blue points to larger ratios for R_15%; please add this information.","section":"Sec. 3.4, Fig. 7"},{"comment":"The sentence 'this SB threshold is over two orders magnitude lower' is missing 'of'; it should read 'over two orders of magnitude lower'.","section":"Sec. 2.1"},{"comment":"The text refers to 'inverted age and metallicity gradients' but the subsequent analysis concerns metallicity only; clarify whether age gradients are used, or restrict the discussion to metallicity gradients.","section":"Sec. 3.3"},{"comment":"References to works 'in prep' such as Rodriguez-Wimberly in prep and Murphy et al. in prep should either be completed with available identifiers or clearly marked as private communications in the reference list.","section":"References"},{"comment":"The phrase 'The increase in overall surface brightness due to the R 156% cut' is confusing even after correcting the typo, because the cut reduces galaxy sizes and therefore increases the mean surface brightness within the edge; please rephrase to avoid ambiguity.","section":"Sec. 3.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a good fit for a specialized astrophysics journal and the numerical experiment is clean. My main concern is external validity rather than internal consistency: the quantitative bias claims depend on the isophotal cut being a faithful proxy for real UFD detection, and on the simulated sample being representative of observed Local Group satellites. These are addressable with additional tests or with appropriately softened conclusions, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nThe short version: this paper does something useful and does it honestly. It shows that if you define the edge of a simulated ultra-faint galaxy by a 32.5 mag/arcsec^2 isophote instead of the usual 15% virial radius, the inferred sizes and stellar velocity dispersions shift toward the observed UFD population—but the Wolf et al. mass estimator gets systematically worse, and for the lowest-mass systems the inferred dark matter halo mass can be off by an order of magnitude. That bias quantification is new, and it is exactly the kind of thing people need to worry about when they use UFDs to constrain dark matter models or galaxy formation thresholds.\n\nThe paper is well executed. The sample is small—19 galaxies from three FIRE-2 runs—but the resolution is high (30 solar masses baryonic), and the comparison between the two edge cuts is clean. The figures are convincing, particularly the mass-size plane shift and the M_est/M_true ratios. The authors are also candid about what their setup does not include: no Milky Way-mass host, so tides could alter the outer stellar halos; a hand-chosen M/L=1; and a single mu=32.5 threshold. They explicitly flag some of these.\n\nThe soft spots are real but not fatal. The biggest one is the mapping from R_SB to actual detection. Observed UFDs are found as overdensities of resolved stars, not by integrated surface brightness, and the paper itself says the comparison is 'a reasonably good approximation' and speculates that it is valid. For the faintest simulated galaxies—a few hundred solar masses, tens of star particles—the isophotal radius is dominated by counting noise. So the quantitative claim about an order-of-magnitude halo mass inflation is a prediction under a particular proxy, not a direct measurement of the bias in real UFDs. That should be stated more prominently, but it doesn't break the internal logic. The analysis code is not released, which is a minor reproducibility gap.\n\nBottom line: this is a solid, useful cautionary paper. It deserves serious refereeing. I would accept it for peer review, and I would cite it when interpreting UFD mass measurements. The main thing I'd ask for in revision is a sharper discussion of how the isophotal proxy maps to real star-count surveys, and some acknowledgment of how sensitive the bias magnitude is to that mapping.","headline":"Clean, honest simulation study showing that ignoring surface brightness limits can bias UFD sizes and Wolf masses by up to an order of magnitude—but the bias magnitude depends on an isophotal proxy that is admittedly approximate.","tokens_in":22037,"tokens_out":2391,"would_cite":true,"duration_ms":25344,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A realistic surface-brightness cut makes simulated ultra-faint galaxies match observations, but the same cut inflates their inferred dark-matter masses by up to tenfold.","keywords":["ultra-faint dwarf galaxies","stellar halos","surface brightness limits","mass estimators","galaxy sizes","dark matter halo masses","FIRE-2 simulations","Local Group dwarf galaxies"],"falsifier":"Deep star-count surveys of known ultra-faint Milky Way satellites reaching surface brightnesses near or below $\\mu_V \\approx 32.5$ mag arcsec$^{-2}$ should either find previously undetected outer stellar populations or fail to find them; if no such extended halos exist, the invisible-halo premise collapses and Wolf-estimator masses recomputed over any recovered stars should not move far from unity.","tokens_in":20860,"feed_emoji":"🔭","tokens_out":11862,"duration_ms":120928,"temperature":0.7,"pith_summary":"This paper argues that a realistic observational detection limit changes what counts as the galaxy in ultra-faint dwarf galaxy studies, and that the choice carries real scientific consequences. Using 19 simulated ultra-faint galaxies (stellar masses roughly 400 to 40,000 solar masses) from FIRE-2 cosmological simulations, it applies a surface brightness threshold of $\\mu_V \\approx 32.5$ mag arcsec$^{-2}$ and recalculates radii, stellar masses, velocity dispersions, metallicities, and dark-matter mass estimates. The cut shrinks the inferred sizes and lowers the inferred velocity dispersions, moving the simulated galaxies closer to observed Milky Way satellites in the mass-size and velocity-dispersion planes. However, the same cut degrades the standard Wolf et al. (2010) mass estimator: estimated masses move farther from the true enclosed masses, and for the lowest-mass galaxies the inferred central dark-matter densities and virial masses are biased upward, in the most extreme cases by about an order of magnitude. A sympathetic reader should care because ultra-faint galaxies are used as probes of dark-matter physics and of the minimum halo mass that can form a galaxy, and both uses are sensitive to this bias.","feed_headline":"Invisible stellar halos inflate ultra-faint galaxy masses tenfold","feed_subtitle":"A realistic brightness cut makes simulated dwarfs match observations, but warps their dark-matter mass estimates.","key_machinery":"The load-bearing mechanism is the contrast between two ways of defining a galaxy's edge in simulations: $R_{15\\%}$, a fixed cut at 15% of the host halo's virial radius, and $R_{\\rm SB}$, the isophotal radius at which the projected stellar surface brightness profile drops below $\\mu_V \\approx 32.5$ mag arcsec$^{-2}$. The second cut is the machine that removes the invisible stellar halo. The Wolf et al. (2010) mass estimator, $M_{\\rm est}^{1/2} \\approx 930\\,(\\sigma_{\\rm los}/{\\rm km\\,s^{-1}})^2\\,((4/3)R_{1/2}/{\\rm pc})\\,M_{\\odot}$, then converts the truncated half-light radius and line-of-sight velocity dispersion into a dynamical mass. The bias arises because both $R_{1/2}$ and $\\sigma_{\\rm los}$ fall when the halo is removed, but the true enclosed mass $M_{\\rm true}^{1/2}$ falls even faster, so the ratio $M_{\\rm est}^{1/2}/M_{\\rm true}^{1/2}$ grows from a median of 1.48 to 2.79.","core_discovery":"The central claim is that ultra-faint galaxies can carry extended, very low-surface-brightness stellar halos that real surveys do not detect, and that ignoring those halos is not a neutral choice. Defined by the conventional simulation edge at 15% of the virial radius, the simulated galaxies are too large, too massive, and too kinematically hot compared with observed ultra-faint dwarfs. Re-defining the edge as the radius where the projected surface brightness falls below $\\mu_V \\approx 32.5$ mag arcsec$^{-2}$ reduces half-light radii by more than half and pushes velocity dispersions below 5 km s$^{-1}$, producing better agreement with observations. But the same operation shrinks the true enclosed mass within the new half-light radius faster than it shrinks the Wolf et al. (2010) mass estimate, so the estimated-to-true mass ratio worsens from a median of 1.48 to 2.79. The consequence is that the very galaxies that look most like observed ultra-faints are the ones whose inferred dark-matter halo masses are most inflated, by up to an order of magnitude in the most extreme cases.","pith_inferences":["If real Milky Way satellites have already been tidally stripped by the host galaxy, their extended halos may be gone, and the sign or size of the bias could differ from what the isolated simulation sample predicts.","A direct observational test is to search for previously undetected, very low-surface-brightness stellar populations around known ultra-faints; detecting them and recomputing masses over the larger radii should lower inferred dark-matter densities.","Deeper co-added surveys should steadily shrink this bias as the effective detection threshold moves below 32.5 mag arcsec$^{-2}$, a trend that future data can verify."],"forward_implications":["Observed ultra-faint dwarfs may be the bright central cores of galaxies that are intrinsically more extended, which would explain why simulations appear too puffy until the low-surface-brightness outskirts are discarded.","Line-of-sight velocity dispersions below 5 km s$^{-1}$ can occur without tidal stripping by a massive host, so cold kinematics alone does not require a Milky Way-mass perturbing galaxy.","Metallicities of observed ultra-faints may be biased slightly high because the missing outer stars are preferentially the most metal-poor, though the surface-brightness cut only partially closes the mass-metallicity gap.","Dynamical masses of ultra-faint galaxies derived with the Wolf estimator should be viewed as upper limits when their stellar halos are not detected, since estimated masses systematically overshoot the true enclosed mass.","Dark-matter constraints drawn from ultra-faint galaxy kinematics, including limits on warm, self-interacting, or fuzzy dark matter, could shift substantially if the halo-mass overestimate is not corrected."],"supporting_citations":[{"why":"Supplies the 19-galaxy FIRE-2 sample and the earlier prediction that ultra-faints have extended, low-surface-brightness stellar populations.","marker":"Wheeler et al. (2019)"},{"why":"Provides the standard mass estimator whose accuracy under the surface-brightness cut is the paper's main target.","marker":"Wolf et al. (2010)"},{"why":"Gives the compilation of observed local ultra-faint galaxies used for comparison in the mass-size, velocity-dispersion, and metallicity planes.","marker":"Pace (2024)"},{"why":"Defines the FIRE-2 star-formation and stellar-feedback physics that generates the simulated galaxies.","marker":"Hopkins et al. (2018a)"},{"why":"Supplies the expectation that lower-mass halos of fixed velocity dispersion have lower surface brightness, which frames the invisibility bias.","marker":"Bullock et al. (2010)"},{"why":"Establishes the baseline roughly 18% accuracy of the Wolf estimator for more massive relaxed FIRE-2 galaxies that the ultra-faint results are compared against.","marker":"Gonzalez-Samaniego et al. (2017)"},{"why":"Shows that mock-image-based half-light radii are about 30% smaller than particle-based ones, supporting the claim that surface-brightness effects shrink inferred sizes.","marker":"Klein et al. (2024)"}],"fun_headline_variants":["Brightness cuts inflate dark matter estimates in ultra-faint dwarfs","Invisible halos skew ultra-faint dwarf mass estimates by 10x","Realistic brightness cuts inflate ultra-faint halo masses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulations are isolated low-mass galaxies without a Milky Way-mass host, so the paper assumes their extended stellar halos are representative of observed Milky Way satellites; if tides from a massive host strip or reshape those halos, the bias could change in magnitude or even direction.","fun_headline_variants_meta":{"raw":{"variants":["Brightness cuts inflate dark matter estimates in ultra-faint dwarfs","Invisible halos skew ultra-faint dwarf mass estimates by 10x","Realistic brightness cuts inflate ultra-faint halo masses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000403,"raw_usage":{"total_tokens":2160,"prompt_tokens":1065,"completion_tokens":1095,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":681,"completion_tokens_details":{"reasoning_tokens":1035}},"tokens_in":681,"tokens_out":1095,"duration_ms":9115,"temperature":1.0,"reasoning_tokens":1035,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:51:17.862063+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deep star-count surveys of known ultra-faint Milky Way satellites reaching surface brightnesses near or below $\\mu_V \\approx 32.5$ mag arcsec$^{-2}$ should either find previously undetected outer stellar populations or fail to find them; if no such extended halos exist, the invisible-halo premise collapses and Wolf-estimator masses recomputed over any recovered stars should not move far from unity.","supporting_citations":[{"cited_title":"S., Stewart, K","cited_arxiv_id":null,"evidence_quote":"Supplies the expectation that lower-mass halos of fixed velocity dispersion have lower surface brightness, which frames the invisibility bias."}],"review_version":1}