{"id":"a0522742-5e5d-4b23-b05e-bf368f749456","arxiv_id":"2412.01901","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Ultra-diffuse galaxies and nearly-UDGs share similar stellar populations and split into two classes marked by globular cluster richness.","lead":"This paper compares ultra-diffuse galaxies (UDGs) with a nearly-UDG control sample using one consistent SED fitting method. It finds that both groups split into two classes, distinguished by globular cluster content, which suggests distinct formation histories.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The two-class clustering result may be largely imposed by the manual zeroing of δdwarf MZR before KMeans; a sensitivity test without that pre-processing is required.","rationale":"The reader's conditional verdict is appropriate, but I would place the load-bearing uncertainty on a different step than the reader's weakest_assumption. The reader flags the NUDGe distance assumption (Section 3), which is a real limitation for roughly a quarter of the NUDGe sample and can affect their derived masses, sizes, and stellar populations. However, that assumption does not touch the UDG-only clustering, which already recovers two classes with a silhouette score of 0.7. The more central vulnerability is the manual zeroing of δdwarf MZR before KMeans. Since this variable is subsequently reported as one of the key discriminators between the two classes, the analysis risks imprinting the two-class structure through the preprocessing. The paper does present supporting evidence: the B24 predecessor found similar classes, the clustering is stable when GC information is excluded, and the refitted B22 data are checked against spectroscopy. These give the two-class picture independent support. Still, the δdwarf MZR preprocessing is a user-defined transformation applied to a variable that is then used to define the physical interpretation of the classes, and no sensitivity test addresses it. Adding such a test is a natural condition for full acceptance. The reader did mention the δdwarf MZR adjustment as a caveat in the rationale, so this is a partial overlap rather than a disagreement. I keep the verdict at CONDITIONAL/UNCHANGED because the concern is not yet demonstrated to overturn the result, but it should be explicitly resolved before the two-class claim is treated as robust.","tokens_in":38134,"tokens_out":5368,"duration_ms":62635,"concrete_test":"Rerun the KMeans clustering on the UDG-only and UDG+NUDGes samples under at least three variants: (i) omit δdwarf MZR entirely; (ii) use the raw δdwarf MZR values without zeroing galaxies inside the MZR scatter; and (iii) replace δdwarf MZR with [M/H] and log M* as separate inputs, or use a uncertainty-aware clustering method such as a Gaussian mixture model. If the same two classes with similar membership and silhouette scores (≳0.5 for UDG-only, ≳0.4 for the combined sample) are recovered in variants (i) and (ii), the concern is resolved. If Class A/B membership changes substantially or the silhouette score drops below about 0.3, the two-class claim is not robust to the pre-processing and should be presented only conditionally.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing weakness is the pre-processing of δdwarf MZR before the KMeans clustering in Section 4.3. The authors set δdwarf MZR = 0 for every galaxy within the scatter of the Simon (2019) mass-metallicity relation because KMeans cannot incorporate uncertainties. This turns a continuous, noisy variable into a distribution with a large spike at exactly zero and tails at negative and positive values. KMeans then uses this transformed variable as one of eleven inputs, and the paper later identifies δdwarf MZR as one of the key factors separating Class A (δ ≈ 0) from Class B (δ < 0). The separation is therefore at least partly built into the input rather than discovered from the data. Because δdwarf MZR is derived from stellar mass and metallicity, and Class B is interpreted as a 'failed galaxy' pathway defined by being metal-poor for its mass, this pre-processing can inflate the silhouette score and artificially stabilise the two-class solution. The paper does not report a run without δdwarf MZR, with the raw residual, or with uncertainties propagated, so the robustness of the central two-class claim to this analytical choice is unquantified. This concern affects both the UDG-only and the UDG+NUDGes clustering, and thus the paper's central claim more directly than the NUDGe distance assumption, which mainly affects the smaller NUDGe subset.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper assembles a sample of 88 ultra-diffuse galaxies (UDGs) from two earlier studies, refits the cluster-dominated subset with the same PROSPECTOR configuration used for the group/field sample, and adds 36 lower-surface-brightness dwarf galaxies ('NUDGes') from the MATLAS survey, all analysed with GALFITM and PROSPECTOR using DECaLS and WISE photometry. After comparing stellar populations, structures, globular cluster (GC) properties and environments, the authors report that NUDGes resemble UDGs in most respects and differ mainly in being smaller and brighter. They then apply KMeans clustering to the UDG-only and UDG+NUDGe samples and recover two classes, which they associate with 'puffy dwarf' and 'failed galaxy' formation pathways. Finally, they use the ratio of GC system mass to stellar mass as a function of environment to argue that these classes cannot simply evolve into one another and may have formed through distinct processes.","tokens_in":38425,"tokens_out":5818,"duration_ms":63912,"significance":"If the conclusions hold, the paper would strengthen the case that GC system properties, rather than size or surface brightness alone, are the physically meaningful separator among low-surface-brightness dwarf galaxies, and that two distinct formation channels (puffy dwarfs and early-quenched/failed galaxies) are present across the UDG/NUDG boundary. The work's main strengths are the homogeneous SED-fitting methodology applied to all 124 galaxies, the refitting of the B22 sample with the same configuration as B24, the use of WISE upper limits to constrain dust, the inclusion of HST-based GC counts, and the direct comparison of three NUDGes with MUSE spectroscopy. These are genuine assets and make the compilation a useful resource. However, the central two-class claim depends on several analysis choices that are not currently tested, most importantly the pre-processing of the mass-metallicity residual before clustering and the inconsistent UDG classification between the CFHT-based labels and the DECaLS photometry actually used.","major_comments":[{"comment":"The pre-processing step in Section 4.3 that assigns δdwarf MZR = 0 to every galaxy within the scatter of the Simon (2019) relation is load-bearing for the central two-class result. Because δdwarf MZR is one of the eleven KMeans inputs and is later quoted as the leading discriminant (Table 1 and Section 5), the separation between Class A (δ ≈ 0) and Class B (δ < 0), and the placement of exactly the five NUDGes below the MZR into Class B, is partly imposed by this transformation rather than discovered. The paper does not report a run with the raw residual, with δdwarf MZR omitted, or with uncertainties propagated through repeated draws. Please add those sensitivity tests and show the silhouette scores and class centroids; if the two-class solution survives, the claim is much stronger.","section":"4.3"},{"comment":"The sample classification is inconsistent with the photometry used in the analysis. Section 3.1 states that UDG/non-UDG labels are kept from the CFHT determination even though all structural parameters are re-measured with shallower DECaLS data, and Section 4.2 admits that many UDGs are smaller than the 1.5 kpc threshold or brighter than the surface-brightness threshold when measured with DECaLS. Since the paper's first conclusion is that NUDGes differ from UDGs 'by definition' in size and surface brightness, the comparison is contaminated by dataset-dependent labels. Please either re-derive UDG/NUDG classifications from the homogeneous DECaLS measurements used throughout, or provide a quantitative cross-tabulation of how many objects move across each boundary and re-run the key comparisons on the consistently defined subsample.","section":"3.1 and 4.2"},{"comment":"The assumed distance for NUDGes without spectroscopic redshifts is the distance of the closest massive galaxy. The paper cites Heesters et al. (2023) that 75% of MATLAS dwarfs are at the host redshift, so roughly a quarter of the NUDGe sample may have distances, and therefore stellar masses, effective radii, ages and metallicities, that are systematically wrong. Because the NUDGes are the new sample and are included in the clustering, the conclusion that UDGs and NUDGes fall into the same two classes needs to be tested against this uncertainty. A simple check would be to repeat the clustering and the comparisons in Figures 4 and 5 using only NUDGes with spectroscopic distances, or to perturb distances by plausible factors and report the range of class memberships.","section":"3 and 3.3"},{"comment":"The SED validation for the NUDGes rests on three MUSE galaxies, and the agreement is not uniform: MATLAS-1400 has SED [M/H] = −0.43 ± 0.63 dex versus [M/H] ≈ −1.2 from spectroscopy, a much larger offset than for the other two objects, and the B22 refit also shows a −0.25 dex median metallicity offset against Ferré-Mateu et al. (2023). Since δdwarf MZR is computed from [M/H] and is a leading clustering input, a metallicity bias of this size is large enough to move galaxies across the MZR scatter boundary and change class membership. Please quantify the sensitivity of the clustering to a ±0.25 dex shift in metallicity, and discuss the MATLAS-1400 discrepancy explicitly rather than attributing it to large uncertainties.","section":"Appendix B2"},{"comment":"The identification of δdwarf MZR, N_GC, and b/a as 'key factors supporting the classification' is circular because these three quantities are among the inputs to KMeans. Finding that an input feature differs between output clusters is expected and does not independently corroborate the puffy dwarf / failed galaxy interpretation. Please rephrase this as a description of which input features drive the separation, and add a validation such as feature permutation or comparison with a classifier trained on a hold-out set; alternatively, show that the same two classes are recovered when the clustering is run on a subset of features or on the independent spectroscopic sample.","section":"4.3"}],"minor_comments":[{"comment":"The text refers to the 'right-hand side of Fig. 3' for the size-luminosity diagram, but in the printed figure the size-luminosity panel is on the left; please correct the cross-reference.","section":"4.2"},{"comment":"The table note lists a 'GALFITM DECaLS i-band magnitude' column, but the table contains g, r, g−r, z, g−z, and WISE columns; the note appears to be a leftover from an earlier version and should be corrected.","section":"Table C1"},{"comment":"The phrase 'dynamic nestled sampling' should read 'dynamic nested sampling'.","section":"3.3"},{"comment":"The field environment conclusion in Section 4.4 rests on four galaxies with median M_GC/M_* = 0.00 ± 0.10%; please add a bootstrap uncertainty or explicitly caution that the field sample is very small.","section":"4.4"},{"comment":"The notation is inconsistent: logρ_N is defined in the appendix text while the main text and figures use logρ_10; please harmonise the symbols.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The observational compilation and homogeneous SED fitting are valuable, and the paper is within the journal's scope. The main issue is not the data but the analysis choices: the δdwarf MZR pre-processing, the CFHT versus DECaLS classification mismatch, and the distance assumption for NUDGes all directly affect the central two-class claim. These are fixable with robustness tests, so I recommend major revision rather than rejection. I do not see a novelty disclosure concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper's main value is the uniform SED analysis of 36 NUDGes plus a reanalysis of the B22 UDG sample, giving a combined sample of 124 LSB dwarfs analyzed consistently. That is real work, and the paper does it carefully. The NUDGe results support the earlier picture that UDGs and NUDGes overlap in most properties, and showing that GC content separates extreme galaxies is a useful practical point. The environment argument about distinct formation paths is suggestive and honestly discussed.\n\nThe soft spots are real, though not fatal. First, the KMeans clustering includes a pre-processed delta_dwarf_MZR where all galaxies within the MZR scatter are set to zero because KMeans cannot handle uncertainties. Later, delta_dwarf_MZR is named one of the key class discriminants. That means the separation into a 'failed galaxy' class that is metal-poor for its mass is part of the input, not purely a discovery. The paper reports no run without delta, with raw residuals, or with uncertainties, so we don't know how much of the two-class split survives that choice. A referee should ask for that sensitivity test.\n\nSecond, the UDG sample is defined by CFHT measurements, but many galaxies do not satisfy the UDG cuts when remeasured with shallower DECaLS data. The authors acknowledge this, but it means the UDG/NUDGe boundary is data-dependent. That limits the interpretation that NUDGes are 'nearly' UDGs.\n\nThird, the distance assumption for NUDGes without spectroscopy (closest massive galaxy) affects maybe a quarter of those 36. The authors cite Heesters et al. (2023) showing 75% agreement, so they know the risk. For about nine objects, masses and sizes could shift, and the NUDGe comparison could change. This is a moderate concern, not a catastrophe.\n\nThe circularity worry is real: the classes are defined using the same properties later listed as discriminants. But the clustering is not trivial—it uses eleven inputs, and the classes do differ in multiple ways, including GC number and axis ratio, which are not derived from delta. So the two-population picture is not invented, but it is less robust than the silhouette scores imply.\n\nWho should read it: UDG specialists and anyone working on LSB dwarf populations. It belongs in the literature with revisions. My recommendation: send to peer review, but require the robustness checks on the clustering input and a clear statement of which assignments would flip if distances or delta are treated differently.","headline":"A careful and useful uniform SED analysis of 124 LSB dwarfs, but the two-class clustering result is less secure than the paper claims because a key input was manually zeroed before clustering.","tokens_in":38999,"tokens_out":2088,"would_cite":true,"duration_ms":23913,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that nearly-UDGs are the same kind of object as ultra-diffuse galaxies—differing only in size and surface brightness—and that both populations split cleanly into a puffy-dwarf class and a failed-galaxy class.","keywords":["ultra-diffuse galaxies","low surface brightness galaxies","nearly-UDGs","globular clusters","galaxy formation","SED fitting","mass-metallicity relation","galaxy clustering"],"falsifier":"Obtain spectroscopic redshifts for all 36 NUDGes; if the true distances move a substantial fraction of them across the UDG size or surface-brightness boundary, recompute their stellar masses and rerun the clustering to see whether the same two classes emerge with the same member galaxies.","tokens_in":37947,"feed_emoji":"🌌","tokens_out":11072,"duration_ms":100221,"temperature":0.7,"pith_summary":"Ultra-diffuse galaxies (UDGs) are defined by size and surface brightness alone, and this paper argues that those two criteria are not enough to capture what makes certain low-surface-brightness galaxies extreme. Comparing 88 UDGs with 36 nearly-UDGs (NUDGes) — dwarf galaxies that fall just outside the UDG cuts — the authors find that NUDGes are statistically indistinguishable from UDGs in stellar populations, star formation histories, shapes, and globular cluster content, differing only in the very properties used to define them. A clustering analysis splits both the UDG-only and the combined sample into the same two classes: a younger, bluer, elongated, globular-cluster-poor 'puffy dwarf' class and an older, rounder, globular-cluster-rich 'failed galaxy' class lying below the classical dwarf mass–metallicity relation. The paper proposes that globular cluster number and cluster mass relative to stellar mass should join size and surface brightness as standard diagnostics for identifying extreme low-surface-brightness galaxies.","feed_headline":"Ultra-diffuse galaxies split into two distinct classes","feed_subtitle":"Nearly-UDGs share the split: one class is a puffy dwarf, the other a failed galaxy rich in globular clusters.","key_machinery":"The load-bearing machinery is a uniform spectral energy distribution (SED) fitting procedure applied to every galaxy, followed by an unsupervised centroid-based clustering algorithm (KMeans) run on eleven scaled properties. The SED fitting yields stellar masses, metallicities, ages, star formation timescales, and dust attenuation for all 124 galaxies under identical assumptions, so differences between samples are not artifacts of heterogeneous methods. The clustering uses the residual $\\delta_{\\rm dwarf\\,MZR}$ (offset from the classical dwarf mass–metallicity relation), globular cluster number $N_{\\rm GC}$, and axis ratio $b/a$ as the strongest discriminators, and a silhouette score selects the number of classes; the same two classes emerge whether NUDGes are included or not. The globular cluster counts and cluster system masses, measured from space-based imaging, play the decisive role in separating the classes and in the environmental argument.","core_discovery":"On the paper's own terms, the central discovery is that the UDG designation is less a physical class than a region of parameter space: galaxies just outside that region (NUDGes) share every stellar population property, structural property, and globular cluster trend with the UDGs inside it. When clustered on stellar mass, color, mass-to-light ratio, age, axis ratio, size, globular cluster number, globular cluster mass fraction, central surface brightness, star formation timescale, and offset from the dwarf mass–metallicity relation, both the 88-UDG sample and the 124-galaxy UDG+NUDGe sample split into two classes with high silhouette scores. Class A matches a 'puffy dwarf' formation path — low mass, blue, young, elongated, GC-poor, following the classical dwarf MZR — while Class B matches a 'failed galaxy' path — massive, red, old, round, GC-rich, below the MZR. The globular cluster system mass relative to stellar mass rises from field to group to cluster, and the paper argues that this monotonic difference, combined with the implausibility of forming or destroying globular clusters during infall, implies that the two classes did not simply evolve into one another but formed through distinct processes.","pith_inferences":["Beyond the paper's claims: if GC content is the sharper separator, a practical test would be to select low-surface-brightness dwarfs by GC richness alone and check whether they reproduce the same two stellar-population groups without any size or surface-brightness cut.","Beyond the paper's claims: the distance assumption for NUDGes without redshifts is the main lever on the result; a spectroscopic campaign that measures distances for all 36 NUDGes would show whether the two classes survive with accurate sizes and masses.","Beyond the paper's claims: galaxy formation simulations that produce puffy dwarfs versus failed galaxies could be run through the same clustering pipeline; if simulated populations do not separate along these axes, the mapping from observed classes to formation paths would need revision.","Beyond the paper's claims: the flat-to-rising color gradients seen here are in tension with simulations predicting declining metallicity gradients for high-spin halos; deeper imaging or spectroscopy could determine whether the gradients trace age rather than metallicity."],"forward_implications":["The standard UDG selection by size and surface brightness mixes two physically distinct populations, so samples built on those cuts alone need to be re-examined.","Globular cluster number and GC system mass fraction should be adopted as standard diagnostics for selecting extreme low-surface-brightness galaxies.","NUDGes should be included in UDG studies when the question is about formation physics rather than about the operational cut.","The two-class split means formation models must explain both a puffy-dwarf channel and a failed-galaxy channel, with the failed-galaxy class concentrated in denser environments.","If the GC-environment trend is real, field GC-poor galaxies are unlikely to become cluster GC-rich galaxies through infall, so environment alone does not transform one class into the other."],"supporting_citations":[{"why":"Supplies the 29 cluster-dominated UDGs and the original SED-fitting results that are refit here with the updated methodology.","marker":"B22"},{"why":"Provides the 59 MATLAS UDGs, the uniform SED-fitting methodology, and the previous two-class clustering result this paper extends.","marker":"B24"},{"why":"Provides the 36 NUDGes and the globular cluster counts for the MATLAS dwarfs.","marker":"Marleau et al. (2024b)"},{"why":"Defines the UDG selection criteria (central surface brightness and half-light radius) that the paper is testing and extending.","marker":"van Dokkum et al. (2015)"},{"why":"Introduces the NUDGe term and the suggestion that nearly-UDGs be considered alongside UDGs.","marker":"Forbes & Gannon (2024)"},{"why":"Supports the distance assumption by finding 75% of MATLAS dwarfs at their host redshift.","marker":"Heesters et al. (2023)"},{"why":"Provides the spectroscopic stellar populations used to validate the SED-fitting results and to connect GC richness to metallicity.","marker":"Ferré-Mateu et al. (2023)"},{"why":"Supplies the classical dwarf mass-metallicity relation whose offset defines the delta-dwarf-MZR clustering input.","marker":"Simon (2019)"},{"why":"Supplies the high-redshift mass-metallicity relation used to identify the early-quenched failed-galaxy candidates.","marker":"Ma et al. (2016)"},{"why":"Supplies the KMeans clustering algorithm used to identify the two classes.","marker":"MacQueen et al. (1967)"}],"fun_headline_variants":["UDGs split into puffy dwarfs and failed galaxies","Nearly-UDGs reveal same two-class split as UDGs","Globular clusters distinguish two UDG formation paths","Ultra-diffuse galaxies: not one type, but two distinct origins"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The analysis assumes that galaxies without spectroscopic redshifts sit at the distance of their nearest massive neighbor, so roughly a quarter of the NUDGe sample could have systematically wrong distances and therefore wrong stellar masses, sizes, and stellar population properties.","fun_headline_variants_meta":{"raw":{"variants":["UDGs split into puffy dwarfs and failed galaxies","Nearly-UDGs reveal same two-class split as UDGs","Globular clusters distinguish two UDG formation paths","Ultra-diffuse galaxies: not one type, but two distinct origins"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000386,"raw_usage":{"total_tokens":2094,"prompt_tokens":1058,"completion_tokens":1036,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":674,"completion_tokens_details":{"reasoning_tokens":964}},"tokens_in":674,"tokens_out":1036,"duration_ms":10190,"temperature":1.0,"reasoning_tokens":964,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:50:38.723322+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain spectroscopic redshifts for all 36 NUDGes; if the true distances move a substantial fraction of them across the UDG size or surface-brightness boundary, recompute their stellar masses and rerun the clustering to see whether the same two classes emerge with the same member galaxies.","supporting_citations":[{"cited_title":"Radial velocities and stellar population properties of 56 MATLAS dwarf galaxies observed with MUSE","cited_arxiv_id":"2305.04593","evidence_quote":"Supports the distance assumption by finding 75% of MATLAS dwarfs at their host redshift."}],"review_version":1}