{"id":"9fd27404-a4ba-48a5-8b49-0f2cf3df3939","arxiv_id":"2412.15860","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"Integrated optical spectra of 2052 ECO star-forming galaxies, modeled as superpositions of many H II regions, reveal low-excitation gas with average metallicity near 0.3 solar and a mass-metallicity relation consistent with direct abundance methods.","lead":"Astronomers modeled the combined light from 2052 nearby star-forming galaxies to recover the range of gas conditions inside them, without needing spatially resolved maps. They showed that the resulting average metallicities line up with independent direct measurements, which supports the use of integrated spectra for studying how galaxies become enriched with heavy elements.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal Z/U/n distributions may not be identifiable from the strong-line set; §4.4 states tracers constrain only averages, so the claimed recovery of internal distributions is prior-dominated. A mock-injection recovery test is needed.","rationale":"The reader's weakest assumption (grid adequacy) is real, but the paper partially addresses it by testing alternate grids (BOND, SFGX) and by comparing with empirical diagnostics and the direct-method MZR. The more load-bearing gap is the identifiability of the distribution hyperparameters: §4.4 concedes that the tracers mostly constrain the average values, yet §§5.2–5.3 interpret the derived boundaries and dispersions as physical internal distributions and even as an evolutionary sequence. Without an injection-recovery test, we cannot know whether the strong-line set contains information on the shape of the internal distributions or only on a luminosity-weighted mean. The LOC-over-1C1S improvement is, by itself, strong evidence for nonuniform conditions, so the broad conclusion is directionally right; what is conditional is the quantitative recovery of the distributions. I therefore agree with the reader's CONDITIONAL verdict and would not change it, but the condition should explicitly include a synthetic recovery validation.","tokens_in":25990,"tokens_out":6461,"duration_ms":57808,"concrete_test":"Build mock galaxies by drawing from a known two-phase or log-normal distribution of (Z, U, n, age) using the very Cloudy/BPASS grid used here; compute the same strong-line fluxes, add noise matched to ECO S/N, and run the same MULTIGRIS LOC power-law inference. Check whether the marginalized posteriors for α_Z, Zmin, Zmax, and ΔZ recover the input values. Also fit the mocks with the 1C2S/2C1S models; if these fit the mock data as well as LOC, the line set cannot constrain the internal distribution shape.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.4 states that the inferred power-law slopes are 'similar on first order for all galaxies' and that 'the observed tracers mostly constrain the average physical parameter value.' If only the averages are constrained, then the boundary hyperparameters pmin,max — and hence the internal metallicity dispersion ΔZ = Zmax−Zmin used in §5.2.1 to infer an evolutionary enrichment sequence — are not data-driven. The LOC model's better PPP relative to 1C1S (§4.2) shows that a single uniform phase underfits, but it does not establish that the specific power-law distribution shape is recovered from the data; a single flexible distribution will trivially fit better. The paper presents no injection-recovery test on synthetic spectra with known multi-component distributions, and its benchmark in §2.1 only compares on-the-fly inference against precomputed LOC grids, not against a true underlying distribution. Consequently, the central claim that integrated spectra carry enough information to recover internal distributions of Z, U, and n is not yet demonstrated. This concern is independent of grid absolute calibration: even with a perfect grid, the strong-line set may constrain only luminosity-weighted averages.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper applies the MULTIGRIS Bayesian inference framework, using combinations of 1D Cloudy photoionization models (LOC distributions with power-law parameter distributions and free boundaries), to fit optical strong-line fluxes in 2052 star-forming galaxies from the volume-limited ECO catalog. It compares single-component 1D models (1C1S) with multicomponent and LOC architectures using posterior predictive p-values, marginal likelihoods, and a fraction-within-3σ metric, and finds that LOC models with free boundaries outperform single 1D models. The authors then infer per-galaxy averages and internal distributions of metallicity, ionization parameter, density, and stellar age; report a weakly bimodal average-metallicity distribution; discuss the internal metallicity dispersion in terms of an evolutionary enrichment sequence; and construct a mass-metallicity relation that they compare with direct-method and strong-line calibrations. A substantial part of the analysis is concerned with the influence of the line set, particularly the decision to exclude the [S II] lines.","tokens_in":26220,"tokens_out":4298,"duration_ms":40743,"significance":"If the inference is reliable, the paper would demonstrate that integrated optical spectra contain enough information to recover not only average ISM conditions but also luminosity-weighted internal distributions of Z, U, and n in galaxies, with important implications for high-redshift unresolved samples. The strength of the paper lies in its explicit model-selection framework (PPP, marginal likelihood, and the decision tree of Sect. 2.3), the public MULTIGRIS code, the use of a volume-limited sample spanning the dwarf regime, and the anchoring of the grid metallicity against the empirical line-ratio diagnostics of Garg et al. (2024). The agreement of the no-[S II] MZR with direct-method determinations over a wide mass range is a valuable external check. However, the central recovery claim and several headline results depend on untested or model-dependent steps, which must be made explicit and quantitatively supported.","major_comments":[{"comment":"The central claim that integrated spectra recover internal distributions of Z, U, and n is not yet demonstrated. Section 4.4 states that the inferred power-law slopes are \"similar on first order for all galaxies\" and that \"the observed tracers mostly constrain the average physical parameter value.\" If only the averages are constrained, the free boundaries pmin,max—and therefore the metallicity dispersion ΔZ = Zmax−Zmin used in §5.2.1 to infer an evolutionary enrichment sequence—may be largely prior-dominated. The benchmark in §2.1 compares on-the-fly inference with precomputed LOC grids, but it does not test whether an unknown input multi-component distribution can be recovered. I request an injection-recovery test: generate synthetic galaxy spectra from known LOC or multi-component distributions, fit them with the same pipeline, and report bias and uncertainty on pmin,max, the power-law slopes, and ΔZ. Without such a test, the abstract's claim that the inference \"predicts non-uniform physical conditions within galaxies\" remains an architectural assumption rather than an empirical result.","section":"Sect. 4.4 / 5.2.1"},{"comment":"The headline numerical results—the ≈0.3 Z⊙ peak and the MZR—are presented mainly for runs that exclude the [S II] lines, while the runs that include [S II] produce a stronger metallicity bimodality and a significantly different high-mass MZR (Figs. 9, 13). Excluding a diagnostic because it is underfit is a post-hoc model choice; the PPP improvement and the [O II] replacement test in §4.1 support the decision, but they do not quantify how much of the interpretation rests on this choice. I ask for a quantitative statement, perhaps in tabular form, of how the inferred average-metallicity PDF, the MZR fit coefficients in Eq. (10), and the bimodality indicators change across the line-set configurations (with [S II], without [S II], with [O II]), together with the corresponding model-selection metrics. This will let the reader assess the robustness of the central results rather than relying on qualitative statements.","section":"Sect. 4.1 / 5.4"},{"comment":"The metallicity bimodality is explicitly attributed to the grid and its abundance patterns: \"We conclude that the Z bimodality is mostly driven by the grid presently used and the underlying abundance patterns.\" The alternative grids are discussed only qualitatively (\"BOND ... does not show any bimodality\"), and the authors themselves conclude that the bimodality, \"if real, is likely not a strong one.\" Since the bimodality appears in the abstract and conclusions as a substantive result, the paper should either downgrade it to a model-dependent suggestion in the abstract and conclusions, or provide a quantitative comparison with BOND and SFGX (e.g., the fraction of models in the secondary peak, or a mixture-model fit to the PDF). As written, the level of support is disproportionate to the prominence of the claim.","section":"Sect. 5.3"}],"minor_comments":[{"comment":"Typo: \"withough\" should be \"without\" in the parenthetical about machine-learning UV magnitude predictions.","section":"Sect. 3.1"},{"comment":"Typo: \"statitistical distributions\" should be \"statistical distributions.\"","section":"Sect. 3.2"},{"comment":"Typo: \"interperation\" should be \"interpretation.\"","section":"Sect. 5.1"},{"comment":"Equation (5) is ambiguous because the text says the average is calculated in log scale but the formula writes p in the numerator and denominator. Please define p as log10 of the physical value (or write the equation for log10 pavg explicitly), and define the corresponding weighted average for the boundaries.","section":"Eq. (5)"},{"comment":"The \"fraction of posterior draws matching the observed values within 3σ\" is used in Fig. 7 and elsewhere but is not formally defined; please give its definition and specify how it is computed from the posterior samples.","section":"Sect. 2.3"},{"comment":"The U–Z interpretation would benefit from stating explicitly whether the comparison curves from Kashino & Inoue (2019) and Ji & Yan (2022) are plotted in the same quantity (log U) and against the same metallicity scale as the inferred Zavg; otherwise the visual agreement could be misleading.","section":"Sect. 5.1 / Fig. 10"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern raised by the skeptical reader is valid and is the main reason for my recommendation: the absence of an injection-recovery test leaves the central 'recovery of internal distributions' claim unsupported. The paper is a good fit for an A&A methods-oriented paper, and the authors have been unusually candid about limitations; the requested tests and quantification of line-set dependence are within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The reader's conditional verdict is about right, but the stress-test note puts the main worry in the right place. Section 4.4 states that the observed tracers mostly constrain the average physical parameter value and that the power-law slopes are similar for all galaxies. If that is true, then the boundaries Zmin/Zmax — and hence the internal dispersion ΔZ that drives the Section 5.2.1 \"evolutionary enrichment sequence\" — are not actually data-constrained. The fact that LOC fits better than a single 1D component shows that a uniform phase underfits, not that a power-law interior is recovered. There is no injection-recovery test on synthetic spectra with known multi-component distributions. So the headline claim, recovering internal distributions of Z, U, and n from integrated spectra, goes beyond what the line set can identify. This concern is independent of grid calibration: even with a perfect grid, the strong lines may only pin down luminosity-weighted averages.\n\nThat said, there is real value here. The free-boundary LOC architecture with power-law age and Z applied to 2052 volume-limited ECO galaxies is new, and the [SII] exclusion analysis is a useful caution. The model-selection metrics (PPP, marginal likelihood, fraction within 3σ) do support LOC over single-1D models, and the comparison of the model grid against Garg et al. empirical diagnostics is a solid check. The agreement of the low-mass MZR with direct-method fits (Indahl et al., Andrews & Martini) is genuinely encouraging, and the paper is unusually honest: it explicitly says results depend on the reference grid and that the Z bimodality is mostly grid-driven.\n\nOther soft spots are more minor. Excluding [SII] post-hoc improves the fit and they test [OII] as a substitute, but the headline results shift with line set; at minimum the paper needs a systematic error budget for line-set choice and grid prescriptions. The MZR polynomial (Eq. 10) has no uncertainties. The 12 hyperparameters are many, but the priors and posteriors appear handled transparently and the degeneracies are not hidden.\n\nWho this is for: anyone working on strong-line ISM modeling or unresolved high-z galaxies will want to read it. It deserves a serious referee. I recommend sending it to peer review with the requirement that the authors either add an injection-recovery test or clearly reframe the claims from \"recovery\" toward \"averages plus a plausible, prior-dependent spread.\" A single synthetic test showing that the power-law boundaries can be recovered would strengthen the central claim considerably.","headline":"A transparent and useful LOC application whose sharpest claim — recovering internal parameter distributions — is not yet demonstrated; the average values and MZR are the robust parts.","tokens_in":26887,"tokens_out":3251,"would_cite":true,"duration_ms":29824,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that the integrated spectrum of a galaxy can be inverted to recover internal distributions of metallicity, ionization parameter, and gas density.","keywords":["integrated spectroscopy","photoionization models","LOC models","mass-metallicity relation","dwarf galaxies","Bayesian inference","ionized gas","ECO survey"],"falsifier":"Take a subsample of ECO star-forming galaxies observed with integral-field spectroscopy, sum the spaxels to form the integrated spectrum, run the LOC inference, and compare the recovered power-law parameters to the actual luminosity-weighted distributions in the resolved maps; if the inferred internal metallicity dispersions (up to roughly 1 dex around solar metallicity) are not present in the resolved maps, the recovery is an artifact of the grid or topology. Similarly, direct electron-temperature metallicity measurements across the sample would settle whether the inferred metallicity scale matches reality.","tokens_in":25723,"feed_emoji":"🌌","tokens_out":5622,"duration_ms":47828,"temperature":0.7,"pith_summary":"The paper tries to establish that integrated galaxy spectra—single unresolved spectra like those coming from large surveys—contain enough information to recover the internal distribution of physical conditions in the ionized gas, not just a single average. The authors model 2052 star-forming galaxies from the volume-limited ECO catalog as sums of 1D photoionization models whose parameters (metallicity, ionization parameter, density, stellar age) follow power-law distributions with free boundaries. They find that these distributed models fit the observed strong lines better than single-component models, that most galaxies are dominated by low-excitation gas near 0.3 solar metallicity, and that the resulting mass–metallicity relation agrees with direct electron-temperature methods. If true, this means unresolved surveys can still probe how gas conditions vary inside galaxies, which matters for interpreting high-redshift and all-sky spectroscopy.","feed_headline":"Integrated spectra reveal galaxies' inner gas spread","feed_subtitle":"How 2,052 star-forming galaxies yield the mass–metallicity relation from unresolved light alone.","key_machinery":"The central object is the LOC (locally optimally emitted clouds) model: a galaxy's integrated line emission is written as the sum over many 1D photoionization clouds, each weighted by a power-law distribution in the physical parameters, with the boundary positions of that distribution as free hyperparameters. The machinery is the MULTIGRIS Bayesian framework, which evaluates such combinations on the fly with sequential Monte Carlo sampling and returns posterior distributions for the average values, the power-law slopes, and the lower and upper boundaries. The LOC equation is $L_{\\rm tot} = \\sum_p \\Phi(p)\\,I(p)\\,\\Delta(p)$, with weights of the form $\\Phi(p) = 10^{\\alpha p}$ inside $[p_{\\rm min}, p_{\\rm max}]$. This construction is what lets the paper claim it recovers distributions rather than single representative values.","core_discovery":"On its own terms, the paper claims that the optical strong lines of a star-forming galaxy are produced by a luminosity-weighted mixture of many ionized clouds with a spread of conditions, and that this spread can be recovered. The reference model treats metallicity, ionization parameter, density, and stellar age as power-law-distributed parameters with free lower and upper bounds; the inference returns probability density functions for these parameters. The average metallicity of the ECO star-forming sample peaks around 0.3 solar, the integrated emission is dominated by low-excitation gas with log U near -3.2, and the average mass–metallicity relation follows direct abundance determinations from the low-metallicity calibrated regime to high-metallicity stacks. The paper also finds that the LOC models outperform single 1D models by every metric, interprets this as evidence that physical conditions within galaxies are nonuniform, and identifies the [SII] lines as a source of systematic bias.","pith_inferences":["If integrated spectra really encode internal distributions, the same LOC machinery could be applied to high-redshift galaxies and large-area surveys that only have one spectrum per object, effectively recovering a 'poor-man's IFU' from unresolved data.","The near-zero power-law slopes for age, density, and metallicity might reflect the fact that line fluxes weight clouds by brightness, so the recovered distributions are luminosity-weighted, not volume-weighted; physical interpretations of the slopes should be treated as such.","A direct test would be to run the same inference on IFU data cubes of ECO-like dwarfs, summing the spaxels into a fake integrated spectrum and comparing the recovered power-law parameters with the actual resolved distribution.","If the [SII] problem is a grid abundance or depletion issue rather than a data issue, similar biases may affect sulfur-based diagnostics such as N2S2 in other strong-line studies."],"forward_implications":["The mass–metallicity relation inferred without [SII] constraints matches direct electron-temperature determinations from low-mass calibrations to high-mass stacks, suggesting strong-line photoionization modeling can yield reliable metallicities across the full mass range.","Because LOC models outperform single 1D models in all metrics, single-'representative' cloud analyses of integrated spectra likely miss real internal spreads of ionization and metallicity.","Most ECO star-forming galaxies are dominated by low-excitation gas near 0.3 solar metallicity, with a tight average ionization parameter around log U near -3.2 and ages peaking near 5 Myr.","The inferred internal metallicity dispersion grows from factors of 2–3 in the most metal-poor galaxies to factors of 5–10 around solar metallicity, implying metal-rich regions enrich faster than metal-poor regions.","The [SII] line doublet, when included, overpredicts [SII] and worsens [NII] and [OI] fits, so current grids and/or SDSS measurements carry a systematic that should be addressed before using [SII] as a constraint."],"supporting_citations":[{"why":"Supplies the 1D Cloudy grid methodology, BPASS SEDs, depletion patterns, and the AGN-fraction subgrid used for all inference.","marker":"Richardson et al. (2022)"},{"why":"Provides the MULTIGRIS statistical framework with sequential Monte Carlo estimation of marginal likelihood and posterior predictive p-values.","marker":"Lebouteiller & Ramambason (2022)"},{"why":"Introduces the LOC hypothesis that emission is dominated by optimally emitting clouds, the basis for the power-law distributions.","marker":"Ferguson et al. (1997)"},{"why":"Supplies the Galactic Concordance abundances that set abundance scalings and the solar standard in the model grid.","marker":"Nicholls et al. (2017)"},{"why":"Provides the BPASS stellar population models used for the ionizing SEDs in the grid.","marker":"Eldridge et al. (2017)"},{"why":"Provides the ECO–SDSS crossmatch, line flux and signal-to-noise selection, star-forming classification, and internal extinction corrections.","marker":"Polimera et al. (2022)"},{"why":"Gives the low-mass direct-method mass–metallicity relation that the inferred relation matches at low stellar masses.","marker":"Indahl et al. (2021)"},{"why":"Gives the stacked direct-method mass–metallicity relation used as the high-mass comparison.","marker":"Andrews & Martini (2013)"},{"why":"Provides empirical strong-line metallicity diagnostics used to cross-check the metallicities from the model grid.","marker":"Garg et al. (2024)"}],"fun_headline_variants":["Unresolved light reveals galaxy gas spread","Integrated spectra expose hidden gas in galaxies","From integrated light: gas spread and metallicity","New mass-metallicity fit from integrated spectra","Galaxy spectra decode gas conditions without imaging"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 1D Cloudy grid built with BPASS stellar SEDs, Nicholls et al. (2017) abundances, a fixed depletion strength, and a metallicity-dependent dust-to-gas ratio maps the observed line ratios onto the true metallicities and densities.","fun_headline_variants_meta":{"raw":{"variants":["Unresolved light reveals galaxy gas spread","Integrated spectra expose hidden gas in galaxies","From integrated light: gas spread and metallicity","New mass-metallicity fit from integrated spectra","Galaxy spectra decode gas conditions without imaging"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":3042,"prompt_tokens":1059,"completion_tokens":1983,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":675,"completion_tokens_details":{"reasoning_tokens":1926}},"tokens_in":675,"tokens_out":1983,"duration_ms":14090,"temperature":1.0,"reasoning_tokens":1926,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T11:00:54.772168+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a subsample of ECO star-forming galaxies observed with integral-field spectroscopy, sum the spaxels to form the integrated spectrum, run the LOC inference, and compare the recovered power-law parameters to the actual luminosity-weighted distributions in the resolved maps; if the inferred internal metallicity dispersions (up to roughly 1 dex around solar metallicity) are not present in the resolved maps, the recovery is an artifact of the grid or topology. Similarly, direct electron-temperature metallicity measurements across the sample would settle whether the inferred metallicity scale matches reality.","supporting_citations":[{"cited_title":"& Ramambason , L","cited_arxiv_id":null,"evidence_quote":"Provides the MULTIGRIS statistical framework with sequential Monte Carlo estimation of marginal likelihood and posterior predictive p-values."},{"cited_title":"C., Sutherland , R","cited_arxiv_id":null,"evidence_quote":"Supplies the Galactic Concordance abundances that set abundance scalings and the solar standard in the model grid."},{"cited_title":"J., Stanway , E","cited_arxiv_id":null,"evidence_quote":"Provides the BPASS stellar population models used for the ionizing SEDs in the grid."},{"cited_title":"J., et al","cited_arxiv_id":null,"evidence_quote":"Gives the low-mass direct-method mass–metallicity relation that the inferred relation matches at low stellar masses."},{"cited_title":"L., et al","cited_arxiv_id":null,"evidence_quote":"Provides empirical strong-line metallicity diagnostics used to cross-check the metallicities from the model grid."}],"review_version":1}