{"id":"21d411ca-731d-4855-b5a4-6eb09fdd97b0","arxiv_id":"2608.12471","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"UMAP embeddings of 13-band JWST and HST photometry preserve enough information to statistically rank five galaxy formation models against observed GOODS-S galaxies, with JAGUAR performing best.","lead":"This paper tests whether a 2D map of galaxy brightness in thirteen telescope filters can distinguish between five computer models of galaxy formation. It finds the map preserves enough information to rank the models, with JAGUAR matching bright galaxies in the GOODS-S field about six times better than the next best model.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never states whether UMAP is fitted to the observed sample alone or jointly to observed plus all model lightcones; if the latter, every S score and the JAGUAR factor ratios depend on the model sample and are not a fixed property of the observations.","rationale":"The reader's weakest_assumption focuses on catalogue comparability, which the paper tests extensively through source-extraction variations and noise scaling. The more decisive gap is the unspecified UMAP training protocol, because it controls the coordinate system in which every score is computed. The paper's own text (Section 2.3 combined sample; Section 3.3 'position of each galaxy') suggests a joint fit, but no statement or code confirms it. If joint, the observed embedding is not fixed, and the S-score ratios are not a well-defined property of the models; this is exactly the kind of condition that should be resolved before the results are used as model constraints. The paper otherwise has considerable strengths: 500 stochastic iterations, hyperparameter sensitivity checks, a degradation layer that gives physical interpretability, and an honest acknowledgement that JAGUAR is partly built from the observed field. The recommended verdict stays CONDITIONAL, matching the reader, with the added condition that the training protocol be specified and demonstrated not to control the ranking.","tokens_in":24952,"tokens_out":11941,"duration_ms":119619,"concrete_test":"Using the released fiducial catalogues and code, rerun the full pipeline under two protocols: (A) fit UMAP on the 4,590 observed galaxies only, then project the 40,802 model galaxies with UMAP's transform; (B) fit UMAP jointly on the combined 45,392-object sample as the manuscript appears to imply. Compare the resulting S_e and S_c values, the model ranking, and the JAGUAR/sc-sam and JAGUAR/sage ratios. If the ranking or the factor-of-6/12 claim changes by more than the quoted 16th-84th percentile ranges, the central differentiation claim is not robust to the training protocol and the manuscript must state which protocol is correct.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper leaves unspecified the most consequential step of the analysis: what data are used to fit the UMAP embedding that underlies Figures 4 and 6. Section 3.3 describes fitting UMAP models and 'recording the position of each galaxy', while Section 2.3 explicitly combines 40,802 simulated and 4,590 observed galaxies into a single sample of 45,392 objects. If the fit is performed on this combined sample, the observed coordinates are not intrinsic to the observations: they are co-optimized with the model galaxies, and the relative number of galaxies per model (sage is said to contribute more than twice as many as jaguar) influences the graph and the resulting embedding. The bin-wise s_i and summary S_e and S_c then measure how well each model reproduces a space that was partially constructed from that same model, not how well it reproduces a fixed observed manifold. The headline factors ('six times as well', 'twelve times as well') are ratios of these embedding-dependent scores. The absence of any statement about fit versus transform, and the absence of the code (Data Availability promises links only after acceptance), makes this a validity and reproducibility gap in the central claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a method for comparing observed galaxy multi-band photometry to theoretical model predictions without SED fitting. The authors reduce 13-band HST/JWST fluxes of 4590 bright (m_F444W<26) galaxies in GOODS-S to 2D UMAP embeddings, add synthetic noise to five public model lightcones (JAGUAR, SC-SAM, SAGE, SPRITZ1, SPRITZ4), and evaluate each model with a bin-wise chi-square-like score S computed with density (euclidean metric) or fraction (cosine metric) in the embedding. They report that JAGUAR reproduces the observed population best (S_e=408), followed by SC-SAM (S_e~2530), SAGE (4730), and SPRITZ, and that rankings are robust to source-extraction and hyperparameter variations. They also analyze discrepant bins with NIRSpec spectra, identify extreme emission-line galaxies as a model failure, and compare UMAP to PCA and to Bayesian SED fitting in speed.","tokens_in":25191,"tokens_out":8985,"duration_ms":79489,"significance":"If the method is validated, it offers a fast, unbiased statistical constraint on galaxy formation models from direct observables, with natural scaling to Euclid/LSST and CAMELS simulation-based inference. The paper's strengths include a careful treatment of observational systematics (45 SE parameter variations, alternative extraction codes, noise scaling), 500 UMAP iterations to quantify stochasticity, and a concrete physical follow-up using spectroscopy. However, the headline claims are currently stronger than the evidence: the JAGUAR comparison is partly circular, and the UMAP embedding procedure is not fully specified. The core methodological idea is promising and the systematics work is unusually thorough, but the central numerical results need clarification and re-analysis before the paper can be accepted.","major_comments":[{"comment":"The manuscript does not state what data are used to fit the UMAP model. Section 2.3 reports a combined sample of 45,392 objects (40,802 simulated and 4,590 observed), and Section 3.3 says the models are fit with positions recorded for every galaxy; if the embedding is trained jointly on simulated and observed galaxies, the observed coordinates are not intrinsic to the observations but depend on the composition and size of the model sample. In that case the bin-wise scores and the factor ratios (e.g., SAGE/JAGUAR ~12) measure each model against a space partly constructed from that model, and the relative model abundances in the fit (SAGE contributes more than twice as many galaxies as JAGUAR) can bias the manifold. Please specify whether the UMAP is fitted to the observed sample alone and models are transformed through that fixed embedding, or fitted jointly; if the latter, re-run with an observed-only fit and show whether the rankings and ratios persist.","section":"§2.3 and §3.3"},{"comment":"The paper states that orientation within the UMAP space can vary across the 500 iterations and that median positions are used for plotting and for reporting s_i. If the coordinate axes are arbitrary rotations or reflections, per-coordinate medians are not well-defined unless the embeddings are aligned (e.g., by Procrustes analysis), and no alignment step is described. Please either describe the alignment used, or restrict the median and percentile reporting to the per-iteration S values and show the maps for a single representative iteration.","section":"§3.3"},{"comment":"JAGUAR's SEDs for massive z<4 galaxies are matched to 3D-HST sources, and GOODS-S is a significant part of 3D-HST; the abstract's 'six times as well' and 'twelve times as well' do not constitute an independent test of JAGUAR. The paper acknowledges this in §4.1, but the abstract and conclusions present the factor ratios without the caveat. Please add the caveat to the abstract and conclusions, or re-frame the headline as demonstrating the method rather than as independent evidence for JAGUAR.","section":"§2.2.2 and Abstract"},{"comment":"The five lightcones differ not only in galaxy-formation physics but also in simulation volume and area, mass resolution, SPS models, dust prescriptions, and the presence or absence of nebular emission (e.g., SAGE lacks photoionisation). Because S_e includes number counts and S_c only removes the normalization, the ranking could be dominated by these catalogue-level differences rather than by the physics the paper discusses, such as SNe feedback efficiency in SAGE or template limitations in SPRITZ. The authors should either add a control analysis that isolates these factors (for example, matching redshift ranges or comparing models processed through the same forward-modelling pipeline), or clearly state that the scores compare the public model catalogues and cannot uniquely attribute the discrepancies to specific physical processes.","section":"§2.2 and §4.1-4.2"}],"minor_comments":[{"comment":"The SC-SAM summary score is reported as S_e = 2532+162-75 in §4.1 but as S_e = 2578+504-210 in the Conclusions; please make these values consistent.","section":"§5 vs §4.1"},{"comment":"Equation (1) is garbled in the displayed formula (an extra 'q' appears in the numerator and before the square root); please check the typesetting of this equation.","section":"Eq. (1)"},{"comment":"The abstract uses m_AB<26 while the text defines the cut as m_F444W<26; please clarify whether the magnitude limit is F444W-specific or a general AB magnitude.","section":"Abstract and §2.3"},{"comment":"The statement that a reduced chi-square of about 3 'confirms excellent replication' is surprising given that only Poisson uncertainties are included; please explain why this value is considered excellent in this context.","section":"§4.1"},{"comment":"The speed comparison compares a full 500-iteration UMAP fit to a single Bagpipes fit, but the relevant operation for a new survey galaxy is transforming through a fixed embedding rather than refitting; please clarify the comparison so the '>100 times faster' claim is not misleading.","section":"§3.3"},{"comment":"The manuscript says links to the fiducial catalogues, embedded positions, and analysis code will be added upon acceptance; for a methods paper, providing the code and catalogues with the submission, or at least a detailed pseudocode for the UMAP fitting and alignment steps, would substantially aid reproducibility.","section":"Data Availability"},{"comment":"The sentence 'A 121 arcmin2 realisation of includes galaxies' is missing a word (presumably JAGUAR); please correct this typo.","section":"§2.2.2"}],"recommendation":"major_revision","confidential_remarks":"The key issue for me is the unspecified UMAP fitting sample and the lack of an alignment description for the 500 iterations; both are fixable in revision, but without them the central numbers are not reproducible. I would not reject the paper: the systematics treatment is genuinely thorough and the method has clear value. I would ask for a re-analysis with an observed-only embedding, a clear statement on JAGUAR's circularity, and correction of the internal score inconsistency before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper you'd want to know about: FLAGS II applies UMAP to direct observables (13-band HST+JWST photometry) from GOODS-S and five public galaxy formation model lightcones, then scores how well each model reproduces the observed distribution in 2D embedding space with a chi2-like statistic S. The punchline is that JAGUAR, a semi-empirical model whose SEDs are partly drawn from observed GOODS-S galaxies, scores \"six times as well\" as SC-SAM, and the paper shows that linear PCA would give different rankings. That ranking-change result is the most interesting thing here, because it has direct consequences for simulation-based inference.\n\nWhat's new and good: This is a serious, careful application of a known method. The systematics work is genuinely thorough—45 source-extraction variations, three extraction codes, noise scaling, UMAP hyperparameters, 500 stochastic iterations—and the conclusion that rankings survive these variations is well supported. The degradation layer (drop one filter, watch S change) is a nice tool for tracing discrepancies to specific bands. The writing is clear, and the paper is upfront that JAGUAR's lead is partly built-in.\n\nSoft spots, in order of severity. First, the paper never states whether UMAP is fit to the observed sample alone or jointly to observed plus all model galaxies. Everything in the text points to a joint fit: a combined sample of 45,392 objects, \"recording the position of each galaxy.\" If that's what was done, the embedding space is co-constructed with the models, so the absolute S values and the \"six times\" factor are properties of the specific model mix, not of the observations. The authors need to say which it is, and ideally fit on observations and transform models, or at least show the embedding is stable across model subsamples. Second, the headline JAGUAR result is acknowledged as circular but the abstract still sells it as a constraint. It's a demonstration, not a measurement. Third, code and data aren't public yet. That's fixable, but it currently blocks reproduction.\n\nNone of this breaks the central methodological claim. The paper is a solid, useful demonstration that UMAP on direct observables can statistically differentiate forward models, and the PCA comparison is a real cautionary result.\n\nWho it's for: anyone doing forward-model comparison or simulation-based inference with large photometric surveys. It deserves a serious referee, but it needs a revision that clarifies the fit-vs-transform question before acceptance. Recommendation: send to peer review, with the fit question and code release as conditions.","headline":"Solid methods paper on UMAP-based direct-observable model comparison; the PCA-vs-UMAP ranking flip is the real result, but the embedding-fit ambiguity and JAGUAR circularity need fixing before the numbers are interpreted as constraints.","tokens_in":25747,"tokens_out":5111,"would_cite":true,"duration_ms":42465,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 2D map of galaxy fluxes separates five galaxy formation models without SED fitting.","keywords":["galaxy formation models","forward modelling","dimensionality reduction","UMAP","JWST photometry","HST photometry","GOODS-S","SED fitting alternative"],"falsifier":"Treat one model's lightcone as if it were the observed catalogue: apply the same synthetic-noise model and $m_{\\mathrm{F444W}}<26$ cut, and rank the other models against it with the identical UMAP and $S$ pipeline. If the model that generated the pseudo-observations does not achieve the lowest score, the ranking is measuring catalogue artefacts rather than model physics; the claim would also be threatened if a different random seed or a 3D UMAP embedding changed the ordering of the five models.","tokens_in":24727,"feed_emoji":"🛰️","tokens_out":7169,"duration_ms":63013,"temperature":0.7,"pith_summary":"This paper aims to show that a two-dimensional map of galaxy fluxes, built with the non-linear dimensionality reduction algorithm UMAP, keeps enough information to tell five galaxy formation models apart directly from telescope photometry, with no SED fitting. The authors compile 4590 bright galaxies in GOODS-S observed in thirteen JWST and HST bands, inject observationally matched noise into five public model lightcones, embed everything in the same 2D space, and score each model with a $\\chi^2$-like statistic. They report that JAGUAR matches the bright population six times as well as SC-SAM and twelve times as well as SAGE, while template-based SPRITZ and SAGE's missing photoionisation produce the largest discrepancies. If true, statistical model constraints become a fast, bias-free step: locating an object in the embedding is more than 100 times quicker than Bayesian SED fitting.","feed_headline":"A 2D map of galaxy colors ranks five formation models","feed_subtitle":"UMAP on JWST/HST fluxes says JAGUAR beats SC-SAM sixfold and SAGE twelvefold on bright GOODS-S galaxies.","key_machinery":"The engine of the argument is UMAP, a non-linear dimensionality reduction algorithm that assumes the high-dimensional data lie uniformly on a locally connected manifold, builds a fuzzy nearest-neighbour graph, and finds a 2D projection that best preserves that graph. Working in raw flux space rather than magnitudes lets non-detections and negative values enter, and a cosine variant of the distance metric isolates SED shape from overall brightness. On top of the embedding sits the score $S=\\sum_i s_i^2$ with $s_i=(q_{\\mathrm{obs}}-q_{\\mathrm{sim}})/\\sqrt{\\sigma_{\\mathrm{obs}}^2+\\sigma_{\\mathrm{sim}}^2}$, computed in 20 by 20 bins; $q$ is sky density $\\rho$ for the euclidean embedding and sample fraction $f$ for the cosine embedding. The embedding converts a 13-dimensional sparse space into a binnable manifold, and the score turns map differences into model rankings. A degradation layer, removing one filter at a time and rerunning, points to which wavelengths drive a model's discrepancies.","core_discovery":"The central claim is that the information needed to differentiate galaxy formation models survives compression to two dimensions. Binning the 13-band flux space directly is impractical: even with deciles per dimension the space is sparsely populated, and a full 13-dimensional grid demands prohibitive memory. UMAP embeds raw fluxes, including negative values, into a horseshoe-shaped 2D manifold whose long axis tracks apparent F444W brightness and whose inner edge and substructure encode SED shape. Comparing observed and model occupancy in 20 by 20 bins with the score $S=\\sum_i s_i^2$, using sky density for a euclidean metric or sample fraction for a cosine metric, yields a stable ranking: JAGUAR scores lowest, and the paper reports it reproduces the bright GOODS-S population six times as well as SC-SAM and twelve times as well as SAGE. The cosine-embedding maps show that template-based SPRITZ cannot fill large regions of the space, and that both SC-SAM and SAGE fail to produce a bin dominated by dust-poor starbursts at $z\\approx2.5$\\,--\\,$3.5$ with [O\\,III] equivalent widths above 750\\,\\AA; the paper argues this points to missing binaries, negligible nebular emission in SAGE, and coarse snapshot cadence in the underlying dark-matter simulation.","pith_inferences":["Because JAGUAR's spectra are partly drawn from GOODS-S itself through the 3D-HST catalogue, its top rank may partly encode the training field; applying the same pipeline to a semi-empirical model calibrated on a different field would separate method from memory.","The cosine-metric score controls for number counts but not for redshift distributions, so the reported SED-shape failures of SC-SAM and SAGE could partly reflect their predicted redshift distributions rather than spectral physics alone.","A direct extension would be to run the same embedding on mock lightcones drawn from each model's own SEDs to calibrate the null distribution of $S$, converting the sixfold and twelvefold ratios into a significance statement.","The method doubles as an outlier and contamination finder, as demonstrated by the separated dusty star-forming galaxy at $z\\approx7.8$ that is mimicked by low-redshift dusty dwarfs in embedding space."],"forward_implications":["Statistical constraints on galaxy formation models can be derived from observer-frame photometry alone, bypassing SED-fitting biases and completeness corrections.","The same pipeline can scale to large surveys such as LSST and Euclid and to simulation suites with many parameter variations, where per-galaxy Bayesian fitting is computationally prohibitive.","Non-linear dimensionality reduction should be preferred over linear PCA for observer-frame model comparison, because PCA artificially lowers the apparent disagreement and can even change the ranking of models.","The method localises physical failures: template-based SED generation and missing photoionisation produce specific unoccupied regions of embedding space that can be traced back to particular filters and galaxy populations.","Rest-frame or grism-based versions of the approach could make the degradation layer substantially more informative for identifying the physics behind model discrepancies."],"supporting_citations":[{"why":"Supplies the UMAP dimensionality reduction algorithm that is the core of the analysis.","marker":"McInnes et al. 2018"},{"why":"Provides the JAGUAR semi-empirical model lightcones and spectra used in the comparison.","marker":"Williams et al. 2018"},{"why":"Provides the SC-SAM GOODS-S lightcone photometry used to evaluate that model.","marker":"Yung et al. 2022"},{"why":"Provides the Theoretical Astrophysical Observatory infrastructure used to generate the SAGE lightcone.","marker":"Bernyk et al. 2016"},{"why":"Provides the SPRITZ template-based lightcones whose model is evaluated.","marker":"Bisigello et al. 2021"},{"why":"Supplies the Bagpipes Bayesian SED fitting code used for the speed comparison.","marker":"Carnall et al. 2018"}],"fun_headline_variants":["UMAP on galaxy flux maps JAGUAR 6x better than SC-SAM, 12x SAGE","UMAP 2D flux map ranks five galaxy models, JAGUAR top","Compress 13 photometric bands to 2D, rank five galaxy formation models","JAGUAR best matches bright galaxies in 2D UMAP of JWST/HST fluxes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The five public model lightcones, after synthetic noise injection and the $m_{\\mathrm{F444W}}<26$ cut, are directly comparable to the observed GOODS-S catalogue, so differences in the score $S$ reflect galaxy formation physics rather than differences in simulation volume, mass resolution, stellar-population templates, dust prescriptions, or SAGE's missing nebular emission.","fun_headline_variants_meta":{"raw":{"variants":["UMAP on galaxy flux maps JAGUAR 6x better than SC-SAM, 12x SAGE","UMAP 2D flux map ranks five galaxy models, JAGUAR top","Compress 13 photometric bands to 2D, rank five galaxy formation models","JAGUAR best matches bright galaxies in 2D UMAP of JWST/HST fluxes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001283,"raw_usage":{"total_tokens":5314,"prompt_tokens":1088,"completion_tokens":4226,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":704,"completion_tokens_details":{"reasoning_tokens":4125}},"tokens_in":704,"tokens_out":4226,"duration_ms":25104,"temperature":1.0,"reasoning_tokens":4125,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:07:49.630729+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Treat one model's lightcone as if it were the observed catalogue: apply the same synthetic-noise model and $m_{\\mathrm{F444W}}<26$ cut, and rank the other models against it with the identical UMAP and $S$ pipeline. If the model that generated the pseudo-observations does not achieve the lowest score, the ranking is measuring catalogue artefacts rather than model physics; the claim would also be threatened if a different random seed or a 3D UMAP embedding changed the ordering of the five models.","supporting_citations":[],"review_version":1}