{"id":"123257e0-fc74-410c-8084-8878ba8e4ba0","arxiv_id":"2509.00150","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"ProMage predicts ProSpect galaxy magnitudes in five HSC bands with ~0.01 magnitude errors and a 10,000x speedup, based on 10 million simulated galaxies.","lead":"ProMage is a neural network that reproduces galaxy magnitudes computed by the ProSpect SED code, with errors below 0.02 magnitudes for 99% of test galaxies. It runs about 10,000 times faster than ProSpect, which could speed up galaxy property inference for next-generation surveys.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stage IV suitability claim rests on in-prior, AGN-free validation; no test of generalization to realistic galaxy distributions.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the reported accuracy is only validated in-distribution, on the same LHS prior used for training, and the leap to Stage IV surveys and AGN-inclusive SEDs is untested. I agree with the CONDITIONAL verdict. The emulator architecture and training details are concrete, and the in-prior accuracy is credible, but the central claim of suitability for next-generation surveys requires either an out-of-prior validation or a clear qualification of the claim. The AGN exclusion is explicitly stated in Section 2, so this is not a hidden flaw but an unsupported extrapolation. The metric inconsistency between per-mille relative accuracy and 0.01-0.02 mag absolute errors is a secondary issue that should be clarified, but it does not by itself invalidate the practical utility of the emulator. No code or trained model is released, which limits reproducibility, but that is not the most load-bearing scientific concern. A realistic validation experiment would settle whether the generalization concern actually lands, so the current CONDITIONAL verdict should stand unchanged.","tokens_in":4426,"tokens_out":4811,"duration_ms":57748,"concrete_test":"Build an out-of-prior validation set by drawing 10^5 galaxy physical properties from a cosmological simulation or a spectroscopic survey (e.g., EAGLE/IllustrisTNG or GAMA/DeepDrill), run ProSpect with AGN turned on, and compare ProMage magnitudes to ProSpect magnitudes in HSC g,r,i,z,y. If the 99th-percentile absolute error exceeds 0.02 mag, the Stage IV suitability claim fails. As a cheaper check, recompute the flux error implied by 0.02 mag (10^{-0.008}-1 ≈ 1.8%); if the abstract's per-mille claim refers to flux, it is off by an order of magnitude, and the metric must be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline accuracy numbers (Section 4) are measured on a 10^6 test set that is a random 10% split of the same Latin-hypercube sample used to generate the training set (Sections 2-3). Because LHS is space-filling and the split is random, every test point lies in a densely sampled region of the training prior; this demonstrates interpolation, not generalization. The paper then states the emulator is 'well suited' for Stage IV surveys and for forward-modelling frameworks such as GalSBI-SPS, but those applications require the emulator to work for the actual distribution of galaxies, which is not the LHS prior. Section 2 explicitly excludes AGN 'for now'; AGN contribute significant light to a large fraction of Stage IV galaxies, especially at high redshift, and ProSpect itself includes AGN emission. No test is shown for AGN-inclusive SEDs, and the claimed extension to new filters is asserted but not demonstrated. A separate metric inconsistency compounds the issue: the abstract claims per-mille relative accuracy, while the reported absolute errors of 0.01-0.02 mag correspond to roughly 1-2% flux errors; the accuracy metric needs a precise definition. The in-prior emulator is plausible, but the survey-scale suitability claim is not yet supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ProMage, a feed-forward neural network that emulates observer- and rest-frame magnitudes computed by the SED package ProSpect in the HSC g,r,i,z,y bands. The network is trained on 10^7 galaxies drawn from a Latin hypercube prior over 11 physical parameters (redshift, SFH, gas, dust), with per-band and per-frame networks. On a held-out 10^6 galaxy test set, the authors report absolute errors <0.02 mag for 99% of sources and <0.01 mag for 95%, a ~10^4 speed-up over ProSpect, and state that the emulator is well suited for Stage IV surveys and forward-modelling frameworks such as GalSBI-SPS.","tokens_in":4786,"tokens_out":3572,"duration_ms":42997,"significance":"If the reported performance is robust, ProMage would be a practically useful component for fast SED-based inference and synthetic catalogue generation: the architecture is simple, per-band training enables parallelization and extension, and the 10^6 source test set is substantial. The speed-up claim is plausible and well aligned with the needs of simulation-based inference. However, the headline accuracy metric is ambiguous and internally inconsistent, and the generalization claims go beyond the evidence presented. The core emulation result is credible as an in-prior interpolation statement, but the survey-scale suitability claim requires additional validation or explicit scoping.","major_comments":[{"comment":"The abstract claims 'per-mille relative accuracy for 99% of sources', but §4 reports absolute errors <0.02 mag for 99% and <0.01 mag for 95%. In flux units, 0.01 mag is ~0.9% and 0.02 mag is ~1.8%, not 0.1% (per-mille). The relative accuracy metric is never defined. This is the headline quantitative claim and must be corrected or redefined (e.g., as median absolute percentage flux error or as a stated magnitude tolerance).","section":"Abstract and §4"},{"comment":"The validation is entirely in-prior: the test set is a random 10% split of the same Latin hypercube sample used for training, and the LHS is space-filling, so every test point lies in a densely sampled region of the training prior. This demonstrates interpolation, not generalization to the actual galaxy distribution. The statement that ProMage is 'well suited' for Stage IV surveys and for GalSBI-SPS (§4) is therefore not yet supported. Provide an out-of-distribution test, e.g., comparing against ProSpect on a realistic mock catalogue or on the posterior draws produced by GalSBI-SPS, or at minimum a leave-out-region test, and discuss the implied accuracy outside the prior.","section":"§3–§4"},{"comment":"AGN are explicitly excluded 'for now' from the training set, yet ProSpect itself includes AGN emission and a large fraction of Stage IV galaxies host AGN, particularly at high redshift. The suitability claim in §4 is unqualified. Either extend the training to AGN-inclusive SEDs, or explicitly scope the claim to non-AGN galaxies and assess the expected impact on Stage IV applications.","section":"§2"},{"comment":"Quantitative results are shown only for the g and i bands; the statement that 'similar performance occurs for the other HSC optical bands' is asserted without data. Since the abstract claims per-mille accuracy 'across the g,r,i,z,y bands', please provide a table or additional panels reporting the 95th and 99th percentile absolute errors for r, z, and y in both observer and rest frames.","section":"Fig. 1 / §4"}],"minor_comments":[{"comment":"The prior range for mpeak is given as '[−(2 + tlb), 13.4 − tlb]' but tlb is never defined in the text. Please define this quantity or replace with an explicit numerical range. Also, there is a typo in 'αSF,,screen'.","section":"Table 1"},{"comment":"The definition of rest-frame magnitudes is not stated. Please specify how the rest-frame band is computed (e.g., filter transmission shifted to rest wavelength, K-correction conventions, and whether the same filter set is used). This is needed for reproducibility.","section":"§2–§3"},{"comment":"The timing statements '10^5 sources in less than half a second' (Introduction) and '10^6 sources in under three seconds' (Results) are roughly consistent but should be accompanied by the exact hardware and timing methodology (e.g., batch size, number of repeated runs) for a fair speed comparison with ProSpect.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"This is a competent emulator paper for a proceedings volume. The central in-prior emulation result is credible, but the accuracy-metric inconsistency and the overreach to Stage IV suitability need to be fixed. The requested revisions are feasible and do not require new conceptual machinery. I see no grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a workmanlike engineering contribution, not a new physical result. ProMage is a per-band feed-forward network that emulates ProSpect observer- and rest-frame magnitudes in HSC g,r,i,z,y, trained on 10^7 Latin hypercube samples. The 10^4 speed-up and sub-0.02 mag errors are plausible and consistent with Figure 1. The authors give concrete architecture and training details, and they cite the relevant prior emulator work (Alsing et al., Thorp et al.) fairly. That earns real credit.\n\nThe soft spots are all about scope and framing. First, the abstract says 'per-mille relative accuracy' but 0.02 mag is about 2% in flux; the metric needs a precise definition and the abstract overstates it. Second, the test set is a random split of the same Latin hypercube used for training, so the headline accuracy demonstrates interpolation, not generalization to the true galaxy distribution. The Stage IV suitability claim would need validation on a realistic population, ideally from a different SED code or from real photometry. AGN are explicitly excluded, and no test is shown for filters beyond g and i, even though r,z,y are claimed by analogy. Third, no code, data, or trained weights are released, so the numbers can't be independently checked.\n\nNone of these are fatal. They are standard issues for an emulator paper, and they are addressable: fix the metric wording, add an out-of-distribution validation set, and release the artifact. The core idea is sound and the engineering is real. This is a useful paper for anyone doing SED-based inference or forward modelling with ProSpect, and it deserves a proper referee rather than a desk rejection. I'd send it to review for a full-length version, with a request for clarification and an independent validation test.","headline":"ProMage is a solid, useful emulator for ProSpect magnitudes, but the per-mille accuracy claim is overstated and the Stage IV suitability rests on in-prior validation only.","tokens_in":5231,"tokens_out":1788,"would_cite":false,"duration_ms":21894,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A feed-forward network emulates ProSpect galaxy magnitudes to <0.02 mag for 99% of test sources, running about 10,000 times faster.","keywords":["galaxies: stellar content","galaxies: fundamental parameters","methods: numerical","galaxies: statistics","neural network emulator","spectral energy distribution","Hyper Suprime-Cam","forward modelling"],"falsifier":"Run ProMage on 10^6 ProSpect realisations sampled from a prior that includes AGN components or extends to z > 5, and check whether more than 1% of emulated magnitudes deviate from ProSpect by more than 0.02 mag; also compare emulated magnitudes against real HSC photometry for a sample of galaxies with independently fitted physical parameters.","tokens_in":4383,"feed_emoji":"🔭","tokens_out":5626,"duration_ms":59703,"temperature":0.7,"pith_summary":"The paper introduces ProMage, a feed-forward neural network that reproduces the observer- and rest-frame galaxy magnitudes computed by the ProSpect spectral energy distribution code. Trained on 10^7 ProSpect runs, it predicts magnitudes from physical galaxy parameters (redshift, star formation history, dust, gas) and reaches absolute errors below 0.02 magnitudes for 99% of a 10^6-galaxy test set across the HSC g, r, i, z, y bands, roughly 10^4 times faster than ProSpect. The authors argue this speed makes feasible large-scale forward-modelling and Bayesian inference of galaxy properties for Stage IV surveys such as Euclid and Rubin-LSST.","feed_headline":"Galaxy magnitudes 10,000x faster with <0.02 mag error","feed_subtitle":"ProMage reproduces ProSpect SED magnitudes to per-mille accuracy on HSC bands, enabling survey-scale inference.","key_machinery":"A feed-forward neural network with five hidden layers (512-256-128-64-32 neurons) mapping 11 physical ProSpect inputs (redshift, star formation history parameters, dust and gas parameters, metallicity) to a single magnitude per band, trained per band with mean squared error loss on raw magnitudes and the Alsing et al. (2020) activation function. The speed-up comes from replacing a full SPS evaluation (tens of milliseconds in ProSpect) with one matrix multiplication chain (microseconds).","core_discovery":"The central claim is that a modest feed-forward network can stand in for a full stellar population synthesis calculation for the purpose of computing photometric magnitudes, achieving per-mille accuracy and a four-orders-of-magnitude speed-up. ProMage predicts each band independently from the same 11 physical inputs, trained per band with a custom activation function. On a held-out test set of 10^6 sources, 99% of emulated magnitudes lie within 0.02 mag of ProSpect's values in both observer and rest frames in all five HSC bands, below typical photometric zero-point uncertainties of current surveys. The paper positions this as a practical enabler of amortised simulation-based inference and MC","pith_inferences":["If the <0.02 mag accuracy holds at the edges of the prior, extending ProMage to galaxies with AGN or different dust geometries would likely require retraining or an added correction term; the paper notes AGN are excluded.","Because the emulator is fast and differentiable, it could be coupled directly to gradient-based inference or used in simulation-based calibration of survey systematics.","The accuracy numbers are reported on a Latin-hypercube test set; real galaxy photometry includes noise and selection effects, so end-to-end validation on observed HSC data would strengthen the generalization claim."],"forward_implications":["ProMage can compute photometric magnitudes for billions of galaxies in minutes on a single CPU, enabling large synthetic survey realisations for weak-lensing redshift calibration.","It makes amortised simulation-based inference practical: networks can be trained over the emulator to invert galaxy properties from observed magnitudes.","Observer- and rest-frame magnitudes in HSC g, r, i, z, y are cheap enough that MCMC chains over SPS parameters become feasible at survey scale.","The per-band, independent design means adding new filters or expanding parameter ranges only requires retraining that band, not the whole model."],"supporting_citations":[{"why":"ProSpect is the generative SED code whose observer- and rest-frame magnitudes ProMage is trained to emulate; it defines the target outputs and input parameter space.","marker":"Robotham et al. (2020)"},{"why":"Provides the neural-network emulation approach and the activation function adopted for ProMage, plus per-band training rationale.","marker":"Alsing et al. (2020)"},{"why":"A comparable SED emulator that supports the per-band architecture choice and serves as a performance reference.","marker":"Thorp et al. (2025)"},{"why":"GalSBI-SPS, the forward-modelling framework in which ProMage is already integrated and the motivating application for speed.","marker":"Tortorelli et al. (2025)"},{"why":"Simulation-based inference framework that ProMage is intended to make scalable for galaxy property inference.","marker":"Cranmer et al. (2020)"},{"why":"Expected photometric precision of Rubin-LSST used as the accuracy benchmark that ProMage's errors are compared against.","marker":"Crenshaw et al. (2024)"},{"why":"Supplies the Progeny single stellar population models underlying the ProSpect SEDs, defining the physics the emulator learns.","marker":"Robotham & Bellstedt (2025)"}],"fun_headline_variants":["ProMage: Galaxy magnitudes 10,000x faster, 0.02 mag accuracy","SED emulation speeds up galaxy magnitudes by 10,000x","Neural net nails galaxy magnitudes to 0.02 mag","Galaxy photometry 10^4x faster with neural emulator","ProSpect magnitudes emulated with milli-mag error"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"Accuracy is only demonstrated on test galaxies drawn from the same Latin hypercube prior used for training, so the emulator's performance on galaxies outside that prior—higher redshifts, AGN-inclusive SEDs, or different filter sets—is untested.","fun_headline_variants_meta":{"raw":{"variants":["ProMage: Galaxy magnitudes 10,000x faster, 0.02 mag accuracy","SED emulation speeds up galaxy magnitudes by 10,000x","Neural net nails galaxy magnitudes to 0.02 mag","Galaxy photometry 10^4x faster with neural emulator","ProSpect magnitudes emulated with milli-mag error"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000193,"raw_usage":{"total_tokens":1145,"prompt_tokens":661,"completion_tokens":484,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":405,"completion_tokens_details":{"reasoning_tokens":389}},"tokens_in":405,"tokens_out":484,"duration_ms":5854,"temperature":1.0,"reasoning_tokens":389,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T13:52:49.786487+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run ProMage on 10^6 ProSpect realisations sampled from a prior that includes AGN components or extends to z > 5, and check whether more than 1% of emulated magnitudes deviate from ProSpect by more than 0.02 mag; also compare emulated magnitudes against real HSC photometry for a sample of galaxies with independently fitted physical parameters.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ProSpect is the generative SED code whose observer- and rest-frame magnitudes ProMage is trained to emulate; it defines the target outputs and input parameter space."},{"cited_title":"2020, The Astrophysical Journal Supplement Series, 249, 1,","cited_arxiv_id":null,"evidence_quote":"Provides the neural-network emulation approach and the activation function adopted for ProMage, plus per-band training rationale."}],"review_version":1}