{"id":"fe50e894-823d-4a49-a4f2-15b9d6d3c4a9","arxiv_id":"2506.08995","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Simulated JWST/JADES observations of model proto-spheroids at z>=1.5 show stellar masses recoverable to 0.1 dex, with accurate photometric redshifts, while far-infrared data remain essential for star formation rates.","lead":"This paper simulates how well the James Webb Space Telescope's JADES survey in GOODS-South will measure the masses, star formation rates, and dust content of distant dusty galaxies that grew into today's massive elliptical galaxies. It matters because it tests whether JWST can find and characterize these galaxies down to masses roughly ten times lower than the pre-JWST limit.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stellar-mass recovery claims are conditional on the simulated SEDs matching the fitting templates; the paper's own SFH test shows order-of-magnitude SFR sensitivity, and no independent validation of mass recovery exists.","rationale":"The reader identified the model-to-model circularity as the weakest assumption; my analysis agrees. The paper's external validations are real but limited: the LF comparison under-predicts the faint end, the colour-colour test shows the model occupies a narrow sub-region of the observed distribution, and the ASTRODEEP comparison is not an independent accuracy test because both observed and simulated samples are fit with the same pipeline. The paper's own SFR sensitivity test is strong evidence that the recovered physical properties depend on the assumed SFH family, and the same fragility is plausible for stellar masses even though the authors do not report an analogous test. The central 0.1 dex dispersions are therefore conditional on the simulated SEDs being representative, and the abstract does overstate the scope by saying 'majority of proto-spheroids' when the best numbers apply to a Herschel-selected sub-sample after outlier removal. These are not grounds for rejection because the paper is transparent, includes useful quantitative forecasts, and the authors flag several limitations themselves. A conditional accept with requests for independent validation and full-sample reporting remains the appropriate verdict, so I recommend no change from the reader's conditional verdict.","tokens_in":24425,"tokens_out":8453,"duration_ms":93018,"concrete_test":"Take the public JADES GOODS-S NIRCam photometry for sources with spectroscopic redshifts (e.g., JADES NIRSpec or existing ALMA/CO redshifts), select massive (log M*_M_sun > 10) star-forming galaxies with S/N in all nine NIRCam bands comparable to the simulated sample, and run the exact EAZY + CIGALE configuration of Section 3.3. Compute the photo-z outlier fraction and the stellar-mass scatter against spectroscopic redshifts and against masses from an independent SED fit (e.g., Prospector or MAGPHYS using full UV-to-FIR coverage). If the observed outlier fraction exceeds about 0.1 or the mass scatter exceeds about 0.3 dex, the simulation-based 0.1 dex claim is not transferable to real data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—JWST photometry yields |dz|/(1+z)<0.15 outlier fraction 0.05 and stellar-mass scatter 0.1–0.2 dex for z≥1.5 proto-spheroids—is a model-recovery result. The simulated SEDs in Sections 2.1 and 2.3 are produced with the same class of physical ingredients that CIGALE fits: BC03 SSPs, energy-balance dust, a smooth delayed-type SFH, and Fritz et al. (2006) AGN templates. This means the fitter is asked to recover parameters within the generative model family. The external checks do not break the loop: the mid-IR luminosity function comparison (Figure 1) under-predicts the low-luminosity end; the NIRCam colour-colour comparison (Figure 2) shows simulated galaxies confined to a narrower region than the observed CEERS population; and the ASTRODEEP comparison (Section 4.3) applies the same EAZY/CIGALE pipeline to both observed and simulated samples, so any systematic template bias cancels and cannot validate absolute recovery accuracy. The paper itself provides the sharpest evidence of fragility in Section 4.2.2: changing the CIGALE SFH prior from sfhdelayed to sfhdelayedbq changes recovered SFRs by an order of magnitude. No equivalent stress test is shown for the headline M* dispersions. If real proto-spheroids have wider SFH, AGN, or dust variations than this single model family, the quoted outlier fractions and mass scatters are optimistic rather than representative.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper uses the Cai et al. (2013) / Mitra et al. (2024) physical model of proto-spheroids to simulate JADES/GOODS-S observations, then applies EAZY and CIGALE to recover photometric redshifts, stellar masses, SFRs, dust luminosities, and dust masses. It reports photo-z outlier fractions of 0.05 with JWST alone and 0.019 with JWST+HST for the parent sample (0.042 and 0.008 for the DSFG sub-sample), stellar-mass dispersions of about 0.2 and 0.14 dex from JWST alone (down to 0.1 dex for the DSFG sample after adding HST and removing outliers), and demonstrates that far-IR photometry is needed to constrain SFRs. It also constructs a NIRCam-selected DSFG sample from the ASTRODEEP-JWST catalog and compares CIGALE-derived masses and SFRs with the simulations, finding broad consistency and claiming that JWST can detect DSFGs down to about 10^10 solar masses.","tokens_in":24731,"tokens_out":5338,"duration_ms":53549,"significance":"The forecasting question is timely, and the paper deserves credit for reporting honest scatter, outlier, and bias metrics, and for clearly separating what JWST can constrain (stellar masses, photo-z) from what it cannot (SFRs without FIR data). The central result, stellar masses recovered to 0.1-0.2 dex, is physically plausible if the generative model spans the real diversity of proto-spheroids. However, the study is a model-to-model recovery test: the simulated SEDs are generated with the same class of physical ingredients and templates that CIGALE fits. The external checks in Figures 1 and 2 and Section 4.3 are suggestive but do not independently validate absolute accuracy, and the paper's own SFH sensitivity test in Section 4.2.2 shows an order-of-magnitude variation in recovered SFRs. With additional robustness tests and an independent observed-sample validation, this would be a useful reference for JWST DSFG studies.","major_comments":[{"comment":"The paper reports that switching CIGALE's SFH module from sfhdelayed to sfhdelayedbq changes recovered SFRs by an order of magnitude. This is presented as a caveat for SFRs, but no analogous stress test is shown for stellar masses, which are the headline quantities of Section 4.2.1 and the conclusions. Because the generative model and the fitting model share the same delayed-SFH family, the quoted 0.1-0.2 dex mass dispersions likely understate sensitivity to template assumptions. Please add a robustness test in which the simulated SEDs are generated with a different SFH, dust attenuation law, or AGN prescription, or in which CIGALE is run with deliberately mismatched templates, and quote the resulting stellar-mass dispersions and biases.","section":"Sec. 4.2.2"},{"comment":"The recovery analysis is closed-loop: the simulated SEDs are built with the da Cunha et al. (2008) energy-balance formalism and Fritz et al. (2006) AGN templates, and CIGALE fits the same formalism. The external comparisons do not break the loop: Figure 1 shows the model under-predicts the low-luminosity end of the mid-IR luminosity functions, Figure 2 shows the simulated population occupies a narrower colour-colour region than the CEERS sources, and Section 4.3 applies the same EAZY/CIGALE pipeline to both observed and simulated samples, so systematic template biases cancel. To support the absolute accuracy claims, please validate against an independent benchmark, for example by fitting observed galaxies with spectroscopic redshifts and comparing CIGALE stellar masses with dynamical masses or with a second SED-fitting code, and state the implied systematic floor on stellar mass.","section":"Secs. 2.1, 2.3, 4.2, 4.3"},{"comment":"The reduced dispersions of 0.15 and 0.1 dex are quoted after removing catastrophic photo-z outliers. In a real survey the true redshift is unknown, so an outlier-removal criterion based on |∆z|/(1+z) cannot be applied. Please specify a practical, observable outlier-rejection strategy, such as posterior-based quality cuts or agreement between multiple photo-z codes, and recompute the dispersions and mean offsets using only that strategy. Otherwise the post-outlier-removal numbers are not directly transferable to the JADES data.","section":"Sec. 4.2.1"}],"minor_comments":[{"comment":"The abstract states photo-z accuracy of at least 95%, while Section 5 says that for 90% of the sources EAZY gave an estimate accurate to better than 15% in (1+z); please reconcile these two numbers.","section":"Abstract / Sec. 5"},{"comment":"The caption describes the galaxies as also detected by HST, but the left-hand panels show JWST-only photometry and the right-hand panels show JWST+HST; please clarify which panels correspond to which data combination.","section":"Fig. 3 caption"},{"comment":"There is a typo in the sentence preceding Equation (2): 'The the distribution' should read 'The distribution'.","section":"Sec. 4.2.3"},{"comment":"The reference list contains Liao et al. (2024) twice with identical bibliographic data; please merge the duplicate entries.","section":"References"},{"comment":"The text refers to the 'GOOD-S field'; the standard abbreviation used elsewhere is GOODS-S, and this should be made consistent.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":"The main risk is epistemological rather than statistical: the forecast is built on the authors' own model and validated with the same pipeline, so the headline scatter values are in-family recovery statistics unless mismatched-template tests are added. This is a normal situation for a forecast paper and not a sign of bad faith. A major revision adding an independent validation step and a robustness test for stellar masses would make the paper suitable for publication. The heavy self-citation of Mitra et al. (2024) is justified because that paper is the direct antecedent of the model used here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"We both know how these forecast papers usually go, so I'll cut to it. This one is better than most and still has one load-bearing caveat: every headline recovery number is a model-to-model exercise. The simulated proto-spheroids are built with Mitra et al.'s upgrade of Cai et al. (2013), using da Cunha et al. (2008) energy-balance SEDs, BC03 SSPs, and Fritz et al. (2006) AGN templates, and then the recovery pipeline is EAZY plus CIGALE with those same physical ingredients. That means the M* scatters of 0.2 dex (parent) and 0.14 dex (DSFG) from JWST alone, and 0.1 dex for DSFG after outlier removal, are essentially in-family estimates. They are not yet validated against independent truth.\n\nThe paper earns credit for being honest about this, and for the work it does put in. The new pieces are the JADES/GOODS-S forecast from the upgraded model, the MIR luminosity function comparison against Ling et al. (2024), the NIRCam colour-colour check against CEERS, and the ASTRODEEP-JWST NIRCam-selected DSFG comparison. The quantitative tables are useful, and the discussion of where photo-z outliers come from is clear. The authors also flag their own biggest failure mode: switching CIGALE's SFH from sfhdelayed to sfhdelayedbq changes recovered SFRs by an order of magnitude. That test is exactly the kind of honesty you want, but it also makes the absence of an equivalent stress test for M* more conspicuous.\n\nThe external checks do not close the circularity. The MIR LFs under-predict the low-luminosity end. The simulated galaxies occupy a narrower region of the NIRCam colour-colour diagram than observed CEERS galaxies. And the ASTRODEEP comparison runs the same EAZY/CIGALE pipeline on both the observed and simulated samples, so any systematic template bias cancels out. That comparison shows the model can reproduce the observed population under the same fitting assumptions; it cannot certify absolute recovery accuracy.\n\nSmaller issues: the abstract's 'up to 0.1 dex' for the majority glosses over the fact that the 0.1 dex figure is the sub-sample after outlier removal; full-sample metrics should be reported more prominently. The dust mass absorption coefficient kappa_0 is not stated, which matters for the dust mass comparison. And the simulated catalogues are not released; they should be, at least as tables of input and recovered properties.\n\nIt is a solid, useful forecast for anyone planning JADES or similar deep NIRCam programs. The central claims are plausible but conditional. I'd send it to review, not desk-reject it, and I'd ask for the stress test on stellar masses, the kappa_0 value, and full-sample numbers.","headline":"Useful, honest JWST survey forecast whose headline mass and redshift accuracies are in-family estimates from the same SED machinery; deserves review with a request for a stellar-mass SFH stress test.","tokens_in":25384,"tokens_out":2818,"would_cite":true,"duration_ms":27190,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"JWST's JADES survey can recover the stellar masses of dusty proto-spheroids at z≥1.5 to about 0.1 dex, while their star formation rates still need far-infrared data.","keywords":["dusty star-forming galaxies","proto-spheroids","JWST","JADES","photometric redshifts","SED fitting","stellar mass","GOODS-South"],"falsifier":"Compare the model-predicted photo-z outlier fractions against real JADES sources with spectroscopic redshifts in GOODS-S: if the fraction with $|\\Delta z|/(1+z)>0.15$ exceeds roughly 10 percent for NIRCam-selected DSFGs, the 5-percent forecast is wrong. Similarly, measure stellar masses for a dozen DSFGs from ALMA dynamics or from a second, independent SED code; if the median offset against CIGALE exceeds about 0.3 dex, the claimed 0.1 dex mass recovery will not hold on real data.","tokens_in":24132,"feed_emoji":"🔭","tokens_out":8822,"duration_ms":88640,"temperature":0.7,"pith_summary":"This paper asks how well the James Webb Space Telescope, as configured for the JADES survey in GOODS-South, can constrain the nature of the dusty star-forming galaxies (DSFGs) that are thought to become today's massive elliptical galaxies. The authors generate a mock catalog of proto-spheroids at $z\\gtrsim1.5$ from a physical model of galaxy and black-hole co-evolution, assign spectral energy distributions to them, and then run the standard redshift and SED-fitting tools on the simulated photometry. They find that JWST photometry alone recovers photometric redshifts with a 5 percent outlier fraction for the parent sample and 4.2 percent for the 250-$\\mu$m-selected DSFG subsample, and that stellar masses come back with 1$\\sigma$ dispersions of 0.2 and 0.14 dex, improving to 0.1 dex for DSFGs when HST data are added. The same exercise shows that star formation rates cannot be trusted from near-infrared data alone and need the far-infrared data from Spitzer and Herschel. If the forecast is right, JWST can weigh the stellar content of proto-spheroids down to about $10^{10}\\,M_\\odot$, while their ongoing star formation will still need far-infrared and submillimetre observations.","feed_headline":"JWST alone can weigh dusty galaxies to 0.1 dex","feed_subtitle":"Simulated JADES photometry recovers proto-spheroid masses accurately; star formation still needs far-infrared data.","key_machinery":"The load-bearing object is the simulated proto-spheroid catalog built from the Cai et al. (2013) co-evolution model as upgraded in Mitra et al. (2024), which links star formation and black-hole accretion in haloes with $11.3\\le\\log(M_{\\rm vir}/M_\\odot)\\le13.3$ virializing at $1.5\\le z_{\\rm vir}\\le8$. Each simulated galaxy gets an SED from the da Cunha et al. (2008) energy-balance formalism plus a smooth-torus AGN component; the survey strategy is applied by imposing 5$\\sigma$ depth limits in the nine NIRCam bands and the ancillary HST, Spitzer, and Herschel bands. The recovery test then feeds this photometry through EAZY, a template-based photometric-redshift code, and CIGALE, an energy-balance SED-fitting code, and compares every recovered quantity to the known input value via $Q_{\\log P}=\\log(P_{\\rm CIGALE}/P_{\\rm input})$. What carries the argument is this closed model-to-model loop: any bias or scatter the fitting codes show against the model's own SEDs is taken as the expected performance on real DSFGs.","core_discovery":"The central claim is that JWST/JADES photometry alone is sufficient to determine photometric redshifts of massive dusty star-forming galaxies at $z\\gtrsim1.5$ with an outlier fraction $f_{\\rm out}=0.05$ for the parent sample and $0.042$ for the DSFG subsample (defined by a 5$\\sigma$ detection at 250 $\\mu$m), and stellar masses with 1$\\sigma$ dispersions of 0.2 and 0.14 dex respectively. When HST photometry is added, the outlier fractions drop to 0.019 and 0.008, and the DSFG stellar-mass dispersion falls to 0.1 dex. The paper further claims JWST can detect DSFGs with stellar masses down to $\\sim10^{10}\\,M_\\odot$, roughly an order of magnitude below what was accessible before JWST, and that a NIRCam colour-selected DSFG catalog drawn from the ASTRODEEP-JWST data matches the simulated population in stellar mass and SFR distributions. By contrast, the authors report that star formation rates recovered from JWST photometry alone have dispersions of about 0.55–0.8 dex, and that changing the assumed star-formation history in CIGALE changes recovered SFRs by about an order of magnitude. The conclusion is that JWST alone can constrain the stellar content of proto-spheroids, while their ongoing star formation and dust properties remain hostage to far-infrared follow-up.","pith_inferences":["The quoted accuracies are probably optimistic lower bounds, because the test only checks whether the fitting codes can recover the values the model put in; real galaxies with richer star-formation histories or more complex dust geometries could degrade the dispersions.","The reported order-of-magnitude sensitivity of SFR to the assumed star-formation history implies that any JWST-only SFR reported for an obscured galaxy should be read as template-dependent, not as a measurement.","A sharp, testable extension of the paper would be to apply the NIRCam colour selection $f_{444}/f_{150}>3.5$ to the full JADES footprint and compare EAZY outlier fractions against the growing spectroscopic sample; if the outlier fraction stays near 5 percent, the model-based forecast is confirmed.","If real DSFGs at $10^{10}\\,M_\\odot$ are as numerous as simulated, the integrated star-formation-rate density at cosmic noon may be higher than currently inferred from submillimetre-selected samples that miss low-mass dusty galaxies."],"forward_implications":["If the recovery statistics hold for real galaxies, JWST photometry alone can map the stellar mass content of DSFGs at $z\\gtrsim1.5$ to about 0.15–0.2 dex, enough to test models of proto-spheroid assembly at cosmic noon.","Photometric redshifts from NIRCam, with outlier fractions of a few percent, mean that DSFG samples can be selected and roughly placed in redshift without far-infrared data, opening the low-mass regime that was Herschel-blind.","The 250-$\\mu$m-selected DSFG sample recovers SFR, dust luminosity, and dust mass with dispersions of about 0.16–0.18, 0.12, and 0.26 dex respectively only when JWST is combined with Spitzer and Herschel, so the far-infrared complement remains essential for star formation.","JWST lowers the detectable stellar-mass threshold for dusty galaxies by roughly an order of magnitude, to about $10^{10}\\,M_\\odot$, so future deep surveys should uncover a populous low-mass DSFG population invisible to Herschel and Spitzer.","The consistency between simulated and ASTRODEEP NIRCam-selected DSFGs supports using the same physical model to predict what future far-infrared and submillimetre facilities will see."],"supporting_citations":[{"why":"Supplies the base physical model of proto-spheroid and AGN co-evolution used to generate the simulated galaxies.","marker":"Z.-Y. Cai et al. (2013)"},{"why":"Upgrades the model with SED formalisms for stellar/dust and AGN emission and provides the halo formation-rate function used to build the catalog.","marker":"D. Mitra et al. (2024)"},{"why":"Defines the JADES survey strategy, filter set, and depths in GOODS-S that set the detection criteria in the simulation.","marker":"D. J. Eisenstein et al. (2023)"},{"why":"Provides EAZY, the template-based code used to estimate photometric redshifts from the simulated JWST/HST photometry.","marker":"G. B. Brammer et al. (2008)"},{"why":"Provides CIGALE, the energy-balance SED fitting code used to recover stellar mass, SFR, dust luminosity, and dust mass.","marker":"M. Boquien et al. (2019)"},{"why":"Supplies the energy-balance dust/SED formalism that generates the model galaxy SEDs and is the reference for dust attenuation.","marker":"E. da Cunha et al. (2008)"},{"why":"Provides the mid-infrared luminosity functions used to validate that the model reproduces observed proto-spheroid populations.","marker":"C.-T. Ling et al. (2024)"},{"why":"Provides the NIRCam colour criteria used to select DSFGs from the ASTRODEEP catalog and from the simulation for comparison.","marker":"A. J. Barger & L. L. Cowie (2023)"},{"why":"Supplies the ASTRODEEP-JWST photometric catalog whose NIRCam-selected DSFGs are compared with the simulated population.","marker":"E. Merlin et al. (2024)"}],"fun_headline_variants":["JWST photometry alone nails redshifts and masses of dusty galaxies","95% accurate photo-z from JWST alone for dusty star-formers","Stellar masses to 0.1 dex with JWST, but SFRs need IR","JWST alone measures dusty galaxy masses, not their starbursts","Proto-spheroid masses from JWST: 0.1 dex accuracy, SF uncertain"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast stands on the premise that the model's spectral templates, star-formation histories, dust attenuation laws, and AGN fractions span the real diversity of proto-spheroids; the pipeline is only tested against the model's own outputs, and the paper itself shows that changing one star-formation history parametrisation changes recovered SFRs by an order of magnitude.","fun_headline_variants_meta":{"raw":{"variants":["JWST photometry alone nails redshifts and masses of dusty galaxies","95% accurate photo-z from JWST alone for dusty star-formers","Stellar masses to 0.1 dex with JWST, but SFRs need IR","JWST alone measures dusty galaxy masses, not their starbursts","Proto-spheroid masses from JWST: 0.1 dex accuracy, SF uncertain"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000542,"raw_usage":{"total_tokens":2714,"prompt_tokens":1180,"completion_tokens":1534,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":796,"completion_tokens_details":{"reasoning_tokens":1445}},"tokens_in":796,"tokens_out":1534,"duration_ms":11106,"temperature":1.0,"reasoning_tokens":1445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:56:19.570751+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the model-predicted photo-z outlier fractions against real JADES sources with spectroscopic redshifts in GOODS-S: if the fraction with $|\\Delta z|/(1+z)>0.15$ exceeds roughly 10 percent for NIRCam-selected DSFGs, the 5-percent forecast is wrong. Similarly, measure stellar masses for a dozen DSFGs from ALMA dynamics or from a second, independent SED code; if the median offset against CIGALE exceeds about 0.3 dex, the claimed 0.1 dex mass recovery will not hold on real data.","supporting_citations":[{"cited_title":"2024, title Euclid view of the dusty star-forming galaxies at ≳ detected in wide area submillimetre surveys , Monthly Notices of the Royal Astronomical Society, 530, 2292","cited_arxiv_id":null,"evidence_quote":"Upgrades the model with SED formalisms for stellar/dust and AGN emission and provides the halo formation-rate function used to build the catalog."},{"cited_title":"Exploring the faintest end of mid-infrared luminosity functions up to $z\\simeq 5$ with the JWST CEERS survey","cited_arxiv_id":"2402.05386","evidence_quote":"Provides the mid-infrared luminosity functions used to validate that the model reproduces observed proto-spheroid populations."},{"cited_title":"2024, title ASTRODEEP-JWST: NIRCam-HST multi-band photometry and redshifts for half a million sources in six extragalactic deep fields , , 691, A240","cited_arxiv_id":null,"evidence_quote":"Supplies the ASTRODEEP-JWST photometric catalog whose NIRCam-selected DSFGs are compared with the simulated population."}],"review_version":1}