{"id":"d6ba273e-d46a-4610-8400-4f115e4d1590","arxiv_id":"2412.08883","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new simulation suite, ESpRESSO, generates Roman grism survey scenes from Hubble imaging and synthetic SEDs, with injected Lyman-alpha galaxies and noise, for testing future spectral extraction tools.","lead":"ESpRESSO is a new software package that creates realistic simulated images of what NASA's Roman Space Telescope will see when it spreads galaxies' light into spectra. It lets astronomers test data-analysis tools before Roman launches, using real Hubble images plus synthetic galaxies.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'high (synthetic) LAE completeness' is asserted but never quantified; without a recovery test on the delivered noisy simulations, the mock survey's central utility is unverified.","rationale":"The reader's weakest assumption was that the optical model is taken from the grism as designed, not as built. That is a real limitation, but it is acknowledged by the authors, cannot be validated before launch, and does not block the paper's stated planning-purpose use. The more immediate load-bearing weakness is the unquantified completeness claim, which is central to the abstract and directly testable with the products the paper says it will release. The paper presents a coherent forward-modeling pipeline with honest caveats, but the 'high (synthetic) LAE completeness' assertion is never operationalized: no definition of completeness, no recovery simulation, no threshold. This matters because the stated purpose of the suite is to develop spectral extraction tools; a suite for that purpose needs a known, measured injection-recovery performance. The secondary area inconsistency between the abstract and Section 4.4(2) also needs resolution before the delivered data can be used as a survey mock. These concerns are addressable by the authors, so they do not change the reader's CONDITIONAL verdict; they do sharpen the conditions that should be attached to acceptance.","tokens_in":19862,"tokens_out":6085,"duration_ms":67388,"concrete_test":"Run an independent 1D/2D spectral extraction (e.g., CUBGRISM or a matched-filter line search) on the released 25-PA, 10 ks noisy images with injected LAEs; compute recovery fraction versus Ly-alpha flux, rest-frame EW, and redshift. If recovery at the claimed depth is not high (e.g., not >80% at the survey limit), the abstract's 'high completeness' claim should be revised or quantified. Also verify the delivered footprint per PA against the 'nine detector' / 'half of eighteen' claim in the abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract is that the 25-PA, 10 ks suite can serve as 'a mock deep Roman grism survey with high (synthetic) LAE completeness.' Section 3.4.1 describes injecting 5,000 LAEs with a single Sersic stamp and a Gaussian-line-plus-power-law SED, and Section 3.5 adds Poisson noise, but nowhere is a completeness fraction computed. No source extraction or line-search pipeline is run on the noisy images; the only demonstration is one z=9.5 LAE shown in isolation in Fig. 7. Completeness depends on noise, foreground crowding, off-order contamination, and the assumed LAE size/EW distribution—all present in the simulation but never tested end-to-end. A forward model can be correct and still fail to deliver high completeness if the injected population is too faint or too crowded, so the headline claim is unsupported as written. A secondary inconsistency: the abstract promises 'half of the eighteen detector array' while Section 4.4(2) states the current simulations cover only ~2-3 SCAs per PA, leaving the delivered survey area ambiguous.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ESpRESSO, a forward-modeling pipeline for Nancy Grace Roman Space Telescope WFI grism observations. The pipeline combines HST/CANDELS F160W COSMOS imaging with 3D-HST/EAZY spectral energy distributions to construct a wavelength-resolved datacube, then applies per-SCA sky-to-detector distortion polynomials, simulates the (0,0), (1,1), and (2,2) grism orders, and adds Poisson noise to produce 10 ks exposures at 25 position angles, with 12 of those having paired positive and negative dithers. Custom sources, including 5,000 synthetic Ly-alpha emitters, can be injected into the foreground scenes. The paper also presents an argument that (0,0)-order artifacts are unlikely to be confused with true emission-line pairs, and a crowding analysis concluding that foreground sources significantly elevate the background for about 10% of grism pixels. The central advertised product is a mock deep Roman grism survey with 'high (synthetic) LAE completeness' for developing spectral extraction tools.","tokens_in":20083,"tokens_out":7303,"duration_ms":74184,"significance":"If the completeness claim were validated, ESpRESSO would be a genuinely useful community resource: it is the first Roman-specific grism forward model described here to include field-angle-dependent distortions, three spectral orders, source injection, and noise within a modular, parameter-file-driven pipeline. The off-order confusion argument in Sec. 4.2 is internally consistent and useful for survey planning, as is the 10% foreground-contamination estimate in Sec. 4.3. The planned public release of simulated grism scenes is a concrete contribution to Roman preparatory work. The paper's main limitation is that its headline claim of high synthetic LAE completeness is asserted rather than demonstrated: no source-extraction or line-detection test is run on the noisy delivered images, so the central utility of the 25-PA suite as a completeness-testing mock survey is unverified. This is fixable with an end-to-end recovery experiment, but it is a load-bearing gap in the current manuscript.","major_comments":[{"comment":"The abstract and conclusions claim that the 25-PA, 10 ks suite provides 'high (synthetic) LAE completeness', but no completeness fraction is ever computed or reported. Section 3.4.1 describes injecting 5,000 LAEs with a single Sersic stamp and a Gaussian-line-plus-power-law SED, and Sec. 3.5 adds Poisson noise, but the paper never runs a source-extraction or line-search pipeline on the resulting noisy images. The only illustration is a single z=9.5 LAE shown in isolation in Fig. 7, which cannot establish completeness against crowding, noise, and off-order contamination. Please add an end-to-end recovery test and report completeness as a function of Ly-alpha flux, equivalent width, and redshift, or revise the abstract and conclusions to claim only that the scenes are suitable for developing extraction tools, not that they have demonstrated high completeness.","section":"Abstract; Sec. 3.4.1; Sec. 3.5; Sec. 5"},{"comment":"The delivered survey area is described inconsistently. The abstract promises 'a simulation suite of half of the eighteen detector array', while Sec. 4.4(2) states that the current simulations have 'a total area coverage of ~2-3 Roman SCAs per position angle, nowhere near the full detector array'. These statements cannot both describe the released products. Please specify exactly how many SCAs are covered per position angle, how the 25 PAs combine into total unique sky area, and reconcile the abstract wording with Sec. 4.4(2).","section":"Abstract vs Sec. 4.4(2)"},{"comment":"The quantitative claim that foreground contamination affects about 10% of grism pixels and that the sky-limited assumption is valid for about 90% of the field is presented without uncertainty or PA-to-PA variation. The caption of Fig. 10 says 'For our simulated grism extra-galactic scene', suggesting a single realization, even though the paper generates 25 PAs and paired dithers. Since this 10% figure is a headline result for survey design, please report the distribution across PAs and dithers, or explicitly state that it is a single-representative-scene estimate with no quoted uncertainty.","section":"Sec. 4.3 and Fig. 10"}],"minor_comments":[{"comment":"The abstract has several wording and punctuation issues that should be corrected: 'nine detector grism observation' is unclear (nine detectors or one detector?), 'which model field angle dependent optical distortions' should read 'which models ...', '12 with analogous positive and negative dithers,' is a sentence fragment, and the exposure-time clause ends with a comma rather than a period.","section":"Abstract and Sec. 1"},{"comment":"The captions contain incomplete placeholder values: 'a bright, broadband star (m_F160W =, spectral type )' and 'an emission line galaxy (ELG; m_F160W =)' have blank magnitude and spectral-type entries. Please fill in the actual values or remove the parentheticals.","section":"Captions of Figs. 5 and 6"},{"comment":"The comparison of input resolutions '0.06 vs. 0.03 mas' is dimensionally wrong for plate scales; the authors presumably mean 0.06 arcsec/pixel versus 0.03 arcsec/pixel (or 60 vs 30 mas/pixel). Please correct the units.","section":"Sec. 4.1"},{"comment":"The delivered simulations use a sky background of 0.8 counts/s per pixel in Sec. 3.5, while Sec. 4.3 uses the Roman technical-report value of 1.3 counts/s per pixel for the crowding threshold. Please clarify which background level the released 10 ks images use and discuss how the difference affects the 10% contamination estimate and completeness expectations.","section":"Sec. 3.5 and Sec. 4.3"},{"comment":"The sentence 'All emission to left the emission line has been attenuated' contains a typo; it should read 'to the left of the emission line'.","section":"Sec. 3.4.1"},{"comment":"The description of the distortion polynomial is unclear: 'We note that t = 0 for all 4th power and nearly all 5th power terms' does not specify which of the 44 coefficients are actually retained. Please list the non-zero monomials or otherwise clarify the structure of the polynomial used for each SCA.","section":"Sec. 3.2"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision rather than rejection because the forward-modeling machinery and the off-order confusion analysis are sound, and the missing completeness validation is a well-defined addition. I would specifically ask the editor to ensure the authors reconcile the abstract's 'half of the eighteen detector array' statement with Sec. 4.4(2)'s '2-3 SCAs per position angle' claim, since this discrepancy affects the advertised data release. The 10% contamination result should also be presented with PA-to-PA spread if it is to inform survey design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"ESpRESSO is a genuine new tool and worth reading if you plan to work on Roman grism extraction. It builds a datacube from CANDELS COSMOS imaging and 3D-HST SEDs, applies the 44-term sky-to-detector distortion polynomial per SCA, simulates the three in-focus orders, injects LAEs, adds Poisson noise, and will release a 25-PA, 10 ks suite on acceptance. The off-order confusion argument in Sec 4.2 is internally consistent, and the comparison with the aXeSIM-based Wold et al. run in Sec 4.1 is a concrete demonstration of what the new code changes. The authors are also unusually upfront about limitations in Sec 4.4: pre-launch grism model, HST PSF without deconvolution, limited template coverage, whole-object SED assignment.\n\nThe soft spots are real but mostly fixable. The abstract's 'high (synthetic) LAE completeness' is the headline claim, yet nowhere is a completeness fraction computed. No source extraction or line-search pipeline is run on the noisy images; the only demonstration is one z=9.5 LAE in isolation in Fig 7. A forward model can be correct and still yield poor completeness if the injected population is too faint or too crowded, so the claim is unsupported as written. A simple recovery test would fix this. Second, the abstract promises 'half of the eighteen detector array' while Sec 4.4(2) says the current simulations cover only ~2-3 SCAs per PA. That discrepancy needs reconciliation before the data release is described as a half-array survey. Third, the 10% crowding estimate has no uncertainty and depends on the assumed sky level and survey depth; a small sensitivity test would help. The as-designed vs as-built grism concern is real but the authors flag it themselves; it will be validated after launch, and for planning purposes it is a reasonable working assumption.\n\nThe math and data handling look sound: flux conservation in Eqs. 4-5, nearest-neighbor assignment with oversampling, and the parameter-file structure make the pipeline reproducible in principle. There is no circularity in the central simulation; the LAE injection follows their earlier recipe, but nothing is fit to force the 10% or confusion conclusions. Self-citation is minor.\n\nThis paper is for the Roman spectroscopy subfield: grism tool developers, LAE survey planners, and anyone writing extraction software before launch. It deserves a serious referee. A moderate revision should add a completeness/recovery demonstration and fix the area-coverage inconsistency, and then it will be a solid methods paper.","headline":"Useful Roman grism forward model, but the abstract's 'high LAE completeness' and 'half the detector array' claims outrun what the paper actually demonstrates.","tokens_in":20662,"tokens_out":2161,"would_cite":true,"duration_ms":22136,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ESpRESSO produces realistic simulated Roman grism exposures from Hubble imaging and model spectra, creating a mock deep survey for testing spectral extraction software before launch.","keywords":["Roman Space Telescope","WFI grism","slitless spectroscopy simulation","forward modeling","Lyman-alpha emitters","source injection","detector distortion","mock survey"],"falsifier":"Take laboratory or early-on-orbit calibration images of a bright star taken through the Roman grism at several field positions, and compare the measured (1,1) trace position at each wavelength with ESpRESSO's Eq. 6–12 prediction; trace residuals larger than roughly one WFI pixel (about 0.11 arcsec) would mean the 44-term distortion model does not match the flight instrument. Separately, measure the blue- and red-edge response across a detector and check whether the 1:9 blue-to-red ratio of the (0,0) artifact and the roughly 10% foreground-contaminated pixel fraction survive in real data.","tokens_in":19655,"feed_emoji":"🔭","tokens_out":9908,"duration_ms":100279,"temperature":0.7,"pith_summary":"ESpRESSO is a forward model of Roman's Wide Field Instrument grism that turns a three-dimensional cube of flux density, built from deep Hubble imaging and per-object model spectra, into synthetic nine-detector grism exposures. The paper's central claim is that these simulated scenes are realistic enough, including field-angle-dependent optical distortions, three spectral orders, source injection of Lyman-alpha emitters, and photon noise, to serve as a mock deep Roman grism survey for developing and testing spectral analysis tools before launch. Using this machinery, the authors also argue that the compact (0,0) order images cannot easily be mistaken for true emission-line pairs, and that bright foreground sources raise the background for about ten percent of grism pixels, so sky-limited assumptions hold for the rest. The result is a public simulation suite spanning half of the 18-detector array at 25 position angles with 10 ks exposures, meant as test data for recovery and de-confliction algorithms.","feed_headline":"Mock Roman grism surveys are ready to stress-test spectral tools","feed_subtitle":"ESpRESSO reproduces detector distortions, three spectral orders, noise, and injected Lyman-alpha emitters across 25 roll angles.","key_machinery":"The central object is ESpRESSO's forward-modeling pipeline for the Wide Field Instrument grism, a dispersing optic whose undeviated wavelength is 1.55 μm and whose three in-focus orders are (0,0), (1,1), and (2,2). The machinery centers on a precomputed flux-density data cube over position and wavelength, driven by a 44-term polynomial that maps sky coordinates and wavelength to detector pixel coordinates for each sensor chip, plus a piecewise response function that includes position-dependent blue- and red-edge cutoffs. This lets the code place every source pixel's flux at the correct dispersed location for each order and each roll angle, so scene crowding, overlap between orders, and contamination statistics emerge directly from the simulation rather than being assumed. It is also what makes the pipeline modular: any of the optical parameters can be swapped for on-orbit measurements after launch, and custom sources enter through the same cube construction, so the same engine produces both the foreground scene and isolated injected-object images.","core_discovery":"On the paper's own terms, the discovery is that Roman WFI grism data can be emulated before launch at pixel-level fidelity by combining the best available imaging and spectra. ESpRESSO builds an (x, y, λ) data cube by assigning each object's model spectrum to the image pixels belonging to that object, then maps sky coordinates to detector pixels with a 44-term polynomial per detector (Eq. 6), applies a wavelength- and position-dependent response with blue- and red-edge cutoffs (Eqs. 9–12), and assigns flux by nearest-neighbor sampling at 3.7x spatial and 11x spectral oversampling. Three in-focus orders, (0,0), (1,1), and (2,2), are produced, along with dithers and roll angles, and photon noise is added assuming a 0.8 counts/s sky background. The paper demonstrates custom source injection with a star, an emission-line galaxy, and 5,000 synthetic Lyman-alpha emitters, and uses the simulated scenes to conclude that the (0,0) artifact spectrum is unlikely to be confused with real line pairs and that roughly 10% of grism pixels are significantly boosted by foreground sources.","pith_inferences":["A reader can extend the 25-position-angle suite into a completeness experiment: inject the same Lyman-alpha emitter at many fluxes and redshifts, run any extraction code, and map recovery rate against position angle and dither to identify which roll angles maximize the clean area.","Because the code's optical parameters are replaceable, the same pipeline is a natural testbed for on-orbit calibration; once flight data exist, re-running with measured parameters would tell how much of the 10% contamination and the 1:9 off-order ratio survive reality.","The 10% foreground-boosted pixel fraction should be read as tied to the depth and density of the input field; a shallower survey or a different line of sight would shift the number, so survey planners may want this calculation repeated for other deep fields as they become available.","The off-order confusion analysis suggests a concrete algorithmic check: run source detection on the released (0,0) plus (1,1) scenes and count how often an artifact is classified as an emission-line pair; the resulting misclassification rate is the quantitative form of the paper's qualitative conclusion."],"forward_implications":["The released 25-position-angle, 10 ks-per-angle suite provides a ready-made benchmark for testing spectral extraction and source recovery software against a deep Roman grism survey, with injected Lyman-alpha emitters whose true redshifts, line fluxes, and continuum levels are known by construction.","The (0,0) order's double-peaked structure, with 29-pixel separation, 319 Å observed separation, and a 1:9 blue-to-red flux ratio, is very unlikely to be mistaken for a real emission-line pair, because known doublets would require physically implausible brightness ratios or be ruled out by other nearby lines.","Foreground contaminants significantly elevate the background for about 10% of grism pixels, implying that sky-limited noise assumptions are valid for roughly 90% of the field and that de-confliction algorithms are needed for the remaining 10%.","Because the optical model is parameterized and replaceable, the same pipeline can be re-run with post-launch calibration data to update the mock observations as the real instrument's performance becomes known."],"supporting_citations":[{"why":"Defines the mock Lyman-alpha emitter population recipe, including luminosity function slope, equivalent-width distribution, and redshift range, that ESpRESSO adopts, and provides the prior aXeSIM foreground used for visual comparison.","marker":"Wold et al. (2023)"},{"why":"Supplies the photometric catalog, best-fit model spectra from template fitting, and pixel-membership maps that link each source's spectrum to its image pixels in the data cube.","marker":"Skelton et al. (2014)"},{"why":"Provides the deep Hubble WFC3/IR F160W mosaics at 30 mas per pixel that serve as the high-resolution imaging basis of the foreground scene.","marker":"Koekemoer et al. (2011)"},{"why":"The 3D-HST program produced the grism-derived catalogs and imaging products to which the catalog and spectral inputs are tied.","marker":"Brammer et al. (2012)"},{"why":"This is the established aXe and aXeSIM forward-modeling software baseline against which ESpRESSO's grism images are compared in the discussion.","marker":"Kümmel et al. (2009)"},{"why":"Provides the intergalactic medium transmission prescriptions for the Lyman-alpha forest and damped Lyman-alpha systems used to attenuate the injected Lyman-alpha emitter continua.","marker":"Inoue et al. (2014)"},{"why":"Gives the filter magnitude conversion method used to identify bright stars and extended galaxies whose spectra needed replacement modeling.","marker":"Tokunaga & Vacca (2005)"}],"fun_headline_variants":["ESpRESSO emulates Roman grism before launch","Mock Roman grism scenes stress spectral tools","Pixel-level Roman spectroscopy simulator","Pre-launch Roman grism emulator for analysis","ESpRESSO: forward-modeling Roman's spectroscopy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The simulation trusts the grism as designed rather than as it will test out in space; if the real instrument bends or focuses light differently at any wavelength or field position, every simulated spectrum position and the paper's crowding and off-order confusion conclusions shift with it.","fun_headline_variants_meta":{"raw":{"variants":["ESpRESSO emulates Roman grism before launch","Mock Roman grism scenes stress spectral tools","Pixel-level Roman spectroscopy simulator","Pre-launch Roman grism emulator for analysis","ESpRESSO: forward-modeling Roman's spectroscopy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000258,"raw_usage":{"total_tokens":1641,"prompt_tokens":1063,"completion_tokens":578,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":505}},"tokens_in":679,"tokens_out":578,"duration_ms":6037,"temperature":1.0,"reasoning_tokens":505,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T17:28:05.816739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take laboratory or early-on-orbit calibration images of a bright star taken through the Roman grism at several field positions, and compare the measured (1,1) trace position at each wavelength with ESpRESSO's Eq. 6–12 prediction; trace residuals larger than roughly one WFI pixel (about 0.11 arcsec) would mean the 44-term distortion model does not match the flight instrument. Separately, measure the blue- and red-edge response across a detector and check whether the 1:9 blue-to-red ratio of the (0,0) artifact and the roughly 10% foreground-contaminated pixel fraction survive in real data.","supporting_citations":[{"cited_title":"Ly$\\alpha$ at Cosmic Dawn with a Simulated Roman Grism Deep Field","cited_arxiv_id":"2305.01562","evidence_quote":"Defines the mock Lyman-alpha emitter population recipe, including luminosity function slope, equivalent-width distribution, and redshift range, that ESpRESSO adopts, and provides the prior aXeSIM foreground used for visual comparison."}],"review_version":1}