{"id":"fae0efda-a738-4da7-80e5-264b2a5c59c3","arxiv_id":"2505.00093","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Star-galaxy spectral differences induce shear biases of about 0.2% in Roman's weak-lensing bands and 2% in the wide filter, exceeding requirements and requiring chromatic corrections.","lead":"This paper measures how differences in the light spectra of stars and galaxies change galaxy-shape measurements for the Roman Space Telescope, finding errors that exceed mission requirements: about 0.2% in the main weak-lensing bands and 2% in the wide filter. It then tests a first-order correction that works for the narrow bands under idealized conditions, and compares two realistic ways to estimate the needed quantities.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline correction claim is a closed-loop test within the galsim.roman PSF model, with coaddition and detector effects deferred, so Roman-specific bias magnitudes remain conditional on that model.","rationale":"The reader identified the same weakest assumption: the reliability of the galsim.roman PSF model, with coaddition explicitly deferred to future work. My stress-test sharpens this into a concrete mechanism: the correction's success is partly guaranteed by using the same PSF model to both generate the images and construct the B1 basis, so model errors cancel in the corrected result. The uncorrected bias magnitude is also model-dependent, scaling with the PSF chromaticity. This concern does not make the paper internally inconsistent; the authors are transparent about the simplifications (no detector effects in Sec. 3.5, coaddition left to future work in Sec. 7), and the code is publicly available. However, it does mean the headline 'within requirements' claim is a proof of principle within an assumed PSF model rather than a demonstration against the true Roman PSF. The concrete test—repeating the analysis on PyIMCOM-coadded native-pixel images with the coadded PSF—would directly address the most important gap and is feasible with existing tools. The paper deserves conditional acceptance with this caveat explicitly carried into the abstract and conclusions, which is consistent with the reader's CONDITIONAL verdict.","tokens_in":32166,"tokens_out":10920,"duration_ms":122156,"concrete_test":"Reproduce the Fig. 3 and Fig. 5 analysis, but simulate images at the native Roman pixel scale (0.11 arcsec/pixel) with dithered undersampled exposures coadded with PyIMCOM (as in OpenUniverse2024), computing B1 from the coadded PSF rather than the oversampled exposure PSF. If the corrected multiplicative bias in any WL band exceeds 3.2e-4, or the uncorrected m shifts by more than 20% relative to the oversampled result, the closed-loop/oversampled setup is load-bearing for the headline claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central numbers—uncorrected m ~0.2% in the WL bands and ~2% in W146, and the first-order correction bringing all WL bands within |m|<3.2e-4—are generated and tested in a closed loop: the same galsim.roman PSF model produces the simulated galaxy images and defines the B1 correction basis (Sec. 5.2.1). Because the correction subtracts a term built from the same PSF model, any wavelength-dependent PSF error (filter coating transmission, charge diffusion, undersampling/coaddition) is shared between simulation and correction and largely cancels in the corrected result. The uncorrected bias itself scales with the PSF chromaticity, so both the bias quantification and the mitigation claim are conditional on galsim.roman accurately representing Roman's true PSF. The paper explicitly omits detector effects (Sec. 3.5) and defers coaddition of undersampled exposures to future work (Sec. 7), yet the abstract presents the correction result without these caveats. A plausible ~10% change in the PSF's wavelength dependence—e.g., from charge diffusion or from coaddition reshaping the PSF-λ relation—could shift the corrected m by a significant fraction of the 3.2e-4 budget, given that the uncorrected bias is ~2e-3.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper quantifies, with Roman-like image simulations, the shear calibration biases induced by the wavelength dependence of the PSF when the PSF is modeled with stars but applied to galaxies. The authors use two extragalactic catalogs (Diffsky and cosmoDC2) with distinct SED libraries, a Kurucz-based stellar catalog, and the galsim.roman PSF model to generate noiseless oversampled postage stamps, measuring shapes with the FPFS/AnaCal estimator. They find uncorrected multiplicative biases of roughly 0.2% in the four Roman WL bands and roughly 2% in the wide filter W146, exceeding both the SRD requirement |m| < 3.2e-4 and the relaxed 1e-3 requirement; additive biases are acceptable in the WL bands but not in W146. The paper then develops a PSF-level correction based on a Taylor expansion of the flux-normalized SED difference into coefficients ΔS_n and precomputed basis images B_n, and shows that with perfect per-galaxy SED knowledge the first-order correction brings all WL bands within the stringent requirement, while W146 is not corrected to requirement. It also tests an analytical color-based estimator and a self-organizing-map estimator for realistic implementation, finding that both reduce biases but with catalog-dependent performance.","tokens_in":32449,"tokens_out":7521,"duration_ms":85187,"significance":"If the results hold, this is a timely and useful contribution for Roman weak lensing: it is the first systematic quantification of chromatic PSF biases from star-galaxy SED differences for the four Roman WL bands and W146, and it proposes a clean, SED-only correction framework that is independent of the shape-measurement method. The numerical work is careful: 10,000 galaxies per catalog, 45-degree rotations to suppress shape noise, three input shears, bootstrap error bars, a convergence check at 1,500 galaxies, and two independent SED libraries that give consistent uncorrected biases. The correction coefficients ΔS_n are derived from the SEDs, not fitted to the measured bias, so there is no parameter-fitted-to-target circularity. The public code and catalogs also make the analysis reproducible. The principal caveat is that the validation is performed entirely within the galsim.roman PSF model: the same model produces the simulated images and the B_n correction basis, so the Roman-specific amplitude of the biases and the demonstrated post-correction accuracy are conditional on that model's fidelity.","major_comments":[{"comment":"The central mitigation claim is validated in a closed loop: both the simulated galaxy images and the B1 basis functions used in Eq. (14) are generated from the same galsim.roman PSF model. A wavelength-dependent error in that model, for example from filter coating uncertainties, charge diffusion, or reshaping by coaddition (effects explicitly deferred in Sec. 3.5 and Sec. 7), would change both the uncorrected bias amplitude and the correction basis. The manuscript therefore does not yet demonstrate that the corrected bias remains below |m| < 3.2e-4 for the actual Roman PSF. Please add a sensitivity test that perturbs the PSF wavelength dependence (e.g., scaling B1 or using an independent PSF model such as WebbPSF) and shows the corrected m stays within budget, and state this closed-loop caveat explicitly in the abstract.","section":"Sec. 5.2.1 and Abstract"},{"comment":"The text states that the SCA-constant approximation for B_n is 'tested later on for B1 and confirmed to hold,' but I could not find this test anywhere in the manuscript: Secs. 5.3 through 6 contain no comparison of the center-of-SCA basis against the basis at the actual galaxy positions. Because Eq. (15) enters every corrected measurement in Fig. 5, please either present the missing test with quantitative residuals, or remove the claim and propagate the approximation uncertainty into the corrected m values.","section":"Sec. 5.2.1, Eq. (15)"},{"comment":"The abstract's statement that 'higher-order terms are necessary for the wide filter' is not backed by a working higher-order implementation in the paper: Sec. 5.3 reports that a second-order polynomial fit produced a higher multiplicative bias than the first-order correction and failed to meet the relaxed requirement in all redshift bins for both catalogs. Please either demonstrate a higher-order correction that actually reduces W146 biases to requirement, or rephrase the conclusion to say that the first-order correction is insufficient for W146 and a successful higher-order correction is not yet demonstrated.","section":"Sec. 5.3 and Abstract"}],"minor_comments":[{"comment":"The choice of 0.0275 arcsec/pixel oversampled scale is described as 'somewhat arbitrary'; since the correction results and their comparison with coadded Roman images depend on pixel scale, please add a brief justification or a test of sensitivity to this choice.","section":"Sec. 3.5"},{"comment":"What is called an 'upper limit' is actually the largest absolute bias across redshift bins, not a statistical upper limit; please rename it to something like 'largest |m| across redshift bins' to avoid over-interpretation.","section":"Sec. 5.4, Table 2"},{"comment":"The statement that N_t/N_f 'resembles something like the color' is imprecise; the color is proportional to log10(N_t/N_f), not to the flux ratio itself, and this distinction affects how one expects the estimator to behave with noise.","section":"Sec. 6.1, Eq. (23)"},{"comment":"The phrase 'fail to exceed the allowed relative error' should read 'exceed the allowed relative error' or 'fail to stay within the allowed relative error'; the current wording states the opposite of what Table 3 shows.","section":"Sec. 6.3, Table 3"},{"comment":"Please define explicitly whether the SEDs in Eq. (3) are in photon units or energy units; the text mentions the necessary λ/hc conversion factor, but the equation as written is ambiguous, and the slope coefficients ΔS1 depend on this convention.","section":"Sec. 2.2, Eq. (3)"},{"comment":"The green curve representing the second-order approximation is not labeled in the legend of the right panel; please add a legend entry or a clear caption description so the reader can identify it.","section":"Fig. 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically sound and well suited to the journal, and the core bias quantification is careful and reproducible. The main risk is that the Roman-specific claims are stronger than what the closed-loop galsim.roman validation can support; I would prioritize (1) a perturbed-PSF sensitivity test for the correction, and (2) settling the missing SCA-constant B1 test. The W146 'higher-order terms are necessary' wording should also be aligned with the reported failure of the second-order implementation. These are fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first Roman-specific quantification of star-galaxy SED chromatic PSF biases, and it is worth taking seriously. The effect is large—multiplicative biases around 0.2% in the four WL bands and 2% in W146, versus the 0.032% SRD budget—and the two independent SED libraries (Diffsky, cosmoDC2) agree well. The B_n PSF basis formalism is a genuinely useful way to separate SED-dependent coefficients from PSF-dependent basis images, and the proof in Appendix B that ΔS0 vanishes for flux-normalized linear SEDs is clean. They also check convergence at ~1500 galaxies, use 45-degree rotations and three input shears, and release code. Credit where due: this is careful work.\n\nSoft spots, in proportion. The headline correction claim is validated in a closed loop: the same galsim.roman PSF model generates the images and defines the B1 basis. So both the measured bias and the corrected residual are conditional on that model being close to Roman's real chromatic PSF. The paper acknowledges this and defers coaddition and detector effects, but the abstract does not carry the caveat—it states the first-order correction brings WL bands within requirements without saying 'assuming perfect per-galaxy SED knowledge and the simulation PSF model.' Second, the realistic implementations are more pessimistic than the abstract suggests: for Diffsky, the analytical color-based estimators and even the SOM fail several redshift bins under the relaxed requirement. The authors are transparent about this (Table 3, Fig 7), so it is not a hidden flaw, but it does mean the practical correction is not yet solved. Third, the 0.2% number itself could shift with a more detailed PSF model, but even a plausible 10-20% change in the PSF wavelength dependence would leave the bias well above the 0.032% requirement. So the central claim—chromatic PSF errors are a major systematic for Roman WL—is robust.\n\nWho is this for? Anyone building the Roman WL pipeline, and anyone working on PSF chromaticity for space-based surveys. The formalism will get reused. It deserves a serious referee; the main things to push on are the closed-loop caveat and the gap between the idealized correction and the realistic estimators.","headline":"Roman-specific chromatic PSF bias is real and large, the B_n correction formalism is useful, but the headline mitigation claim is idealized and needs its caveats in the abstract.","tokens_in":33002,"tokens_out":1843,"would_cite":true,"duration_ms":18604,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Chromatic PSF errors from star-galaxy color differences bias Roman shear measurements by 0.2% in the four weak-lensing bands and 2% in the wide filter, above mission limits; a first-order PSF-level correction restores the WL bands.","keywords":["weak gravitational lensing","chromatic PSF","point spread function","Roman Space Telescope","shear calibration","SED differences","shape measurement","self-organizing maps"],"falsifier":"An on-orbit measurement of Roman's PSF size versus stellar color in each band would settle the matter: if the observed slope of PSF FWHM with color differs from the model's prediction by more than the sub-percent level required to keep $|m|$ below $3.2\\times 10^{-4}$, the simulated bias amplitudes and the first-order correction basis are both in doubt. Repeating the end-to-end simulation with an independent PSF model, generated without using the model that also defines the correction basis, would test whether the $0.2\\%$/$2\\%$ biases and the correction success are an artifact of model reuse.","tokens_in":31976,"feed_emoji":"🔭","tokens_out":11606,"duration_ms":98943,"temperature":0.7,"pith_summary":"This paper argues that the wavelength dependence of Roman's point-spread function (PSF) makes the spectral energy distribution (SED) difference between stars and galaxies a serious systematic for weak-lensing shear. Using Roman-like image simulations, it measures multiplicative shear biases of about $0.2\\%$ in the four weak-lensing bands Y106, J129, H158, and F184, and about $2\\%$ in the wide filter W146, both above the mission requirement $|m| < 3.2\\times 10^{-4}$; individual redshift bins reach $0.4$-$0.9\\%$ and $3$-$6\\%$. The paper then develops a PSF-level correction that expands the star-galaxy SED difference in a Taylor series about each filter's effective wavelength, and shows that with perfect SED knowledge the first-order term brings the four WL bands within the strict requirement, while W146 requires higher-order terms. It also tests two practical estimators of the needed SED slope, an analytical color-based estimator and self-organizing maps, finding both reduce biases but with sensitivity to the SED library used for training. A careful reader would care because the result makes chromatic correction a required component of Roman's shear calibration and offers a concrete framework for supplying it.","feed_headline":"Chromatic PSF biases Roman shear 0.2-2%, above budget","feed_subtitle":"Star-galaxy color differences exceed Roman's shear budget; a first-order PSF fix restores the four WL bands.","key_machinery":"The load-bearing object is a Taylor-expansion basis for the chromatic PSF error. The effective PSF difference between a star and a galaxy is written as\n$$\\$\\Delta$\\mathrm{PSF}_{\\rm eff}(x,y)=\\sum_n \\$\\Delta$ S_n\\, B_n(x,y),$$\nwhere $\\Delta S_n$ is the difference between flux-normalized SED Taylor coefficients about the filter's effective wavelength $\\lambda_0$, and $B_n(x,y)=\\int \\mathrm{PSF}(x,y,\\lambda)\\,F(\\lambda)\\,(\\lambda-\\lambda_0)^n\\,d\\lambda$ is an SED-independent image depending only on the PSF model and filter throughput. Flux normalization makes $\\Delta S_0$ vanish for linear SEDs, so the first-order term $\\Delta S_1 B_1$ carries nearly all of the bias in the narrow WL bands. The factorization is what carries the argument: $B_n$ can be precomputed from a trusted PSF model, leaving only the scalar coefficient $\\Delta S_1$ to be estimated from galaxy SEDs, and the correction is applied at the PSF level so it is independent of the shape-measurement method.","core_discovery":"The central claim is that chromatic PSF mismatch, caused by systematically different SEDs for the stars used to model the PSF and the galaxies whose shapes are measured, is a dominant systematic for Roman weak-lensing shear. Averaged over all simulated galaxies, the multiplicative bias is roughly $0.2\\%$ in every WL band and $2\\%$ in W146, an order of magnitude larger in the wide filter; additive bias is acceptable in the WL bands but exceeds the systematic budget in W146. Because a $1\\%$ multiplicative bias maps to roughly $1.5\\%$ bias in the cosmological parameter $S_8$, these amplitudes are cosmologically significant. The paper's constructive result is that when the PSF model and filter transmission are known, the chromatic PSF error factorizes into SED-dependent coefficients and SED-independent basis images, and the first-order term corrects the WL bands to within the strictest requirement when each galaxy's SED is known exactly; the wide filter resists first-order correction, and a second-order polynomial version performs unstably.","pith_inferences":["Beyond the paper: the same Taylor-basis formalism transfers to any diffraction-limited NIR survey with a well-characterized PSF model; the key input is the accuracy of $\\mathrm{PSF}(x,y,\\lambda)$, so the structure of the method is portable even though the paper only demonstrates it for Roman.","Beyond the paper: the wide-filter failure implies that a decision to use W146 for Roman weak lensing should be gated on on-orbit measurements of the PSF size-color relation; the paper's simulations are self-consistent, so real-data chromaticity remains untested.","Beyond the paper: the SOM experiments suggest that spectroscopic training samples for Roman should be built with explicit coverage of high-redshift SED space; the demonstrated degradation under cross-library training makes SED-space completeness a testable design requirement.","Beyond the paper: a direct follow-up is to test the correction on coadded Roman images rather than oversampled individual exposures; the paper leaves coaddition to future work, and coaddition can alter the effective PSF chromaticity."],"forward_implications":["If the central claim is correct, Roman's weak-lensing analysis must apply a chromatic PSF correction before shape measurement; leaving the effect uncorrected exceeds the SRD multiplicative-bias requirement by roughly a factor of six in the WL bands.","With perfect per-galaxy SED information, the first-order correction satisfies the strictest multiplicative requirement in Y106, J129, H158, and F184, meaning the dominant remaining uncertainty shifts to how well SED slopes can be estimated from photometry.","In W146, even a perfect first-order correction leaves residual multiplicative bias, and a straightforward second-order polynomial correction is unstable; this argues that using the wide filter for shear requires either a higher-order correction or acceptance of larger systematics.","Ensemble-averaged corrections can fail in individual redshift bins, while redshift-bin-averaged corrections satisfy a relaxed requirement; therefore tomographic analyses need per-redshift-bin calibration.","Galaxy color gradients contribute at most about $10^{-4}$ to multiplicative bias in the WL bands, so the dominant chromatic effect is the star-galaxy SED difference rather than internal color gradients."],"supporting_citations":[{"why":"It established that star-galaxy SED differences bias shear in diffraction-limited surveys and introduced color-matching and template-fitting mitigation.","marker":"Cypriano et al. (2010)"},{"why":"It extended template-fitting and machine-learning estimates of the effective PSF size, informing the correction design and its sensitivity to photometric errors.","marker":"Eriksen & Hoekstra (2018)"},{"why":"It supplied the wavelength-dependent Roman PSF model used both to generate the image simulations and to compute the correction basis.","marker":"Kannawadi et al. (2016)"},{"why":"It provided the image-simulation software used to render galaxy and star postage stamps and the nonphysical basis images.","marker":"Rowe et al. (2015)"},{"why":"It supplied the shapelet-based shear estimator and response formalism used to measure multiplicative and additive biases.","marker":"Li & Mandelbaum (2023)"},{"why":"It introduced the Fourier Power Function Shapelets estimator on which the shape measurement rests.","marker":"Li et al. (2018, 2022)"},{"why":"It provided the primary extragalactic catalog and the stellar catalog with SEDs used in the main simulations.","marker":"OpenUniverse et al. (2025)"},{"why":"It provided the alternative cosmoDC2 extragalactic catalog used to test the sensitivity of the biases and corrections to the SED library.","marker":"Korytov et al. (2019)"},{"why":"It demonstrated that PSF-level perturbations can correct SED-difference biases, the strategy adapted here for Roman.","marker":"Meyers & Burchat (2015a)"},{"why":"It generated the catalog-level photometric noise used to build realistic observed magnitudes and colors for the correction tests.","marker":"Crenshaw et al. (2024)"}],"fun_headline_variants":["Chromatic PSF biases Roman shear beyond requirements","Star-galaxy color mismatch breaks Roman shear budget","First-order PSF fix restores Roman WL bands","SED differences cause chromatic shear bias in Roman","Roman weak lensing demands chromatic PSF corrections"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the simulated wavelength-dependent Roman PSF used to create the images is a faithful model of the real Roman PSF, because the same model provides both the biased images and the correction basis; if actual filter coatings, charge diffusion, or coaddition of undersampled exposures change the PSF chromaticity, both the measured biases and the correction performance would shift.","fun_headline_variants_meta":{"raw":{"variants":["Chromatic PSF biases Roman shear beyond requirements","Star-galaxy color mismatch breaks Roman shear budget","First-order PSF fix restores Roman WL bands","SED differences cause chromatic shear bias in Roman","Roman weak lensing demands chromatic PSF corrections"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000334,"raw_usage":{"total_tokens":1938,"prompt_tokens":1114,"completion_tokens":824,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":730,"completion_tokens_details":{"reasoning_tokens":752}},"tokens_in":730,"tokens_out":824,"duration_ms":8188,"temperature":1.0,"reasoning_tokens":752,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:51:44.395133+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An on-orbit measurement of Roman's PSF size versus stellar color in each band would settle the matter: if the observed slope of PSF FWHM with color differs from the model's prediction by more than the sub-percent level required to keep $|m|$ below $3.2\\times 10^{-4}$, the simulated bias amplitudes and the first-order correction basis are both in doubt. Repeating the end-to-end simulation with an independent PSF model, generated without using the model that also defines the correction basis, would test whether the $0.2\\%$/$2\\%$ biases and the correction success are an artifact of model reuse.","supporting_citations":[],"review_version":1}