{"id":"06f8d515-c9dc-41f3-8f07-fc9624151269","arxiv_id":"2501.05303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":9,"one_line_summary":"MOST could measure strong-lens time delays of Type Ia supernovae to within a few hours at a 2-day cadence, based on simulated CSST-discovered systems.","lead":"This paper simulates follow-up monitoring of gravitationally lensed Type Ia supernovae with the Muztagh-Ata 1.93m telescope and predicts hour-level time-delay measurement errors. If the forecast holds, the telescope could add an independent, supernova-based route to the Hubble constant and the Hubble tension debate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Early-detection requirement (1–3 days, §2.3) is inconsistent with the quoted CSST ~80-day cadence; no evidence shows hour-level delays persist if discovery is weeks after explosion.","rationale":"The paper is a technical forecast of time-delay measurement precision for lensed SNe Ia with MOST. Its machinery is standard: W7 explosion model, SEDONA radiative transfer, SIE lensing, GPU microlensing maps, and SNTD/SALT2 fitting. The pipeline is described in enough detail to be re-implemented, and the light-curve validation against SN2011fe provides some support. The central claim, however, is conditional on the discovery timeline. The selection criterion in §2.3 explicitly requires detection within 1–3 days of explosion, and the paper identifies CSST as the discovery survey. Yet the quoted CSST cadence is ~80 days in both WFS and DFS, so a supernova can only be caught in the very first few days by chance, with probability roughly the duty cycle of the rise relative to the cadence (a few percent at most). No such probability or mitigation is presented. If the first image is discovered weeks after explosion, the rising light curve is missing, which degrades the constraint on the reference time t0 for that image and hence the measured time delay. The paper does not quantify this degradation, so the advertised 'few hours' error and 'below one hour' bias are not established for the actual discovery scenario. The reader's weakest assumption correctly identifies this as the most load-bearing issue. I also note that the time-delay error forecast is based on only two example systems, but that is secondary because the paper presents them as demonstrations rather than a full population study. The proposed truncation test directly settles whether delayed discovery destroys the forecast; if the errors remain small even when the first image is caught at day 20, the concern is mitigated, but without that test the forecast overreaches. The verdict should remain CONDITIONAL, as the reader already conditioned on this assumption; no further adjustment is needed.","tokens_in":22008,"tokens_out":8979,"duration_ms":91310,"concrete_test":"Re-run the SNTD light-curve fitting for the two example systems after truncating each first-image light curve to start at T_det = 2, 5, 10, 20, and 40 days after explosion, keeping the 2-day MOST cadence and the same noise model. If the median time-delay error stays below ~6 hours for T_det = 20 days, the early-detection assumption is not load-bearing; if it exceeds ~1 day, the headline forecast fails under realistic CSST discovery timing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The forecast's central claim (Abstract; §4) depends on selection criterion 1 in §2.3, which requires glSNe Ia to be detected within 1–3 days of explosion. The paper names CSST as the discovery survey, but the CSST strategy it quotes (§2.3, from Ref. [35]) visits each WFS sightline about twice over ten years, with ~80-day cadence. This makes detection within 1–3 days of explosion effectively impossible; a supernova would typically be discovered weeks after explosion. The paper does not show how time-delay fitting degrades when the first-image light curve starts at day 5, 10, or 20 instead of day 1–3. Missing the rising part of the first image widens the uncertainty on its fitted reference time t0, directly inflating the delay error toward days rather than hours. Because the abstract quotes 'few hours' precision and '<1 hour' bias as general capabilities, the applicability of the forecast to the CSST+MOST program is unsubstantiated. The paper mentions ZTF/WFST as alternative discoverers (Conclusion), but all forecast numbers—SNR, detection rate, errors—assume the CSST early-detection scenario.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper forecasts the precision with which the Muztagh-Ata 1.93m Synergy Telescope (MOST) can measure strong-lensing time delays of gravitationally lensed Type Ia supernovae discovered by CSST. The simulation chain combines an SIE strong-lensing population model, W7/SEDONA supernova spectral time series, microlensing magnification maps computed with a GPU ray-shooting code, and light-curve fitting with SNTD/SALT2. For one quadruple-image system and one double-image system, the authors report time-delay fitting errors of a few hours and biases typically below one hour with a 2-day cadence, and they extrapolate an annual detection rate of 2 quadruple and 14 double systems over 4000 square degrees. The paper frames these results as demonstrating MOST's capability to support independent cosmography via glSNe Ia.","tokens_in":22346,"tokens_out":2400,"duration_ms":25636,"significance":"If the central forecast holds, the paper makes a useful contribution by showing that a relatively small ground-based telescope can reach the few-hour time-delay precision needed for glSNe Ia cosmography, complementing the glQSO-based H0 measurements. The main strengths are the realistic end-to-end simulation: the use of the W7 model with SEDONA, explicit microlensing magnification maps, the achromatic-phase analysis following Goldstein et al., and the use of the actual MOST site parameters. The treatment of microlensing-induced color scatter and the recommendation to choose reference images with minimal chromatic contamination are well motivated. However, the headline precision numbers rest on only two example systems, and the selection criteria include a strong assumption about early detection that is not reconciled with the quoted CSST cadence. The significance is therefore currently conditional on those assumptions being quantified and relaxed.","major_comments":[{"comment":"The forecast requires that each glSNe Ia be detected within 1 to 3 days of explosion, but the paper itself quotes the CSST cadence as approximately 80 days, with each WFS sightline visited only about twice in a decade. Since CSST is the stated discovery survey, this makes detection within 1–3 days effectively impossible for most systems. The paper does not show how the time-delay error and bias degrade when the first image is discovered at day 5, 10, or 20 after explosion, when the rising part of the light curve is partially or entirely missed. Because the fitted reference time t0 is anchored by the light-curve rise, later discovery directly widens the inferred delay uncertainty and could change the conclusion from hours to days. This is a load-bearing assumption and must be addressed, either by modeling realistic CSST discovery epochs or by explicitly restricting the forecast to alternative discovery surveys and showing the corresponding precision.","section":"§2.3 (selection criterion 1) and §2 (CSST cadence)"},{"comment":"The central claim that 'the time delay errors are typically around a few hours, with biases generally being under one hour' is based on fits to only one quadruple-image system and one double-image system. The paper does not quote the distribution of errors or biases across the 14 double and 2 quadruple systems that it predicts per year, nor does it show how the quoted values depend on image magnification, microlensing realization, or source/lens redshift. Given that the abstract and conclusion generalize these two examples to a capability claim, the authors should provide a population-level error and bias distribution, ideally from a bootstrap over their simulated catalog, and report the median and scatter rather than a single example.","section":"§4 and Table 2"},{"comment":"The paper attributes the reported sub-hour biases to microlensing, but there is no control run without microlensing. To establish that the bias is caused by microlensing and to quantify its magnitude, the authors should fit the same light curves with the microlensing magnification set to unity and compare the resulting time-delay offsets. Without this control, the statement that 'the time delay bias caused by microlensing will not significantly impact the systematic accuracy' is not directly supported by the presented comparisons.","section":"§4 and §3.1"},{"comment":"The annual detection rate of 2 quadruple and 14 double systems is derived under the early-detection selection criterion, and the paper explicitly states that it 'omits actual weather, observing strategies, and other potential influencing factors.' This limitation is acknowledged, but the rate is quoted in the abstract and conclusion without the caveat. Since the rate is a secondary result and not the main forecast, this should be reworded to make the conditional nature of the rate explicit wherever it is cited.","section":"§2.3 (detection rate)"}],"minor_comments":[{"comment":"The telescope name is spelled inconsistently: 'Muztage-Ata' in the title and abstract versus 'Muztagh-Ata' throughout the body. The spelling should be unified.","section":"Title and abstract"},{"comment":"The caption lists the bands as 'Sloan u, g, i, r, and z' but the text then refers to 'z, i, r, g, and u' offsets. Please reorder the band names for consistency with the plotted curves and offsets.","section":"Figure 2 caption"},{"comment":"The supernova rate normalization is described as yielding 'approximately 1,105 normal Type Ia supernovae' per square degree per year, but the text does not define the units of the integrand in Eq. (9) clearly. A brief statement of how this number is obtained from the stated parameters would help reproducibility.","section":"§2.3, Eq. (9)"},{"comment":"The notation Δt_{i-r} = t_i - t_r is used for image pairs, but in Table 2 the pairs are labeled as '1-2', '1-3', etc. It would be clearer to define r as a fixed reference image and then list the delays relative to that reference, or to define the pair notation explicitly.","section":"§4, Eq. (13)"},{"comment":"The statement that 'CCSNe are brighter than SNe Ia' is too broad: while some core-collapse supernovae (e.g., Type IIP at peak) can be brighter than SNe Ia, others are fainter. This generalization should be qualified or rephrased.","section":"§5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid simulation study with a clear and potentially useful forecast, but the central quantitative claim currently rests on a small number of hand-picked examples and on an early-detection assumption that is inconsistent with the quoted CSST cadence. I believe the issues are fixable within the manuscript's scope: the authors can add a population-level error distribution, include a no-microlensing control, and rerun the fitting for discovery epochs of a few days to a few weeks after explosion. If those additions still support the few-hour precision claim, the paper would be suitable for publication. I would not recommend rejection because the underlying simulation chain is sound and the limitation statements in the text indicate the authors are aware of some of these caveats."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The new content is a concrete forecast: with a 2-day cadence, MOST could measure glSNe Ia time delays to a few hours with biases under an hour, and should catch about 16 systems per year from 4000 square degrees. That's a useful planning number. The weak spot is the load-bearing early-detection criterion: the paper requires each supernova to be found within 1–3 days of explosion, but names CSST as the discoverer and quotes an ~80-day CSST cadence. A supernova would typically be found weeks after explosion, and the paper never tests how the time-delay errors degrade when the first light curve starts at day 5, 10, or 20. That is a real inconsistency, not a nitpick.\n\nWhat the paper does well: the simulation pipeline is standard and described in enough detail to be re-implemented. W7/SEDONA gives spectra, SIE lenses plus realistic VDF and shear, microlensing maps, and SALT2 fits with SNTD. The two example systems in Table 2 show small biases (≤0.04 d) and few-hour errors, which is plausible and consistent with earlier HOLISMOKES work.\n\nThe main soft spots, in proportion. The headline errors come from one quadruple and one double system, yet the abstract states 'typically a few hours' as if it were general. There is no no-microlensing run, so the claimed bias suppression is not isolated—though the small biases in the examples are at least reassuring. The detection rate of 2+14 per year is conditional on the early-detection criterion, so if discovery delays are realistic, both the rate and the precision forecast are optimistic. The paper mentions ZTF/WFST as alternative discoverers in the conclusion, but none of the simulations use them.\n\nThe stress-test note holds up: the cadence mismatch is in the manuscript itself. The fix is straightforward—add a discovery-delay parameter and rerun the fits—but until that's done, the abstract overstates what the CSST+MOST program can deliver.\n\nBottom line: this is a serious, technically competent forecast paper, but the central claim needs a sensitivity analysis. Send it to peer review and ask the authors to test later discovery and add a no-microlensing control. I'd cite it after that. Reading group: maybe, it's a good case study in how a simulation's selection function can drive the headline result.","headline":"Useful forecast for MOST glSNe Ia, but the early-detection assumption clashes with the quoted CSST cadence and the few-hour errors come from two hand-picked systems; worth refereeing.","tokens_in":22849,"tokens_out":3635,"would_cite":true,"duration_ms":33817,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["97.60.Bw","98.62.Sb","95.75.De"],"model":"deepseek-v4-flash","headline":"With a 2-day cadence, the Muztagh-Ata 1.93m Synergy Telescope can measure gravitationally lensed Type Ia supernova time delays to within a few hours, with bias typically below one hour.","keywords":["strong gravitational lensing","type Ia supernovae","time delay","Hubble constant","Muztagh-Ata 1.93m telescope","microlensing","light curve fitting","CSST survey"],"falsifier":"Compare the CSST observing schedule (or a fiducial simulation of it) against the 1-3 day early-detection criterion: if the probability of catching a lensed SN Ia's first image within that window is near zero, the simulated sample does not represent what CSST will actually deliver and the hour-level error forecast would not transfer to a real MOST campaign.","tokens_in":21850,"feed_emoji":"🔭","tokens_out":9186,"duration_ms":76780,"temperature":0.7,"pith_summary":"The paper forecasts how precisely the Muztagh-Ata 1.93m Synergy Telescope (MOST) can measure the gravitational time delay between multiple images of strongly lensed Type Ia supernovae (glSNe Ia). It simulates the whole chain—SNe Ia spectra from a standard explosion model, a lens population based on CSST survey forecasts, microlensing magnification maps, realistic MOST photometry, and SALT2 template fitting—and finds that with a 2-day cadence the time delay error is only a few hours, with bias typically below one hour. Because glSNe Ia have smooth, well-understood light curves, they avoid the variability and microlensing problems that hamper lensed-quasar delays, so hour-level precision on even a handful of systems would provide an independent check in the Hubble-tension debate. The paper concludes that MOST is well suited to deliver such measurements, complementing quasar-based H0 programs.","feed_headline":"Lensed supernova delays measurable to a few hours on 1.93-m telescope","feed_subtitle":"Two-day cadence on MOST keeps bias under one hour, an independent route to Hubble-constant checks.","key_machinery":"The argument is carried by a simulated observation pipeline. The W7 Chandrasekhar-mass explosion model, processed by the SEDONA Monte Carlo radiative-transfer code, produces time-dependent SN Ia spectra; these are projected onto microlensing magnification maps built from a singular-isothermal-ellipsoid strong-lens population (calibrated to CSST forecasts) plus a stellar field following a standard stellar initial mass function, using a GPU ray-shooting code. The microlensed light curves are sampled with realistic MOST photometry (300s times 9 exposures per epoch, 2-day cadence, 0.82 arcsecond seeing) and fit with the SALT2 template using SNTD. The load-bearing element is the achromatic phase of SN Ia color curves: near peak brightness the band-to-band specific-intensity ratio is roughly constant across the projected supernova disk, so microlensing changes total flux but not color, letting fits of the early light curves return time delays with hour-level errors and sub-hour bias.","core_discovery":"The central claim is that MOST can measure the relative time delays of glSNe Ia discovered by CSST with hour-level accuracy despite microlensing contamination. Building on the fact that the specific-intensity profile of a Type Ia supernova near peak is nearly constant across its projected disk, so that color curves stay achromatic until roughly day 50, the simulations show microlensing scatters the measured delays by only a few tenths of a day. In the two worked examples, one quadruple-image and one double-image system, the fitting errors across image pairs are typically a few hours and the biases are below one hour. The same population model predicts about 2 quadruple and 14 double systems per year over 4000 square degrees observable by MOST. On the paper's own terms, glSNe Ia time delays measured this way are precise and accurate enough to support independent cosmography and to help adjudicate the Hubble tension.","pith_inferences":["The forecast depends on catching the first image within 1-3 days of explosion, while CSST's roughly 80-day revisit cadence makes such early detection far from guaranteed; a realistic early-warning or target-of-opportunity scheme is a testable prerequisite.","Because the achromatic phase lasts roughly 50 days, the 2-day cadence could probably be relaxed or traded for more epochs per night without losing hour-level precision, freeing telescope time for other programs.","Hour-level delays, combined with lens modeling, suggest that a handful of MOST-monitored glSNe Ia could move time-delay cosmography toward the 1% H0 goal, though the paper itself stops at the time-delay measurement rather than the full cosmological inference."],"forward_implications":["With 2-day cadence, MOST achieves time-delay errors of only a few hours on glSNe Ia, so no denser monitoring is needed for this precision.","Microlensing-induced bias stays below one hour, so systematic accuracy of glSNe Ia time-delay cosmography is not limited by microlensing.","The forecast rate of about 2 quadruple and 14 double systems per year gives a concrete target list for a dedicated MOST monitoring program.","The method extends to brighter core-collapse lensed supernovae and to systems discovered by ZTF or WFST, broadening the sample beyond CSST."],"supporting_citations":[{"why":"Supplies the mock lens catalog and SN rate forecasts from which the population model draws lens and source parameters.","marker":"[35]"},{"why":"Establishes the color-curve and achromatic-phase method used to keep microlensing bias in time-delay fits small.","marker":"[42]"},{"why":"Provides the W7 explosion model whose spectra and light curves seed the simulated SNe Ia.","marker":"[45]"},{"why":"SEDONA Monte Carlo radiative-transfer code used to compute time-dependent SN Ia spectra.","marker":"[46]"},{"why":"Supplies the strong-lensing simulation code used to generate the multiple-image lens systems.","marker":"[48]"},{"why":"GPU ray-shooting code used to build microlensing magnification maps.","marker":"[49]"},{"why":"Provides the SNTD fitting code used to extract time delays from the simulated light curves.","marker":"[51]"},{"why":"Gives the SALT2 template model used for light-curve fitting.","marker":"[74]"},{"why":"Supplies the measured seeing and sky brightness used to compute MOST photometric signal-to-noise ratios.","marker":"[37]"}],"fun_headline_variants":["Hour-level lensed supernova delays on 1.93-m telescope","MOST pins lensed supernova delays to few hours","Lensed supernovae: hour-accurate delays on 1.93-m scope","1.93-m telescope measures lensed supernova delays in hours"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forecast assumes that the first image of each lensed supernova is discovered within 1 to 3 days of explosion, but the planned CSST survey revisits a given field only about every 80 days and no other early-trigger mechanism is supplied.","fun_headline_variants_meta":{"raw":{"variants":["Hour-level lensed supernova delays on 1.93-m telescope","MOST pins lensed supernova delays to few hours","Lensed supernovae: hour-accurate delays on 1.93-m scope","1.93-m telescope measures lensed supernova delays in hours"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000238,"raw_usage":{"total_tokens":1569,"prompt_tokens":1062,"completion_tokens":507,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":678,"completion_tokens_details":{"reasoning_tokens":428}},"tokens_in":678,"tokens_out":507,"duration_ms":4795,"temperature":1.0,"reasoning_tokens":428,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T21:12:59.538256+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the CSST observing schedule (or a fiducial simulation of it) against the 1-3 day early-detection criterion: if the probability of catching a lensed SN Ia's first image within that window is near zero, the simulated sample does not represent what CSST will actually deliver and the hour-level error forecast would not transfer to a real MOST campaign.","supporting_citations":[{"cited_title":"Forecast of strongly lensed supernovae rates in the China Space Station Telescope surveys","cited_arxiv_id":"2407.10470","evidence_quote":"Supplies the mock lens catalog and SN rate forecasts from which the population model draws lens and source parameters."},{"cited_title":"The BOSS Emission-Line Lens Survey. IV. : Smooth Lens Models for the BELLS GALLERY Sample","cited_arxiv_id":"1608.08707","evidence_quote":"Supplies the strong-lensing simulation code used to generate the multiple-image lens systems."},{"cited_title":"An Improved GPU-based Ray-shooting Code for Gravitational Microlensing","cited_arxiv_id":null,"evidence_quote":"GPU ray-shooting code used to build microlensing magnification maps."},{"cited_title":"Site testing at Muztagh-ata site II: seeing statistics","cited_arxiv_id":null,"evidence_quote":"Supplies the measured seeing and sky brightness used to compute MOST photometric signal-to-noise ratios."}],"review_version":1}