{"id":"5d069719-b192-4f85-a9c7-69f9e49a9faf","arxiv_id":"2506.05161","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Hourglass simulates the Roman High-Latitude Time-Domain survey catalog, forecasting about 21,700 Type Ia supernovae, 39,000 core-collapse supernovae, and 64,000 total transients with photometry and prism spectra.","lead":"The authors release Hourglass, a simulated catalog of the time-domain sky that NASA's Roman Space Telescope is expected to see in its High-Latitude Time-Domain survey. The simulation includes about 64,000 transients, 11 million photometric observations, and 500,000 prism spectra, and is offered as a public testbed for planning and machine-learning classifiers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Catalog yields inherit the paper's own rate uncertainties—KN simulated at 5×, PISN at 2× an alternative model, SN Ia extrapolated to z=3—without propagated error bars; headline counts need sensitivity bands.","rationale":"The reader's conditional verdict is the right one. The rate sensitivity is not a flaw in the simulation mechanics—the authors are transparent and release inputs—but it is the least secure link between 'simulation output' and 'expected Roman catalog.' The 11,000 Å red edge is a real secondary limitation, especially for prism spectra of non-SN Ia classes at z<0.6, where more than half of the prism bandpass is unsimulated; this should be caveated in the spectra claims. Neither issue invalidates the resource; both argue for conditional acceptance with requested sensitivity bands and SED-extension plans. I agree with the reader's weakest_assumption, so no verdict change.","tokens_in":20056,"tokens_out":9431,"duration_ms":113211,"concrete_test":"Re-run the released SNANA/PIPPIN inputs for a 3-point rate grid holding survey and SED models fixed: (1) SN Ia rate at Strolger et al. (2020) central and ±1σ for z>1; (2) KN rate at the true Abbott et al. (2021) value (i.e., 1/5 of the released simulation rate) and at 3× that value; (3) PISN rate from Pan et al. (2012) in place of Eq. (7). Compare Table 4 counts. If SN Ia counts move by more than the Poisson counting uncertainty and KN/PISN counts move by factors of order 2, the catalog should be released with rate-error bands and the abstract rounded with those bands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The Hourglass catalog is a convolution of SED templates, survey geometry, and volumetric rates, and the headline counts are directly proportional to the rates for each class. For the rarest transients the paper's own text brackets the counts by factors of 2–5: KNe are simulated at five times the Abbott et al. (2021) rate (Sec. 2.1.8), whose uncertainty is stated as a factor of ~3 (up to 4 in Table 4); the PISN rate (Eq. 7) is roughly twice the Pan et al. (2012) value at z=1 (Sec. 2.1.9); and the SN Ia rate (Eq. 1) is extrapolated to z=3 with no error bar even though the majority of the 21,700 SNe Ia are at z>1 and the measured z>1 rate is described as significantly uncertain. The paper acknowledges the dependence (Sec. 4) but does not propagate it, so a user of the released catalog cannot separate rate-driven variations from survey-design effects. This is the load-bearing condition: if the high-redshift SN Ia rate is wrong at the level implied by its source, the headline 21,700 and 'cosmologically useful' 19,000 counts shift by tens of percent, and KN/PISN counts shift by factors. The simulation remains a valid planning tool, but the 'expected' numbers require uncertainty bands.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the Hourglass simulation, an end-to-end forward simulation of the extragalactic time-domain catalog expected from the Roman High-Latitude Time-Domain Core Community Survey under the current reference design. The authors combine rest-frame SED templates and volumetric rate functions for ten transient classes (SNe Ia, SNIa-91bg, SNe Iax, CCSNe, SLSNe-I, TDEs, ILOTs, KNe, PISNe, and AGN) with the SNANA/PIPPIN simulation pipeline, a two-tier photometric survey (wide RZYJ, deep YJHF), prism spectroscopy, and a documented WFI noise model. The headline results are predicted catalog yields (approximately 21,700 SNe Ia, 39,000 CCSNe, 70 SLSNe-I, 39 TDEs, 14 KNe at five times the fiducial rate, 15 PISNe, and 139 AGN), a public data release of photometry, prism spectra, and object metadata, and a demonstration that the SCONE convolutional classifier reaches roughly 94% accuracy and 98% precision for photometrically classifying SNe Ia in these simulations.","tokens_in":20279,"tokens_out":6344,"duration_ms":68206,"significance":"If the adopted rates and SED models are accepted, this is a valuable community planning resource. Its strengths include a public data release with input files and catalog products, built on the well-tested SNANA and PIPPIN infrastructure, a forward-propagated simulation with no parameters fit to the output, and unusually transparent documentation of the survey geometry, instrument characteristics, and selection cuts. The paper also provides the first public simulated Roman prism spectral time series for several non-Ia classes, which is useful for pipeline development. The central weakness is that the headline yields are point predictions that inherit large, explicitly acknowledged input-rate uncertainties without any propagation, so the catalog cannot currently separate rate-driven variations from survey-design effects. The SCONE classifier results are a useful in silico benchmark, not a prediction of real-data performance, and the paper mostly frames them appropriately.","major_comments":[{"comment":"The headline yields in Table 4 are quoted as point values even though several input rates carry large, explicitly acknowledged uncertainties and are extrapolated beyond their measured ranges. Section 2.1.1 assumes the Strolger et al. (2020) SN Ia rate is valid to z=3 even though the majority of the 21,700 simulated SNe Ia lie at z>1 where that source reports significant uncertainty; Section 2.1.8 simulates KNe at five times the Abbott et al. (2021) rate, whose stated uncertainty is a factor of about three (up to four); Section 2.1.9 adopts a PISN rate that at z=1 is nearly twice the Pan et al. (2012) value. Section 4 acknowledges that \"all of these results are highly dependent on the assumed rates,\" but Table 4 and the released catalog contain no uncertainty bands or sensitivity variants. Because the yields are linearly proportional to the adopted volumetric rates, users cannot separate rate-driven variation from survey-design effects. I request a sensitivity analysis (for example, rerunning with the upper and lower rate bounds for SNe Ia, KNe, and PISNe) and propagation of those bounds into Table 4 and the abstract's headline numbers.","section":"§2.2, Table 1, §3.2"},{"comment":"Several non-Ia SED models (CCSN, SLSN-I, TDE, ILOT, and PISN) have a rest-frame red edge of 11,000 Å (Table 1), while the Roman prism is transmissive to 18,000 Å and the deep tier includes H and F filters. The paper notes the J-band consequence in Section 2.2, but the abstract's claim of \"the first realistic simulations of non-Type Ia supernovae spectral-time series data\" is stronger than what these truncated SEDs can support, since at low redshift the reddest prism bins and the H/F photometry are not actually simulated for those classes. Please either extend the SEDs into the near-infrared (as was done for SNe Ia and SNIa-91bg), or explicitly quantify which fraction of objects and which wavelength bins are affected, and soften the \"realistic\" claim accordingly.","section":"§2.2, §3, Table 4"},{"comment":"The host-galaxy catalog limits the simulation to z<3, and the paper states that \"significantly higher redshift PISN and SLSN will be visible\" and that Hourglass only captures the low-redshift tail of these populations. Nevertheless, Table 4 lists PISN (15) and SLSN-I (70) as detected counts without flagging them as truncated lower limits in the table itself. Since Moriya et al. (2022) estimate on the order of 100 PISN at z>5, the z=3 cutoff is not a negligible edge effect for exactly the classes that appear in the abstract as \"possibly pair-instability supernovae.\" Please mark these entries as lower limits and, if feasible, add an estimate of the unmodeled high-redshift contribution.","section":"§2.2, §3, Table 4"}],"minor_comments":[{"comment":"There is a missing space in \"theRoman High-Latitude Time-Domain Core Community Survey\" in the abstract.","section":"Abstract"},{"comment":"The AGN volumetric rate is written as \"1.0−3 Mpc−3\"; this appears to be a formatting error for 1.0×10^-3 Mpc^-3 and should be corrected.","section":"§2.1.10"},{"comment":"The phrase \"A volumetric rate for PISN presented stated in Briel et al. (2022)\" contains a doubled verb; please revise.","section":"§2.1.9"},{"comment":"The column description for \"mw_ebv\" contains the typo \"ling-of-sight\" instead of \"line-of-sight.\"","section":"Table 5"},{"comment":"The detection threshold is described as \"two observations with S/N >5\" in Section 2.2 but as \"a S/N at max of >5\" in the abstract; please clarify whether the requirement is two single-epoch detections at any phase, two detections near peak, or something else.","section":"§2.2 and Abstract"},{"comment":"The horizontal axis of Figure 9 ends at z=2.5, while the text discusses a drop-off at z>2.5 and the simulation extends to z≈3; extending the axis would make the claimed high-redshift degradation visible.","section":"Figure 9"},{"comment":"The SCONE precision and recall are measured on a test set generated with the same simulation machinery used for training. This is a legitimate in silico validation, but the caption and text should more prominently state that these numbers quantify simulation-to-simulation transfer, not expected performance on real Roman data.","section":"§3.3"},{"comment":"The header \"ZPAVG\" is a nonstandard abbreviation; consider defining it in the table caption or using the explicit expression from Equation (8).","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The stress-test concern about missing rate uncertainties is well-founded and is the primary basis for the major revision; the circularity concern does not land, because the simulation is a forward propagation of externally measured inputs. The SED red-edge and z=3 truncation issues are secondary but should be addressed as part of the revision. The paper is otherwise a solid, transparent contribution that should be publishable after the requested sensitivity and labeling changes."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Hourglass simulation is a real contribution: the first public, community-oriented simulated catalog for the Roman High-Latitude Time-Domain survey with this breadth of transient classes, prism spectral time series, and a classification test. It is the kind of resource the community needs before launch, and the authors have done the work to make it usable—data, input files, and documentation are all released.\n\nWhat is new is the catalog itself, not the machinery. SNANA and PIPPIN are well-established, and the paper honestly states that it is extending earlier Roman simulations (Hounsell et al. 2018; Rose et al. 2021) with ten transient classes, updated SEDs and rates, prism spectra, and a SCONE demonstration. Those extensions matter: for the first time someone can generate realistic Roman-like light curves and prism spectra for non-SN Ia events without building the pipeline from scratch.\n\nThe paper does several things well. It is transparent about inputs, selection cuts, and the noise model. It uses publicly available SED and rate models rather than fitting anything to the output, so there is no circularity. It explicitly flags limitations: rate uncertainties, the 11,000 Angstrom SED red edge, the z=3 host-galaxy cutoff, and the known gaps in high-redshift rate measurements. The SCONE test is presented as a demonstration, not a claim that the classifier is flight-ready.\n\nNow the soft spots. The headline counts are directly proportional to adopted volumetric rates, several of which are shaky. The SN Ia rate is extrapolated to z=3 with no error bar even though the source (Strolger et al. 2020) is significantly uncertain above z=1; the KN rate is simulated at five times the already uncertain Abbott et al. rate; the PISN rate is roughly twice the Pan et al. value at z=1. The paper acknowledges this in Section 4 but does not propagate the uncertainties into the reported counts. A user of the catalog cannot tell how much of the 21,700 SNe Ia is rate-driven versus survey-driven. That is a genuine limitation, but it is a limitation of the input data more than of the simulation itself. The SED red edge is a minor issue, and the authors say so.\n\nThe stress-test concern about rate sensitivity holds up. The central argument—that Hourglass is a useful planning and training resource—still stands. It is not a measurement paper; it is a public tool with honest caveats. I would like to see sensitivity bands on the headline numbers and a note on how to rescale the catalog for alternative rates, but that is minor relative to the overall value.\n\nThe paper deserves a serious referee. It is the kind of work that helps the community prepare for Roman, and the data release is a public good. I would cite it if I were working on Roman survey strategy or transient classification. For the reading group, it is worth a slot to discuss how to validate such large simulation efforts.\n\nRecommendation: send it to peer review, accept with minor revisions asking for quantified rate sensitivities on the headline counts.","headline":"Hourglass is a genuinely useful planning and training catalog for Roman time-domain work; the headline counts need rate-uncertainty bands, but that does not undercut its main value.","tokens_in":20947,"tokens_out":1905,"would_cite":true,"duration_ms":25224,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hourglass simulation predicts 64,000 transients for Roman survey","keywords":["Surveys","Catalogs","Time domain astronomy","Space telescopes","Astronomical simulations","Type Ia supernovae","Core-collapse supernovae"],"falsifier":"Measure the volumetric SN Ia rate in the 2<z<3 range — for instance with JWST slitless spectroscopy over a modest field — and compare it to the (1+z)^(-0.1) extrapolation used here; a rate significantly below that curve would falsify the predicted ~21,700 SNe Ia and the ~19,000 cosmologically useful subset proportionally.","tokens_in":19770,"feed_emoji":"🔭","tokens_out":10200,"duration_ms":99847,"temperature":0.7,"pith_summary":"The Hourglass simulation is built to establish what the Roman Space Telescope's High-Latitude Time-Domain Core Community Survey should actually see: it forwards ten extragalactic transient models through the reference survey design and reports a predicted catalog of about 64,000 transients, 11 million photometric observations, and 500,000 prism spectra. The headline counts are approximately 21,700 Type Ia supernovae, 39,000 core-collapse supernovae, and smaller numbers of superluminous supernovae, tidal disruption events, kilonovae, and pair-instability supernovae, assuming a loose two-epoch S/N>5 detection threshold. A sympathetic reader would care because these numbers define the scale of Roman's time-domain data products years before launch, giving the community a concrete basis to plan analysis pipelines, estimate selection effects for cosmology, and train classification algorithms. As a first demonstration, the paper shows that the SCONE classifier recovers Type Ia supernovae with ~98% precision out to z>2, and it releases the simulated photometry, spectra, and input files for others to use.","feed_headline":"Hourglass simulation predicts 64,000 transients for Roman survey","feed_subtitle":"Predicted haul: 21,700 Type Ia supernovae, 39,000 core-collapse, 500,000 spectra for community use.","key_machinery":"The machinery is a forward-modeling pipeline built on the SNANA simulation library, run through the PIPPIN pipeline manager, which turns rest-frame spectral-temporal energy distribution (SED) models into observed Roman light curves and prism spectra. Each of the ten transient classes is injected using a specific SED library — SALT3-NIR for SNe Ia, Vincenzi et al. (2019) models for core-collapse supernovae, MOSFiT-generated templates for superluminous supernovae, tidal disruption events, ILOTs, and pair-instability supernovae, the Bulla (2019) model for kilonovae, and the ELAsTiCc damped random walk model for AGN — with volumetric rates drawn from the literature and host galaxies from the 3DHST catalog. The pipeline then applies the survey geometry, filter transmissions, exposure times, PSF noise-equivalent areas, read noise, sky noise, and a 0.15 mag zero-point scatter, and retains objects with two epochs above S/N=5.","core_discovery":"The central claim is that, under the current design-reference survey — a wide tier of 19.04 $deg^{2}$ in R/Z/Y/J, a deep tier of 4.20 $deg^{2}$ in Y/J/H/F, five-day cadence, two-year baseline, with roughly a fifth of the area covered by the R~100 prism — Roman will produce a science-independent time-domain catalog of over 64,000 transients detected at S/N>5 in two epochs. The simulation counts 21,700 Type Ia supernovae (about 19,000 with S/N at maximum above 10, the paper's proxy for cosmologically useful), 39,000 core-collapse supernovae, 1,300 SN1991bg-like events, 1,300 Type Iax supernovae, about 70 superluminous supernovae, 39 tidal disruption events, 35 intermediate-luminosity optical transients, 15 pair-instability supernovae, and 139 active galactic nuclei, with 14 kilonovae recovered from a simulation injected at five times the assumed true rate. The paper further claims these simulations are realistic enough to train machine-learning classifiers, and demonstrates this with a SCONE model that reaches 94% accuracy and 98% precision separating SNe Ia from contaminants; it also presents the first realistic prism spectral time series for non-Ia transients.","pith_inferences":["Because every count scales linearly with the assumed volumetric rate, the Hourglass catalog can double as a forecast to be tested: the first year of real Roman data should resolve whether the z=2–3 extrapolation of the SN Ia and CCSN rates holds, and a single Roman kilonova would directly test the factor-of-five scaled injection rate.","The 11,000 Å red edge of many SED models means the deep tier's H and F bands are effectively untested for low-redshift objects; extending the SED libraries redward would likely change predicted colors and classification performance in those bands.","The same simulation structure could be extended to variable stars, Galactic sources, and rarer exotic transients to build a complete 'everything that varies' catalog for Roman, which would be the natural basis for alert-broker and anomaly-detection development.","SCONE's precision drops noticeably beyond z≈2.5, suggesting that the highest-redshift SNe Ia — the most valuable for cosmology — are the hardest to classify photometrically; training on prism spectra or adding host-galaxy information is a natural next test."],"forward_implications":["Roman's High-Latitude Time-Domain survey should deliver roughly an order of magnitude more Type Ia supernovae than current cosmological samples, with most of the ~19,000 cosmologically useful events lying above z=1.","Core-collapse supernovae outnumber SNe Ia by almost two to one, making them the dominant source of contamination for photometric SN Ia classification and a key input for contamination studies.","Rare transients will be detected in small but meaningful numbers — about 70 SLSNe, 39 TDEs, 15 PISNe, and ~3 kilonovae at the true rate — enough to begin constraining their rates and physics.","The released simulated photometry, prism spectra, and input files give the community a testbed for survey optimization, classifier training, and selection-effect modeling before launch.","Adopting the newer survey recommendations (larger deep tier, interweaving cadence, pilot and extended surveys) would raise the projected yield by about 25–30%, beyond 100,000 transients."],"supporting_citations":[{"why":"Defines the reference survey design (tiers, filters, exposure times, cadence) that the simulation propagates through.","marker":"Rose et al. (2021)"},{"why":"Provides the SNANA simulation engine used to generate fluxes, light curves, and prism spectra.","marker":"Kessler et al. (2009a)"},{"why":"Supplies the SALT3-NIR spectral model used for Type Ia supernovae, the anchor population.","marker":"Pierel et al. (2022)"},{"why":"Supplies the 65 core-collapse supernova spectral models and their host-galaxy correlations.","marker":"Vincenzi et al. (2019)"},{"why":"MOSFiT generates the SED templates for SLSNe, TDEs, ILOTs, and PISNe.","marker":"Guillochon et al. (2018)"},{"why":"Provides the Type Ia supernova volumetric rate used to set the sample size.","marker":"Strolger et al. (2020)"},{"why":"Provides the core-collapse supernova volumetric rate used to set the sample size.","marker":"Strolger et al. (2015)"},{"why":"Provides the kilonova rate, which the simulation injects at five times this value to beat Poisson noise.","marker":"Abbott et al. (2021)"},{"why":"Provides the pair-instability supernova rate, fit with a polynomial for this simulation.","marker":"Briel et al. (2022)"},{"why":"SCONE is the convolutional network whose photometric classification performance is demonstrated on the simulated light curves.","marker":"Qu et al. (2021)"}],"fun_headline_variants":["Hourglass sim shapes Roman's 64,000-transient catalog","Roman's time-domain survey: 64k transients predicted","Hourglass catalog: 21,700 Type Ia supernovae for Roman","Roman alert: 64,000 transients in two-year survey","Simulation predicts Roman net: 64k transients, 500k spectra"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The predicted catalog sizes scale linearly with input volumetric rates that are either extrapolated far beyond their measured redshift range (SN Ia and CCSN rates assumed valid to z=3) or uncertain by factors of 2–3 (kilonovae, pair-instability supernovae), so if those rates are wrong the headline counts change proportionally.","fun_headline_variants_meta":{"raw":{"variants":["Hourglass sim shapes Roman's 64,000-transient catalog","Roman's time-domain survey: 64k transients predicted","Hourglass catalog: 21,700 Type Ia supernovae for Roman","Roman alert: 64,000 transients in two-year survey","Simulation predicts Roman net: 64k transients, 500k spectra"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001229,"raw_usage":{"total_tokens":5135,"prompt_tokens":1118,"completion_tokens":4017,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":734,"completion_tokens_details":{"reasoning_tokens":3922}},"tokens_in":734,"tokens_out":4017,"duration_ms":39059,"temperature":1.0,"reasoning_tokens":3922,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:23:10.903307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the volumetric SN Ia rate in the 2<z<3 range — for instance with JWST slitless spectroscopy over a modest field — and compare it to the (1+z)^(-0.1) extrapolation used here; a rate significantly below that curve would falsify the predicted ~21,700 SNe Ia and the ~19,000 cosmologically useful subset proportionally.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the SALT3-NIR spectral model used for Type Ia supernovae, the anchor population."},{"cited_title":"E., et al","cited_arxiv_id":null,"evidence_quote":"Supplies the 65 core-collapse supernova spectral models and their host-galaxy correlations."},{"cited_title":"A., et al","cited_arxiv_id":null,"evidence_quote":"MOSFiT generates the SED templates for SLSNe, TDEs, ILOTs, and PISNe."},{"cited_title":"D., Abraham, S., et al","cited_arxiv_id":null,"evidence_quote":"Provides the kilonova rate, which the simulation injects at five times this value to beat Poisson noise."},{"cited_title":"M., Eldridge, J","cited_arxiv_id":null,"evidence_quote":"Provides the pair-instability supernova rate, fit with a polynomial for this simulation."}],"review_version":1}