{"id":"e9c5e284-e9ef-4e4d-955b-947014352709","arxiv_id":"2501.16311","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"TiDES simulations forecast over 140,000 cosmologically useful type Ia supernovae and a statistical-only dark energy equation-of-state constraint roughly 10 times tighter than DES-SN5YR.","lead":"This paper presents TiDES, a planned spectroscopic follow-up survey on the 4MOST telescope that will classify live supernovae, measure host galaxy redshifts, and monitor active galactic nuclei in the Rubin Observatory era. It uses simulated LSST and 4MOST observations to forecast sample sizes and cosmological constraints, including a statistical-only sub-2 percent measurement of the dark energy equation of state.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The §5 assumption of perfect photometric typing for the 131,000-host SN sample is the load-bearing link; realistic contamination could bias distances and invalidate the sub-2% w claim.","rationale":"The paper is a well-structured survey forecast. The simulation pipeline (snana + 4MOST simulator + ETC) is described in detail and the sample numbers are plausible. The central claim, however, is the σ(w)=0.012 constraint. I examined the assumptions in §5 that feed this number. The most consequential is the perfect photometric typing for the ~131,000-host SN Ia sample, as this sample provides ~90% of the statistical weight. The paper states this assumption explicitly but does not assess its violation. Since the simulation includes non-Ia classes, a straightforward extension would be to pass the mocked light curves through a photometric classifier and include the output in the cosmology fit. The low-z sample timing is also a caveat (only available at end of 10-year LSST), but it is self-admitted and affects statistical power rather than introducing a bias. I therefore agree with the reader's identification of the weakest assumption. The requested test (contamination-inclusive forecast) is concrete and would settle whether the sub-2% claim is robust. The verdict CONDITIONAL is appropriate; no change to the reader's verdict.","tokens_in":31612,"tokens_out":6159,"duration_ms":61876,"concrete_test":"Rerun the §5.1–5.3 cosmology pipeline on the simulated LSST light curves with a realistic photometric classifier (e.g., a machine-learning classifier trained on the Table 1 classes) replacing the perfect-typing assumption; include the resulting misclassified non-Ia events in the Hubble diagram and re-estimate BBC corrections and σ(w). If the recovered w is biased by more than ~0.012 or the statistical uncertainty degrades beyond 2%, the headline claim fails. As a minimal check, inject a flat contamination rate f∈{0.5%,1%,2%,5%} into the photometric sample with distance biases drawn from the contaminating classes and recompute the w constraint.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (σ(w)=0.012, sub-2%, §5.3) is built on the 'LSST+TiDES photometric SN sample' of ~131,000 SNe Ia, which is an order of magnitude larger than the spectroscopically confirmed sample. Section 5 explicitly assumes 'perfect photometric typing based on the light curves' for these objects (and perfect redshifts for all hosts reaching the SSC). The simulation includes non-Ia classes (Table 1), but the cosmology pipeline removes them perfectly, so the forecast contains no contamination term. Realistic photometric classifiers for LSST SNe achieve a few percent contamination, dominated by core-collapse SNe at z≳0.3 and by 91bg/Iax at low z. If even ~1% of the 131k hosts are misclassified, the Hubble diagram contains ~1,300 non-Ia events whose inferred distances are biased, and the BBC 1D corrections (which are themselves derived from a simulation with zero contamination) cannot remove that bias. The paper does not quantify this; the abstract also drops the 'statistical-only' caveat that limits the claim. Without a contamination-inclusive forecast, the sub-2% precision is an upper limit that may not be achievable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TiDES, the 4MOST Time Domain Extragalactic Survey, and uses end-to-end catalogue-level simulations of LSST and 4MOST to forecast the survey's yields and cosmological reach. Three programmes are described: TiDES-Live (spectroscopy of live transients), TiDES-Hosts (host-galaxy redshifts for photometrically classified SNe), and TiDES-RM (reverberation mapping of AGN in the LSST deep drilling fields). The cosmology forecast combines ~12,600 spectroscopically confirmed SNe Ia with ~131,000 photometric SNe Ia having host-galaxy redshifts and an external low-z sample of 2,400 SNe, then applies SALT2 light-curve fits, BBC bias corrections, and wfit to derive constraints. In a flat wCDM model with a CMB prior, the authors report a statistical-only uncertainty of sigma(w)=0.012, a factor of 10 smaller than their DES-SN5YR comparison, and a w0-wa Figure of Merit of 85 (99 with a DESI-BAO-Y1-like prior).","tokens_in":31872,"tokens_out":4704,"duration_ms":48243,"significance":"If the forecast is realized, TiDES would deliver a Hubble diagram of order 143,000 SNe Ia, an order-of-magnitude leap over current SN Ia samples. The paper's strengths are its transparent simulation chain, the use of the 4MOST facility simulator and exposure-time calculator to produce mock spectra, the explicit inclusion of multiple non-Ia transient classes in the simulation, and the clear statement in Section 5.3 that the quoted constraints are statistical-only. The central claim, however, rests on several assumptions that are acknowledged but not stress-tested, most importantly perfect photometric typing of the 131,000-host sample that dominates the cosmological forecast. The abstract also presents the headline precision without the statistical-only caveat. The result is a useful survey description and a defensible first forecast, but the sub-2% w claim is not yet established as robust against realistic classification contamination.","major_comments":[{"comment":"The assumption of perfect photometric typing for the ~131,000 host-sample SNe Ia is load-bearing. The forecast pipeline in Sections 5.1-5.3 applies SALT2 quality cuts and BBC bias corrections but includes no contamination term: the simulated non-Ia populations of Table 1 are removed perfectly by construction, as stated in the opening of Section 5. Realistic photometric classifiers have a few per cent contamination, dominated by core-collapse SNe at z>0.3 and by 91bg/Iax events at low z; even 1% contamination of the 131,000-host sample would inject ~1,300 non-Ia events with biased distance estimates that the 1D BBC corrections, estimated from the same contamination-free simulation, cannot remove. The paper should include a contamination-inclusive forecast or an analytic degradation estimate before claiming sub-2% w precision; as written, the claim is an upper limit under a perfect-classification scenario.","section":"§5 (first paragraph) and §5.2"},{"comment":"The abstract's claim of being 'capable of a sub-2 per cent measurement of the equation-of-state of dark energy' omits the qualifier that this is a statistical-only uncertainty, even though Section 5.3 and the final bullet of Section 6 explicitly describe the result as 'statistical-only precision.' Section 1 itself notes that current SN samples have a systematic error budget comparable to the statistical uncertainty. The abstract and the first paragraph of Section 6 should state the statistical-only caveat, otherwise the headline claim is materially stronger than the forecast supports.","section":"Abstract and §6"},{"comment":"The forecast includes an external low-z sample of 2,400 SNe, but the text immediately notes that this sample size 'will only be achieved towards the end of 10-years of LSST operations, even though TiDES is expected to conclude after the first five years of LSST.' The contours in Figures 12 and 13 therefore do not represent a standalone 5-year TiDES projection; they assume a post-TiDES external sample whose availability is itself a scheduling assumption. The paper should present this explicitly as a scenario with an external low-z sample and, ideally, show how the constraints degrade without it.","section":"§5 (external low-z sample)"}],"minor_comments":[{"comment":"There is a typo in the sentence 'we assume every galaxy reaching our SSC has had a redshift sucessfully measured': 'sucessfully' should be 'successfully'.","section":"§4.2"},{"comment":"The code name 'wfit.exe' should be formatted as a code/software name without the '.exe' suffix, or the suffix should be explained if it is intentional.","section":"§5.3"},{"comment":"The table caption contains a stray space in 'T able 1'; it should read 'Table 1'.","section":"Table 1 caption"},{"comment":"In the Section 6 bullet, 'equations-of-state parameter' should be 'equation-of-state parameter', and 'at-least 143 000 objects' should be 'at least 143 000 objects'; in Figure 12, 'dotted-dashed contours' is more conventionally rendered as 'dot-dashed contours'.","section":"§6 and Figure 12 caption"}],"recommendation":"major_revision","confidential_remarks":"The paper is a survey description and forecast, and its main contribution is the realistic simulation framework. The gap between the unqualified abstract claim and the statistical-only, perfect-typing forecast is the main substantive issue. I would not accept without either a contamination-inclusive forecast or a re-scoped claim that clearly labels the result as an idealized upper limit. The external low-z sample timing also needs to be presented as a scenario rather than a baseline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Alice,\n\nQuick take: this is a solid, transparent survey-design and forecast paper. If you want to know what TiDES will deliver, this is the reference. The main thing to know is that the headline \"sub-2 per cent w\" is a statistical-only, idealized forecast: it assumes perfect photometric typing for the ~131,000 host-galaxy SNe and perfect host redshifts, and it does not include systematics. The paper is explicit about this in Section 5, but the abstract drops the \"statistical-only\" caveat, which overstates the claim.\n\nWhat's new: the end-to-end simulation chain (LSST OpSim -> SNANA -> 4MOST Facility Simulator -> mock spectra -> BBC cosmology) applied to TiDES' actual target selection and fibre allocation. The sample forecasts—30k live SN spectra, 131k host SNe Ia, 700–1000 AGN with RM—are specific and quantified. The cosmological forecast is a factor ~10 tighter than DES-SN5YR statistical-only, which is the right comparison if you take the assumptions at face value.\n\nWhat it does well: the assumptions are mostly stated, the pipeline is described in enough detail to reproduce, and the population numbers are internally consistent. The use of PLAsTiCC and V19 templates, the explicit selection function, and the SSC for live and host spectra all reflect real survey-design thinking.\n\nSoft spots, in order of importance. First, the perfect-photometric-typing assumption is load-bearing: the photometric host sample is ~10x the spectroscopic sample. A few percent contamination from core-collapse or 91bg/Iax SNe would bias distances, and the BBC corrections—derived from the same zero-contamination simulation—can't remove that bias. The stress-test note is right that this is not quantified. Second, the forecast is statistical-only; the paper itself says systematics are comparable to statistical in current analyses, so the factor-10 claim is an upper bound, not an expectation. Third, the host spectroscopic redshift success is assumed perfect for all galaxies reaching the SSC; real spectra with weak features will fail. These are all stated as assumptions, so it's not a hidden flaw—but together they mean the sub-2% number should be read as \"idealized statistical precision\", not an achievable measurement.\n\nWho it's for: anyone planning transient follow-up with 4MOST, LSST DESC members, and SN cosmology forecasters. It's a legitimate survey-design paper, not a measurement. It deserves serious review; the referee should push for a contamination-inclusive forecast and a qualification in the abstract, but the design itself is sound.\n\nRecommendation: send it to peer review. It's a useful, citable survey-design reference, and the assumptions can be tightened in revision.","headline":"A transparent, well-built survey-design forecast; the sub-2% w headline is statistical-only and assumes perfect photometric typing, but the paper states both assumptions clearly.","tokens_in":32532,"tokens_out":2177,"would_cite":true,"duration_ms":21963,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The planned TiDES survey could build a 143,000-supernova Hubble diagram and pin dark energy's equation-of-state parameter to 1.2 per cent.","keywords":["surveys","type Ia supernovae","dark energy","cosmology","spectroscopic follow-up","reverberation mapping","active galactic nuclei","4MOST"],"falsifier":"A concrete calculation that would settle the claim: pass the simulated LSST light curves through a realistic photometric classifier with a measured confusion matrix, re-run the same bias-correction and cosmological fit, and compare the recovered uncertainty on $w$ with the paper's 0.012; if realistic contamination pushes $\\sigma(w)$ above 2 per cent, the headline forecast depends on the perfect-typing assumption.","tokens_in":31379,"feed_emoji":"🔭","tokens_out":10616,"duration_ms":92595,"temperature":0.7,"pith_summary":"TiDES is a planned five-year spectroscopic follow-up survey on the 4MOST facility, designed to turn the Rubin Observatory's LSST transient stream into the largest ever sample of cosmologically useful type Ia supernovae. The paper simulates both surveys end to end and argues that TiDES will assemble a Hubble diagram of at least 143,000 SNe Ia—about 12,600 with live TiDES spectra and roughly 131,000 photometrically typed SNe Ia with host-galaxy redshifts—and measure the dark-energy equation-of-state parameter $w$ to a statistical-only precision of $\\sigma(w)=0.012$ in a flat $w$CDM model with a CMB prior. That is a factor of ten tighter than the DES-SN5YR forecast the paper uses as its benchmark. The same survey is projected to obtain more than 30,000 live transient spectra, covering core-collapse, superluminous, and rare fast/faint classes, and to run a reverberation-mapping campaign on 700–1,000 AGN that extends an independent Hubble diagram to $z\\sim2.5$. A sympathetic reader would care because the paper is a test of whether a modest, fixed fibre allocation can unlock the cosmological and astrophysical potential of the LSST alert stream.","feed_headline":"TiDES forecast: 143,000 supernovae, dark energy to 1.2%","feed_subtitle":"Simulations of the 4MOST and Rubin surveys predict a Hubble diagram ten times tighter than today's best supernova sample.","key_machinery":"The argument is carried by an end-to-end catalogue-level simulation pipeline rather than by a single analytic identity. LSST OpSim observing patterns are converted into SNANA inputs, which generate realistic light curves from spectral templates and a host-galaxy library; a TiDES selection function mimics real-time triggering of live transients and faded-host targets; the 4MOST Facility Simulator (4FS) allocates fibres and schedules a mock five-year survey alongside other 4MOST programmes; and mock spectra are produced with the 4MOST exposure-time calculator and judged against per-survey Spectral Success Criteria—mean SNR of 5 per 15 Å for live SN classification, SNR of 3 per Å for host-galaxy redshifts, and SNR of 10 per 15 Å for AGN. For cosmology, the surviving SNe Ia are fit with SALT2, distances are estimated with the Tripp formula, selection effects are corrected with BBC, and parameters are extracted with the wfit minimiser in SNANA. This chain converts fibre-hour accounting into a predicted $\\sigma(w)$.","core_discovery":"The paper's central claim is that the TiDES fibre budget—roughly 30–35 low-resolution fibres per 4MOST pointing, about 250,000 fibre-hours over five years—is sufficient to create a SN Ia sample an order of magnitude larger than today's best surveys. In the simulations, the live-transient programme yields about 18,000 SNe Ia with spectra at SNR per 15 Å of at least 3 (12,600 of them pass the SALT2 'cosmologically useful' cuts), while the host-galaxy programme yields about 131,000 photometrically identified SNe Ia with host spectroscopic redshifts. Combining these with a low-redshift sample of 2,400 SNe Ia and a Planck prior, the paper forecasts $\\sigma(w)=0.012$ for a flat $w$CDM model, quoting this as a sub-2 per cent measurement and a factor-of-ten improvement over a DES-SN5YR-like simulation; for a flat $w_0w_a$CDM model the figure of merit is 85, or 99 when combined with a DESI BAO Y1 prior, and is 15 times larger than the DES-SN5YR figure of merit. The paper also claims this one survey will enlarge the spectroscopically confirmed transient population by an order of magnitude, including over 9,000 core-collapse SNe and about 3,000 hydrogen-poor superluminous SNe, and will deliver one of the largest reverberation-mapped AGN samples over $0.1<z<2.5$, with 700–1,000 targets across the four deep fields.","pith_inferences":["The paper's headline forecast treats the 131,000-object host-redshift sample as perfectly typed from light curves; a natural next calculation is to inject a realistic classifier confusion matrix into the BBC pipeline and quantify how much of $\\sigma(w)=0.012$ is carried by that assumption.","Once real data arrive, the overlap between TiDES-Live and TiDES-Hosts provides a built-in empirical check of photometric typing purity—spectroscopically classified SNe within the host-selected sample can be used to measure contamination directly and re-weight the Hubble diagram.","The AGN programme's yield depends on maintaining the simulated 14-day cadence and season lengths in the deep fields; if the actual 4MOST schedule delivers shorter baselines, the number of recovered lags and the reach of the $z\\sim2.5$ Hubble diagram will shrink accordingly.","Because other 4MOST surveys will observe millions of galaxies, many SN hosts will acquire redshifts without TiDES fibre-hours, so the effective host sample may be larger or cheaper than simulated, which would justify re-optimising the fibre split among the three TiDES programmes."],"forward_implications":["If the forecast holds, SN Ia cosmology becomes statistics-limited far below current systematic budgets, shifting the field's bottleneck to photometric typing purity, bias corrections, and distance-estimator systematics.","The roughly 12,600 live SN Ia spectra will form a large, homogeneous training set for photometric classifiers, potentially extending the cosmological sample beyond TiDES hosts to the full LSST SN Ia population.","The order-of-magnitude expansion of spectroscopically confirmed transients, including rare classes such as calcium-strong transients and tidal disruption events, would sharpen measured rates and map the luminosity–timescale plane far more completely than current samples.","The AGN reverberation-mapping sample provides a second, independent standardisable candle out to $z\\sim2.5$, allowing a cross-check of the SN Ia dark-energy measurement and dynamical supermassive black hole masses at cosmic noon.","Host-galaxy spectroscopy will tie SN Ia distances to environmental properties such as stellar mass, metallicity, and star-formation rate, enabling tests of whether environment-dependent corrections reduce Hubble residuals."],"supporting_citations":[{"why":"Defines the 4MOST facility—fibre numbers, field of view, and low-resolution spectrographs—that set every TiDES yield and fibre-hour budget.","marker":"R. S. de Jong et al. 2019"},{"why":"Provides the SNANA simulation code used to generate realistic SN light curves from survey metadata and spectral templates.","marker":"R. Kessler et al. 2009"},{"why":"Supplies the PLAsTiCC simulation framework and SED templates for peculiar SNe, TDEs, CaST, and SLSNe used in the mock transient populations.","marker":"R. Kessler et al. 2019a"},{"why":"Updates the core-collapse SN templates and host-galaxy association prescriptions used in the SNANA simulations.","marker":"M. Vincenzi et al. 2019"},{"why":"Provides the volumetric SN Ia rate evolution used to normalize the number of simulated type Ia supernovae.","marker":"C. Frohmaier et al. 2019"},{"why":"Supplies the BBC bias-correction method used to build the selection-corrected Hubble diagram and derive cosmological constraints.","marker":"R. Kessler et al. 2019b"},{"why":"Defines the DES-SN5YR simulation and analysis that serves as the benchmark for the claimed factor-of-ten improvement in $w$.","marker":"DES Collaboration et al. 2024"},{"why":"Provides the simulated low-redshift Foundation SN sample included in the TiDES cosmology forecast.","marker":"D. O. Jones et al. 2019"},{"why":"Sets the DESC science-requirement forecast methodology and the external low-$z$ sample strategy adopted in the cosmology analysis.","marker":"The LSST Dark Energy Science Collaboration et al. 2018"},{"why":"Supplies the probabilistic fibre-assignment algorithm used by the 4MOST Facility Simulator to allocate TiDES targets.","marker":"E. Tempel et al. 2020a"}],"fun_headline_variants":["TiDES: 143k supernovae, 1.2% dark energy","TiDES sims: 10x tighter dark energy from 143k SNe","TiDES: sub-2% dark energy from 143k supernovae","TiDES: 250k fiber hours, 143k SNe, 1.2% dark energy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The 131,000-host supernova sample—the bulk of the Hubble diagram—is assumed to be perfectly classified as standard type Ia supernovae from their light curves alone; if even a small fraction of those events are actually a different kind of explosion, the distance estimates and the sub-2 per cent dark-energy claim would not hold.","fun_headline_variants_meta":{"raw":{"variants":["TiDES: 143k supernovae, 1.2% dark energy","TiDES sims: 10x tighter dark energy from 143k SNe","TiDES: sub-2% dark energy from 143k supernovae","TiDES: 250k fiber hours, 143k SNe, 1.2% dark energy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001963,"raw_usage":{"total_tokens":7769,"prompt_tokens":1139,"completion_tokens":6630,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":755,"completion_tokens_details":{"reasoning_tokens":6535}},"tokens_in":755,"tokens_out":6630,"duration_ms":46267,"temperature":1.0,"reasoning_tokens":6535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T13:32:26.468732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete calculation that would settle the claim: pass the simulated LSST light curves through a realistic photometric classifier with a measured confusion matrix, re-run the same bias-correction and cosmological fit, and compare the recovered uncertainty on $w$ with the paper's 0.012; if realistic contamination pushes $\\sigma(w)$ above 2 per cent, the headline forecast depends on the perfect-typing assumption.","supporting_citations":[],"review_version":1}