{"id":"242ae52d-99f2-4dbc-aa7b-3de22efdf1f8","arxiv_id":"2507.08858","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Zero-shot time series foundation models can improve split conformal prediction intervals when training data is scarce, because nearly all data can be reserved for calibration.","lead":"This paper tests whether time series foundation models like Chronos, Lag-Llama, and TimesFM give better uncertainty intervals than classical forecasting models when used with conformal prediction. It finds that in data-scarce settings these pretrained models often produce tighter and more accurate prediction intervals because they need little or no training data, leaving more data for calibration.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Zero-shot premise unverified: TSFM pretraining corpora likely include public benchmarks such as M3, so the limited-data advantage may reflect memorization rather than reliability.","rationale":"The reader's exchangeability concern is legitimate and already acknowledged by the authors, but the pretraining-overlap question is more fundamental: if the benchmark datasets appear in TSFM training corpora, the comparison is not zero-shot at all. This would invalidate the abstract's attribution of reliability to superior predictive accuracy rather than to prior exposure. The M3 Monthly experiments are the sharpest test because series length 66 is the most data-constrained case, exactly the regime where the paper claims TSFMs shine. This concern is concrete and checkable. I do not change the reader's CONDITIONAL verdict because a clean audit could resolve it either way; if overlap is confirmed, the affected claims would need to be dropped or re-run on genuinely held-out data.","tokens_in":19330,"tokens_out":11875,"duration_ms":150190,"concrete_test":"Inspect the official pretraining data manifests/appendices of Chronos, TimesFM, and Lag-Llama for ERCOT, NN5, and M3. For any overlapping benchmark, rerun the affected experiments (Tables 5-7) with that dataset removed and report MCR, MSIW, and MASE on the remaining non-overlap data. If the TSFM advantage persists, the zero-shot claim survives; if M3 or NN5 removal erases it, the central conclusion is an artifact of data leakage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the assertion in Section 2 that the TSFMs were evaluated \"zero-shot\" and \"have never seen the benchmarking datasets.\" The paper offers no verification, and this is the most load-bearing premise: public pretraining documentation for Chronos lists M3 (and other open benchmarks) among its training corpora, and TimesFM/Lag-Llama are also trained on broad public data. If M3 Monthly—the most data-scarce setting (66 points per series)—overlaps with pretraining, then Table 6's TSFM advantage on M3 reflects memorization, not zero-shot reliability. The abstract's causal claim (\"more reliable ... thanks to their superior predictive accuracy\") would not follow for genuinely unseen limited data. This is a falsifiable correctness risk that can be settled by inspecting data manifests and rerunning after exclusion.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript compares time-series foundation models (Chronos, Chronos-Bolt, TimesFM, TimesFM 2.0, Lag-Llama) against classical baselines (Naive, Seasonal Naive, StatisticalEnsemble light, LightGBM with three train/calibration splits) in a split conformal prediction framework. Using four public datasets (ERCOT, NN5 Daily, NN5 Weekly, M3 Monthly) at short, medium, and long horizons, it reports mean coverage rate (MCR), mean scaled interval width (MSIW), and mean absolute scaled error (MASE). The central claim is that TSFMs give more reliable conformalized prediction intervals in data-constrained settings because they require almost no training data, leaving more data for calibration, and that this advantage grows as data become scarcer. The paper also includes a discussion of limitations, an appendix with detailed result tables, and states that all reproduction code is available.","tokens_in":19498,"tokens_out":5151,"duration_ms":55580,"significance":"If the central claim were fully supported, the practical recommendation—prefer zero-shot TSFMs as base forecasters in split conformal prediction when data are scarce—would be immediately useful and actionable. The paper is transparent about its limitations, ships code, and evaluates multiple frequencies, horizons, and model families. However, the evidence as presented does not yet establish the claim. Coverage shortfalls, the acknowledged violation of exchangeability, and the unverified zero-shot premise leave the observed advantage confounded. The paper does make a worthwhile contribution by identifying the data-allocation trade-off in conformal prediction with foundation models, but the supporting experiments need strengthening before the conclusions can be accepted.","major_comments":[{"comment":"Several TSFM configurations fail to reach the 90% target coverage (e.g., Chronos 82.7%, ChronosBolt 86.3%, TimesFM2 85.8% in ERCOT-8760-S; Chronos 81.7%, ChronosBolt 83.8%, TimesFM 87.4%, TimesFM2 83.2% in ERCOT-8760-M; Chronos 83.9%, TimesFM2 82.8% in ERCOT-8760-L), yet the text concludes in favor of TSFMs 'for the overall prediction.' The claim of 'more reliable' intervals requires an explicit criterion that accounts for coverage shortfalls, and reporting point MCR without standard errors or confidence intervals (20 windows are aggregated without variance) is insufficient to distinguish genuine reliability from calibration luck.","section":"Section 4.4, Table 5"},{"comment":"The authors state that the data do not satisfy the exchangeability criterion of SCP: 'We argue that the coverage rate was not met because the data do not satisfy the exchangeability criterion of SCP' and 'our data do not meet the exchangeability criteria that guarantee marginal coverage.' Under non-exchangeability, the finite-sample marginal coverage guarantee of split conformal prediction is void, so coverage differences between models cannot be interpreted as reliability differences. The paper should either use conformal methods designed for time series (e.g., ACI, EnbPI, CQR) or justify why SCP remains valid under the rolling-window protocol used in Section 4.2.","section":"Sections 4.4 and 6"},{"comment":"The claim that the TSFMs 'have never seen the benchmarking datasets' is load-bearing for the zero-shot interpretation, but it is not verified. Public pretraining corpora for Chronos and TimesFM include broad open benchmarks such as M3; if M3 Monthly (66 points per series) overlapped with pretraining, the TSFM advantage on M3 reported in Table 6 would reflect memorization rather than zero-shot reliability. The authors should inspect the models' data manifests and rerun the analysis after excluding any overlapping pretraining datasets.","section":"Section 2 and Section 4.1"},{"comment":"The experimental design confounds model class with calibration-set size. TSFMs are allocated nearly all observations to calibration (e.g., 8248/8760 points in ERCOT-8760 and 1720/2232 points in ERCOT-2232), while LGBM and StatisticalEnsemble light must reserve 20–80% of data for training. Consequently, the second claimed advantage—'the calibration process is more stable because more data are used for calibration'—is partly true by construction and does not demonstrate an intrinsic property of TSFMs. A controlled comparison with equal calibration sizes across model classes, or with classical models trained on a separate external dataset, is needed to support the causal claim in the abstract.","section":"Tables 3 and 4"}],"minor_comments":[{"comment":"There are numerous typographical errors and inconsistencies, including 'zero-sot' (Section 1), 'resons' (Section 2), 'againt' (Section 4.2), 'primarly' (Section 3), 'montly' (Section 4.5), 'Morever' (Abstract), 'M ISW' (Section 4.3.2), and the inconsistent notation 'maenaive,j' vs. 'M AEnaive,j' in Eq. (8).","section":"Throughout"},{"comment":"The definition of the adjusted miscoverage rate 1−α̂ = ⌈(|Cal|+1)(1−α)⌉/|Cal| should be clarified: as written, it appears to set the effective miscoverage level rather than the quantile index. The relationship to the standard SCP quantile (e.g., the ⌈(|Cal|+1)(1−α)⌉-th smallest conformity score) should be spelled out.","section":"Eq. (1)"},{"comment":"The StatisticalEnsemble light model is omitted from the ERCOT-8760 experiments 'due to time constraints,' but this is only mentioned in the prose; the tables mark the entries with 'x.' This omission should be stated more prominently in the experimental setup and in the table captions.","section":"Section 4.4"},{"comment":"Cross-references to appendix tables are imprecise: Section 4.4 refers to 'Table 5 in the Appendix' and Section 4.5 refers to 'Table 6 in Appendix,' but the tables are in Appendix C and the numbering depends on the final layout. Please use automatic cross-references or explicit appendix labels.","section":"Appendix C"},{"comment":"The reference to 'Angelopoulos and Bate' should be updated to the published version or cite the arXiv identifier consistently; the 'Nixla' entry is incomplete and appears to contain a typo.","section":"References"},{"comment":"The acknowledged limitation about context-length choice is important; a small sensitivity analysis varying the context/calibration split would help quantify how much the conclusions depend on this choice.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant and timely question, but the strength of the conclusions currently exceeds the evidence. The two most serious issues—the unverified no-pretraining-overlap premise and the use of SCP under non-exchangeability—are both empirically checkable and could be fixed within the scope of the manuscript. The data-allocation confound is also fixable by rerunning with matched calibration sizes. If the authors can address these, the paper would be a useful practical contribution. I would advise the editor to request a revision rather than reject, because the core idea is sound and the required changes are well-defined."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is the first side-by-side comparison I've seen of split conformal prediction running on Chronos, TimesFM, Lag-Llama plus classical baselines; that alone makes it useful for applied forecasting under data scarcity. Second, the central claim — TSFMs give more reliable conformalized intervals when data is limited because zero-shot prediction frees data for calibration — is plausible but not yet established. The biggest risk is the unverified zero-shot premise.\n\nWhere credit is due: the experimental setup is transparent. The configuration tables give exact train/calibration splits per model per dataset, and the rolling-window calibration scheme is clearly described. The authors also openly concede in Sections 4.4 and 6 that the data violate the exchangeability assumption SCP relies on; they know the finite-sample coverage guarantee does not strictly hold. The repeated finding that TSFMs beat the baselines on MASE while producing narrower intervals is a real empirical signal, and the gap widening as data shrinks is consistent with the tables.\n\nSoft spots, in order of severity. First, the zero-shot premise is load-bearing and unverified. Section 2 asserts the models \"have never seen the benchmarking datasets,\" with no evidence. M3 Monthly — the most data-scarce case where TSFMs look best — is a public benchmark that plausibly sits in Chronos's and TimesFM's pretraining corpora. If so, the M3 advantage is memorization, not zero-shot generalization. This is falsifiable by checking data manifests and rerunning after exclusion, but as it stands the abstract's causal claim (\"more reliable ... thanks to their superior predictive accuracy\") overreaches. Second, the reliability claim is undercut by the coverage rates themselves: Chronos lands at 82.7% on ERCOT-8760-S, and several other TSFM settings miss the 90% target. With the exchangeability violation conceded, those coverage numbers cannot carry the weight the conclusions place on them. The honest reading is \"tighter intervals with often-but-not-always adequate coverage.\" Some of the comparison also builds the advantage in — TSFMs get nearly all data for calibration while LGBM must reserve training data — and while the paper is upfront about it, the headline numbers mix model accuracy with data allocation. Third, no variance or confidence intervals are reported on any aggregate; the twenty ERCOT windows are averaged bare, and NN5/M3 are single passes. Near the coverage boundary, those differences could be noise. Minor: the code-availability statement points to GitHub with no link, and typos like \"zero-sot\" suggest a rushed final pass.\n\nWho this is for: practitioners doing uncertainty quantification in data-scarce forecasting, and people benchmarking TSFMs. It deserves serious peer review — the question is timely and the comparison is genuinely new. My recommendation: send it to review, but make acceptance contingent on verifying the zero-shot claim and either fixing the coverage shortfalls or softening the reliability language.","headline":"Useful first empirical pass at conformalizing time series foundation models, but the unverified zero-shot premise and coverage shortfalls mean the reliability claim is not yet established.","tokens_in":19989,"tokens_out":5984,"would_cite":false,"duration_ms":56693,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that when data are limited, zero-shot time-series foundation models give more reliable conformalized prediction intervals than classic forecasting models because their predictive accuracy lets nearly all available data go…","keywords":["time series foundation models","zero-shot forecasting","conformal prediction","split conformal prediction","prediction intervals","data scarcity","uncertainty quantification","coverage rate"],"falsifier":"Take a single non-stationary series (for example the hourly ERCOT data) and compute SCP intervals using a random split of the historical window versus a time-ordered split of the same window; if the empirical coverage differs substantially between the two splits, exchangeability fails and the reported model rankings in coverage cannot be attributed to the models themselves.","tokens_in":19154,"feed_emoji":"🎯","tokens_out":9256,"duration_ms":104571,"temperature":0.7,"pith_summary":"The paper sets out to show that time-series foundation models (TSFMs), which forecast in zero-shot mode without training on the target data, are the better base model for split conformal prediction when data are scarce. Because TSFMs need no training set, almost every available data point can be used for calibration, so the conformal interval rests on a large, stable calibration set. Classic models such as LightGBM or statistical ensembles must split scarce data between training and calibration, leaving a small calibration set and forcing a trade-off between accuracy and reliable intervals. On hourly, daily, weekly, and monthly public datasets, the paper finds that TSFMs deliver tighter intervals and better coverage, and that the advantage grows as data shrink.","feed_headline":"Zero-shot forecast models sharpen conformal intervals on scarce data","feed_subtitle":"Foundation models skip training, freeing nearly all data for calibration and tightening prediction intervals.","key_machinery":"The central object is Split Conformal Prediction (SCP) with absolute-residual conformity scores and a user-set miscoverage rate, which splits each series into a context and a calibration set, computes the empirical quantile of the calibration residuals, and applies that quantile as a symmetric interval width around point forecasts. What carries the paper's claim is the interaction between zero-shot forecasting and this calibration split: TSFMs need only a context window (32 to 512 points), so the remaining historical data, up to roughly 94% of the series, is free for calibration, while classical learners like LightGBM give up 20% to 80% of the same data for training and lag creation. The paper uses a rolling-window scheme for TSFM calibration and reports local (per-series) quantiles as the headline results, with global quantiles deferred to the appendix.","core_discovery":"On its own terms, the paper's claim is that in a split conformal prediction (SCP) framework with a target 90% coverage, zero-shot TSFMs—Chronos, Chronos-Bolt, TimesFM, TimesFM 2.0, and Lag-Llama—produce prediction intervals that are both narrower (lower MSIW) and better calibrated (closer to MCR around 90%) than statistical and gradient-boosting baselines when the historical data available for training and calibration is limited. The mechanism the paper identifies is compositional: the zero-shot property removes the training-set requirement, so a TSFM uses a short context and assigns the remaining data (for example 1720 to 8248 points on ERCOT, or 791 minus context minus horizon on NN5 Daily) entirely to calibration, whereas LightGBM must consume 20% to 80% of the same data for fitting and for lag creation. As data shrink, the paper argues, the calibration quantile estimate becomes imprecise and classic models degrade, so the gap is most pronounced in the smallest-data settings. The paper reports the effect holds across horizons and data frequencies, and that larger TSFMs performed better, suggesting the scaling rule has not plateaued.","pith_inferences":["If the result holds, a testable extension is replacing SCP with a time-series-aware conformal method, such as adaptive conformal inference or ensemble batch prediction intervals, on top of the same TSFM forecasts; the paper's own exchangeability caveat suggests coverage gains could be more honest or even larger there.","The context-length versus calibration-fraction trade-off the paper leaves open could be quantified by sweeping the context size against the calibration fraction; TSFMs may tolerate much shorter contexts than the 128- or 512-point windows used here, and the optimal split is likely data-dependent.","The paper does not probe the extreme tail of scarcity, where a TSFM's context window (as low as 32 points) consumes most of the available history; in that regime a classical model trained on the full series could plausibly win, and mapping that threshold would sharpen the practical guidance."],"forward_implications":["In data-scarce forecasting tasks, TSFMs become a natural default base model for split conformal prediction, since they free nearly all data for calibration.","The claimed advantage grows as available data shrinks, so the regime where classic models need the most training data is exactly where TSFM-based intervals should be preferred.","Because TSFM intervals are built on larger calibration sets, the quantile threshold is more stable across rolling windows, reducing variability in interval width.","Improved point accuracy of TSFMs translates into narrower intervals at the same coverage level, since the conformal width is driven by the magnitude of calibration residuals.","Larger TSFMs (more parameters) gave better results, implying further scaling could widen the gap over classical methods in conformal prediction settings."],"supporting_citations":[{"why":"Defines the split conformal prediction procedure with absolute residual calibration that the whole experimental setup uses.","marker":"(Lei et al., 2017)"},{"why":"Grounds conformal prediction and the exchangeability-based coverage guarantee that the paper invokes and then concedes is violated.","marker":"(Shafer and V ovk, 2007)"},{"why":"Supplies the gentle introduction to conformal prediction and the miscoverage definition the paper follows.","marker":"(Angelopoulos and Bate, 2021)"},{"why":"Supplies Chronos, one of the three zero-shot TSFMs whose conformalized intervals are benchmarked.","marker":"(Ansari et al., 2024)"},{"why":"Supplies TimesFM, another zero-shot TSFM used in the comparison.","marker":"(Das et al., 2024)"},{"why":"Supplies Lag-Llama, the compact zero-shot TSFM included to study the effect of model size.","marker":"(Rasul et al., 2024)"},{"why":"Supplies LightGBM, the gradient boosting baseline representing the classical methods that need training data.","marker":"(Ke et al., 2017)"},{"why":"Supplies the statistical ensemble whose components form the StatisticalEnsemble light baseline.","marker":"(Petropoulos and Svetunkov, 2018)"}],"fun_headline_variants":["Zero-shot models sharpen conformal intervals on scarce data","Skip training, shrink intervals: TSFMs in low-data conformal","Foundation models tighten conformal prediction when data is limited","Scarce data? Zero-shot forecast models give sharper conformal intervals"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The coverage guarantee of split conformal prediction rests on the calibration residuals being exchangeable with the residuals that appear at prediction time, and the paper states plainly that its own time-series data do not satisfy this exchangeability, so differences in coverage across models may reflect the violation rather than genuine reliability.","fun_headline_variants_meta":{"raw":{"variants":["Zero-shot models sharpen conformal intervals on scarce data","Skip training, shrink intervals: TSFMs in low-data conformal","Foundation models tighten conformal prediction when data is limited","Scarce data? Zero-shot forecast models give sharper conformal intervals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1464,"prompt_tokens":953,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":440}},"tokens_in":569,"tokens_out":511,"duration_ms":5878,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:55:20.025844+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a single non-stationary series (for example the hourly ERCOT data) and compute SCP intervals using a random split of the historical window versus a time-ordered split of the same window; if the empirical coverage differs substantially between the two splits, exchangeability fails and the reported model rankings in coverage cannot be attributed to the models themselves.","supporting_citations":[],"review_version":1}