Pith. sign in

REVIEW 22 cited by

GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.10393 v2 pith:5OX6G6IZ submitted 2024-10-14 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelsseriestimefoundationbenchmarkevaluationforecastinggift-eval
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Time series foundation models excel in zero-shot forecasting, handling diverse tasks without explicit training. However, the advancement of these models has been hindered by the lack of comprehensive benchmarks. To address this gap, we introduce the General Time Series Forecasting Model Evaluation, GIFT-Eval, a pioneering benchmark aimed at promoting evaluation across diverse datasets. GIFT-Eval encompasses 23 datasets over 144,000 time series and 177 million data points, spanning seven domains, 10 frequencies, multivariate inputs, and prediction lengths ranging from short to long-term forecasts. To facilitate the effective pretraining and evaluation of foundation models, we also provide a non-leaking pretraining dataset containing approximately 230 billion data points. Additionally, we provide a comprehensive analysis of 17 baselines, which includes statistical models, deep learning models, and foundation models. We discuss each model in the context of various benchmark characteristics and offer a qualitative analysis that spans both deep learning and foundation models. We believe the insights from this analysis, along with access to this new standard zero-shot time series forecasting benchmark, will guide future developments in time series foundation models. Code, data, and the leaderboard can be found at https://github.com/SalesforceAIResearch/gift-eval .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis

    cs.AI 2025-10 conditional novelty 7.0 of 10

    TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.

  2. Into the ORBIT for Time Series: Training Regimes for Foundation Models

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A training regime that controls data exposure, context lengths, and prediction horizons yields state-of-the-art zero-shot time series forecasting with a simple encoder-only Transformer.

  3. ReasonCast: Towards Explainable Time Series Forecasting with Reasoning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    A fine-tuned LLM that states its reasoning, then its forecast, in one response beats specialized forecasters on five synthetic time series patterns.

  4. A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Transformers—especially a standard encoder-decoder—yield the lowest hourly load forecast errors across TSO, low-voltage feeder, and client-level datasets, with 6.6–10.7% error reduction over the best non-Transformer baseline.

  5. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.

  6. RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    A curated 142-billion-point real-world multivariate time series corpus improves zero-shot forecasting when combined with existing synthetic and univariate pretraining data across four foundation models.

  7. Forecasting Realized Volatility with Time Series Foundation Models: A Comparison with Econometric Benchmarks

    q-fin.ST 2026-07 accept novelty 6.0 of 10

    Zero-shot time series foundation models largely fail to beat econometric benchmarks for realized volatility forecasting, with only TTM achieving a narrow, calibration-driven edge.

  8. Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

    stat.AP 2026-07 conditional novelty 6.0 of 10

    Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.

  9. OpenMHC: Accelerating the Science of Wearable Foundation Models

    cs.LG 2026-06 conditional novelty 6.0 of 10

    OpenMHC contributes the largest open-access consumer wearable dataset to date (67M hours, 11,894 participants), a standardized three-track benchmark, and the first open implementations of Apple WBM and Google LSM-2.

  10. FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    Foundation models outperform dataset-specific machine learning in energy time series forecasting across 54 datasets in 9 categories.

  11. TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning

    eess.SP 2026-04 unverdicted novelty 6.0 of 10

    TimeRFT fine-tunes time-series foundation models with step-wise reward signals and difficulty-filtered data, beating supervised fine-tuning on eight benchmarks across data regimes.

  12. Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables

    cs.LG 2026-03 conditional novelty 6.0 of 10

    ApolloPFN trains a time-aware prior-data fitted network on synthetic time series with exogenous variables and outperforms existing zero-shot forecasters on M5 and electricity price benchmarks.

  13. Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Time-series foundation models are better calibrated than ARIMA and N-BEATS baselines on the tested datasets, and their calibration does not show the systematic overconfidence seen in image and language models.

  14. ARIES: Relation Assessment and Model Recommendation for Deep Time Series Forecasting

    cs.LG 2025-09 conditional novelty 6.0 of 10

    ARIES shows that deep forecasting models have consistent performance preferences tied to time series properties, and uses those preferences to recommend models for new datasets.

  15. BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A balanced sampling strategy over statistically characterized time series patterns lets universal forecasting models train on 78 billion tokens instead of 419 billion, with equal or better zero-shot accuracy.

  16. MoTime: A Dataset Suite for Multimodal Time Series Forecasting

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.

  17. Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Forecasting with a non-causal encoder-only model improves fixed-horizon accuracy when the model is asked to output extra future tokens, an effect the authors call delayed chain-of-thought.

  18. Investigating Compositional Reasoning in Time Series Foundation Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    On a benchmark where models train on Fourier components and test on their sums, patch-based Transformers and residual MLP architectures show the strongest compositional generalization, while most standard transformers...

  19. Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting

    stat.ML 2026-07 reject novelty 5.0 of 10

    A two-stage forecaster (SPA trend extraction + LoRA-fine-tuned residual Transformer) that the paper claims beats prior models by 6.56% MASE, though the claim is not robust to its own extended baseline tables.

  20. Foundation models for time series forecasting: Application in conformal prediction

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Zero-shot time series foundation models can improve split conformal prediction intervals when training data is scarce, because nearly all data can be reserved for calibration.

  21. From Vector Autoregressions to AI-based Time Series Forecasting: A Review

    econ.EM 2026-07 unverdicted novelty 4.0 of 10

    AI forecasting methods are flexible generalizations of the classical VAR's conditional forecast distribution, gaining adaptability and scale but losing ready-made inference, identification, and structural interpretation.

  22. ModRWKV: Transformer Multimodality in Linear Time

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A linear RNN backbone (RWKV7) with lightweight adapters can handle vision, speech, and time-series inputs, with competitive vision and speech results but unreliable time-series evaluation.

Pith tools