REVIEW 22 cited by
GIFT-Eval: A Benchmark For General Time Series Forecasting Model Evaluation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Time series foundation models excel in zero-shot forecasting, handling diverse tasks without explicit training. However, the advancement of these models has been hindered by the lack of comprehensive benchmarks. To address this gap, we introduce the General Time Series Forecasting Model Evaluation, GIFT-Eval, a pioneering benchmark aimed at promoting evaluation across diverse datasets. GIFT-Eval encompasses 23 datasets over 144,000 time series and 177 million data points, spanning seven domains, 10 frequencies, multivariate inputs, and prediction lengths ranging from short to long-term forecasts. To facilitate the effective pretraining and evaluation of foundation models, we also provide a non-leaking pretraining dataset containing approximately 230 billion data points. Additionally, we provide a comprehensive analysis of 17 baselines, which includes statistical models, deep learning models, and foundation models. We discuss each model in the context of various benchmark characteristics and offer a qualitative analysis that spans both deep learning and foundation models. We believe the insights from this analysis, along with access to this new standard zero-shot time series forecasting benchmark, will guide future developments in time series foundation models. Code, data, and the leaderboard can be found at https://github.com/SalesforceAIResearch/gift-eval .
Forward citations
Cited by 22 Pith papers
-
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.
-
Into the ORBIT for Time Series: Training Regimes for Foundation Models
A training regime that controls data exposure, context lengths, and prediction horizons yields state-of-the-art zero-shot time series forecasting with a simple encoder-only Transformer.
-
ReasonCast: Towards Explainable Time Series Forecasting with Reasoning
A fine-tuned LLM that states its reasoning, then its forecast, in one response beats specialized forecasters on five synthetic time series patterns.
-
A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods
Transformers—especially a standard encoder-decoder—yield the lowest hourly load forecast errors across TSO, low-voltage feeder, and client-level datasets, with 6.6–10.7% error reduction over the best non-Transformer baseline.
-
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.
-
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
A curated 142-billion-point real-world multivariate time series corpus improves zero-shot forecasting when combined with existing synthetic and univariate pretraining data across four foundation models.
-
Forecasting Realized Volatility with Time Series Foundation Models: A Comparison with Econometric Benchmarks
Zero-shot time series foundation models largely fail to beat econometric benchmarks for realized volatility forecasting, with only TTM achieving a narrow, calibration-driven edge.
-
Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule
Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.
-
OpenMHC: Accelerating the Science of Wearable Foundation Models
OpenMHC contributes the largest open-access consumer wearable dataset to date (67M hours, 11,894 participants), a standardized three-track benchmark, and the first open implementations of Apple WBM and Google LSM-2.
-
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
Foundation models outperform dataset-specific machine learning in energy time series forecasting across 54 datasets in 9 categories.
-
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning
TimeRFT fine-tunes time-series foundation models with step-wise reward signals and difficulty-filtered data, beating supervised fine-tuning on eight benchmarks across data regimes.
-
Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables
ApolloPFN trains a time-aware prior-data fitted network on synthetic time series with exogenous variables and outperforms existing zero-shot forecasters on M5 and electricity price benchmarks.
-
Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?
Time-series foundation models are better calibrated than ARIMA and N-BEATS baselines on the tested datasets, and their calibration does not show the systematic overconfidence seen in image and language models.
-
ARIES: Relation Assessment and Model Recommendation for Deep Time Series Forecasting
ARIES shows that deep forecasting models have consistent performance preferences tied to time series properties, and uses those preferences to recommend models for new datasets.
-
BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
A balanced sampling strategy over statistically characterized time series patterns lets universal forecasting models train on 78 billion tokens instead of 419 billion, with equal or better zero-shot accuracy.
-
MoTime: A Dataset Suite for Multimodal Time Series Forecasting
MoTime provides a large multimodal forecasting benchmark and shows that external text or images can improve forecasts in some datasets, especially cold-start and sparse settings, though gains are inconsistent.
-
Output Scaling: YingLong-Delayed Chain of Thought in a Large Pretrained Time Series Forecasting Model
Forecasting with a non-causal encoder-only model improves fixed-horizon accuracy when the model is asked to output extra future tokens, an effect the authors call delayed chain-of-thought.
-
Investigating Compositional Reasoning in Time Series Foundation Models
On a benchmark where models train on Fourier components and test on their sums, patch-based Transformers and residual MLP architectures show the strongest compositional generalization, while most standard transformers...
-
Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting
A two-stage forecaster (SPA trend extraction + LoRA-fine-tuned residual Transformer) that the paper claims beats prior models by 6.56% MASE, though the claim is not robust to its own extended baseline tables.
-
Foundation models for time series forecasting: Application in conformal prediction
Zero-shot time series foundation models can improve split conformal prediction intervals when training data is scarce, because nearly all data can be reserved for calibration.
-
From Vector Autoregressions to AI-based Time Series Forecasting: A Review
AI forecasting methods are flexible generalizations of the classical VAR's conditional forecast distribution, gaining adaptability and scale but losing ready-made inference, identification, and structural interpretation.
-
ModRWKV: Transformer Multimodality in Linear Time
A linear RNN backbone (RWKV7) with lightweight adapters can handle vision, speech, and time-series inputs, with competitive vision and speech results but unreliable time-series evaluation.
Discussion (0). Continue with ORCID to comment.