REVIEW 8 cited by
TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Time series are generated in diverse domains such as economic, traffic, health, and energy, where forecasting of future values has numerous important applications. Not surprisingly, many forecasting methods are being proposed. To ensure progress, it is essential to be able to study and compare such methods empirically in a comprehensive and reliable manner. To achieve this, we propose TFB, an automated benchmark for Time Series Forecasting (TSF) methods. TFB advances the state-of-the-art by addressing shortcomings related to datasets, comparison methods, and evaluation pipelines: 1) insufficient coverage of data domains, 2) stereotype bias against traditional methods, and 3) inconsistent and inflexible pipelines. To achieve better domain coverage, we include datasets from 10 different domains: traffic, electricity, energy, the environment, nature, economic, stock markets, banking, health, and the web. We also provide a time series characterization to ensure that the selected datasets are comprehensive. To remove biases against some methods, we include a diverse range of methods, including statistical learning, machine learning, and deep learning methods, and we also support a variety of evaluation strategies and metrics to ensure a more comprehensive evaluations of different methods. To support the integration of different methods into the benchmark and enable fair comparisons, TFB features a flexible and scalable pipeline that eliminates biases. Next, we employ TFB to perform a thorough evaluation of 21 Univariate Time Series Forecasting (UTSF) methods on 8,068 univariate time series and 14 Multivariate Time Series Forecasting (MTSF) methods on 25 datasets. The benchmark code and data are available at https://github.com/decisionintelligence/TFB. We have also launched an online time series leaderboard: https://decisionintelligence.github.io/OpenTS/OpenTS-Bench/.
Forward citations
Cited by 8 Pith papers
-
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.
-
The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting
Bidirectional joint-attention, complete forecasting aggregation, and direct mapping form the most effective Transformer design for long-term time series forecasting.
-
BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models
A balanced sampling strategy over statistically characterized time series patterns lets universal forecasting models train on 78 billion tokens instead of 419 billion, with equal or better zero-shot accuracy.
-
GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting
Adaptive soft routing among three complementary linear bases via a channel-horizon-phase gate yields competitive multivariate forecasting accuracy with a small, interpretable model.
-
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
CastFSR improves context-aware time series forecasting by combining a fast data-driven forecast prior, slow LLM-driven contextual reasoning, and reflective validation, outperforming most baselines on public benchmarks.
-
MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning
MemCast claims LLM time-series forecasting improves when retrieval from a hierarchical memory of patterns, wisdom, and laws conditions reasoning, but the reported gains depend on a test-label-rewarded confidence update.
-
DisMS-TS: Eliminating Redundant Multi-Scale Features for Time Series Classification
A multi-scale time series classifier that disentangles scale-shared and scale-specific features and reports improved accuracy on six benchmarks.
-
ModRWKV: Transformer Multimodality in Linear Time
A linear RNN backbone (RWKV7) with lightweight adapters can handle vision, speech, and time-series inputs, with competitive vision and speech results but unreliable time-series evaluation.
Discussion (0). Continue with ORCID to comment.