Pith. sign in

REVIEW 8 cited by

TFB: Towards Comprehensive and Fair Benchmarking of Time Series Forecasting Methods

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.20150 v4 pith:MOJA3QA7 submitted 2024-03-29 cs.LG cs.AIcs.CY

classification cs.LGcs.AIcs.CY
keywords methodsseriestimeforecastingcomprehensivedatasetsbenchmarkdifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Time series are generated in diverse domains such as economic, traffic, health, and energy, where forecasting of future values has numerous important applications. Not surprisingly, many forecasting methods are being proposed. To ensure progress, it is essential to be able to study and compare such methods empirically in a comprehensive and reliable manner. To achieve this, we propose TFB, an automated benchmark for Time Series Forecasting (TSF) methods. TFB advances the state-of-the-art by addressing shortcomings related to datasets, comparison methods, and evaluation pipelines: 1) insufficient coverage of data domains, 2) stereotype bias against traditional methods, and 3) inconsistent and inflexible pipelines. To achieve better domain coverage, we include datasets from 10 different domains: traffic, electricity, energy, the environment, nature, economic, stock markets, banking, health, and the web. We also provide a time series characterization to ensure that the selected datasets are comprehensive. To remove biases against some methods, we include a diverse range of methods, including statistical learning, machine learning, and deep learning methods, and we also support a variety of evaluation strategies and metrics to ensure a more comprehensive evaluations of different methods. To support the integration of different methods into the benchmark and enable fair comparisons, TFB features a flexible and scalable pipeline that eliminates biases. Next, we employ TFB to perform a thorough evaluation of 21 Univariate Time Series Forecasting (UTSF) methods on 8,068 univariate time series and 14 Multivariate Time Series Forecasting (MTSF) methods on 25 datasets. The benchmark code and data are available at https://github.com/decisionintelligence/TFB. We have also launched an online time series leaderboard: https://decisionintelligence.github.io/OpenTS/OpenTS-Bench/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.

  2. The Power of Architecture: Deep Dive into Transformer Architectures for Long-Term Time Series Forecasting

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Bidirectional joint-attention, complete forecasting aggregation, and direct mapping form the most effective Transformer design for long-term time series forecasting.

  3. BLAST: Balanced Sampling Time Series Corpus for Universal Forecasting Models

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A balanced sampling strategy over statistically characterized time series patterns lets universal forecasting models train on 78 billion tokens instead of 419 billion, with equal or better zero-shot accuracy.

  4. GatedLinear: Adaptive Routing of Complementary Linear Bases for Time Series Forecasting

    cs.LG 2026-07 conditional novelty 5.5 of 10

    Adaptive soft routing among three complementary linear bases via a channel-horizon-phase gate yields competitive multivariate forecasting accuracy with a small, interpretable model.

  5. CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting

    cs.AI 2026-08 conditional novelty 5.0 of 10

    CastFSR improves context-aware time series forecasting by combining a fast data-driven forecast prior, slow LLM-driven contextual reasoning, and reflective validation, outperforming most baselines on public benchmarks.

  6. MemCast: Memory-Driven Time Series Forecasting with Experience-Conditioned Reasoning

    cs.LG 2026-02 reject novelty 5.0 of 10

    MemCast claims LLM time-series forecasting improves when retrieval from a hierarchical memory of patterns, wisdom, and laws conditions reasoning, but the reported gains depend on a test-label-rewarded confidence update.

  7. DisMS-TS: Eliminating Redundant Multi-Scale Features for Time Series Classification

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A multi-scale time series classifier that disentangles scale-shared and scale-specific features and reports improved accuracy on six benchmarks.

  8. ModRWKV: Transformer Multimodality in Linear Time

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A linear RNN backbone (RWKV7) with lightweight adapters can handle vision, speech, and time-series inputs, with competitive vision and speech results but unreliable time-series evaluation.

Pith tools