Pith. sign in

REVIEW 7 cited by

Timer: Generative Pre-trained Transformers Are Large Time Series Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02368 v3 pith:3VKDMKLW submitted 2024-02-04 cs.LG stat.ML

classification cs.LGstat.ML
keywords modelstimeserieslargedeepgenerativesmalldatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning has contributed remarkably to the advancement of time series analysis. Still, deep models can encounter performance bottlenecks in real-world data-scarce scenarios, which can be concealed due to the performance saturation with small models on current benchmarks. Meanwhile, large models have demonstrated great powers in these scenarios through large-scale pre-training. Continuous progress has been achieved with the emergence of large language models, exhibiting unprecedented abilities such as few-shot generalization, scalability, and task generality, which are however absent in small deep models. To change the status quo of training scenario-specific small models from scratch, this paper aims at the early development of large time series models (LTSM). During pre-training, we curate large-scale datasets with up to 1 billion time points, unify heterogeneous time series into single-series sequence (S3) format, and develop the GPT-style architecture toward LTSMs. To meet diverse application needs, we convert forecasting, imputation, and anomaly detection of time series into a unified generative task. The outcome of this study is a Time Series Transformer (Timer), which is generative pre-trained by next token prediction and adapted to various downstream tasks with promising capabilities as an LTSM. Code and datasets are available at: https://github.com/thuml/Large-Time-Series-Model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis

    cs.AI 2025-10 conditional novelty 7.0 of 10

    TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.

  2. LightGTS: A Lightweight General Time Series Forecasting Model

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A lightweight time series foundation model using period-aligned patches and parallel decoding reports zero-shot and full-shot accuracy on nine benchmarks comparable to much larger models.

  3. Fusing Large Language Models with Temporal Transformers for Time Series Forecasting

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A gated fusion of GPT-2 semantic features and a PatchTST-style Transformer encoder improves average MSE/MAE slightly on ETT, Weather, and ILI, while losing to PatchTST on four of the six datasets.

  4. Large Causal Models for Temporal Causal Discovery

    cs.LG 2026-02 conditional novelty 4.0 of 10

    A transformer pretrained on a large mixed corpus of synthetic and simulated realistic time series can discover lagged causal graphs zero-shot on datasets up to 12 variables, outperforming several classical baselines.

  5. Foundation Model Empowered Synesthesia of Machines (SoM): AI-native Intelligent Multi-Modal Sensing-Communication Integration

    eess.SP 2025-06 conditional novelty 4.0 of 10

    The paper proposes a systematic classification and two roadmaps for using foundation models (LLMs and wireless foundation models) to design Synesthesia of Machines systems for 6G, with preliminary case-study evidence ...

  6. ModRWKV: Transformer Multimodality in Linear Time

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A linear RNN backbone (RWKV7) with lightweight adapters can handle vision, speech, and time-series inputs, with competitive vision and speech results but unreliable time-series evaluation.

  7. Scaling Transformers for Time Series Forecasting: Do Pretrained Large Models Outperform Small-Scale Alternatives?

    cs.LG 2025-06 reject novelty 3.0 of 10

    LLM4TS_FS achieves the best MSE on four of seven long-term datasets, but the claimed broad advantage of pre-trained large models over small transformers is not consistent across all benchmarks.

Pith tools