Pith. sign in

REVIEW 36 cited by

Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.08278 v3 pith:75ZQCHZC submitted 2023-10-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords foundationmodelsseriestimeforecastinglag-llamacapabilitiesdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Over the past years, foundation models have caused a paradigm shift in machine learning due to their unprecedented capabilities for zero-shot and few-shot generalization. However, despite the success of foundation models in modalities such as natural language processing and computer vision, the development of foundation models for time series forecasting has lagged behind. We present Lag-Llama, a general-purpose foundation model for univariate probabilistic time series forecasting based on a decoder-only transformer architecture that uses lags as covariates. Lag-Llama is pretrained on a large corpus of diverse time series data from several domains, and demonstrates strong zero-shot generalization capabilities compared to a wide range of forecasting models on downstream datasets across domains. Moreover, when fine-tuned on relatively small fractions of such previously unseen datasets, Lag-Llama achieves state-of-the-art performance, outperforming prior deep learning approaches, emerging as the best general-purpose model on average. Lag-Llama serves as a strong contender to the current state-of-art in time series forecasting and paves the way for future advancements in foundation models tailored to time series data.

Discussion (0). Sign in to comment.

Forward citations

Cited by 36 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Expert-Guided Forecast Editing for Time-Series Foundation Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    DEFT edits frozen time-series foundation-model forecasts by exploiting model samples and searching over trend/seasonal components, improving forecast quality under small expert-query budgets.

  2. UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control

    cs.LG 2026-06 unverdicted novelty 7.0 of 10

    UC-Search is a model-agnostic test-time wrapper that adds feasibility-automaton search and uncertainty-based risk adjustment to produce better delayed constrained control than CEM, MPPI, and risk-random baselines on p...

  3. Byte Pair Encoding for Efficient Time Series Forecasting

    cs.LG 2025-05 conditional novelty 7.0 of 10

    A byte-pair-encoding tokenizer that converts repeated temporal motifs into single tokens improves zero-shot forecasting accuracy and speed over sample-wise and patch-based methods.

  4. CENTILE: A Telemetry Foundation Model Evaluated by the Decisions It Drives

    cs.NI 2026-08 conditional novelty 6.0 of 10

    One pretrained telemetry model, CENTILE, improves both HPC backfilling and ISP capacity provisioning decisions under replay, with zero-shot transfer across months and domains.

  5. LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models

    cs.LG 2026-07 conditional novelty 6.0 of 10

    LAFP uses MCTS so a frozen TSFM proposes forecast pieces and an LLM ranks and judges them against text, improving text-conditioned forecasts without retraining.

  6. LiFT-MPC: Language-in-the-Loop Feedback Tuning of Cost Previews for MPC

    math.OC 2026-07 conditional novelty 6.0 of 10

    Language-informed residual corrections of cost previews, tuned online by a delayed control-performance loss, improve MPC economic performance with a proven regret-style bound.

  7. Residual-Guided Multi-Resolution Refinement of Foundation Models: A Case Study in Drought Forecasting

    cs.LG 2026-07 conditional novelty 6.0 of 10

    A residual-guided, coarse-to-fine inference wrapper consistently improves frozen time-series foundation models for monthly drought-index forecasting, cutting one-month-ahead MSE by up to 18.9%.

  8. A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Transformers—especially a standard encoder-decoder—yield the lowest hourly load forecast errors across TSO, low-voltage feeder, and client-level datasets, with 6.6–10.7% error reduction over the best non-Transformer baseline.

  9. When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Time-series foundation models are unconditionally better than classical methods on 15/30 datasets, lose early on 6, and a n_train<700 + seasonality rule resolves 10 deployment decisions without training.

  10. Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule

    stat.AP 2026-07 conditional novelty 6.0 of 10

    Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.

  11. Time-Aware Prior Fitted Networks for Zero-Shot Forecasting with Exogenous Variables

    cs.LG 2026-03 conditional novelty 6.0 of 10

    ApolloPFN trains a time-aware prior-data fitted network on synthetic time series with exogenous variables and outperforms existing zero-shot forecasters on M5 and electricity price benchmarks.

  12. DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks

    cs.LG 2026-02 conditional novelty 6.0 of 10

    A fine-tuned GraphPFN foundation model forecasts DeFi exposure networks and stress-test losses, beating learned baselines everywhere and persistence on link statistics, though not on average stress-test error.

  13. Beyond Accuracy: Are Time Series Foundation Models Well-Calibrated?

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Time-series foundation models are better calibrated than ARIMA and N-BEATS baselines on the tested datasets, and their calibration does not show the systematic overconfidence seen in image and language models.

  14. STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements

    cs.NE 2025-08 unverdicted novelty 6.0 of 10

    A deep learning architecture that turns raw eye-tracking coordinates into spatial tokens for a time-series foundation model to predict consumer choices.

  15. Hallucination Detection and Mitigation with Diffusion in Multi-Variate Time-Series Foundation Models

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Pre-trained multivariate time-series imputation models frequently return values that violate known relations between variables, and a diffusion-based score can detect and filter these errors.

  16. Time Series Foundation Models for Multivariate Financial Time Series Forecasting

    q-fin.GN 2025-07 reject novelty 6.0 of 10

    Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...

  17. Decomposing the Time Series Forecasting Pipeline: A Modular Approach for Time Series Representation, Information Extraction, and Projection

    cs.AI 2025-07 conditional novelty 6.0 of 10

    REP-Net, a modular pipeline of representation, memory, and projection modules, achieves competitive forecasting accuracy on seven multivariate benchmarks with lower computational cost.

  18. Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A foundation model of wearable behavioral data outperforms simple baselines and complements a PPG sensor model across 57 health detection tasks.

  19. Time Series Representations for Classification Lie Hidden in Pretrained Vision Transformers

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Frozen vision transformers, applied to image representations of time series, produce classification features that outperform or match time series foundation models on UCR and UEA benchmarks.

  20. Towards a Foundation Model for Communication Systems

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A single pre-trained transformer can forecast and interpolate multiple wireless channel features (rank, precoder, Doppler, delay) on simulated 5G NR data.

  21. Hopformer: Homogeneity-Pursuit Transformer for Time Series Forecasting

    stat.ML 2026-07 reject novelty 5.0 of 10

    A two-stage forecaster (SPA trend extraction + LoRA-fine-tuned residual Transformer) that the paper claims beats prior models by 6.56% MASE, though the claim is not robust to its own extended baseline tables.

  22. Lightweight Wrappers for Adapting Time Series Foundation Models to Regional Drought Forecasting

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Inference-time wrappers that add multi-resolution residual corrections or block-bootstrap averaging to frozen time-series foundation models reduce MSE for regional one-month-ahead SPEI forecasts.

  23. VAIOM: Continuous-Input, Discrete-Output Decoder-Only Financial Sequence Modeling

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A continuous-input, categorical-output decoder-only Transformer improves held-out one-hour FX return likelihood over single-bar LightGBM and statistical baselines.

  24. Towards Interpretable Time Series Foundation Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    After fine-tuning on 180 synthetic mean-reverting series annotated by a large multimodal model, small Qwen models can describe trend direction, noise intensity, and extremum location in natural language.

  25. Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Pretrained T5 language weights give a persistent validation-loss advantage over random initialization for low-data time series forecasting, and the advantage does not vanish within the training budget.

  26. When can isotropy help adapt LLMs' next word prediction to numerical domains?

    cs.CL 2025-05 reject novelty 5.0 of 10

    Using a log-linear model and Jacobian analysis, the paper claims isotropy in LLM hidden embeddings stabilizes the softmax partition function and improves time-series forecasting, with illustrative experiments on five ...

  27. A Novel Hybrid Approach to Contraceptive Demand Forecasting: Integrating Point Predictions with Probabilistic Distributions

    cs.LG 2025-02 reject novelty 5.0 of 10

    A hybrid quantile-averaging method that pins a probabilistic forecast to an expert point forecast is reported to improve contraceptive demand forecasts, but the headline results are internally inconsistent.

  28. Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail

    cs.LG 2026-07 conditional novelty 4.0 of 10

    A smooth-baseline-plus-sparse-shock decomposition lowers forecast variance and safety stock but increases stockout costs, reducing total inventory cost only when holding costs exceed ~20% of stockout costs.

  29. From Vector Autoregressions to AI-based Time Series Forecasting: A Review

    econ.EM 2026-07 unverdicted novelty 4.0 of 10

    AI forecasting methods are flexible generalizations of the classical VAR's conditional forecast distribution, gaining adaptability and scale but losing ready-made inference, identification, and structural interpretation.

  30. On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating

    cs.LG 2025-08 conditional novelty 4.0 of 10

    On four public datasets, Gradient Boosting with hand-built features beat Chronos, Llama, and ARIMA on most accuracy metrics, while Chronos only led on financial sMAPE.

  31. Foundation Models for Demand Forecasting via Dual-Strategy Ensembling

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A dual ensemble of hierarchical partitions and diverse backbones improves foundation-model sales forecasts on M5 and three external datasets, though the zero-shot protocol is under-specified.

  32. Towards Foundation Auto-Encoders for Time-Series Anomaly Detection

    cs.LG 2025-07 conditional novelty 4.0 of 10

    A univariate VAE with dilated convolutions is proposed as a simple 'foundation' model for time-series anomaly detection, with preliminary zero-shot experiments on two datasets.

  33. Delayformer: spatiotemporal transformation for predicting high-dimensional dynamics

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Delayformer embeds each time series variable into a Hankel matrix, processes the matrices as images with a shared ViT, and predicts all variables in parallel.

  34. Foundation Models for Clean Energy Forecasting: A Comprehensive Review

    eess.SY 2025-07 conditional novelty 3.0 of 10

    A survey of foundation model methods, data, and open problems for renewable energy forecasting, built from roughly 218 cited works.

  35. Human in the Loop Adaptive Optimization for Improved Time Series Forecasting

    cs.LG 2025-05 reject novelty 3.0 of 10

    The core idea is standard forecast recalibration, and the reported experiments show mixed, sometimes negative, results with internal table errors, so the claim of consistent improvement fails.

  36. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

Pith tools