REVIEW 19 cited by
Unified Training of Universal Time Series Forecasting Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Deep learning for time series forecasting has traditionally operated within a one-model-per-dataset framework, limiting its potential to leverage the game-changing impact of large pre-trained models. The concept of universal forecasting, emerging from pre-training on a vast collection of time series datasets, envisions a single Large Time Series Model capable of addressing diverse downstream forecasting tasks. However, constructing such a model poses unique challenges specific to time series data: i) cross-frequency learning, ii) accommodating an arbitrary number of variates for multivariate time series, and iii) addressing the varying distributional properties inherent in large-scale data. To address these challenges, we present novel enhancements to the conventional time series Transformer architecture, resulting in our proposed Masked Encoder-based Universal Time Series Forecasting Transformer (Moirai). Trained on our newly introduced Large-scale Open Time Series Archive (LOTSA) featuring over 27B observations across nine domains, Moirai achieves competitive or superior performance as a zero-shot forecaster when compared to full-shot models. Code, data, and model weights can be found at https://github.com/SalesforceAIResearch/uni2ts.
Forward citations
Cited by 19 Pith papers
-
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.
-
Byte Pair Encoding for Efficient Time Series Forecasting
A byte-pair-encoding tokenizer that converts repeated temporal motifs into single tokens improves zero-shot forecasting accuracy and speed over sample-wise and patch-based methods.
-
A Benchmark for Electrical Load Forecasting Across Grid Levels: Time-Series Transformers Outperform Established Methods
Transformers—especially a standard encoder-decoder—yield the lowest hourly load forecast errors across TSO, low-voltage feeder, and client-level datasets, with 6.6–10.7% error reduction over the best non-Transformer baseline.
-
Trend strength predicts when generative foundation models win: a power-controlled benchmark, a mechanism, and an actionable selection rule
Zero-shot Chronos wins time-series benchmarks by under-extrapolating trend, and trend strength computed before forecasting predicts when it will beat classical models.
-
WindFM: An Open-Source Foundation Model for Zero-Shot Wind Power Forecasting
A wind-specific foundation model, WindFM, uses hierarchical tokenization and autoregressive pre-training on the NREL WIND Toolkit to achieve state-of-the-art zero-shot wind power forecasts.
-
STARE: Predicting Decision Making Based on Spatio-Temporal Eye Movements
A deep learning architecture that turns raw eye-tracking coordinates into spatial tokens for a time-series foundation model to predict consumer choices.
-
Time Series Foundation Models for Multivariate Financial Time Series Forecasting
Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...
-
Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions
A foundation model of wearable behavioral data outperforms simple baselines and complements a PPG sensor model across 57 health detection tasks.
-
LightGTS: A Lightweight General Time Series Forecasting Model
A lightweight time series foundation model using period-aligned patches and parallel decoding reports zero-shot and full-shot accuracy on nine benchmarks comparable to much larger models.
-
Modular Foundation Models for Time-Series Perception in Digital Twins
A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.
-
FLDmamba: Integrating Fourier and Laplace Transform Decomposition with Mamba for Enhanced Time Series Prediction
FLDmamba combines a learnable Fourier filter on Mamba's step size with a damped-sinusoid output layer and reports superior long-term forecasting accuracy on standard benchmarks.
-
Towards Interpretable Time Series Foundation Models
After fine-tuning on 180 synthetic mean-reverting series annotated by a large multimodal model, small Qwen models can describe trend direction, noise intensity, and extremum location in natural language.
-
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
Pretrained T5 language weights give a persistent validation-loss advantage over random initialization for low-data time series forecasting, and the advantage does not vanish within the training budget.
-
Benchmarking Pre-Trained Time Series Models for Electricity Price Forecasting
No time series foundation model statistically outperforms the biseasonal MSTL model in most European day-ahead electricity price markets in 2024, though Chronos-Bolt and Time-MoE match traditional methods.
-
Probabilistic Forecasting for Building Energy Systems using Time-Series Foundation Models
Fine-tuned time-series foundation models, especially Chronos with LoRA, outperform trained-from-scratch deep forecasters on multi-signal building energy forecasting with limited data.
-
When can isotropy help adapt LLMs' next word prediction to numerical domains?
Using a log-linear model and Jacobian analysis, the paper claims isotropy in LLM hidden embeddings stabilizes the softmax partition function and improves time-series forecasting, with illustrative experiments on five ...
-
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
QuantFlow combines inverted embeddings, bidirectional Mamba decoders, quantile regression, TSMixup, and three-round FedAvg to report competitive one-step MSE on ETTm1/Weather and retain accuracy across 20 non-IID clients.
-
Contextual Deconvolution for Variance-Stable Demand Sensing: Kernel-Modulated Operators in Promotional Retail
A smooth-baseline-plus-sparse-shock decomposition lowers forecast variance and safety stock but increases stockout costs, reducing total inventory cost only when holding costs exceed ~20% of stockout costs.
-
On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating
On four public datasets, Gradient Boosting with hand-built features beat Chronos, Llama, and ARIMA on most accuracy metrics, while Chronos only led on financial sMAPE.
Discussion (0). Sign in to comment.