Pith. sign in

REVIEW 1 cited by

Multi-Task Time Series Forecasting With Shared Attention

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2101.09645 v1 pith:GVLUD46V submitted 2021-01-24 cs.LG stat.ML

classification cs.LGstat.ML
keywords forecastingseriestimemulti-tasktasksattentionmodelsoutperform
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Time series forecasting is a key component in many industrial and business decision processes and recurrent neural network (RNN) based models have achieved impressive progress on various time series forecasting tasks. However, most of the existing methods focus on single-task forecasting problems by learning separately based on limited supervised objectives, which often suffer from insufficient training instances. As the Transformer architecture and other attention-based models have demonstrated its great capability of capturing long term dependency, we propose two self-attention based sharing schemes for multi-task time series forecasting which can train jointly across multiple tasks. We augment a sequence of paralleled Transformer encoders with an external public multi-head attention function, which is updated by all data of all tasks. Experiments on a number of real-world multi-task time series forecasting tasks show that our proposed architectures can not only outperform the state-of-the-art single-task forecasting baselines but also outperform the RNN-based multi-task forecasting method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mamba Adaptive Anomaly Transformer with association discrepancy for time series

    cs.LG 2025-02 conditional novelty 5.0 of 10

    MAAT integrates sparse attention and a Mamba state-space block into the Anomaly Transformer, reporting small F1 improvements over Anomaly Transformer and DCdetector on seven datasets.

Pith tools