REVIEW 2 cited by
Rethinking Spatio-Temporal Transformer for Traffic Prediction:Multi-level Multi-view Augmented Learning Framework
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Traffic prediction is a challenging spatio-temporal forecasting problem that involves highly complex spatio-temporal correlations. This paper proposes a Multi-level Multi-view Augmented Spatio-temporal Transformer (LVSTformer) for traffic prediction. The model aims to capture spatial dependencies from three different levels: local geographic, global semantic, and pivotal nodes, along with long- and short-term temporal dependencies. Specifically, we design three spatial augmented views to delve into the spatial information from the perspectives of local, global, and pivotal nodes. By combining three spatial augmented views with three parallel spatial self-attention mechanisms, the model can comprehensively captures spatial dependencies at different levels. We design a gated temporal self-attention mechanism to effectively capture long- and short-term temporal dependencies. Furthermore, a spatio-temporal context broadcasting module is introduced between two spatio-temporal layers to ensure a well-distributed allocation of attention scores, alleviating overfitting and information loss, and enhancing the generalization ability and robustness of the model. A comprehensive set of experiments is conducted on six well-known traffic benchmarks, the experimental results demonstrate that LVSTformer achieves state-of-the-art performance compared to competing baselines, with the maximum improvement reaching up to 4.32%.
Forward citations
Cited by 2 Pith papers
-
Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?
Uniform full-range mean broadcasting matches standard spatial attention on six traffic benchmarks (0.14% mean MAE gap) while cutting node mixing cost from O(N²) to O(N).
-
UrbanMind: Urban Dynamics Prediction with Multifaceted Spatial-Temporal Large Language Models
UrbanMind combines a multifaceted masked autoencoder, semantic prompting, and test-time adaptation in an LLM to forecast traffic speed, inflow, and demand, reporting lower MAE and RMSE than baselines in three cities.
Discussion (0). Continue with ORCID to comment.