Pith. sign in

REVIEW 3 major objections 7 minor 92 references

Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting

T0 review · 3 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read DOST hooks a lightweight per-location adapter and a weekly awake-hibernate update cycle onto an offline traffic forecaster, cutting forecast error by 12.89% across four real-world urban datasets.

desk verdict DOST is a well-engineered online adapter for urban ST forecasting, but the paper underspecifies the baseline protocol, which could trivialize the headline gains. read the letter →

arxiv 2411.15893 v1 pith:HDNLVK67 submitted 2024-11-24 cs.LG cs.AI

classification cs.LGcs.AI
keywords onlinecontinuallearningspatio-temporalforecastingdistributionshiftconceptdrifttraffictaxidemandpredictionawake-hibernatestreamingmemory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that urban spatio-temporal data—taxi demand, traffic speed—does not just vary from moment to moment; its underlying distribution drifts gradually over weeks and months, and the drift differs by location. DOST is an online continual-learning wrapper that bolts onto an existing spatio-temporal forecaster and retrains only a tiny location-specific adapter, on a fixed awake-then-hibernate schedule aligned to weekly periodicity, to track those shifts at low compute cost. The authors report that DOST beats thirteen offline and online baselines on four real-world datasets, with online forecasts at an average of about 0.1 seconds and a 12.89% average error reduction. If this holds, the value is practical: cities can keep existing forecasting models accurate in production without expensive full re-training.

What carries the argument

The Variable-Independent Adapter (VIA) is a set of $N$ lightweight MLP sub-adapters, one per urban location, each with a residual skip connection, inserted before the spatio-temporal module; only these per-location adapters are fine-tuned online. The Awake-Hibernate (AH) learning strategy alternates an awake phase of length $L_a$ (set to one week) with a hibernate phase of length $L_h = \lambda L_a$ (with $\lambda=1$), and the Streaming Memory Update (SMU) mechanism maintains a reservoir-sampled Streaming Memory Buffer of the most recent AH cycle and fine-tunes the adapter on a tiny randomly drawn Episodic Memory to capture recent patterns while preventing catastrophic forgetting.

What would settle it

Feed DOST a stream trained on normal traffic that then contains a single sudden, permanent large shift—for instance a bridge closure that changes a corridor's demand pattern overnight—and measure MAE during the hibernate week; if errors spike sharply relative to the awake week, the weekly-periodicity assumption is the limiting mechanism, whereas no spike would suggest the schedule is not actually driving the reported gains.

Watch

Extended reading notes

Core claim

The paper's central claim is that urban spatio-temporal distributions shift gradually and location-specifically, and that a fixed alternating schedule of one awake week of fine-tuning followed by one hibernate week of frozen parameters can track those shifts better than immediate per-sample updates or full-model fine-tuning. The proposed DOST framework attaches a small Variable-Independent Adapter (VIA) to an existing spatio-temporal network, updates only that adapter during awake phases, and uses a Streaming Memory Update (SMU) mechanism to sample a tiny episodic memory from a reservoir buffer so that adaptation avoids catastrophic forgetting. On Chicago taxi demand, Singapore taxi demand, METR-LA traffic speed, and PEMS-BAY traffic speed, DOST reports lower MAE, RMSE, and WMAPE than all thirteen baselines, with an average forecast-error reduction of 12.89% and online inference at about 0.1 seconds per forecast.

Load-bearing premise

The whole schedule rests on the assumption that urban distributions shift gradually and repeat weekly, so a fixed cycle of one awake week and one frozen hibernate week is enough to track the drift; abrupt or non-periodic changes like an accident, a weather extreme, or a policy change would freeze all adaptation for up to a week.

Editorial extensions

If this is right

  • Offline spatio-temporal forecasters such as STGCN, MTGNN, and GWNet can be upgraded to online drift-tracking by adding the VIA and AH strategy, as the paper's strategy-integration experiments on Singapore-T show.
  • Because only the adapter is updated, per-forecast compute stays around 0.1 seconds, making the approach feasible for city-wide real-time deployment.
  • A 12.89% average error reduction across both region-based and road-based datasets implies directly better taxi-demand and traffic-speed predictions under real-world streaming conditions.
  • The weekly periodicity prior could be reused to schedule model updates for other urban data streams with gradual drift, such as crowd flow or energy demand, without full re-training.
  • The fixed awake-hibernate cycle bounds the computational cost of continual learning, since the hibernate phase only updates the memory buffer and performs inference.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The awake-hibernate schedule likely generalizes to other spatio-temporal domains with strong weekly periodicity, but abrupt non-periodic events such as storms, incidents, or policy changes would probably break it because adaptation freezes for up to a week.
  • The VIA's variable-independent design suggests a broader pattern: separate a shared spatial model from per-variable drift adapters, a modularity that could transfer to multivariate time-series forecasting beyond urban spatio-temporal data.
  • A testable extension is to make the awake-hibernate cycle adaptive, adjusting $L_a$ and $\lambda$ based on drift detection, since the fixed cycle cannot react mid-hibernate to a shift.
  • The SMU's choice to exclude the very latest sample from the episodic memory is a deliberate bias; an ablation comparing that choice against always including the latest sample would quantify the trade-off between stability and responsiveness.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 7 minor

Summary. The paper proposes DOST, an online continual learning framework for urban spatio-temporal forecasting. DOST combines a Variable-Independent Adapter (VIA) with per-location sub-adapters, an Awake-Hibernate (AH) strategy that alternates between active adapter fine-tuning and frozen phases, and a Streaming Memory Update (SMU) mechanism with a reservoir-sampled buffer. Experiments on Chicago-T, Singapore-T, METR-LA, and PEMS-BAY compare DOST with 13 baselines and report lower MAE, RMSE, and WMAPE, an average error reduction of 12.89%, sub-0.1-second per-sample inference, and ablations isolating VIA, AH, SMU, and memory reset.

Significance. If the empirical claims hold, DOST would be a practically useful online adaptation layer for existing ST forecasters: VIA is parameter-efficient, the AH schedule reduces update frequency, and the component ablations show that each design choice contributes. The paper reports five-seed repetitions with standard deviations and significance markers, and it tests the framework on multiple ST backbones (STGCN, MTGNN, GWNET), which are strengths. However, two protocol issues—an unspecified baseline update schedule and the use of validation samples to seed the streaming memory buffer—currently prevent the headline superiority claim from being accepted at face value. No public code or data is provided, which further limits verification.

major comments (3)
  1. [§4.1.3, Table 2] The paper never states whether the 13 baselines are trained only on the warm-up split and then frozen, periodically retrained, or given any online adaptation during the 6/8 online span. If STGCN, GWNET, AGCRN, MTGNN, GMSDR, PDFormer, REVIN, PatchTST, and DLinear are frozen, then DOST's online data access alone could explain a large part of the reported 12.89% average improvement, independent of VIA, AH, or SMU. Please specify the exact protocol for every baseline and justify it; if baselines are frozen, include an online-adapted variant or otherwise separate the effect of streaming access from the effect of the proposed components.
  2. [Algorithm 1 (lines 3–4), §3.2.2] Validation samples from D_val are inserted into the SMB before the online phase, and during awake phases the SMU samples from this buffer to compute gradients for the adapter (Algorithm 1 lines 12–15). This means the model is updated on validation data before and during the reported test phase, so the test errors in Table 2 are not produced on fully held-out data, and early stopping based on the validation split is compromised. The warm-up protocol should be changed to use only training data in the memory buffer, or the validation set should be shown to be disjoint and non-influential; Table 2 should then be rerun and the significance claims rechecked.
  3. [§3.1.3, Fig. 6] The awake-hibernate schedule assumes weekly periodicity and sets L_a to one full week and lambda=1; no independent validation of this periodicity assumption is given beyond the illustrative KDE plots in Figure 1. For non-periodic or abrupt shifts such as incidents, weather extremes, or policy changes, the hibernate phase can freeze adaptation for up to a week, which is a material limitation on the paper's general claim that urban ST distributions 'typically' shift gradually. Please provide a sensitivity analysis with L_a and lambda varied over non-weekly values, or add a trigger-based awake decision, and temper the general claim accordingly.
minor comments (7)
  1. [Table 2] OneNet's WMAPE entry '9.14%±0.06%%' contains a double percent sign, and the statistical test for the ‡ marker is not described; please specify whether the tests are paired by seed, by time step, or by dataset and whether any multiple-comparison correction is applied.
  2. [Abstract, §4.2.1] The 12.89% average error reduction is not defined; please state the aggregation formula over datasets and metrics and report the per-dataset and per-metric reductions that lead to this average.
  3. [Table 4] The column 'Total Inference Time' appears to include adaptation time for online methods such as DOST, FSNet, and OneNet; the caption should state explicitly whether forward pass and update time are both included, so the 'inference' nomenclature is not misleading.
  4. [§3.2.1, Eq. (6)] The modulo notation 'tau . 0 (mod L_ah)' is nonstandard; please use a standard congruence symbol such as tau ≡ 0 (mod L_ah).
  5. [Figure 2] The caption says the Memory Placeholder is omitted, but several memory/buffer boxes appear in the figure; a clearer legend or annotated mapping to Algorithm 1 would help readers understand the data flow.
  6. [§4.3] The ablation study reports only PEMS-BAY; reporting ablations on at least one region-based dataset would strengthen the claim that VIA and SMU benefit both road-based and region-based urban ST data.
  7. [Reproducibility] No link to code or data is provided; given the protocol ambiguities above, public code and a precise configuration file would substantially aid verification.

Circularity Check

0 steps flagged · score 2.0 of 10

No substantive circularity; DOST's central claims rest on external benchmark comparisons and ablations, with only minor non-load-bearing self-citations.

full rationale

The paper does not derive its headline result from its own assumptions. DOST's superiority claim is an empirical comparison against 13 external baselines on four real-world datasets (Table 2), and the component contributions are tested through ablations (Figure 5), strategy studies (Table 5), and hyperparameter analyses (Figure 6). No equation in Sections 3.1-3.3 is defined in terms of the outcome it is later said to predict, and no fitted parameter is renamed as a prediction. The only self-citations are [60] and [61]. Reference [60] is cited alongside external reference [53] to support the weekly-periodicity premise of the awake-hibernate schedule, and Figure 1 provides direct empirical evidence of week-to-week distribution similarity; it is a supporting citation, not the load-bearing justification. Reference [61] appears only in related-work positioning. The awake-hibernate assumption of gradual, weekly-periodic drift is a stated modeling assumption rather than a circularly derived conclusion; the fact that it is validated only on the same four datasets is a limitation or generalization concern, not a circularity. The reviewer concern about whether offline baselines are frozen or updated during the online phase is a benchmarking-fairness question and does not make the derivation circular; even if true, it would affect the interpretation of Table 2, not the logical dependency of the method on its inputs. Accordingly, the score reflects only the presence of minor self-citation that is not load-bearing.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central empirical claim rests on several domain assumptions rather than on a derived theory. The method assumes urban distributions drift gradually and repeat weekly, that each location's shift can be modeled independently by a small adapter, that the spatial backbone stays valid while only adapters are updated, and that a memory reset to the latest awake-hibernate cycle keeps the most relevant samples. These assumptions are plausible for the tested traffic datasets but are not tested outside the four benchmarks. Hyperparameters lambda, M, M_e, d_m, and awake length are selected experimentally on PEMS-BAY and then fixed across datasets. No new physical or conceptual entities are introduced; the adapter, memory buffer, and awake decider are standard ML components.

free parameters (6)
  • AH parameter lambda = 1
    Chosen by sweeping lambda on PEMS-BAY (Figure 6b); controls the ratio of hibernate to awake phase. Applied to all datasets without per-dataset tuning.
  • SMB size M = 1000
    Selected from a sweep on PEMS-BAY (Figure 6c) over {100, 300, 500, 1000, 1500}.
  • Episodic memory size M_e = 8
    Selected from a sweep on PEMS-BAY (Figure 6d) over {0, 4, 8, 16, 32}.
  • VIA bottleneck dimension d_m = 4
    Selected from a sweep on PEMS-BAY (Figure 6a); d_m=8 improves accuracy but adds about 128K parameters.
  • Awake phase length L_a = 672 for Chicago-T and Singapore-T; 2016 for METR-LA and PEMS-BAY
    Set to one week of time intervals, embedding the weekly-periodicity assumption into the schedule.
  • Warm-up/online split ratio = 2:6
    Protocol choice affecting how much history is seen before online adaptation; no sensitivity analysis is reported.
assumptions (5)
  • domain assumption Urban ST data distributions drift gradually over time and exhibit weekly periodic patterns.
    Sections 1 and 3.1.3; this justifies fixed awake and hibernate phases aligned to a week. Abrupt or non-periodic shifts would break the schedule.
  • domain assumption Location-specific distribution shifts can be modeled independently by one small adapter per location, without cross-location sharing of shift information.
    Section 3.1.2, VIA uses N independent sub-adapters; no mechanism shares shift knowledge across locations.
  • domain assumption Samples from the most recent AH cycle are the most relevant for adaptation; older samples can be discarded when the SMB is reset.
    Section 3.2.1 states 'distant past data can become irrelevant' and the SMB is reset at the start of each hibernate phase.
  • domain assumption The spatial-temporal backbone remains valid during online shifts, so only the VIA adapter needs updating.
    Section 3.2.3 fine-tunes only adapters in awake phases; if the spatial correlation structure itself shifts, the frozen backbone may fail.
  • domain assumption External factors, specifically date and time, are sufficient to schedule adaptation; the awake decider is a fixed periodic schedule, not a learned policy.
    Section 3.1.3 sets L_a and L_h as fixed proportions of a week; there is no explicit shift detection.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting." pith.science (2026). https://pith.science/paper/HDNLVK67

@misc{pith2026241115893,
  author       = {Pith},
  title        = {Pith review of: Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HDNLVK67}},
  note         = {Machine review of arXiv:2411.15893}
}
read the original abstract

Urban spatio-temporal (ST) forecasting is crucial for various urban applications such as intelligent scheduling and trip planning. Previous studies focus on modeling ST correlations among urban locations in offline settings, which often neglect the non-stationary nature of urban ST data, particularly, distribution shifts over time. This oversight can lead to degraded performance in real-world scenarios. In this paper, we first analyze the distribution shifts in urban ST data, and then introduce DOST, a novel online continual learning framework tailored for ST data characteristics. DOST employs an adaptive ST network equipped with a variable-independent adapter to address the unique distribution shifts at each urban location dynamically. Further, to accommodate the gradual nature of these shifts, we also develop an awake-hibernate learning strategy that intermittently fine-tunes the adapter during the online phase to reduce computational overhead. This strategy integrates a streaming memory update mechanism designed for urban ST sequential data, enabling effective network adaptation to new patterns while preventing catastrophic forgetting. Experimental results confirm DOST's superiority over state-of-the-art models on four real-world datasets, providing online forecasts within an average of 0.1 seconds and achieving a 12.89% reduction in forecast errors compared to baseline models.

Figures

Figures reproduced from arXiv: 2411.15893 by the authors.

Figure 1
Figure 1. An illustration of Chicago’s taxi demand distribu [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of DOST, which employs two strategies: (a) Adaptive spatio-temporal (ST) network for online learning, where modules are represented in three colors: gray for traditional modules, green for the adapter, and yellow for the awake decider. (b) Awake-Hibernate (AH) learning strategy, which alternates network updates between awake and hibernate phases. During the awake phase, the adapter is fine-tuned using the S… view at source ↗
Figure 3
Figure 3. Overview of Variable-Independent Adapter (VIA), [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: This mechanism fine-tunes the network with a tiny [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Ablation study of DOST on PEMS-BAY dataset. [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Effects of hyperparameters on PEMS-BAY dataset. [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

92 extracted references · 55 canonical work pages

  1. [1]

    Sheng-hai An, Byung-Hyug Lee, and Dong-Ryeol Shin. 2011. A survey of intelli- gent transportation systems. In 2011 third international conference on computa- tional intelligence, communication systems and networks . IEEE, 332–337

  2. [2]

    Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting. Advances in neural information processing systems 33 (2020), 17804–17815

  3. [3]

    Peter J Brockwell, Peter J Brockwell, Richard A Davis, and Richard A Davis. 2016. Introduction to time series and forecasting . Springer

  4. [4]

    Pietro Buzzega, Matteo Boschini, Angelo Porrello, Davide Abati, and Simone Calderara. 2020. Dark experience for general continual learning: a strong, simple baseline. Advances in neural information processing systems 33 (2020), 15920– 15930

  5. [5]

    Fabio Cermelli, Dario Fontanel, Antonio Tavera, Marco Ciccone, and Barbara Caputo. 2022. Incremental learning in semantic segmentation from image la- bels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4371–4381

  6. [6]

    Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajan- than, Puneet K Dokania, Philip HS Torr, and Marc’Aurelio Ranzato. 2019. On tiny episodic memories in continual learning. arXiv preprint arXiv:1902.10486 (2019)

  7. [7]

    Mario Cools, Elke Moons, and Geert Wets. 2010. Assessing the impact of weather on traffic intensity. Weather, Climate, and Society 2, 1 (2010), 60–68

  8. [8]

    Carlos Oliveira Cruz and Joaquim Miranda Sarmento. 2021. The impact of COVID-19 on highway traffic and management: The case study of an operator perspective. Sustainability 13, 9 (2021), 5320

Show all 92 references
  1. [9]

    Yousef-Awwad Daraghmi and Motaz Daadoo. 2015. Improved dynamic route guidance based on holt-winters-taylor method for traffic flow prediction. (2015)

  2. [10]

    Marcos VO de Assis, Luiz F Carvalho, Joel JPC Rodrigues, and Mario Lemes Proença. 2013. Holt-winters statistical forecasting and aco metaheuristic for traffic characterization. In 2013 IEEE International Conference on Communications (ICC). IEEE, 2524–2528

  3. [11]

    Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Gregory Slabaugh, and Tinne Tuytelaars. 2021. A continual learning survey: Defying forgetting in classification tasks. IEEE transactions on pattern analysis and machine intelligence 44, 7 (20...

  4. [12]

    Jinliang Deng, Xiusi Chen, Renhe Jiang, Xuan Song, and Ivor W Tsang. 2021. St-norm: Spatial and temporal normalization for multi-variate time series fore- casting. In Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining. 269–278

  5. [13]

    Arthur Douillard, Yifu Chen, Arnaud Dapogny, and Matthieu Cord. 2021. Plop: Learning without forgetting for continual semantic segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 4040–4050

  6. [14]

    Wenying Duan, Xiaoxi He, Zimu Zhou, Lothar Thiele, and Hong Rao. 2023. Localised Adaptive Spatial-Temporal Graph Neural Network. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining

  7. [15]

    Ziquan Fang, Lu Pan, Lu Chen, Yuntao Du, and Yunjun Gao. 2021. MDTP: A multi-source deep traffic prediction framework over spatio-temporal trajectory data. Proceedings of the VLDB Endowment 14, 8 (2021), 1289–1297

  8. [16]

    Kaiqun Fu, Taoran Ji, Liang Zhao, and Chang-Tien Lu. 2019. Titan: A spatiotem- poral feature learning framework for traffic incident duration prediction. In Proceedings of the 27th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems. 329–338

  9. [17]

    Yuan Gao and Dorota Glowacka. 2016. Deep gate recurrent neural network. In Asian conference on machine learning . PMLR, 350–365

  10. [18]

    Nuwan Gunasekara, Bernhard Pfahringer, Heitor Murilo Gomes, and Albert Bifet. 2023. Survey on online streaming continual learning. In Proceedings of the Thirty-Second International Joint Conference on Artificial Intelligence, IJCAI-23 . 6628–6637

  11. [19]

    Jindong Han, Weijia Zhang, Hao Liu, Tao Tao, Naiqiang Tan, and Hui Xiong

  12. [20]

    Liangzhe Han, Ruixing Zhang, Leilei Sun, Bowen Du, Yanjie Fu, and Tongyu Zhu. 2023. Generic and dynamic graph representation learning for crowd flow modeling. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 37. 4293–4301

  13. [21]

    Md Yousuf Harun, Jhair Gallardo, Tyler L Hayes, Ronald Kemker, and Christopher Kanan. 2023. SIESTA: Efficient Online Continual Learning with Sleep. arXiv preprint arXiv:2303.10725 (2023)

  14. [22]

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778

  15. [23]

    Sepp Hochreiter and Jürgen Schmidhuber. 1997. Long short-term memory.Neural computation 9, 8 (1997), 1735–1780

  16. [24]

    Wei-Chiang Hong, Ping-Feng Pai, Shun-Lin Yang, and Robert Theng. 2006. High- way traffic forecasting by support vector regression model with tabu search algorithms. In The 2006 IEEE International Joint Conference on Neural Network Proceedings. IEEE, 1617–1621

  17. [25]

    Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-efficient transfer learning for NLP. In International Conference on Machine Learning. PMLR, 2790–2799

  18. [26]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large Language Models. In The Tenth International Conference on Learning Representa- tions, 2022, Virtual Event, April 25-29, 2022 ...

  19. [27]

    Jiahao Ji, Jingyuan Wang, Zhe Jiang, Jingtian Ma, and Hu Zhang. 2020. Inter- pretable spatiotemporal deep learning model for traffic flow prediction based on potential energy fields. In 2020 IEEE International Conference on Data Mining (ICDM). IEEE, 1076–1081

  20. [28]

    Jiawei Jiang, Chengkai Han, Wayne Xin Zhao, and Jingyuan Wang. 2023. PDFormer: Propagation Delay-aware Dynamic Long-range Transformer for Traffic Flow Prediction. In AAAI. AAAI Press

  21. [29]

    Guangyin Jin, Yuxuan Liang, Yuchen Fang, Zezhi Shao, Jincai Huang, Junbo Zhang, and Yu Zheng. 2023. Spatio-temporal graph neural networks for predictive learning in urban computing: A survey. IEEE Transactions on Knowledge and Data Engineering (2023)

  22. [30]

    Taesung Kim, Jinhee Kim, Yunwon Tae, Cheonbok Park, Jang-Ho Choi, and Jaegul Choo. 2021. Reversible instance normalization for accurate time-series forecasting against distribution shift. In International Conference on Learning Representations

  23. [31]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. 2015. Adam: A Method for Stochastic Opti- mization. In 3rd International Conference on Learning Representations, 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings , Yoshua Bengio and Yann LeCun (Eds.)

  24. [32]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In International Conference on Learning Repre- sentations

  25. [33]

    Yaguang Li, Kun Fu, Zheng Wang, Cyrus Shahabi, Jieping Ye, and Yan Liu. 2018. Multi-task representation learning for travel time estimation. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. 1695–1704

  26. [34]

    Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In 6th International Conference on Learning Representations, 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedi...

  27. [35]

    Zhonghang Li, Lianghao Xia, Yong Xu, and Chao Huang. 2024. FlashST: A Simple and Universal Prompt-Tuning Framework for Traffic Prediction. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net

  28. [36]

    Dachuan Liu, Jin Wang, Shuo Shang, and Peng Han. 2022. Msdr: Multi-step dependency relation networks for spatial temporal forecasting. In Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1042–1050

  29. [37]

    Hangchen Liu, Zheng Dong, Renhe Jiang, Jiewen Deng, Jinliang Deng, Quan- jun Chen, and Xuan Song. 2023. Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of the 32nd ACM International Conference on Information and Knowled...

  30. [38]

    David Lopez-Paz and Marc’Aurelio Ranzato. 2017. Gradient episodic memory for continual learning. Advances in neural information processing systems 30 (2017)

  31. [39]

    Ilya Loshchilov and Frank Hutter. 2019. Decoupled Weight Decay Regularization. In 7th International Conference on Learning Representations, 2019, New Orleans, Chengxin Wang, Gary Tan, Swagato Barman Roy, and Beng Chin Ooi LA, USA, May 6-9, 2019 . OpenReview.net

  32. [40]

    Jie Lu, Anjin Liu, Fan Dong, Feng Gu, Joao Gama, and Guangquan Zhang. 2018. Learning under concept drift: A review. IEEE transactions on knowledge and data engineering 31, 12 (2018), 2346–2363

  33. [41]

    Hao Miao, Yan Zhao, Chenjuan Guo, Bin Yang, Kai Zheng, Feiteng Huang, Jiandong Xie, and Christian S Jensen. 2024. A unified replay-based continuous learning framework for spatio-temporal prediction on streaming data. arXiv preprint arXiv:2404.14999 (2024)

  34. [42]

    Truong Thao Nguyen, François Trahay, Jens Domke, Aleksandr Drozd, Emil Vatai, Jianwei Liao, Mohamed Wahib, and Balazs Gerofi. 2022. Why globally re-shuffle? Revisiting data shuffling in large scale deep learning. In 2022 IEEE International Parallel and Distributed Processing S...

  35. [43]

    Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam

    Yuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, and Jayant Kalagnanam. 2023. A Time Series is Worth 64 Words: Long-term Forecasting with Transformers. In The Eleventh International Conference on Learning Representations, 2023, Kigali, Rwanda, May 1-5, 2023 . OpenReview.net

  36. [44]

    Judea Pearl et al. 2000. Models, reasoning and inference. Cambridge, UK: Cam- bridgeUniversityPress 19, 2 (2000)

  37. [45]

    Jonas Pfeiffer, Andreas Rücklé, Clifton Poth, Aishwarya Kamath, Ivan Vulic, Sebastian Ruder, Kyunghyun Cho, and Iryna Gurevych. 2020. AdapterHub: A Framework for Adapting Transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: S...

  38. [46]

    Jonas Pfeiffer, Ivan Vulic, Iryna Gurevych, and Sebastian Ruder. 2020. MAD-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer. InProceed- ings of the 2020 Conference on Empirical Methods in Natural Language Processing, EMNLP 2020, Online, November 16-20, 2020 ...

  39. [47]

    Quang Pham, Chenghao Liu, Doyen Sahoo, and Steven C. H. Hoi. 2023. Learning Fast and Slow for Online Time Series Forecasting. In The Eleventh International Conference on Learning Representations, 2023, Kigali, Rwanda, May 1-5, 2023

  40. [48]

    Xinwu Qian and Satish V Ukkusuri. 2015. Spatial variation of the urban taxi ridership using GPS data. Applied geography 59 (2015), 31–42

  41. [49]

    Bin Ran and David Boyce. 2012. Modeling dynamic transportation networks: an intelligent transportation system oriented approach . Springer Science & Business Media

  42. [50]

    Zezhi Shao, Fei Wang, Yongjun Xu, Wei Wei, Chengqing Yu, Zhao Zhang, Di Yao, Guangyin Jin, Xin Cao, Gao Cong, et al. 2023. Exploring Progress in Multivari- ate Time Series Forecasting: Comprehensive Benchmarking and Heterogeneity Analysis. arXiv preprint arXiv:2310.06119 (2023)

  43. [51]

    Zezhi Shao, Zhao Zhang, Fei Wang, Wei Wei, and Yongjun Xu. 2022. Spatial- temporal identity: A simple yet effective baseline for multivariate time series forecasting. InProceedings of the 31st ACM International Conference on Information & Knowledge Management. 4454–4458

  44. [52]

    Zezhi Shao, Zhao Zhang, Wei Wei, Fei Wang, Yongjun Xu, Xin Cao, and Chris- tian S. Jensen. 2022. Decoupled dynamic spatial-temporal graph neural network for traffic forecasting. Proc. VLDB Endow. 15, 11 (jul 2022), 2733–2746

  45. [53]

    Hongzhi Shi and Yong Li. 2018. Discovering periodic patterns for large scale mo- bile traffic data: Method and applications.IEEE Transactions on Mobile Computing 17, 10 (2018), 2266–2278

  46. [54]

    Chen-Hui Song, Xi Xiao, Bin Zhang, and Shu-Tao Xia. 2023. Follow the Will of the Market: A Context-Informed Drift-Aware Method for Stock Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management. 2311–2320

  47. [55]

    Yongxin Tong, Yuqiang Chen, Zimu Zhou, Lei Chen, Jie Wang, Qiang Yang, Jieping Ye, and Weifeng Lv. 2017. The simpler the better: a unified approach to predicting original taxi demands based on large-scale online platforms. In Proceedings of the 23rd ACM SIGKDD international co...

  48. [56]

    Quang Thanh Tran, Zhihua Ma, Hengchao Li, Li Hao, and Quang Khai Trinh

  49. [57]

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)

  50. [58]

    Jeffrey S Vitter. 1985. Random sampling with a reservoir. ACM Transactions on Mathematical Software (TOMS) 11, 1 (1985), 37–57

  51. [59]

    Binwu Wang, Yudong Zhang, Xu Wang, Pengkun Wang, Zhengyang Zhou, Lei Bai, and Yang Wang. 2023. Pattern expansion and consolidation on evolving graphs for continual traffic prediction. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 2223–2232

  52. [60]

    Chengxin Wang, Yuxuan Liang, and Gary Tan. 2022. Periodic residual learning for crowd flow forecasting. In Proceedings of the 30th International Conference on Advances in Geographic Information Systems . 1–10

  53. [61]

    Chengxin Wang, Yuxuan Liang, and Gary Tan. 2024. CityCAN: Causal Attention Network for Citywide Spatio-Temporal Forecasting. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining . 702–711

  54. [62]

    Jingyuan Wang, Jiawei Jiang, Wenjun Jiang, Chao Li, and Wayne Xin Zhao

  55. [63]

    Kuo Wang, LingBo Liu, Yang Liu, GuanBin Li, Fan Zhou, and Liang Lin. 2023. Urban regional function guided traffic flow prediction. Information Sciences 634 (2023), 308–320

  56. [64]

    Leye Wang, Di Chai, Xuanzhe Liu, Liyue Chen, and Kai Chen. 2023. Exploring the Generalizability of Spatio-Temporal Traffic Prediction: Meta-Modeling and an Analytic Framework. IEEE Transactions on Knowledge and Data Engineering 35, 4 (2023), 3870–3884

  57. [65]

    Dali Wei and Hongchao Liu. 2013. An adaptive-margin support vector regression for short-term traffic flow forecast. Journal of Intelligent Transportation Systems 17, 4 (2013), 317–327

  58. [66]

    Qingsong Wen, Weiqi Chen, Liang Sun, Zhang Zhang, Liang Wang, Rong Jin, Tieniu Tan, et al. 2024. Onenet: Enhancing time series forecasting models under concept drift by online ensembling. Advances in Neural Information Processing Systems 36 (2024)

  59. [67]

    Billy M Williams and Lester A Hoel. 2003. Modeling and forecasting vehicular traffic flow as a seasonal ARIMA process: Theoretical basis and empirical results. Journal of transportation engineering 129, 6 (2003), 664–672

  60. [68]

    Xinle Wu, Dalin Zhang, Chenjuan Guo, Chaoyang He, Bin Yang, and Christian S Jensen. 2021. AutoCTS: Automated correlated time series forecasting.Proceedings of the VLDB Endowment 15, 4 (2021), 971–983

  61. [69]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining . 753–763

  62. [70]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph WaveNet for Deep Spatial-Temporal Graph Modeling. InProceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI 2019, Macao, China, August 10-16, 2019 . ijcai...

  63. [71]

    Mingxing Xu, Wenrui Dai, Chunmiao Liu, Xing Gao, Weiyao Lin, Guo-Jun Qi, and Hongkai Xiong. 2020. Spatial-temporal transformer networks for traffic flow forecasting. arXiv preprint arXiv:2001.02908 (2020)

  64. [72]

    Huaxiu Yao, Fei Wu, Jintao Ke, Xianfeng Tang, Yitian Jia, Siyu Lu, Pinghua Gong, Jieping Ye, and Zhenhui Li. 2018. Deep multi-view spatial-temporal network for taxi demand prediction. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32

  65. [73]

    Bing Yu, Haoteng Yin, and Zhanxing Zhu. 2018. Spatio-Temporal Graph Con- volutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of the Twenty-Seventh International Joint Conference on Artificial Intelligence, 2018, July 13-19, 2018, Stockholm, S...

  66. [74]

    Haitao Yu and Zhong-Ren Peng. 2019. Exploring the spatial variation of rides- ourcing demand and its relationship to built environment and socioeconomic factors with the geographically weighted Poisson regression.Journal of Transport Geography 75 (2019), 147–163

  67. [75]

    Le Yu, Leilei Sun, Bowen Du, and Weifeng Lv. 2023. Towards better dynamic graph learning: New architecture and unified library. Advances in Neural Information Processing Systems 36 (2023), 67686–67700

  68. [76]

    Ailing Zeng, Muxi Chen, Lei Zhang, and Qiang Xu. 2023. Are transformers effective for time series forecasting?. In Proceedings of the AAAI conference on artificial intelligence, Vol. 327. 11121–11128

  69. [77]

    Renrui Zhang, Jiaming Han, Aojun Zhou, Xiangfei Hu, Shilin Yan, Pan Lu, Hong- sheng Li, Peng Gao, and Yu Qiao. 2023. Llama-adapter: Efficient fine-tuning of language models with zero-init attention. arXiv preprint arXiv:2303.16199 (2023)

  70. [78]

    Xiyue Zhang, Chao Huang, Yong Xu, Lianghao Xia, Peng Dai, Liefeng Bo, Junbo Zhang, and Yu Zheng. 2021. Traffic Flow Forecasting with Spatial-Temporal Graph Diffusion Network. In Thirty-Fifth AAAI Conference on Artificial Intelli- gence. AAAI Press, 15008–15015

  71. [79]

    Xin Zhang, Yanhua Li, Xun Zhou, Oren Mangoubi, Ziming Zhang, Vincent Filardi, and Jun Luo. 2021. Dac-ml: domain adaptable continuous meta-learning for urban dynamics prediction. In 2021 IEEE International Conference on Data Mining (ICDM). IEEE, 906–915

  72. [80]

    Yingxue Zhang, Yanhua Li, Xun Zhou, Jun Luo, and Zhi-Li Zhang. 2022. Ur- ban traffic dynamics prediction—a continuous spatial-temporal meta-learning approach. ACM Transactions on Intelligent Systems and Technology (TIST) 13, 2 (2022), 1–19

  73. [81]

    Zijian Zhang, Ze Huang, Zhiwei Hu, Xiangyu Zhao, Wanyu Wang, Zitao Liu, Junbo Zhang, S Joe Qin, and Hongwei Zhao. 2023. MLPST: MLP is All You Need for Spatio-Temporal Prediction. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management . ...

  74. [82]

    Lifan Zhao, Shuming Kong, and Yanyan Shen. 2023. DoubleAdapt: A Meta- learning Approach to Incremental Learning for Stock Trend Forecasting. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. 3492–3503

  75. [83]

    Wei Zhao, Shiqi Zhang, Bei Wang, and Bing Zhou. 2023. Spatio-temporal causal graph attention network for traffic flow prediction in intelligent transportation Distribution-aware Online Continual Learning for Urban Spatio-Temporal Forecasting systems. PeerJ Computer Science 9 (...

  76. [84]

    Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020. Gman: A graph multi-attention network for traffic prediction. In Proceedings of the AAAI conference on artificial intelligence , Vol. 34. 1234–1241

  77. [85]

    Fan Zhou, Qing Yang, Kunpeng Zhang, Goce Trajcevski, Ting Zhong, and Ashfaq Khokhar. 2020. Reinforced spatiotemporal attentive graph neural networks for traffic forecasting. IEEE Internet of Things Journal 7, 7 (2020), 6414–6428

  78. [86]

    Tian Zhou, Peisong Niu, Xue Wang, Liang Sun, and Rong Jin. 2023. One Fits All: Power General Time Series Analysis by Pretrained LM. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orlea...

  79. [87]

    Yirong Zhou, Jun Li, Hao Chen, Ye Wu, Jiangjiang Wu, and Luo Chen. 2020. A spatiotemporal attention mechanism-based model for multi-step citywide passenger demand prediction. Information Sciences 513 (2020), 372–385

  80. [88]

    Zhengyang Zhou, Qihe Huang, Kuo Yang, Kun Wang, Xu Wang, Yudong Zhang, Yuxuan Liang, and Yang Wang. 2023. Maintaining the Status Quo: Capturing Invariant Relations for OOD Spatiotemporal Learning. (2023)

  81. [89]

    Martin Zinkevich. 2003. Online convex programming and generalized infinitesi- mal gradient ascent. ICML’03

  82. [2015]

    International Journal of Communications, Network and System Sciences 8, 4 (2015)

    A multiplicative seasonal ARIMA/GARCH model in EVN traffic prediction. International Journal of Communications, Network and System Sciences 8, 4 (2015)

  83. [2021]

    In Proceedings of the 29th International Conference on Advances in Geographic Information Systems

    Libcity: An open library for traffic prediction. In Proceedings of the 29th International Conference on Advances in Geographic Information Systems . 145– 148

  84. [2024]

    Proceedings of the VLDB Endowment 17, 5 (2024), 1081–1090

    BigST: Linear Complexity Spatio-Temporal Graph Neural Network for Traffic Forecasting on Large-Scale Road Networks. Proceedings of the VLDB Endowment 17, 5 (2024), 1081–1090

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.