Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read PatchSTG claims that dynamic spatial attention for traffic forecasting can be made sub-quadratic by partitioning irregularly placed sensors into balanced patches with a leaf KDTree, and reports state-of-the-art accuracy with up to 10x…

desk verdict PatchSTG has a genuinely new KDTree-based patching idea and strong accuracy results on LargeST, but its efficiency analysis omits a per-batch padding cost that could eat into the claimed speedup. read the letter →

arxiv 2412.09972 v2 pith:BE4XAIDE submitted 2024-12-13 cs.LG cs.AI

classification cs.LGcs.AI
keywords trafficforecastingspatio-temporalgraphneuralnetworksTransformerKDTreeirregularspatialpatchingdynamicmodelinglarge-scaledatamanagement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

PatchSTG aims to make dynamic spatial modeling in traffic forecasting practical at city scale. The paper argues that the quadratic cost of point-to-point attention is what blocks large-scale use, and that this cost can be cut by first partitioning irregularly placed road sensors into balanced, non-overlapping patches with a leaf KDTree and then running attention only within and across patches. On four large-scale datasets (up to 8,600 sensors), the resulting Transformer is claimed to match or beat existing dynamic and static spatial models in accuracy while training up to 10 times faster and using up to 4 times less GPU memory.

What carries the argument

The machinery is irregular spatial patching driven by a leaf KDTree. A leaf KDTree partitions irregularly distributed sensor points into leaf nodes of capacity at most $C$, using median hyperplanes so all points live in leaves; padding and backtracking then merge these leaves into occupancy-equal, non-overlapped patches of $P$ points each. This reduces the set of points involved in dynamic attention from $N$ to $M$ (with the paper asserting $M\approx N$), and depth/breadth attention on the patched tensor converts the quadratic point-pair calculation into two smaller attention passes, one local and one global.

What would settle it

Profile the CA training run with the leaf-KDTree construction and cosine-similarity padding included in the per-epoch timer and excluded from it; if the padding step accounts for a significant share of wall-clock time, the reported 10x speedup is not the cost of the full method. A simpler check is to compute $M$ from the reported $R=512$, $C=3$, and $N_p$ for CA and compare it with $N=8{,}600$.

Watch

Extended reading notes

Core claim

The central claim is that full dynamic spatial dependency modeling does not require every sensor to attend to every other sensor; it can be reorganized through spatial data management. PatchSTG builds a leaf KDTree over sensor coordinates, stores all points in leaves, re-indexes them by breadth-first search, pads unfilled leaves with the most temporally similar points from other leaves, and merges leaves of the same subtree into patches of equal size. The encoder then alternates depth attention (local, within a patch) and breadth attention (global, across same-index positions in different patches), giving $O(\max(P,R)Md)$ complexity with $M\approx N$ instead of $O(N^2 d)$. The paper reports state-of-the-art MAE/RMSE/MAPE on SD, GBA, GLA, and CA, with up to $10\times$ training speed and $4\times$ memory advantage over dynamic spatial baselines.

Load-bearing premise

The near-linear efficiency claim rests on the padded point count staying close to the real sensor count and on the similarity-based padding step costing little; on the CA dataset the reported hyperparameters force $M\approx 12{,}288$ against $N=8{,}600$, and that padding cost is never measured in the speed comparison.

Editorial extensions

If this is right

  • If PatchSTG's results hold, dynamic point-to-point spatial modeling becomes usable on networks with thousands of sensors, where quadratic baselines either run out of memory or take days to train.
  • The same patched-attention design can produce interpretable spatial partitions and explicit per-point correlation maps, which linear and low-rank efficient methods cannot provide.
  • Because breadth attention is applied per index across patches rather than fusing patch representations, global correlations retain per-point heterogeneity, supporting the paper's claim of fidelity.
  • The efficiency gain is large enough that forecasting on the 8,600-sensor CA dataset completes in about 14 hours of total training time with a batch size of 32 on one 48 GB GPU, compared with baselines that need larger batches or fail to run.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: The complexity analysis is incomplete because Equation 5 computes a cosine-similarity matrix across all input sensors, an $O(N^2 H)$ operation per batch that is absent from the efficiency comparison; including it could shrink or erase the reported speedup on the largest datasets.
  • Editorial inference: The same irregular-patching recipe should transfer to other irregularly placed sensor arrays, such as air quality monitors or weather stations; a natural test is whether depth attention's local bias still helps when the spatial correlation length is much longer than typical patch size.
  • Editorial inference: The hyperparameter search shows the best patch count grows with dataset size (16, 16, 64, 512 for SD, GBA, GLA, CA), suggesting a data-dependent scaling rule could make the method nearly parameter-free in practice.
  • Editorial inference: A crossover point likely exists: as $N$ grows, the $O(N^2 H)$ padding cost may overtake the savings from patched attention, so the method's advantage may be largest in a mid-scale regime.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes PatchSTG, a Transformer framework for large-scale traffic forecasting that reduces the cost of dynamic spatial attention by partitioning irregularly distributed sensors into balanced, non-overlapping patches using a leaf KDTree. Depth and breadth attention are then applied within and across patches to capture local and global spatial dependencies. On four LargeST datasets (SD, GBA, GLA, CA), PatchSTG consistently outperforms ten baselines on MAE, RMSE, and MAPE, and the paper reports up to 10x training speed and 4x memory reduction over dynamic spatial modeling baselines.

Significance. If the efficiency claims are substantiated, the paper makes a useful contribution: it provides a practical way to apply full dynamic spatial attention at city scale (8,600 sensors) with subquadratic complexity, and the KDTree-based patching offers a spatially interpretable alternative to clustering or low-rank approximations. The empirical evaluation is thorough: four public large-scale datasets, ten baselines, standard metrics, and extensive ablations. The code is publicly available, which is a strength. The paper's primary novelty, however, is the efficiency argument, and that argument currently rests on an incomplete complexity analysis and an unaccounted preprocessing/padding cost.

major comments (3)
  1. [§4.2 and §4.5] The expression for the padded size, M = C×2^{⌊log(N)⌋−log(C)} ≥ N, is not valid as written. With the paper's own hyperparameters, for CA (N=8600, C=3, R=512) the required number of points per patch is P=24, giving M=12,288 ≈ 1.43N; for GBA (N=2352, C=2, R=16), P=256, giving M=4096 ≈ 1.74N; for SD (N=716, C=2, R=16), M=1024 ≈ 1.43N. Only GLA is close to the claimed M≈N. Since §4.5 bases its complexity comparison on M≈N, the efficiency analysis overstates the benefit of patching. The authors should correct the formula, report the actual M/N ratios, and revisit the asymptotic and constant-factor claims.
  2. [§4.2, Eq. (5)] The padding step in Eq. (5) computes CosSim(X, X^T) to select padding points for unfull leaf nodes. If this computation is executed inside the training loop for each batch, it adds O(N^2 H) work per batch; for CA with H=12, this is roughly 8.9×10^8 multiply-adds per batch, which is comparable to the attention cost and would substantially erode the speedups reported in Table 5. The paper does not state whether the padding indices are precomputed once (e.g., on the training set) or recomputed per batch, and §4.5 omits this cost entirely. This is a load-bearing gap: the efficiency advantage is the paper's primary claim, and the authors must specify the placement of Eq. (5) and include its cost in the complexity and runtime analyses.
  3. [§5.4, Table 5] The headline 'up to 10× and 4× improvements in speed and memory' is drawn from total training time on GLA (41h vs 4h) and batch size on CA (32 vs 8), respectively. The per-epoch training time improvements are 5× on GLA (1483s vs 295s) and 3.7× on CA (3641s vs 981s). The paper should state explicitly which quantity is being reported in each claim, and the abstract and conclusions should avoid conflating total training time, per-epoch time, and batch size.
minor comments (6)
  1. [Throughout] There are several typos and infelicitous phrases: 'breath first searching' should be 'breadth first searching'; 'Similarity' should be 'Similarly'; 'intepretable' should be 'interpretable'; 'equaled' should be 'equal'; and 'non-overlap patches' should be 'non-overlapping patches'.
  2. [§4.2] The definition of the number of points per patch uses N_p, but the relationship between P, R, C, and N_p is not written explicitly in the equations; please clarify that P = C × N_p and that N_p must be a power of 2.
  3. [§4.5] The sentence 'which requires less time than quadratic dynamic spatial modeling methods because P≪N, R≪N, and M≈N' is imprecise: even with a correct M, the comparison of asymptotic complexity should be stated as a ratio or with explicit big-O expressions for both depth and breadth attention.
  4. [§5.3] The ablation variant 'w/o FGGC' is listed in the bullet points and in Table 4, but the subsequent discussion does not clearly reference it by name; consider explicitly mapping each variant to its findings.
  5. [§5.4, Table 5] The 'Improvements' row would be clearer if it indicated the baseline used for the ratios (e.g., STWave, the only dynamic baseline that runs on all datasets) and the direction of each ratio (PatchSTG value divided by baseline value).
  6. [§4.2] The paper claims the padded patches are 'non-overlapped' and 'fidelity' (no information loss), but padding duplicates points from other leaf nodes, so the same original point may appear both in its own patch and as padding in another patch. The paper should clarify how this duplication affects the fidelity claim and whether it influences the attention computation.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the central result is an empirically benchmarked architecture; the disputed M≈N and padding-cost points are efficiency-accounting concerns, not derivation-by-construction.

full rationale

PatchSTG's central claims are a new irregular spatial patching algorithm (leaf KDTree plus cosine-similarity padding and backtracking merging), a dual depth/breadth attention encoder, and state-of-the-art results on four LargeST datasets. These are presented as an empirical contribution with held-out train/validation/test splits (Sec. 5.1.1), and the performance numbers in Table 3 are benchmark comparisons, not quantities fitted to the test targets. The complexity statement in Sec. 4.5, 'the dominant complexity of PatchSTG is O(max(P,R)Md)... because P<<N, R<<N, and M≈N,' is an asymptotic claim; whether M≈N holds under the reported hyperparameters (e.g., CA: M=12288 vs N=8600) is a correctness or precision issue in the efficiency analysis, not a circular reduction. Similarly, Eq. 5 pads unfull leaf nodes using CosSim(X, X^T); this uses the observed historical input, not the future target, and while the paper does not state whether this padding is precomputed offline or per batch, the associated cost is an unaccounted overhead rather than a self-referential derivation. Self-citations appear mainly as baselines (STWave, Lastjomer, STID) and standard embedding components, and the benchmark measurements do not depend on those citations being accepted; no uniqueness theorem or fitted-parameter-as-prediction pattern is present. Accordingly, no circular step can be exhibited, and the score is 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on no new physical entities. The free parameters are standard architectural hyperparameters, though several are tuned per dataset. The main ad hoc assumption is that M is approximately N, which the paper's own hyperparameters violate.

free parameters (5)
  • patch count R = 16 (SD), 16 (GBA), 64 (GLA), 512 (CA)
    Selected per dataset from {8,...,512}; controls attention complexity and receptive field.
  • leaf node capacity C = 3 (CA), 2 (GLA), 2 (GBA), 2 (SD)
    Determines leaf KDTree granularity and padding amount.
  • number of encoder layers = 5
    Tuned from {1,3,5,7}; best at 5.
  • input embedding dimension = 128 (SD, GBA), 64 (GLA, CA)
    Tuned from {32,64,128,256}.
  • embedding dims for day-of-week, timeslice, spatial = 32 each
    Fixed without sensitivity analysis.
assumptions (5)
  • standard math KDTree construction and median splitting are correct and balanced for the sensor locations
    Used in Section 4.2 to build the leaf KDTree.
  • domain assumption Geographic proximity (latitude and longitude) is a valid proxy for functional traffic correlation
    Basis for the KDTree patch locality; supported only indirectly by the ablation study.
  • domain assumption Points with the same patch index capture meaningful global spatial dependencies
    Core of breadth attention; no theoretical justification beyond empirical results.
  • ad hoc to paper The padded representation M is approximately equal to N
    Stated in Section 4.5 but contradicted by the reported R and C choices, which yield M up to 1.74N.
  • domain assumption Cosine similarity of raw input time series is a suitable criterion for choosing padding points
    Used in Eq. 5; no sensitivity analysis for this choice.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective." pith.science (2026). https://pith.science/paper/BE4XAIDE

@misc{pith2026241209972,
  author       = {Pith},
  title        = {Pith review of: Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BE4XAIDE}},
  note         = {Machine review of arXiv:2412.09972}
}
abstract

Road traffic forecasting is crucial in real-world intelligent transportation scenarios like traffic dispatching and path planning in city management and personal traveling. Spatio-temporal graph neural networks (STGNNs) stand out as the mainstream solution in this task. Nevertheless, the quadratic complexity of remarkable dynamic spatial modeling-based STGNNs has become the bottleneck over large-scale traffic data. From the spatial data management perspective, we present a novel Transformer framework called PatchSTG to efficiently and dynamically model spatial dependencies for large-scale traffic forecasting with interpretability and fidelity. Specifically, we design a novel irregular spatial patching to reduce the number of points involved in the dynamic calculation of Transformer. The irregular spatial patching first utilizes the leaf K-dimensional tree (KDTree) to recursively partition irregularly distributed traffic points into leaf nodes with a small capacity, and then merges leaf nodes belonging to the same subtree into occupancy-equaled and non-overlapped patches through padding and backtracking. Based on the patched data, depth and breadth attention are used interchangeably in the encoder to dynamically learn local and global spatial knowledge from points in a patch and points with the same index of patches. Experimental results on four real world large-scale traffic datasets show that our PatchSTG achieves train speed and memory utilization improvements up to $10\times$ and $4\times$ with the state-of-the-art performance.

Figures

Figures reproduced from arXiv: 2412.09972 by the authors.

Figure 1
Figure 1. Sketch of our PatchSTG. We first partition the [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The workflow of our PatchSTG. the preceding 𝐻 time slices and the location of points. 𝑌ˆ = 𝑓𝜃 (𝑋, 𝐿𝑎𝑡, 𝐿𝑛𝑔) (1) where 𝑌ˆ ∈ R 𝐹 ×𝑁 is the predicted traffic in the future, which will be used to compare with the ground truth 𝑌 ∈ R 𝐹 ×𝑁 . 𝐿𝑎𝑡 ∈ R 𝑁 and 𝐿𝑛𝑔 ∈ R 𝑁 denote the real-world latitude and longitude of points. Moreover, the function 𝑓𝜃 (·) indicates a data-driven forecasting model parameterized by 𝜃. 4 METHODOLOG… view at source ↗
Figure 3
Figure 3. (a)-(b): left part draws the spatial partition using the original and leaf KDTree, and the right part is the corresponding [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Structure of depth and breadth attention. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Leaf KDTree on the GLA dataset. (a) Correlations on 23𝑡ℎ (b) Correlations on 56𝑡ℎ [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Learned patch-level correlations. points with indices 23 and 56 in the patch. As shown in [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HieraMix: A Hierarchical MLP-Mixer for Large-Scale Traffic Forecasting

    cs.LG 2025-11 conditional novelty 6.0 of 10

    HSTMixer, a hierarchical all-MLP mixer with adaptive region mixing, reports state-of-the-art MAE/RMSE/MAPE on four LargeST datasets at 15-minute resolution.

Reference graph

Works this paper leans on

69 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting. In Proceedings of NeurIPS. 17804–17815

  2. [2]

    Srinivasa Ravi Chandra and Haitham Al-Deek. 2009. Predictions of freeway traffic speeds and volumes using vector autoregressive models. Journal of Intelligent Transportation Systems (2009), 53–72

  3. [3]

    Jeongwhan Choi, Hwangyong Choi, Jeehyun Hwang, and Noseong Park. 2022. Graph neural controlled differential equations for traffic forecasting. In Proceed- ings of AAAI. 6367–6374

  4. [4]

    Yue Cui, Kai Zheng, Dingshan Cui, Jiandong Xie, Liwei Deng, Feiteng Huang, and Xiaofang Zhou. 2021. METRO: a generic graph neural network framework for multivariate time series forecasting. In Proceedings of VLDB. 224–236

  5. [5]

    Mark De Berg, Otfried Cheong, Marc Van Kreveld, and Mark Overmars. 2008. Orthogonal range searching: Querying a database. Computational Geometry: Algorithms and Applications (2008), 95–120

  6. [6]

    Liwei Deng, Tianfu Wang, Yan Zhao, and Kai Zheng. 2024. MILLION: A General Multi-Objective Framework with Controllable Risk for Portfolio Management. arXiv preprint arXiv:2412.03038 (2024)

  7. [7]

    Liwei Deng, Yan Zhao, Yue Cui, Yuyang Xia, Jin Chen, and Kai Zheng. 2024. Task Recommendation in Spatial Crowdsourcing: A Trade-Off Between Diversity and Coverage. In Proceedings of ICDE. 276–288

  8. [8]

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of ICLR. 1–22

Show all 69 references
  1. [9]

    Wenying Duan, Xiaoxi He, Zimu Zhou, Lothar Thiele, and Hong Rao. 2023. Localised adaptive spatial-temporal graph neural network. In Proceedings of SIGKDD. 448–458

  2. [10]

    Yuchen Fang, Haiyong Luo, Fang Zhao, Poly ZH Sun, Yanjun Qin, Liang Zeng, Bo Hui, and Chenxing Wang. 2024. CDGNet: A Cross-Time Dynamic Graph- Based Deep Learning Model for Vehicle-Based Traffic Speed Forecasting. IEEE Transactions on Intelligent Vehicles (2024)

  3. [11]

    Yuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao, Bingbing Xu, Liang Zeng, and Chenxing Wang. 2023. When spatio-temporal meet wavelets: Disentangled traffic forecasting via efficient spectral graph attention networks. In Proceedings of ICDE. 517–529

  4. [12]

    Yuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao, and Kai Zheng. 2023. STWave+: A Multi-Scale Efficient Spectral Graph Attention Network With Long- Term Trends for Disentangled Traffic Flow Forecasting. IEEE Transactions on Knowledge and Data Engineering (2023), 2671–2685

  5. [13]

    Yuchen Fang, Fang Zhao, Yanjun Qin, Haiyong Luo, and Chenxing Wang. 2022. Learning all dynamics: Traffic forecasting via locality-aware spatio-temporal joint transformer. IEEE Transactions on Intelligent Transportation Systems (2022), 23433–23446

  6. [14]

    Zheng Fang, Qingqing Long, Guojie Song, and Kunqing Xie. 2021. Spatial- temporal graph ode networks for traffic flow forecasting. In Proceedings of SIGKDD. 364–373

  7. [15]

    Han Gao, Xu Han, Jiaoyang Huang, Jian-Xun Wang, and Liping Liu. 2022. Patchgt: Transformer over non-trainable clusters for learning graph representations. In Proceedings of LoG. 1–27

  8. [16]

    Ge Guo, Wei Yuan, Jinyuan Liu, Yisheng Lv, and Wei Liu. 2021. Traffic forecast- ing via dilated temporal convolution with peak-sensitive loss. IEEE Intelligent Transportation Systems Magazine (2021), 48–57

  9. [17]

    Shengnan Guo, Youfang Lin, Ning Feng, Chao Song, and Huaiyu Wan. 2019. Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. In Proceedings of AAAI. 922–929

  10. [18]

    Shengnan Guo, Youfang Lin, Letian Gong, Chenyu Wang, Zeyu Zhou, Zekai Shen, Yiheng Huang, and Huaiyu Wan. 2023. Self-supervised spatial-temporal bottle- neck attentive network for efficient long-term traffic forecasting. In Proceedings of ICDE. 1585–1596

  11. [19]

    Antonin Guttman. 1984. R-trees: A dynamic index structure for spatial searching. In Proceedings of SIGMOD. 47–57

  12. [20]

    Jindong Han, Weijia Zhang, Hao Liu, Tao Tao, Naiqiang Tan, and Hui Xiong. 2024. BigST: Linear Complexity Spatio-Temporal Graph Neural Network for Traffic Forecasting on Large-Scale Road Networks. In Proceedings of VLDB. 1081–1090

  13. [21]

    Liangzhe Han, Bowen Du, Leilei Sun, Yanjie Fu, Yisheng Lv, and Hui Xiong

  14. [22]

    Xiaoxin He, Bryan Hooi, Thomas Laurent, Adam Perold, Yann LeCun, and Xavier Bresson. 2023. A generalization of vit/mlp-mixer to graphs. In Proceedings of ICML. 12724–12745

  15. [23]

    Bo Hui, Da Yan, Haiquan Chen, and Wei-Shinn Ku. 2021. Trajectory waveNet: A trajectory-based model for traffic forecasting. InProceedings of ICDM. 1114–1119

  16. [24]

    Jiawei Jiang, Chengkai Han, Wayne Xin Zhao, and Jingyuan Wang. 2023. Pdformer: Propagation delay-aware dynamic long-range transformer for traffic flow prediction. In Proceedings of AAAI. 4365–4373

  17. [25]

    Xinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang, Yichen Zhu, Ruizhe Zhang, Yuchen Fang, Xu Chu, Junfeng Zhao, and Yasha Wang. 2024. RAGraph: A General Retrieval-Augmented Graph Learning Framework. arXiv preprint arXiv:2410.23855 (2024)

  18. [26]

    Xinke Jiang, Dingyi Zhuang, Xianghui Zhang, Hao Chen, Jiayuan Luo, and Xiaowei Gao. 2023. Uncertainty quantification via spatial-temporal tweedie model for zero-inflated and long-tail travel demand prediction. In Proceedings of CIKM. 3983–3987

  19. [27]

    Guangyin Jin, Yuxuan Liang, Yuchen Fang, Zezhi Shao, Jincai Huang, Junbo Zhang, and Yu Zheng. 2023. Spatio-temporal graph neural networks for predictive learning in urban computing: A survey. IEEE Transactions on Knowledge and Data Engineering (2023)

  20. [28]

    S Vasantha Kumar and Lelitha Vanajakshi. 2015. Short-term traffic flow predic- tion using seasonal ARIMA model with limited input data. European Transport Research Review (2015), 1–9

  21. [29]

    Siqi Lai, Zhao Xu, Weijia Zhang, Hao Liu, and Hui Xiong. 2023. Large language models as traffic signal control agents: Capacity and opportunity. arXiv preprint arXiv:2312.16044 (2023)

  22. [30]

    Zhichen Lai, Dalin Zhang, Huan Li, Christian S Jensen, Hua Lu, and Yan Zhao

  23. [31]

    Shiyong Lan, Yitong Ma, Weikang Huang, Wenwu Wang, Hongyu Yang, and Pyang Li. 2022. Dstagnn: Dynamic spatial-temporal aware graph neural network for traffic flow forecasting. In Proceedings of ICML. 11906–11917

  24. [32]

    Fuxian Li, Jie Feng, Huan Yan, Guangyin Jin, Fan Yang, Funing Sun, Depeng Jin, and Yong Li. 2023. Dynamic graph convolutional recurrent network for traffic prediction: Benchmark and solution. ACM Transactions on Knowledge Discovery from Data (2023), 1–21

  25. [33]

    Rongfan Li, Ting Zhong, Xinke Jiang, Goce Trajcevski, Jin Wu, and Fan Zhou

  26. [34]

    Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In Proceedings of ICLR

  27. [35]

    Zhonghang Li, Lianghao Xia, Jiabin Tang, Yong Xu, Lei Shi, Long Xia, Dawei Yin, and Chao Huang. 2024. Urbangpt: Spatio-temporal large language models. In Proceedings of SIGKDD. 1–11

  28. [36]

    Yuxuan Liang, Kun Ouyang, Junkai Sun, Yiwei Wang, Junbo Zhang, Yu Zheng, David Rosenblum, and Roger Zimmermann. 2021. Fine-grained urban flow prediction. In Proceedings of WWW . 1833–1845

  29. [37]

    Yuxuan Liang, Yutong Xia, Songyu Ke, Yiwei Wang, Qingsong Wen, Junbo Zhang, Yu Zheng, and Roger Zimmermann. 2023. Airformer: Predicting nationwide air quality in china with transformers. In Proceedings of AAAI. 14329–14337

  30. [38]

    Dachuan Liu, Jin Wang, Shuo Shang, and Peng Han. 2022. Msdr: Multi-step dependency relation networks for spatial temporal forecasting. In Proceedings of SIGKDD. 1042–1050

  31. [39]

    Hangchen Liu, Zheng Dong, Renhe Jiang, Jiewen Deng, Jinliang Deng, Quanjun Chen, and Xuan Song. 2023. Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of CIKM. 4125–4129

  32. [40]

    Shuncheng Liu, Xu Chen, Ziniu Wu, Liwei Deng, Han Su, and Kai Zheng. 2022. HeGA: heterogeneous graph aggregation network for trajectory prediction in high-density traffic. In Proceedings of CIKM. 1319–1328

  33. [41]

    Shuncheng Liu, Yuyang Xia, Xu Chen, Jiandong Xie, Han Su, and Kai Zheng. 2023. Impact-aware maneuver decision with enhanced perception for autonomous vehicle. In Proceedings of ICDE. 3255–3268

  34. [42]

    Xu Liu, Yuxuan Liang, Chao Huang, Hengchang Hu, Yushi Cao, Bryan Hooi, and Roger Zimmermann. 2024. Reinventing Node-centric Traffic Forecasting for Improved Accuracy and Efficiency. In Proceedings of ECML-PKDD. 21–38

  35. [43]

    Xu Liu, Yutong Xia, Yuxuan Liang, Junfeng Hu, Yiwei Wang, Lei Bai, Chao Huang, Zhenguang Liu, Bryan Hooi, and Roger Zimmermann. 2024. Largest: A benchmark dataset for large-scale traffic forecasting. In Proceedings of NeurIPS

  36. [44]

    Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of ICCV. 10012–10022

  37. [45]

    Jiayuan Luo, Wentao Zhang, Yuchen Fang, Xiaowei Gao, Dingyi Zhuang, Hao Chen, and Xinke Jiang. 2024. Timeseries suppliers allocation risk optimization via deep black litterman model. arXiv preprint arXiv:2401.17350 (2024)

  38. [46]

    Zhongjian Lv, Jiajie Xu, Kai Zheng, Hongzhi Yin, Pengpeng Zhao, and Xiaofang Zhou. 2018. Lc-rnn: A deep learning model for traffic speed prediction.. In Proceedings of IJCAI. 27

  39. [47]

    Qian Ma, Zijian Zhang, Xiangyu Zhao, Haoliang Li, Hongwei Zhao, Yiqi Wang, Zitao Liu, and Wanyu Wang. 2023. Rethinking sensors modeling: Hierarchical information enhanced traffic forecasting. In Proceedings of CIKM. 1756–1765

  40. [48]

    Chunghyun Park, Yoonwoo Jeong, Minsu Cho, and Jaesik Park. 2022. Fast point transformer. In Proceedings of CVPR. 16949–16958

  41. [49]

    Cheonbok Park, Chunggi Lee, Hyojin Bahng, Yunwon Tae, Seungmin Jin, Kihwan Kim, Sungahn Ko, and Jaegul Choo. 2020. ST-GRAT: A novel spatio-temporal graph attention networks for accurately forecasting dynamically changing road speed. In Proceedings of CIKM. 1215–1224. Efficient...

  42. [50]

    Zezhi Shao, Zhao Zhang, Fei Wang, Wei Wei, and Yongjun Xu. 2022. Spatial- temporal identity: A simple yet effective baseline for multivariate time series forecasting. In Proceedings of CIKM. 4454–4458

  43. [51]

    Zezhi Shao, Zhao Zhang, Wei Wei, Fei Wang, Yongjun Xu, Xin Cao, and Chris- tian S Jensen. 2022. Decoupled dynamic spatial-temporal graph neural network for traffic forecasting. In Proceedings of VLDB. 2733–2746

  44. [52]

    Robert F Sproull. 1991. Refinements to nearest-neighbor searching in k- dimensional trees. Algorithmica (1991), 579–589

  45. [53]

    Qian Sun, Rui Zha, Le Zhang, Jingbo Zhou, Yu Mei, Zhiling Li, and Hui Xiong

  46. [54]

    Peng-Shuai Wang. 2023. Octformer: Octree-based transformers for 3d point clouds. ACM Transactions on Graphics (TOG) (2023), 1–11

  47. [55]

    Tianfu Wang, Liwei Deng, Chao Wang, Jianxun Lian, Yue Yan, Nicholas Jing Yuan, Qi Zhang, and Hui Xiong. 2024. COMET: NFT Price Prediction with Wallet Profiling. In Proceedings of SIGKDD. 5893–5904

  48. [56]

    Xu Wang, Lianliang Chen, Hongbo Zhang, Pengkun Wang, Zhengyang Zhou, and Yang Wang. 2023. A Multi-graph Fusion Based Spatiotemporal Dynamic Learning Framework. In Proceedings of WSDM. 294–302

  49. [57]

    Haomin Wen, Youfang Lin, Yutong Xia, Huaiyu Wan, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang. 2023. Diffstg: Probabilistic spatio-temporal graph forecasting with denoising diffusion models. In Proceedings of SIGSPATIAL. 1–12

  50. [58]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of SIGKDD. 753–763

  51. [59]

    Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph wavenet for deep spatial-temporal graph modeling. InProceedings of IJCAI. 1907–1913

  52. [60]

    Chin-Chia Michael Yeh, Yujie Fan, Xin Dai, Uday Singh Saini, Vivian Lai, Prince Osei Aboagye, Junpeng Wang, Huiyuan Chen, Yan Zheng, Zhongfang Zhuang, et al. 2024. RPMixer: Shaking Up Time Series Forecasting with Random Projections for Large Spatial-Temporal Data. In Proceedin...

  53. [61]

    Bing Yu, Haoteng Yin, and Zhanxing Zhu. 2018. Spatio-Temporal Graph Con- volutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of IJCAI. 3634–3640

  54. [62]

    Xingtong Yu, Yuan Fang, Zemin Liu, Yuxia Wu, Zhihao Wen, Jianyuan Bo, Xin- ming Zhang, and Steven CH Hoi. 2024. Few-shot learning on graphs: from meta-learning to pre-training and prompting. arXiv preprint arXiv:2402.01440 (2024)

  55. [63]

    Xingtong Yu, Zemin Liu, Yuan Fang, and Xinming Zhang. 2023. Learning to count isomorphisms with graph neural networks. In Proceedings of AAAI. 4845–4853

  56. [64]

    Xingtong Yu, Zhenghao Liu, Yuan Fang, and Xinming Zhang. 2024. DyG- Prompt: Learning Feature and Time Prompts on Dynamic Graphs. arXiv preprint arXiv:2405.13937 (2024)

  57. [65]

    Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020. Gman: A graph multi-attention network for traffic prediction. In Proceedings of AAAI. 1234–1241

  58. [2021]

    In Proceedings of SIGKDD

    Dynamic and multi-faceted spatio-temporal deep learning for traffic speed forecasting. In Proceedings of SIGKDD. 547–555

  59. [2022]

    In Proceedings of SIGKDD

    Mining spatio-temporal relations via self-paced graph contrastive learning. In Proceedings of SIGKDD. 936–944

  60. [2023]

    In Proceedings of SIGMOD

    Lightcts: A lightweight framework for correlated time series forecasting. In Proceedings of SIGMOD. 1–26

  61. [2024]

    In Proceedings of SIGKDD

    CrossLight: Offline-to-Online Reinforcement Learning for Cross-City Traffic Signal Control. In Proceedings of SIGKDD. 2765–2774

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.