REVIEW 3 major objections 6 minor 1 cited by
Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read PatchSTG claims that dynamic spatial attention for traffic forecasting can be made sub-quadratic by partitioning irregularly placed sensors into balanced patches with a leaf KDTree, and reports state-of-the-art accuracy with up to 10x…
desk verdict PatchSTG has a genuinely new KDTree-based patching idea and strong accuracy results on LargeST, but its efficiency analysis omits a per-batch padding cost that could eat into the claimed speedup. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is irregular spatial patching driven by a leaf KDTree. A leaf KDTree partitions irregularly distributed sensor points into leaf nodes of capacity at most $C$, using median hyperplanes so all points live in leaves; padding and backtracking then merge these leaves into occupancy-equal, non-overlapped patches of $P$ points each. This reduces the set of points involved in dynamic attention from $N$ to $M$ (with the paper asserting $M\approx N$), and depth/breadth attention on the patched tensor converts the quadratic point-pair calculation into two smaller attention passes, one local and one global.
What would settle it
Profile the CA training run with the leaf-KDTree construction and cosine-similarity padding included in the per-epoch timer and excluded from it; if the padding step accounts for a significant share of wall-clock time, the reported 10x speedup is not the cost of the full method. A simpler check is to compute $M$ from the reported $R=512$, $C=3$, and $N_p$ for CA and compare it with $N=8{,}600$.
Extended reading notes
Core claim
The central claim is that full dynamic spatial dependency modeling does not require every sensor to attend to every other sensor; it can be reorganized through spatial data management. PatchSTG builds a leaf KDTree over sensor coordinates, stores all points in leaves, re-indexes them by breadth-first search, pads unfilled leaves with the most temporally similar points from other leaves, and merges leaves of the same subtree into patches of equal size. The encoder then alternates depth attention (local, within a patch) and breadth attention (global, across same-index positions in different patches), giving $O(\max(P,R)Md)$ complexity with $M\approx N$ instead of $O(N^2 d)$. The paper reports state-of-the-art MAE/RMSE/MAPE on SD, GBA, GLA, and CA, with up to $10\times$ training speed and $4\times$ memory advantage over dynamic spatial baselines.
Load-bearing premise
The near-linear efficiency claim rests on the padded point count staying close to the real sensor count and on the similarity-based padding step costing little; on the CA dataset the reported hyperparameters force $M\approx 12{,}288$ against $N=8{,}600$, and that padding cost is never measured in the speed comparison.
Editorial extensions
If this is right
- If PatchSTG's results hold, dynamic point-to-point spatial modeling becomes usable on networks with thousands of sensors, where quadratic baselines either run out of memory or take days to train.
- The same patched-attention design can produce interpretable spatial partitions and explicit per-point correlation maps, which linear and low-rank efficient methods cannot provide.
- Because breadth attention is applied per index across patches rather than fusing patch representations, global correlations retain per-point heterogeneity, supporting the paper's claim of fidelity.
- The efficiency gain is large enough that forecasting on the 8,600-sensor CA dataset completes in about 14 hours of total training time with a batch size of 32 on one 48 GB GPU, compared with baselines that need larger batches or fail to run.
Reading between the lines
- Editorial inference: The complexity analysis is incomplete because Equation 5 computes a cosine-similarity matrix across all input sensors, an $O(N^2 H)$ operation per batch that is absent from the efficiency comparison; including it could shrink or erase the reported speedup on the largest datasets.
- Editorial inference: The same irregular-patching recipe should transfer to other irregularly placed sensor arrays, such as air quality monitors or weather stations; a natural test is whether depth attention's local bias still helps when the spatial correlation length is much longer than typical patch size.
- Editorial inference: The hyperparameter search shows the best patch count grows with dataset size (16, 16, 64, 512 for SD, GBA, GLA, CA), suggesting a data-dependent scaling rule could make the method nearly parameter-free in practice.
- Editorial inference: A crossover point likely exists: as $N$ grows, the $O(N^2 H)$ padding cost may overtake the savings from patched attention, so the method's advantage may be largest in a mid-scale regime.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PatchSTG, a Transformer framework for large-scale traffic forecasting that reduces the cost of dynamic spatial attention by partitioning irregularly distributed sensors into balanced, non-overlapping patches using a leaf KDTree. Depth and breadth attention are then applied within and across patches to capture local and global spatial dependencies. On four LargeST datasets (SD, GBA, GLA, CA), PatchSTG consistently outperforms ten baselines on MAE, RMSE, and MAPE, and the paper reports up to 10x training speed and 4x memory reduction over dynamic spatial modeling baselines.
Significance. If the efficiency claims are substantiated, the paper makes a useful contribution: it provides a practical way to apply full dynamic spatial attention at city scale (8,600 sensors) with subquadratic complexity, and the KDTree-based patching offers a spatially interpretable alternative to clustering or low-rank approximations. The empirical evaluation is thorough: four public large-scale datasets, ten baselines, standard metrics, and extensive ablations. The code is publicly available, which is a strength. The paper's primary novelty, however, is the efficiency argument, and that argument currently rests on an incomplete complexity analysis and an unaccounted preprocessing/padding cost.
major comments (3)
- [§4.2 and §4.5] The expression for the padded size, M = C×2^{⌊log(N)⌋−log(C)} ≥ N, is not valid as written. With the paper's own hyperparameters, for CA (N=8600, C=3, R=512) the required number of points per patch is P=24, giving M=12,288 ≈ 1.43N; for GBA (N=2352, C=2, R=16), P=256, giving M=4096 ≈ 1.74N; for SD (N=716, C=2, R=16), M=1024 ≈ 1.43N. Only GLA is close to the claimed M≈N. Since §4.5 bases its complexity comparison on M≈N, the efficiency analysis overstates the benefit of patching. The authors should correct the formula, report the actual M/N ratios, and revisit the asymptotic and constant-factor claims.
- [§4.2, Eq. (5)] The padding step in Eq. (5) computes CosSim(X, X^T) to select padding points for unfull leaf nodes. If this computation is executed inside the training loop for each batch, it adds O(N^2 H) work per batch; for CA with H=12, this is roughly 8.9×10^8 multiply-adds per batch, which is comparable to the attention cost and would substantially erode the speedups reported in Table 5. The paper does not state whether the padding indices are precomputed once (e.g., on the training set) or recomputed per batch, and §4.5 omits this cost entirely. This is a load-bearing gap: the efficiency advantage is the paper's primary claim, and the authors must specify the placement of Eq. (5) and include its cost in the complexity and runtime analyses.
- [§5.4, Table 5] The headline 'up to 10× and 4× improvements in speed and memory' is drawn from total training time on GLA (41h vs 4h) and batch size on CA (32 vs 8), respectively. The per-epoch training time improvements are 5× on GLA (1483s vs 295s) and 3.7× on CA (3641s vs 981s). The paper should state explicitly which quantity is being reported in each claim, and the abstract and conclusions should avoid conflating total training time, per-epoch time, and batch size.
minor comments (6)
- [Throughout] There are several typos and infelicitous phrases: 'breath first searching' should be 'breadth first searching'; 'Similarity' should be 'Similarly'; 'intepretable' should be 'interpretable'; 'equaled' should be 'equal'; and 'non-overlap patches' should be 'non-overlapping patches'.
- [§4.2] The definition of the number of points per patch uses N_p, but the relationship between P, R, C, and N_p is not written explicitly in the equations; please clarify that P = C × N_p and that N_p must be a power of 2.
- [§4.5] The sentence 'which requires less time than quadratic dynamic spatial modeling methods because P≪N, R≪N, and M≈N' is imprecise: even with a correct M, the comparison of asymptotic complexity should be stated as a ratio or with explicit big-O expressions for both depth and breadth attention.
- [§5.3] The ablation variant 'w/o FGGC' is listed in the bullet points and in Table 4, but the subsequent discussion does not clearly reference it by name; consider explicitly mapping each variant to its findings.
- [§5.4, Table 5] The 'Improvements' row would be clearer if it indicated the baseline used for the ratios (e.g., STWave, the only dynamic baseline that runs on all datasets) and the direction of each ratio (PatchSTG value divided by baseline value).
- [§4.2] The paper claims the padded patches are 'non-overlapped' and 'fidelity' (no information loss), but padding duplicates points from other leaf nodes, so the same original point may appear both in its own patch and as padding in another patch. The paper should clarify how this duplication affects the fidelity claim and whether it influences the attention computation.
Circularity Check
No circularity found: the central result is an empirically benchmarked architecture; the disputed M≈N and padding-cost points are efficiency-accounting concerns, not derivation-by-construction.
full rationale
PatchSTG's central claims are a new irregular spatial patching algorithm (leaf KDTree plus cosine-similarity padding and backtracking merging), a dual depth/breadth attention encoder, and state-of-the-art results on four LargeST datasets. These are presented as an empirical contribution with held-out train/validation/test splits (Sec. 5.1.1), and the performance numbers in Table 3 are benchmark comparisons, not quantities fitted to the test targets. The complexity statement in Sec. 4.5, 'the dominant complexity of PatchSTG is O(max(P,R)Md)... because P<<N, R<<N, and M≈N,' is an asymptotic claim; whether M≈N holds under the reported hyperparameters (e.g., CA: M=12288 vs N=8600) is a correctness or precision issue in the efficiency analysis, not a circular reduction. Similarly, Eq. 5 pads unfull leaf nodes using CosSim(X, X^T); this uses the observed historical input, not the future target, and while the paper does not state whether this padding is precomputed offline or per batch, the associated cost is an unaccounted overhead rather than a self-referential derivation. Self-citations appear mainly as baselines (STWave, Lastjomer, STID) and standard embedding components, and the benchmark measurements do not depend on those citations being accepted; no uniqueness theorem or fitted-parameter-as-prediction pattern is present. Accordingly, no circular step can be exhibited, and the score is 0.
Assumptions & free parameters
free parameters (5)
- patch count R =
16 (SD), 16 (GBA), 64 (GLA), 512 (CA)
- leaf node capacity C =
3 (CA), 2 (GLA), 2 (GBA), 2 (SD)
- number of encoder layers =
5
- input embedding dimension =
128 (SD, GBA), 64 (GLA, CA)
- embedding dims for day-of-week, timeslice, spatial =
32 each
assumptions (5)
- standard math KDTree construction and median splitting are correct and balanced for the sensor locations
- domain assumption Geographic proximity (latitude and longitude) is a valid proxy for functional traffic correlation
- domain assumption Points with the same patch index capture meaningful global spatial dependencies
- ad hoc to paper The padded representation M is approximately equal to N
- domain assumption Cosine similarity of raw input time series is a suitable criterion for choosing padding points
Cite this review
Pith. "Pith review of Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective." pith.science (2026). https://pith.science/paper/BE4XAIDE
@misc{pith2026241209972,
author = {Pith},
title = {Pith review of: Efficient Large-Scale Traffic Forecasting with Transformers: A Spatial Data Management Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/BE4XAIDE}},
note = {Machine review of arXiv:2412.09972}
}
abstract
Road traffic forecasting is crucial in real-world intelligent transportation scenarios like traffic dispatching and path planning in city management and personal traveling. Spatio-temporal graph neural networks (STGNNs) stand out as the mainstream solution in this task. Nevertheless, the quadratic complexity of remarkable dynamic spatial modeling-based STGNNs has become the bottleneck over large-scale traffic data. From the spatial data management perspective, we present a novel Transformer framework called PatchSTG to efficiently and dynamically model spatial dependencies for large-scale traffic forecasting with interpretability and fidelity. Specifically, we design a novel irregular spatial patching to reduce the number of points involved in the dynamic calculation of Transformer. The irregular spatial patching first utilizes the leaf K-dimensional tree (KDTree) to recursively partition irregularly distributed traffic points into leaf nodes with a small capacity, and then merges leaf nodes belonging to the same subtree into occupancy-equaled and non-overlapped patches through padding and backtracking. Based on the patched data, depth and breadth attention are used interchangeably in the encoder to dynamically learn local and global spatial knowledge from points in a patch and points with the same index of patches. Experimental results on four real world large-scale traffic datasets show that our PatchSTG achieves train speed and memory utilization improvements up to $10\times$ and $4\times$ with the state-of-the-art performance.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
HieraMix: A Hierarchical MLP-Mixer for Large-Scale Traffic Forecasting
HSTMixer, a hierarchical all-MLP mixer with adaptive region mixing, reports state-of-the-art MAE/RMSE/MAPE on four LargeST datasets at 15-minute resolution.
Reference graph
Works this paper leans on
-
[1]
Lei Bai, Lina Yao, Can Li, Xianzhi Wang, and Can Wang. 2020. Adaptive graph convolutional recurrent network for traffic forecasting. In Proceedings of NeurIPS. 17804–17815
work page 2020
-
[2]
Srinivasa Ravi Chandra and Haitham Al-Deek. 2009. Predictions of freeway traffic speeds and volumes using vector autoregressive models. Journal of Intelligent Transportation Systems (2009), 53–72
work page 2009
-
[3]
Jeongwhan Choi, Hwangyong Choi, Jeehyun Hwang, and Noseong Park. 2022. Graph neural controlled differential equations for traffic forecasting. In Proceed- ings of AAAI. 6367–6374
work page 2022
-
[4]
Yue Cui, Kai Zheng, Dingshan Cui, Jiandong Xie, Liwei Deng, Feiteng Huang, and Xiaofang Zhou. 2021. METRO: a generic graph neural network framework for multivariate time series forecasting. In Proceedings of VLDB. 224–236
work page 2021
-
[5]
Mark De Berg, Otfried Cheong, Marc Van Kreveld, and Mark Overmars. 2008. Orthogonal range searching: Querying a database. Computational Geometry: Algorithms and Applications (2008), 95–120
work page 2008
-
[6]
Liwei Deng, Tianfu Wang, Yan Zhao, and Kai Zheng. 2024. MILLION: A General Multi-Objective Framework with Controllable Risk for Portfolio Management. arXiv preprint arXiv:2412.03038 (2024)
arXiv 2024
-
[7]
Liwei Deng, Yan Zhao, Yue Cui, Yuyang Xia, Jin Chen, and Kai Zheng. 2024. Task Recommendation in Spatial Crowdsourcing: A Trade-Off Between Diversity and Coverage. In Proceedings of ICDE. 276–288
work page 2024
-
[8]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Proceedings of ICLR. 1–22
work page 2020
Show all 69 references
-
[9]
Wenying Duan, Xiaoxi He, Zimu Zhou, Lothar Thiele, and Hong Rao. 2023. Localised adaptive spatial-temporal graph neural network. In Proceedings of SIGKDD. 448–458
2023
-
[10]
Yuchen Fang, Haiyong Luo, Fang Zhao, Poly ZH Sun, Yanjun Qin, Liang Zeng, Bo Hui, and Chenxing Wang. 2024. CDGNet: A Cross-Time Dynamic Graph- Based Deep Learning Model for Vehicle-Based Traffic Speed Forecasting. IEEE Transactions on Intelligent Vehicles (2024)
2024
-
[11]
Yuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao, Bingbing Xu, Liang Zeng, and Chenxing Wang. 2023. When spatio-temporal meet wavelets: Disentangled traffic forecasting via efficient spectral graph attention networks. In Proceedings of ICDE. 517–529
2023
-
[12]
Yuchen Fang, Yanjun Qin, Haiyong Luo, Fang Zhao, and Kai Zheng. 2023. STWave+: A Multi-Scale Efficient Spectral Graph Attention Network With Long- Term Trends for Disentangled Traffic Flow Forecasting. IEEE Transactions on Knowledge and Data Engineering (2023), 2671–2685
2023
-
[13]
Yuchen Fang, Fang Zhao, Yanjun Qin, Haiyong Luo, and Chenxing Wang. 2022. Learning all dynamics: Traffic forecasting via locality-aware spatio-temporal joint transformer. IEEE Transactions on Intelligent Transportation Systems (2022), 23433–23446
2022
-
[14]
Zheng Fang, Qingqing Long, Guojie Song, and Kunqing Xie. 2021. Spatial- temporal graph ode networks for traffic flow forecasting. In Proceedings of SIGKDD. 364–373
2021
-
[15]
Han Gao, Xu Han, Jiaoyang Huang, Jian-Xun Wang, and Liping Liu. 2022. Patchgt: Transformer over non-trainable clusters for learning graph representations. In Proceedings of LoG. 1–27
2022
-
[16]
Ge Guo, Wei Yuan, Jinyuan Liu, Yisheng Lv, and Wei Liu. 2021. Traffic forecast- ing via dilated temporal convolution with peak-sensitive loss. IEEE Intelligent Transportation Systems Magazine (2021), 48–57
2021
-
[17]
Shengnan Guo, Youfang Lin, Ning Feng, Chao Song, and Huaiyu Wan. 2019. Attention based spatial-temporal graph convolutional networks for traffic flow forecasting. In Proceedings of AAAI. 922–929
2019
-
[18]
Shengnan Guo, Youfang Lin, Letian Gong, Chenyu Wang, Zeyu Zhou, Zekai Shen, Yiheng Huang, and Huaiyu Wan. 2023. Self-supervised spatial-temporal bottle- neck attentive network for efficient long-term traffic forecasting. In Proceedings of ICDE. 1585–1596
2023
-
[19]
Antonin Guttman. 1984. R-trees: A dynamic index structure for spatial searching. In Proceedings of SIGMOD. 47–57
1984
-
[20]
Jindong Han, Weijia Zhang, Hao Liu, Tao Tao, Naiqiang Tan, and Hui Xiong. 2024. BigST: Linear Complexity Spatio-Temporal Graph Neural Network for Traffic Forecasting on Large-Scale Road Networks. In Proceedings of VLDB. 1081–1090
2024
-
[21]
Liangzhe Han, Bowen Du, Leilei Sun, Yanjie Fu, Yisheng Lv, and Hui Xiong
-
[22]
Xiaoxin He, Bryan Hooi, Thomas Laurent, Adam Perold, Yann LeCun, and Xavier Bresson. 2023. A generalization of vit/mlp-mixer to graphs. In Proceedings of ICML. 12724–12745
2023
-
[23]
Bo Hui, Da Yan, Haiquan Chen, and Wei-Shinn Ku. 2021. Trajectory waveNet: A trajectory-based model for traffic forecasting. InProceedings of ICDM. 1114–1119
2021
-
[24]
Jiawei Jiang, Chengkai Han, Wayne Xin Zhao, and Jingyuan Wang. 2023. Pdformer: Propagation delay-aware dynamic long-range transformer for traffic flow prediction. In Proceedings of AAAI. 4365–4373
2023
-
[25]
Xinke Jiang, Rihong Qiu, Yongxin Xu, Wentao Zhang, Yichen Zhu, Ruizhe Zhang, Yuchen Fang, Xu Chu, Junfeng Zhao, and Yasha Wang. 2024. RAGraph: A General Retrieval-Augmented Graph Learning Framework. arXiv preprint arXiv:2410.23855 (2024)
2024 arXiv
-
[26]
Xinke Jiang, Dingyi Zhuang, Xianghui Zhang, Hao Chen, Jiayuan Luo, and Xiaowei Gao. 2023. Uncertainty quantification via spatial-temporal tweedie model for zero-inflated and long-tail travel demand prediction. In Proceedings of CIKM. 3983–3987
2023
-
[27]
Guangyin Jin, Yuxuan Liang, Yuchen Fang, Zezhi Shao, Jincai Huang, Junbo Zhang, and Yu Zheng. 2023. Spatio-temporal graph neural networks for predictive learning in urban computing: A survey. IEEE Transactions on Knowledge and Data Engineering (2023)
2023
-
[28]
S Vasantha Kumar and Lelitha Vanajakshi. 2015. Short-term traffic flow predic- tion using seasonal ARIMA model with limited input data. European Transport Research Review (2015), 1–9
2015
-
[29]
Siqi Lai, Zhao Xu, Weijia Zhang, Hao Liu, and Hui Xiong. 2023. Large language models as traffic signal control agents: Capacity and opportunity. arXiv preprint arXiv:2312.16044 (2023)
2023 arXiv
-
[30]
Zhichen Lai, Dalin Zhang, Huan Li, Christian S Jensen, Hua Lu, and Yan Zhao
-
[31]
Shiyong Lan, Yitong Ma, Weikang Huang, Wenwu Wang, Hongyu Yang, and Pyang Li. 2022. Dstagnn: Dynamic spatial-temporal aware graph neural network for traffic flow forecasting. In Proceedings of ICML. 11906–11917
2022
-
[32]
Fuxian Li, Jie Feng, Huan Yan, Guangyin Jin, Fan Yang, Funing Sun, Depeng Jin, and Yong Li. 2023. Dynamic graph convolutional recurrent network for traffic prediction: Benchmark and solution. ACM Transactions on Knowledge Discovery from Data (2023), 1–21
2023
-
[33]
Rongfan Li, Ting Zhong, Xinke Jiang, Goce Trajcevski, Jin Wu, and Fan Zhou
-
[34]
Yaguang Li, Rose Yu, Cyrus Shahabi, and Yan Liu. 2018. Diffusion Convolutional Recurrent Neural Network: Data-Driven Traffic Forecasting. In Proceedings of ICLR
2018
-
[35]
Zhonghang Li, Lianghao Xia, Jiabin Tang, Yong Xu, Lei Shi, Long Xia, Dawei Yin, and Chao Huang. 2024. Urbangpt: Spatio-temporal large language models. In Proceedings of SIGKDD. 1–11
2024
-
[36]
Yuxuan Liang, Kun Ouyang, Junkai Sun, Yiwei Wang, Junbo Zhang, Yu Zheng, David Rosenblum, and Roger Zimmermann. 2021. Fine-grained urban flow prediction. In Proceedings of WWW . 1833–1845
2021
-
[37]
Yuxuan Liang, Yutong Xia, Songyu Ke, Yiwei Wang, Qingsong Wen, Junbo Zhang, Yu Zheng, and Roger Zimmermann. 2023. Airformer: Predicting nationwide air quality in china with transformers. In Proceedings of AAAI. 14329–14337
2023
-
[38]
Dachuan Liu, Jin Wang, Shuo Shang, and Peng Han. 2022. Msdr: Multi-step dependency relation networks for spatial temporal forecasting. In Proceedings of SIGKDD. 1042–1050
2022
-
[39]
Hangchen Liu, Zheng Dong, Renhe Jiang, Jiewen Deng, Jinliang Deng, Quanjun Chen, and Xuan Song. 2023. Spatio-temporal adaptive embedding makes vanilla transformer sota for traffic forecasting. In Proceedings of CIKM. 4125–4129
2023
-
[40]
Shuncheng Liu, Xu Chen, Ziniu Wu, Liwei Deng, Han Su, and Kai Zheng. 2022. HeGA: heterogeneous graph aggregation network for trajectory prediction in high-density traffic. In Proceedings of CIKM. 1319–1328
2022
-
[41]
Shuncheng Liu, Yuyang Xia, Xu Chen, Jiandong Xie, Han Su, and Kai Zheng. 2023. Impact-aware maneuver decision with enhanced perception for autonomous vehicle. In Proceedings of ICDE. 3255–3268
2023
-
[42]
Xu Liu, Yuxuan Liang, Chao Huang, Hengchang Hu, Yushi Cao, Bryan Hooi, and Roger Zimmermann. 2024. Reinventing Node-centric Traffic Forecasting for Improved Accuracy and Efficiency. In Proceedings of ECML-PKDD. 21–38
2024
-
[43]
Xu Liu, Yutong Xia, Yuxuan Liang, Junfeng Hu, Yiwei Wang, Lei Bai, Chao Huang, Zhenguang Liu, Bryan Hooi, and Roger Zimmermann. 2024. Largest: A benchmark dataset for large-scale traffic forecasting. In Proceedings of NeurIPS
2024
-
[44]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin transformer: Hierarchical vision transformer using shifted windows. In Proceedings of ICCV. 10012–10022
2021
-
[45]
Jiayuan Luo, Wentao Zhang, Yuchen Fang, Xiaowei Gao, Dingyi Zhuang, Hao Chen, and Xinke Jiang. 2024. Timeseries suppliers allocation risk optimization via deep black litterman model. arXiv preprint arXiv:2401.17350 (2024)
2024 arXiv
-
[46]
Zhongjian Lv, Jiajie Xu, Kai Zheng, Hongzhi Yin, Pengpeng Zhao, and Xiaofang Zhou. 2018. Lc-rnn: A deep learning model for traffic speed prediction.. In Proceedings of IJCAI. 27
2018
-
[47]
Qian Ma, Zijian Zhang, Xiangyu Zhao, Haoliang Li, Hongwei Zhao, Yiqi Wang, Zitao Liu, and Wanyu Wang. 2023. Rethinking sensors modeling: Hierarchical information enhanced traffic forecasting. In Proceedings of CIKM. 1756–1765
2023
-
[48]
Chunghyun Park, Yoonwoo Jeong, Minsu Cho, and Jaesik Park. 2022. Fast point transformer. In Proceedings of CVPR. 16949–16958
2022
-
[49]
Cheonbok Park, Chunggi Lee, Hyojin Bahng, Yunwon Tae, Seungmin Jin, Kihwan Kim, Sungahn Ko, and Jaegul Choo. 2020. ST-GRAT: A novel spatio-temporal graph attention networks for accurately forecasting dynamically changing road speed. In Proceedings of CIKM. 1215–1224. Efficient...
2020
-
[50]
Zezhi Shao, Zhao Zhang, Fei Wang, Wei Wei, and Yongjun Xu. 2022. Spatial- temporal identity: A simple yet effective baseline for multivariate time series forecasting. In Proceedings of CIKM. 4454–4458
2022
-
[51]
Zezhi Shao, Zhao Zhang, Wei Wei, Fei Wang, Yongjun Xu, Xin Cao, and Chris- tian S Jensen. 2022. Decoupled dynamic spatial-temporal graph neural network for traffic forecasting. In Proceedings of VLDB. 2733–2746
2022
-
[52]
Robert F Sproull. 1991. Refinements to nearest-neighbor searching in k- dimensional trees. Algorithmica (1991), 579–589
1991
-
[53]
Qian Sun, Rui Zha, Le Zhang, Jingbo Zhou, Yu Mei, Zhiling Li, and Hui Xiong
-
[54]
Peng-Shuai Wang. 2023. Octformer: Octree-based transformers for 3d point clouds. ACM Transactions on Graphics (TOG) (2023), 1–11
2023
-
[55]
Tianfu Wang, Liwei Deng, Chao Wang, Jianxun Lian, Yue Yan, Nicholas Jing Yuan, Qi Zhang, and Hui Xiong. 2024. COMET: NFT Price Prediction with Wallet Profiling. In Proceedings of SIGKDD. 5893–5904
2024
-
[56]
Xu Wang, Lianliang Chen, Hongbo Zhang, Pengkun Wang, Zhengyang Zhou, and Yang Wang. 2023. A Multi-graph Fusion Based Spatiotemporal Dynamic Learning Framework. In Proceedings of WSDM. 294–302
2023
-
[57]
Haomin Wen, Youfang Lin, Yutong Xia, Huaiyu Wan, Qingsong Wen, Roger Zimmermann, and Yuxuan Liang. 2023. Diffstg: Probabilistic spatio-temporal graph forecasting with denoising diffusion models. In Proceedings of SIGSPATIAL. 1–12
2023
-
[58]
Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, Xiaojun Chang, and Chengqi Zhang. 2020. Connecting the dots: Multivariate time series forecasting with graph neural networks. In Proceedings of SIGKDD. 753–763
2020
-
[59]
Zonghan Wu, Shirui Pan, Guodong Long, Jing Jiang, and Chengqi Zhang. 2019. Graph wavenet for deep spatial-temporal graph modeling. InProceedings of IJCAI. 1907–1913
2019
-
[60]
Chin-Chia Michael Yeh, Yujie Fan, Xin Dai, Uday Singh Saini, Vivian Lai, Prince Osei Aboagye, Junpeng Wang, Huiyuan Chen, Yan Zheng, Zhongfang Zhuang, et al. 2024. RPMixer: Shaking Up Time Series Forecasting with Random Projections for Large Spatial-Temporal Data. In Proceedin...
2024
-
[61]
Bing Yu, Haoteng Yin, and Zhanxing Zhu. 2018. Spatio-Temporal Graph Con- volutional Networks: A Deep Learning Framework for Traffic Forecasting. In Proceedings of IJCAI. 3634–3640
2018
-
[62]
Xingtong Yu, Yuan Fang, Zemin Liu, Yuxia Wu, Zhihao Wen, Jianyuan Bo, Xin- ming Zhang, and Steven CH Hoi. 2024. Few-shot learning on graphs: from meta-learning to pre-training and prompting. arXiv preprint arXiv:2402.01440 (2024)
2024 arXiv
-
[63]
Xingtong Yu, Zemin Liu, Yuan Fang, and Xinming Zhang. 2023. Learning to count isomorphisms with graph neural networks. In Proceedings of AAAI. 4845–4853
2023
-
[64]
Xingtong Yu, Zhenghao Liu, Yuan Fang, and Xinming Zhang. 2024. DyG- Prompt: Learning Feature and Time Prompts on Dynamic Graphs. arXiv preprint arXiv:2405.13937 (2024)
2024 arXiv
-
[65]
Chuanpan Zheng, Xiaoliang Fan, Cheng Wang, and Jianzhong Qi. 2020. Gman: A graph multi-attention network for traffic prediction. In Proceedings of AAAI. 1234–1241
2020
-
[2021]
In Proceedings of SIGKDD
Dynamic and multi-faceted spatio-temporal deep learning for traffic speed forecasting. In Proceedings of SIGKDD. 547–555
-
[2022]
In Proceedings of SIGKDD
Mining spatio-temporal relations via self-paced graph contrastive learning. In Proceedings of SIGKDD. 936–944
-
[2023]
In Proceedings of SIGMOD
Lightcts: A lightweight framework for correlated time series forecasting. In Proceedings of SIGMOD. 1–26
-
[2024]
In Proceedings of SIGKDD
CrossLight: Offline-to-Online Reinforcement Learning for Cross-City Traffic Signal Control. In Proceedings of SIGKDD. 2765–2774
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.