REVIEW 5 major objections 6 minor 62 references
SFi-Former: Sparse Flow Induced Attention for Graph Transformer
T0 review · 5 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Sparse attention from l1-regularized network-flow energy minimization makes graph transformers select relevant nodes, and the paper reports top results on long-range graph benchmarks with smaller generalization gaps.
desk verdict A clean energy-based attention reformulation with strong LRGB results, but the row-sum penalty issue and overbroad SOTA claim need fixing before it is fully convincing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the SFi-attention pattern, obtained by solving a regularized network-flow energy minimization per attention head: $$\min_{\boldsymbol{Z}}\ \tfrac12\mathrm{Tr}\big((\boldsymbol{R}^h\circ\boldsymbol{Z})\boldsymbol{Z}^T\big) + \$\lambda$\|\boldsymbol{F}^h\circ\boldsymbol{Z}\|_{1,1} + \tfrac{\$\alpha$}{2}\|\boldsymbol{Z}\mathbf{1}_n - \mathbf{1}_n\|$_2^{2}$,$$ where $\boldsymbol{R}^h$ is a learnable resistance matrix built from softmax of negative scaled query-key products, $\boldsymbol{F}^h$ is a learnable friction matrix acting as a node-wise noise filter, and the $\ell_1$ term drives small flows to zero. Each row of the optimal flow $\boldsymbol{Z}^*$ plays the role of an attention distribution; because the constraint $\boldsymbol{Z}\mathbf{1}_n=\mathbf{1}_n$ is replaced by a quadratic penalty with coefficient $\alpha$, the rows are only approximately normalized. The flow is computed by proximal-gradient iteration with a two-point spectral step-size rule, and the resulting sparse pattern enters the residual update of Eq. (13), where the normalized adjacency $\tilde{\boldsymbol{A}}$ is added to $\gamma\,\mathrm{SFi\text{-}ATT}_h(\boldsymbol{X})$ before feature mixing.
What would settle it
Measure $\|\boldsymbol{Z}^*\mathbf{1}_n - \mathbf{1}_n\|_2$ or the maximum absolute row-sum deviation on trained SFi-Former heads across the long-range datasets; if rows deviate substantially from one and imposing exact row normalization changes test scores materially, then the sparse attention mechanism itself is not what produces the reported gains.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that the dense attention of a graph transformer can be re-derived as the minimizer of a quadratic flow energy on a complete graph, and that adding an $\ell_1$ penalty on the flows, with learnable per-node frictions, turns that minimizer into a genuinely sparse attention pattern. Plugging this pattern into the residual adjacency-enhanced update $$\boldsymbol{X}^{(k+1)} = \boldsymbol{X}^{(k)} + (1+\gamma)^{-1}\sum_{h}\big[\tilde{\boldsymbol{A}} + \gamma\,\mathrm{SFi\text{-}ATT}_h(\boldsymbol{X}^{(k)})\big]\boldsymbol{X}^{(k)}\boldsymbol{W}^h_V\boldsymbol{W}^h_O$$ yields a graph transformer that the paper reports reaches best reported numbers on most Long Range Graph Benchmark datasets and competitive numbers on the standard graph benchmark suite, with consistently smaller train-test gaps than a dense-attention counterpart. The paper reads this as evidence that selective aggregation is a useful inductive bias: irrelevant nodes can be ignored rather than weakly averaged into every representation.
Load-bearing premise
The load-bearing premise is that the computed sparse flows are valid attention weights: the optimization only enforces the row-sum-to-one constraint indirectly through a quadratic penalty with coefficient $0.1$, and the paper never reports how far the rows actually deviate from summing to one.
Editorial extensions
If this is right
- On image-derived long-range datasets such as PascalVOC-SP and COCO-SP, the sparse mechanism produces the largest gains, which the paper attributes to many background superpixels needing no interaction; about 20% of attention entries are driven to zero there.
- Across PascalVOC-SP, Peptides-Func, and Peptides-Struct, the train-test gap is consistently smaller than for a dense-attention baseline with the same backbone, implying sparsity acts as a regularizer rather than a speed-up.
- Setting $\lambda=0$ recovers dense attention inside the same energy framework, so standard self-attention becomes a special case of the flow model, giving a unified derivation and a flexible template for other attention designs.
- The adjacency-enhanced residual term alone is already competitive, and sparsity adds further improvement on most benchmarks, so the two components are complementary rather than redundant.
- On the long-range benchmark the model reports leading results, while on the standard graph benchmark it is competitive but not uniformly best, consistent with sparsity helping most when many nodes are task-irrelevant.
Reading between the lines
- A direct testable extension is to replace the quadratic penalty on row sums with exact row normalization or a hard flow-conservation constraint; if gains persist, the story is about sparsity, and if they vanish, it is about scale miscalibration.
- The learned friction matrix $\boldsymbol{F}^h$ could serve as a per-node importance map; an interpretability study could check whether nodes that retain high flow are the semantically salient ones in superpixel graphs.
- Because the energy function can be defined on non-complete graph topologies, a direction the paper flags, SFi-attention could be applied to $k$-hop or expander graphs, bringing the same selectivity with reduced computation.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SFi-Former, a graph transformer whose attention is obtained by minimizing a network-flow energy with l1-norm regularization, yielding sparse attention patterns. The sparse flow attention is combined with a residual adjacency-enhanced update inside the GraphGPS framework. The authors report competitive results on the GNN Benchmark datasets, state-of-the-art performance on several Long Range Graph Benchmark datasets, and smaller train-test gaps relative to GraphGPS, which they interpret as evidence of reduced overfitting.
Significance. If the mechanism performs as claimed, the paper contributes a flexible energy-based framework for attention that includes standard softmax attention as a special case and introduces learnable friction terms for adaptive sparsity. The conceptual connection between network flows and attention is interesting, and the authors provide code and extensive benchmarking on standard graph datasets. However, the central empirical claims rest on the correctness of the proximal solver and on the normalized behavior of the resulting attention matrix; both need to be verified. At present, the significance is moderate: the framework is promising, but the paper does not yet establish that the observed gains come from sparse adaptive attention rather than from an artifact of unnormalized or incorrectly computed flows.
major comments (5)
- [Section 3.3, Eq. (12)] The proximal update as written is not the soft-threshold operator defined in Section 3.2: it computes sign(Y) max(|Y - tλF|, 0), whereas the correct operator is sign(Y) max(|Y| - tλF, 0). For |Y| < tλF, the expression as written produces a positive value instead of zero. Consequently, the iterates do not minimize the penalized objective in Eq. (11), and the sparsity statistics reported in Section 5.1 may not reflect the claimed optimum.
- [Section 3.3, Eq. (11); Tables 5 and 6] No diagnostic is provided for the row sums of the computed Z*. Since Eq. (11) replaces the hard constraint Z 1 = 1 by a quadratic penalty with α = 0.1, rows of Z* may deviate substantially from unit sum. If a row collapses to zero, Eq. (13) suppresses the global aggregation for that query rather than renormalizing a sparse attention distribution. The paper should report the distribution of row sums, or apply an explicit normalization step, to substantiate the claim that SFi-Former implements adaptive sparse attention.
- [Section 5.4, Figure 3] The generalization-gap evidence is confounded: Figure 3 compares SFi-Former only against GraphGPS, which lacks both the adjacency enhancement and the sparse attention mechanism. The ablation in Table 3 shows that the adjacency component alone contributes substantial performance, so the smaller train-test gap cannot be attributed to sparsity without a comparison against DFi-Former (or another dense-attention model that also includes the adjacency enhancement).
- [Abstract and Section 5.1, Table 1] The claim of SOTA performance on LRGB is not supported by the reported numbers. On PCQM-Contact, SFi-Former achieves MRR 0.3516 while Exphormer achieves 0.3637 and the authors' own DFi-Former achieves 0.3765; on COCO-SP, SFi-Former (0.3801) is below DFi-Former (0.3974). The abstract and Section 5.1 should be qualified to say 'competitive or state-of-the-art on some LRGB datasets.'
- [Section 3.3 and Supplementary A.2] The convergence guarantee stated after Eq. (12) is not connected to the actual iteration: the BB step size t^(k) is not constrained to satisfy t^(k) ≤ (||R^h|| + α√n)^{-1}, and the BB formula in Eq. (12) lacks the standard squared norm in the denominator. In addition, the supplementary proof that ||R^h|| ≤ 1 via Perron-Frobenius is not generally true for the spectral norm of a row-stochastic matrix. The convergence statement should be corrected, or the step size should be explicitly bounded.
minor comments (6)
- [Section 3.1, Eq. (5)] The constraint 'Z 1_n − 1_n = 0_n' uses 1_n to denote both the vector of ones and the scalar 1; this notation should be disambiguated.
- [Section 3.1] The phrase 'r_i ∝ exp(−q_s^T k_i / sqrt(d_k))^2' is ambiguous because the exponent appears to be a superscript on the entire expression; please clarify whether the square is part of the definition.
- [Section 6] The conclusion contains a typo: 'attetion' should be 'attention.'
- [Tables 1 and 2] The tables state that the first, second, and third best results are highlighted, but the highlighting is not visible in the manuscript text; please ensure the formatting is clear.
- [Abstract] The 'Click here for codes' link is not a working URL; please provide a repository link.
- [Eq. (12)] The notation '∇(k)_Z H' should be defined as the gradient of H evaluated at Z^(k) to avoid confusion.
Circularity Check
The softmax-recovery step in Sec. 3.1 is by construction, but it is transparent and not load-bearing; the paper's empirical claims rest on external benchmarks.
-
self definitional
[Section 3.1, after Eq. (6)]
"If we identify the resistance as 𝑟𝑖∝ exp(−𝒒𝑇𝑠 𝒌𝑖/ √ 𝑑𝑘)2, we observe that the optimal network flow is 𝑧∗ 𝑖 = exp(𝒒𝑇 𝑠 𝒌𝑖/√𝑑𝑘)˝𝑛 𝑗=1 exp(𝒒𝑇𝑠 𝒌𝑗/√𝑑𝑘) , which is the attention score from the query node 𝒗𝑠 to a key node 𝒗𝑖 in the standard self-attention mechanism."
The recovery of standard softmax attention is forced by the parameterization: the resistance r_i is chosen as the reciprocal of the softmax numerator (up to sign/temperature), so the optimal flow z*_i = (1/r_i)/Σ(1/r_j) algebraically reduces to exp(q·k/√d)/Σ exp(q·k/√d). The energy minimization therefore does not independently derive softmax; it encodes softmax as a special case by definition. This is transparently presented as an observation, and it is used only to motivate the framework, not to generate any fitted prediction, so its circular character is minor and not load-bearing for the central claims.
full rationale
The paper's main derivation chain has one clearly by-construction step: Section 3.1 defines resistance in terms of the attention logits and then recovers softmax attention as the energy-minimizing flow. This is a self-definitional special case, not an independent first-principles result. However, the central contribution, SFi-attention, is not built on this identity as a prediction: the sparse patterns come from the l1-penalized flow problem in Eq. (10)-(12), with learnable resistances and frictions, solved by an iterative proximal method. The dense DFi-Former is a λ=0 ablation, not a fitted target, and the LRGB and GNN benchmark numbers are held-out test results against standard baselines, so no parameter is fit to the reported metric and then renamed as a prediction. There are no self-citations in the reference list, and no load-bearing uniqueness theorem from prior work by the same authors is invoked. The finite-penalty issue with α=0.1 and possible row-sum drift is a correctness/robustness concern rather than circularity: even if the rows of Z* are far from summing to one, that would be an implementation artifact of the penalty method, not evidence that the conclusion was assumed in the premise. Overall, the circularity is limited to the transparent softmax-recovery special case, which does not compromise the empirical evaluation.
Assumptions & free parameters
free parameters (4)
- lambda (l1 sparsity coefficient) =
1.0
- alpha (flow conservation penalty) =
0.1
- gamma (adjacency/attention balance)
- Proximal iteration count or tolerance =
not stated
assumptions (4)
- standard math The penalized energy in Eq. (11) is convex and the proximal gradient method in Eq. (12) converges to its global minimum under a Lipschitz step-size bound.
- standard math The softmax parameterization keeps R^h and F^h in (0,1).
- domain assumption The finite-penalty solution with alpha=0.1 approximates the constrained attention normalization accurately enough.
- domain assumption Sparse selective aggregation reduces overfitting in graph transformers, by analogy with LASSO-style shrinkage.
invented entities (1)
-
Learnable friction field F^h
Cite this review
Pith. "Pith review of SFi-Former: Sparse Flow Induced Attention for Graph Transformer." pith.science (2026). https://pith.science/paper/LAR2C6NS
@misc{pith2026250420666,
author = {Pith},
title = {Pith review of: SFi-Former: Sparse Flow Induced Attention for Graph Transformer},
year = {2026},
howpublished = {\url{https://pith.science/paper/LAR2C6NS}},
note = {Machine review of arXiv:2504.20666}
}
read the original abstract
Graph Transformers (GTs) have demonstrated superior performance compared to traditional message-passing graph neural networks in many studies, especially in processing graph data with long-range dependencies. However, GTs tend to suffer from weak inductive bias, overfitting and over-globalizing problems due to the dense attention. In this paper, we introduce SFi-attention, a novel attention mechanism designed to learn sparse pattern by minimizing an energy function based on network flows with l1-norm regularization, to relieve those issues caused by dense attention. Furthermore, SFi-Former is accordingly devised which can leverage the sparse attention pattern of SFi-attention to generate sparse network flows beyond adjacency matrix of graph data. Specifically, SFi-Former aggregates features selectively from other nodes through flexible adaptation of the sparse attention, leading to a more robust model. We validate our SFi-Former on various graph datasets, especially those graph data exhibiting long-range dependencies. Experimental results show that our SFi-Former obtains competitive performance on GNN Benchmark datasets and SOTA performance on LongRange Graph Benchmark (LRGB) datasets. Additionally, our model gives rise to smaller generalization gaps, which indicates that it is less prone to over-fitting. Click here for codes.
Figures
Reference graph
Works this paper leans on
-
[1]
Ralph Abboud, Radoslav Dimitrov, and Ismail Ilkan Ceylan. 2022. Shortest path networks for graph property prediction. InLearning on Graphs Conference. PMLR, 5–1
work page 2022
-
[2]
Uri Alon and Eran Yahav. 2020. On the bottleneck of graph neural networks and its practical implications. arXiv preprint arXiv:2006.05205 (2020)
arXiv 2020
-
[3]
Jonathan Barzilai and Jonathan M Borwein. 1988. Two-point step size gradient methods. IMA journal of numerical analysis 8, 1 (1988), 141–148
work page 1988
-
[4]
Ali Behrouz and Farnoosh Hashemi. 2024. Graph mamba: Towards learning on graphs with state space models. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 119–130
work page 2024
-
[5]
Cristian Bodnar, Fabrizio Frasca, Nina Otter, Yuguang Wang, Pietro Liò, Guido F Montufar, and Michael Bronstein. 2021. Weisfeiler and Lehman go cellular: CW networks. In Advances in Neural Information Processing Systems (NeurIPS) . 2625–2640
work page 2021
-
[6]
Giorgos Bouritsas, Fabrizio Frasca, Stefanos P Zafeiriou, and Michael Bronstein
-
[7]
Xavier Bresson and Thomas Laurent. 2017. Residual gated graph convnets. arXiv preprint arXiv:1711.07553 (2017)
arXiv 2017
-
[8]
Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković
Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. 2021. Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges. arXiv:2104.13478 [cs.LG] https://arxiv.org/abs/2104.13478
arXiv 2021
Show all 62 references
-
[9]
Fabio Catania, Micol Spitale, and Franca Garzotto. 2023. Conversational agents in therapeutic interventions for neurodevelopmental disorders: a survey. Comput. Surveys 55, 10 (2023), 1–34
2023
-
[10]
Ben Chamberlain, James Rowbottom, Maria I Gorinova, Michael Bronstein, Stefan Webb, and Emanuele Rossi. 2021. GRAND: Graph Neural Diffusion. InProceedings of the 38th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139), Marina Meil...
2021
-
[11]
Dexiong Chen, Leslie O’Bray, and Karsten Borgwardt. 2022. Structure-aware transformer for graph representation learning. In Proceedings of the 39th Interna- tional Conference on Machine Learning (ICML)
2022
-
[12]
Jinsong Chen, Kaiyuan Gao, Gaichao Li, and Kun He. 2022. NAGphormer: A tokenized graph transformer for node classification in large graphs.arXiv preprint arXiv:2206.04910 (2022)
2022 arXiv
-
[13]
Qi Chen, Yifei Wang, Yisen Wang, Jiansheng Yang, and Zhouchen Lin. 2022. Optimization-induced graph implicit nonlinear diffusion. In International Confer- ence on Machine Learning . PMLR, 3648–3661
2022
-
[14]
Krzysztof Marcin Choromanski, Valerii Likhosherstov, David Dohan, Xingyou Song, Andreea Gane, Tamás Sarlós, Peter Hawkins, Jared Quincy Davis, Afroz Mohiuddin, Lukasz Kaiser, David Benjamin Belanger, Lucy J Colwell, and Adrian Weller. 2021. Rethinking attention with performers...
2021
-
[15]
Correia, Vlad Niculae, and André F
Gonçalo M. Correia, Vlad Niculae, and André F. T. Martins. 2019. Adaptively Sparse Transformers. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) ,...
2019 doi
-
[16]
Michaël Defferrard, Xavier Bresson, and Pierre Vandergheynst. 2016. Con- volutional Neural Networks on Graphs with Fast Localized Spectral Fil- tering. In Advances in Neural Information Processing Systems , D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Eds.), Vol....
2016
-
[17]
Doyle and J.L
P.G. Doyle and J.L. Snell. 1984.Random walks and electric networks. Mathematical Association of America
1984
-
[18]
Vijay Prakash Dwivedi and Xavier Bresson. 2021. A Generalization of Trans- former Networks to Graphs.AAAI Workshop on Deep Learning on Graphs: Methods and Applications (2021)
2021
-
[19]
Vijay Prakash Dwivedi, Chaitanya K Joshi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. 2023. Benchmarking graph neural networks. Journal of Machine Learning Research 24, 43 (2023), 1–48
2023
-
[20]
Vijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio, and Xavier Bresson. 2022. Graph neural networks with learnable structural and positional representations. InInternational Conference on Learning Representations (ICLR)
2022
-
[21]
Vijay Prakash Dwivedi, Ladislav Rampášek, Mikhail Galkin, Ali Parviz, Guy Wolf, Anh Tuan Luu, and Dominique Beaini. 2022. Long range graph benchmark. In Neural Information Processing Systems (NeurIPS 2022), Track on Datasets and Benchmarks
2022
-
[22]
Stéphane d’Ascoli, Hugo Touvron, Matthew L Leavitt, Ari S Morcos, Giulio Biroli, and Levent Sagun. 2021. Convit: Improving Vision Transformers with Soft Convolutional Inductive Biases. In International Conference on Machine Learning . PMLR, 2286–2296
2021
-
[23]
Quentin Fournier, Gaétan Marceau Caron, and Daniel Aloise. 2023. A practical survey on faster and lighter transformers. Comput. Surveys 55, 14s (2023), 1–40
2023
-
[24]
Guoji Fu, Mohammed Haroon Dupty, Yanfei Dong, and Lee Wee Sun. 2023. Im- plicit graph neural diffusion based on constrained Dirichlet energy minimization. arXiv preprint arXiv:2308.03306 (2023)
2023 arXiv
-
[25]
Jianyuan Guo, Kai Han, Han Wu, Chang Xu, Yehui Tang, Chunjing Xu, and Yunhe Wang. 2021. CMT: Convolutional neural networks meet vision transformers. arXiv preprint arXiv:2107.06263 (2021)
2021 arXiv
-
[26]
Will Hamilton, Zhitao Ying, and Jure Leskovec. 2017. Inductive representation learning on large graphs. Advances in neural information processing systems 30 (2017)
2017
-
[27]
Andi Han, Dai Shi, Lequan Lin, and Junbin Gao. 2023. From continuous dy- namics to graph neural networks: Neural diffusion and beyond. arXiv preprint arXiv:2310.10121 (2023)
2023 arXiv
-
[28]
Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. 2022. A survey on vision transformer. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
-
[29]
Trevor Hastie, Robert Tibshirani, and Jerome Friedman. 2009. The Elements of Statistical Learning. Springer New York. doi:10.1007/978-0-387-84858-7
2009 doi
-
[30]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
2016
-
[31]
Katikapalli Subramanyam Kalyan, Ajit Rajasekharan, and Sivanesan Sangeetha
-
[32]
Thomas N Kipf and Max Welling. 2016. Semi-supervised classification with graph convolutional networks. arXiv preprint arXiv:1609.02907 (2016)
2016 arXiv
-
[33]
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The efficient transformer. In 8th International Conference on Learning Representations (ICLR)
2020
-
[34]
Rik Koncel-Kedziorski, Dhanush Bekal, Yi Luan, Mirella Lapata, and Hannaneh Hajishirzi. 2019. Text generation from knowledge graphs with graph transformers. arXiv preprint arXiv:1904.02342 (2019)
2019 arXiv
-
[35]
Devin Kreuzer, Dominique Beaini, Will Hamilton, Vincent Létourneau, and Pru- dencio Tossou. 2021. Rethinking Graph Transformers with Spectral Attention. In Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P.S. Liang, and J. Wortman Va...
2021
-
[36]
Hamilton, Vincent Létourneau, and Prudencio Tossou
Devin Kreuzer, Dominique Beaini, William L. Hamilton, Vincent Létourneau, and Prudencio Tossou. 2021. Rethinking graph transformers with spectral attention. ICMR ’25, June 30–July 3, 2025, Chicago, IL, USA. Zhonghao Li, Ji Shi, Xinming Zhang, Miao Zhang, and Bo Li In Advances ...
2021
-
[37]
Chaoliu Li, Lianghao Xia, Xubin Ren, Yaowen Ye, Yong Xu, and Chao Huang. 2023. Graph Transformer for Recommendation. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (Taipei, Taiwan) (SIGIR ’23). Association for...
2023
-
[38]
Xiaorui Liu, Wei Jin, Yao Ma, Yaxin Li, Hua Liu, Yiqi Wang, Ming Yan, and Jiliang Tang. 2021. Elastic graph neural networks. InInternational Conference on Machine Learning. PMLR, 6837–6849
2021
-
[39]
Ilya Makarov, Dmitrii Kiselev, Nikita Nikitinsky, and Lovro Subelj. 2021. Survey on graph embeddings and their applications to machine learning problems on graphs. PeerJ Computer Science 7 (2021), e357
2021
-
[40]
Joshua Mitton, Hans M Senn, Klaas Wynne, and Roderick Murray-Smith. 2021. A graph vae and graph transformer approach to generating molecular graphs. arXiv preprint arXiv:2104.04345 (2021)
2021 arXiv
-
[41]
Luis Müller, Mikhail Galkin, Christopher Morris, and Ladislav Rampášek. 2024. Attending to Graph Transformers. Transactions on Machine Learning Research (2024). https://openreview.net/forum?id=HhbqHBBrfZ
2024
-
[42]
Kenta Oono and Taiji Suzuki. 2019. Graph neural networks exponentially lose expressive power for node classification. arXiv preprint arXiv:1905.10947 (2019)
2019 arXiv
-
[43]
Neal Parikh and Stephen Boyd. 2014. Proximal Algorithms. Found. Trends Optim. 1, 3 (Jan. 2014), 127–239. doi:10.1561/2400000003
2014 doi
-
[44]
Lukas Rampasek, Mikhail Galkin, Vijay P Dwivedi, Anh Tuan Luu, Giacomo Wolf, and Dominique Beaini. 2022. Recipe for a general, powerful, scalable graph transformer. CoRR abs/2205.12454 (2022)
2022 arXiv
-
[45]
Ladislav Rampášek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu, Guy Wolf, and Dominique Beaini. 2022. Recipe for a General, Powerful, Scalable Graph Transformer. In Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. C...
2022
-
[46]
Patrick Rebeschini and Sekhar Tatikonda. 2019. A new approach to Laplacian solvers and flow problems. Journal of Machine Learning Research 20, 36 (2019), 1–37
2019
-
[47]
Raif Rustamov and James Klosowski. 2018. Interpretable graph-based semi- supervised learning via flows. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32
2018
-
[48]
Suther- land, and Ali Kemal Sinop
Hamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J. Suther- land, and Ali Kemal Sinop. 2023. Exphormer: Sparse Transformers for Graphs. arXiv:2303.06147 [cs.LG] https://arxiv.org/abs/2303.06147
2023 arXiv
-
[49]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. In Advances in Neural Information Processing Systems , Vol. 30. Curran Associates, Inc
2017
-
[50]
Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Lio, and Yoshua Bengio. 2017. Graph attention networks. arXiv preprint arXiv:1710.10903 (2017)
2017 arXiv
-
[51]
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang. 2024. Graph-mamba: Towards long-range graph sequence modeling with selective state spaces. arXiv preprint arXiv:2402.00789 (2024)
2024 arXiv
-
[52]
John Wright and Yi Ma. 2022. High-Dimensional Data Analysis with Low- Dimensional Models: Principles, Computation, and Applications . Cambridge Uni- versity Press
2022
-
[53]
Qitian Wu, Chenxiao Yang, Wentao Zhao, Yixuan He, David Wipf, and Junchi Yan
-
[54]
Yujie Xing, Xiao Wang, Yibo Li, Hai Huang, and Chuan Shi. 2024. Less is More: on the Over-Globalizing Problem in Graph Transformers. In Forty-first International Conference on Machine Learning. https://openreview.net/forum?id=uKmcyyrZae
2024
-
[55]
Keyulu Xu, Weihua Hu, Jure Leskovec, and Stefanie Jegelka. 2018. How powerful are graph neural networks? arXiv preprint arXiv:1810.00826 (2018)
2018 arXiv
-
[56]
Yang Ye and Shihao Ji. 2021. Sparse graph attention networks. IEEE Transactions on Knowledge and Data Engineering 35, 1 (2021), 905–916
2021
-
[57]
Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu. 2021. Do transformers really perform badly for graph representation?. In Advances in Neural Information Processing Systems (NeurIPS)
2021
-
[58]
Manzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie, Chris Alberti, Santiago Ontañón, Philip Pham, Anirudh Ravula, Qifan Wang, Li Yang, and Amr Ahmed. 2020. Big Bird: Transformers for longer sequences. In Advances in Neural Information Processing Systems (NeurIPS)
2020
-
[59]
search direction,
Si Zhang, Hanghang Tong, Jiejun Xu, and Ross Maciejewski. 2019. Graph convo- lutional networks: a comprehensive review. Computational Social Networks 6, 1 (2019), 1–23. SFi-Former: Sparse Flow Induced Attention for Graph Transformer ICMR ’25, June 30–July 3, 2025, Chicago, IL,...
2019
-
[2021]
arXiv preprint arXiv:2108.05542 (2021)
Ammus: A survey of transformer-based pretrained models in natural language processing. arXiv preprint arXiv:2108.05542 (2021)
2021 arXiv
-
[2022]
IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
Improving graph neural network expressivity via subgraph isomorphism counting. IEEE Transactions on Pattern Analysis and Machine Intelligence (2022)
2022
-
[2023]
arXiv preprint arXiv:2301.09474 (2023)
Difformer: Scalable (graph) transformers induced by energy constrained diffusion. arXiv preprint arXiv:2301.09474 (2023)
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.