REVIEW 27 cited by
A Generalization of Transformer Networks to Graphs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a generalization of transformer neural network architecture for arbitrary graphs. The original transformer was designed for Natural Language Processing (NLP), which operates on fully connected graphs representing all connections between the words in a sequence. Such architecture does not leverage the graph connectivity inductive bias, and can perform poorly when the graph topology is important and has not been encoded into the node features. We introduce a graph transformer with four new properties compared to the standard model. First, the attention mechanism is a function of the neighborhood connectivity for each node in the graph. Second, the positional encoding is represented by the Laplacian eigenvectors, which naturally generalize the sinusoidal positional encodings often used in NLP. Third, the layer normalization is replaced by a batch normalization layer, which provides faster training and better generalization performance. Finally, the architecture is extended to edge feature representation, which can be critical to tasks s.a. chemistry (bond type) or link prediction (entity relationship in knowledge graphs). Numerical experiments on a graph benchmark demonstrate the performance of the proposed graph transformer architecture. This work closes the gap between the original transformer, which was designed for the limited case of line graphs, and graph neural networks, that can work with arbitrary graphs. As our architecture is simple and generic, we believe it can be used as a black box for future applications that wish to consider transformer and graphs.
Forward citations
Cited by 27 Pith papers
-
Hybrid Lagrangian-Eulerian Model for Lagrangian Fluid Simulation
A hybrid Lagrangian-Eulerian graph neural simulator with adaptive downsampling and cross-attention achieves state-of-the-art accuracy and rollout stability on particle-based fluid benchmarks.
-
GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models
A demand-initialization layer that embeds travel demand on virtual links instead of node features lets GNN traffic-flow surrogates transfer across city networks with minimal fine-tuning.
-
Node4All: Learning Node Representation Beyond Datasets
A single graph encoder pretrained only on synthetic graphs produces node representations that rank 5th among 21 per-dataset-optimized baselines on 25 benchmarks without any per-dataset tuning.
-
Eigenbasis-Independent Learnable Spectral Positional Encodings for Directed Graphs via Hermitian Block Krylov Subspaces
Learnable spectral positional encodings for directed graphs are computed as gauge-invariant matrix functions of the magnetic operator via block Krylov subspaces, achieving O(log(1/ε)) approximation with sparse matrix-...
-
Toward Manifest Relationality in Transformers via Symmetry Reduction
Transformer attention and parameter optimization can be rewritten on symmetry-reduced relational variables (Gram matrices and invariant parameter composites), removing coordinate redundancies by construction.
-
Text2Structure3D: Graph-Based Generative Modeling of Equilibrium Structures with Diffusion Transformers
A text-conditioned latent diffusion model over structural graphs generates funicular and truss bridge designs that are post-processed into static equilibrium.
-
A Hierarchical Quantized Tokenization Framework for Task-Adaptive Graph Representation Learning
QUIET is a hierarchical RVQ-based graph tokenizer with a learned level-weighting gate; it improves several benchmarks but not consistently against the strongest baselines.
-
CoBAD: Modeling Collective Behaviors for Human Mobility Anomaly Detection
CoBAD detects collective mobility anomalies (unexpected co-occurrence and absence) by pre-training a two-stage attention model over collective event sequences and event graphs.
-
Player-Team Heterogeneous Interaction Graph Transformer for Soccer Outcome Prediction
HIGFormer predicts soccer match outcomes by jointly modeling player-player event interactions and team-team historical win rates with a heterogeneous graph transformer and graph convolution network.
-
AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research
AblationBench is a new benchmark for testing AI planning of ablation experiments, and it shows that current language models recover only a minority of the human reference ablations.
-
NN-Former: Rethinking Graph Structure in Neural Architecture Representation
NN-Former improves neural accuracy and latency prediction by using attention masks over sibling nodes in the architecture graph.
-
HEIST: A Graph Foundation Model for Spatial Transcriptomics and Proteomics Data
HEIST pretrains a hierarchical graph transformer on 22.3 million cells so that one frozen encoder can annotate cell types, cluster cells, impute missing genes, and predict clinical outcomes in spatial transcriptomics ...
-
Learnable Spatial-Temporal Positional Encoding for Link Prediction
L-STEP learns time-evolving positional encodings for graph nodes via a learnable spectral filter and predicts links with MLPs only, matching or beating attention-based baselines on 13 temporal datasets.
-
Distributed Control of Network Systems in the Space of Stabilizing Graph Neural Network Policies
A graph-neural-network policy parameterization guarantees closed-loop stability by construction and transfers from small to large networks without retraining.
-
Asynchronous Message Passing for Addressing Oversquashing in Graph Neural Networks
CAMP updates nodes in centrality-ranked batches to spread information across GNN layers and claims to reduce oversquashing without rewiring, but the proof and evidence are not convincing.
-
Towards Interpretable Drug-Drug Interaction Prediction: A Graph-Based Approach with Molecular and Network-Level Explanations
MolecBioNet, a graph neural network that treats drug pairs as unified entities with knowledge graph and molecular substructure views, reports state-of-the-art accuracy, F1, and PR-AUC on the Ryu and DrugBank DDI benchmarks.
-
Few-shot Learning on AMS Circuits and Its Application to Parasitic Capacitance Prediction
A few-shot graph-pretraining pipeline, built from subgraph sampling and a hybrid graph transformer, predicts parasitic coupling capacitance on unseen AMS circuits with substantially lower error than prior graph baselines.
-
Density-aware Walks for Coordinated Campaign Detection
Density-aware random-walk embeddings improve coordinated campaign detection accuracy on the LEN Twitter graph dataset.
-
GITO: Graph-Informed Transformer Operator for Learning Complex Partial Differential Equations
GITO, a graph-informed transformer operator, reports lower relative L2 errors than existing transformer-based neural operators on Navier-Stokes, heat conduction, and airfoil benchmark datasets.
-
Demystifying Topological Message-Passing with Relational Structures: A Case Study on Oversquashing in Simplicial Message-Passing
Simplicial message passing can be analyzed for oversquashing by collapsing its relational structure into an influence graph and applying graph-theoretic sensitivity, curvature, and rewiring tools.
-
On Preserving Geometrical Invariance for Superpixel Image Classification using Graph Transformer
A GraphGPS-style transformer on SLIC RAGs with mean-centered centroids reaches ~80.2% CIFAR-10 accuracy, matching ShapeGNN without boundary-point features and with better low-data stability.
-
STGAtt: A Spatial-Temporal Unified Graph Attention Network for Traffic Flow Forecasting
STGAtt claims a unified spatial-temporal graph attention model with a neighborhood signal-exchanging mechanism beats state-of-the-art traffic forecasters on PEMS-BAY and SHMetro.
-
MPFSR-Enhanced GNNs: Spectral Graph Neural Networks Enhancement Through Learnable Multiple-Parameter Graph Fractional Fourier Transforms
The paper proposes multiple-parameter graph fractional Fourier transforms and a plug-in module that lets spectral GNNs learn per-frequency transform orders, with modest node-classification gains.
-
Commute Networks as a Signature of Urban Socioeconomic Performance: Evaluating Mobility Structures with Deep Learning Models
Commute networks alone predict median household income at the census-tract level across 12 U.S. cities, with an end-to-end GNN+VNN pipeline outperforming a 311-feature baseline in most cities.
-
Transformers are Graph Neural Networks
Transformer self-attention is message passing on a complete graph, making Transformers a special case of graph neural networks.
-
Graph-of-Causal Evolution: Challenging Chain-of-Model for Reasoning
GoCE swaps CoM's chain structure for a differentiable causal graph and reports accuracy gains on CLUTRR, CLadder, EX-FEVER, and CausalQA, but the evidence is sandbox-generated and unauditable.
-
Leveraging Manifold Embeddings for Enhanced Graph Transformer Representations and Learning
R-SGFormer pairs GraphMoRE embeddings with SGFormer, but its own tables show the full model fails to consistently beat SGFormer or GraphMoRE baselines.
Discussion (0). Sign in to comment.