Pith. sign in

REVIEW 4 cited by

Do Transformers Really Perform Bad for Graph Representation?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.05234 v5 pith:DKGX6GQ5 submitted 2021-06-09 cs.LG cs.AI

classification cs.LGcs.AI
keywords graphgraphormerencodingrepresentationstructuraltransformerarchitectureinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The Transformer architecture has become a dominant choice in many domains, such as natural language processing and computer vision. Yet, it has not achieved competitive performance on popular leaderboards of graph-level prediction compared to mainstream GNN variants. Therefore, it remains a mystery how Transformers could perform well for graph representation learning. In this paper, we solve this mystery by presenting Graphormer, which is built upon the standard Transformer architecture, and could attain excellent results on a broad range of graph representation learning tasks, especially on the recent OGB Large-Scale Challenge. Our key insight to utilizing Transformer in the graph is the necessity of effectively encoding the structural information of a graph into the model. To this end, we propose several simple yet effective structural encoding methods to help Graphormer better model graph-structured data. Besides, we mathematically characterize the expressive power of Graphormer and exhibit that with our ways of encoding the structural information of graphs, many popular GNN variants could be covered as the special cases of Graphormer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Graph Neural Networks for the Graphical Bootstrap

    hep-th 2026-07 conditional novelty 6.0 of 10

    GNNs and graph transformers classify vanishing coefficients on millions of N=4 SYM f-graphs, generalizing to larger n with 99.996% ROC AUC and pruning up to 85.5% of redundant d-graphs.

  2. ReDiSC: A Reparameterized Masked Diffusion Model for Scalable Node Classification with Structured Predictions

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A reparameterized masked diffusion model with variational EM gives scalable structured node classification, matching or beating GNN, label propagation, and continuous diffusion baselines.

  3. Polaritonic Machine Learning for Graph-based Data Analysis

    cond-mat.dis-nn 2025-07 conditional novelty 6.0 of 10

    Simulated polariton condensate lattices act as physics-based feature generators for CNNs and improve classification of cliques and asymmetries in point clouds over raw point images in three synthetic tasks.

  4. QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization

    cs.AI 2026-07 conditional novelty 5.0 of 10

    QLPO resamples GRPO training groups to favor short correct and long incorrect responses, cutting reasoning length substantially while keeping accuracy roughly unchanged.

Pith tools