Pith. sign in

REVIEW 2 cited by

Less is More: on the Over-Globalizing Problem in Graph Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.01102 v2 pith:JMFTLPPH submitted 2024-05-02 cs.LG cs.AI

classification cs.LGcs.AI
keywords graphnodesattentionglobalinformationmechanismover-globalizingproblem
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Graph Transformer, due to its global attention mechanism, has emerged as a new tool in dealing with graph-structured data. It is well recognized that the global attention mechanism considers a wider receptive field in a fully connected graph, leading many to believe that useful information can be extracted from all the nodes. In this paper, we challenge this belief: does the globalizing property always benefit Graph Transformers? We reveal the over-globalizing problem in Graph Transformer by presenting both empirical evidence and theoretical analysis, i.e., the current attention mechanism overly focuses on those distant nodes, while the near nodes, which actually contain most of the useful information, are relatively weakened. Then we propose a novel Bi-Level Global Graph Transformer with Collaborative Training (CoBFormer), including the inter-cluster and intra-cluster Transformers, to prevent the over-globalizing problem while keeping the ability to extract valuable information from distant nodes. Moreover, the collaborative training is proposed to improve the model's generalization ability with a theoretical guarantee. Extensive experiments on various graphs well validate the effectiveness of our proposed CoBFormer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Generative Adaptive Replay Continual Learning Model for Temporal Knowledge Graph Reasoning

    cs.IR 2025-06 conditional novelty 6.0 of 10

    DGAR uses diffusion-generated, model-guided historical entity distributions with adaptive replay to reduce catastrophic forgetting in temporal knowledge graph reasoning.

  2. OpenGT: A Comprehensive Benchmark For Graph Transformers

    cs.LG 2025-06 conditional novelty 5.0 of 10

    OpenGT benchmarks 16 graph models on 14 datasets, finding graph transformers excel on heterophilous graphs, though several observations are not robustly supported.

Pith tools