Pith. sign in

REVIEW 3 major objections 5 minor 75 references

Graph Neural Networks Need Cluster-Normalize-Activate Modules

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Cluster-Normalize-Activate modules stop node features from collapsing in deep graph networks, and the paper reports accuracy gains over state-of-the-art baselines on most benchmark tasks.

desk verdict Useful plug-and-play module with credible controlled gains, but the SOTA claims are undermined by missing split details and leaderboard cherry-picking. read the letter →

arxiv 2412.04064 v1 pith:WZNIAJWH submitted 2024-12-05 cs.LG cs.AI

classification cs.LGcs.AI
keywords graphneuralnetworksoversmoothingcluster-normalize-activatelearnableactivationsrationalnodeclassificationmessagepassing
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Graph neural networks lose accuracy when made deep because node features converge to a common fixed point, a failure known as oversmoothing. This paper proposes Cluster-Normalize-Activate (CNA), a plug-and-play module that replaces the activation function in the update step of any message-passing GNN. CNA clusters node features with k-means, normalizes each cluster separately, then applies a distinct learned rational activation to each cluster, effectively creating per-cluster super nodes. The authors report that CNA improves node-classification accuracy on 8 of 11 benchmark datasets, reaching 94.18% on Cora and 95.75% on CiteSeer, lowers regression error, improves graph classification, and uses substantially fewer parameters than competing architectures. If the claims hold, CNA is a general and parameter-efficient remedy to oversmoothing that makes deeper GNNs practical.

What carries the argument

The central object is the CNA module itself, inserted after the aggregate-and-update computation of any message-passing GNN. Step 1, Cluster, applies hard k-means to the node feature vectors, yielding K clusters that the paper calls super nodes. Step 2, Normalize, standardizes features within each cluster using per-cluster, per-dimension means and variances, without an affine transform. Step 3, Activate, applies a separate learnable rational activation $R(x) = P(x)/Q(x)$ to each cluster, where rationals are smoothly differentiable universal approximators whose unbounded, non-Lipschitz form breaks the assumptions of existing oversmoothing proofs. The module's work is to give different groups of nodes different learned transformations, preserving discriminative information as layers deepen.

What would settle it

Train a GCN with CNA on Cora while fixing the cluster assignment once at initialization instead of re-clustering at every layer. If fixed clusters match the full CNA accuracy, the per-layer clustering step is not what prevents oversmoothing; if accuracy collapses, the claim that fresh cluster assignments carry the benefit is confirmed.

Watch

Extended reading notes

Core claim

The paper argues that oversmoothing can be countered by making the node-feature update adaptive per group of nodes rather than applying one fixed nonlinearity to all nodes. At each layer, node features are partitioned into K clusters by hard k-means; the features in each cluster are then normalized separately per dimension, and each cluster is passed through its own learnable rational activation function. The paper claims this Cluster-Normalize-Activate sequence keeps representations from collapsing to a single point, allows training networks with up to 96 layers without the usual accuracy drop, and outperforms state-of-the-art methods on a majority of node classification, graph classification, and node regression benchmarks while requiring fewer learnable parameters.

Load-bearing premise

The method's gains rest on the assumption that hard k-means cluster assignments at each layer are stable enough that normalizing and activating each cluster separately produces useful, trainable gradients; the paper does not analyze how the non-differentiable, changing assignment affects training.

Editorial extensions

If this is right

  • CNA preserves accuracy in node classification at depths up to 96 layers, where ReLU networks collapse and even linearized GNNs degrade.
  • CNA is a drop-in activation replacement for GCN, GAT, GraphSAGE, TransformerConv, and Dir-GNN, so its benefit is not tied to one architecture.
  • Against the 11-dataset Papers with Code node-classification leaderboard, CNA achieves the best reported accuracy on 8 of 11 datasets, including 94.18% on Cora and 95.75% on CiteSeer.
  • CNA also reduces normalized mean squared error on Chameleon and Squirrel node regression and improves graph-level classification on Mutag, Enzymes, and Proteins.
  • Models with CNA reach higher or comparable accuracy with far fewer learnable parameters than competing deep GNNs, such as 74.64% on ogbn-arxiv with 389.2k parameters.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not analyze the stability of the k-means assignments across epochs; if clusters are noisy, part of the gain could come from normalization and rational activations alone, not from semantically meaningful super nodes.
  • A direct testable extension is to replace hard k-means with differentiable soft clustering; comparable accuracy would suggest the hard assignment is not load-bearing, while a drop would confirm that it is.
  • Because attention layers behave like message passing on complete graphs, CNA-style per-cluster normalization and activation may transfer to Transformers, where token representations similarly converge with depth.
  • Cluster quality itself, for example the silhouette score on learned features, could be measured and correlated with accuracy gains; the paper does not report whether better clusters yield better generalization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes Cluster-Normalize-Activate (CNA), a plug-and-play module for message-passing GNNs that replaces the usual activation function with three steps: hard k-means clustering of node features, per-cluster normalization, and per-cluster learnable rational activations. The authors claim that CNA limits oversmoothing, improves accuracy over the state of the art on node classification and property prediction, reduces parameter counts, and benefits graph-level classification and node regression. The experimental sections report gains over backbone architectures (Tables 1, 2, 3), a leaderboard comparison (Table 4), an ogbn-arxiv parameter-efficiency study (Table 5), and an ablation (Table 6). A short theoretical section (Section 3.4) argues that existing oversmoothing proofs do not apply to CNA.

Significance. If the empirical claims are confirmed, CNA is a simple, architecture-agnostic component with broad potential: it gives consistent gains over its own backbones in controlled experiments, appears to permit much deeper GNNs without the usual performance collapse (Figure 4), and may yield parameter savings (Table 5). The paper ships code and provides a fairly wide evaluation across node classification, node regression, graph classification, and a large OGB dataset. However, the strongest claim in the abstract—state-of-the-art accuracy on Cora and CiteSeer—is not yet established because the evaluation protocol for the leaderboard comparison is incompletely specified, and the theoretical discussion contains a concrete error in its extremal-case analysis. The core ablation in Table 6 is a real strength: it shows all three components are needed for the effect on both Cora and ogbn-arxiv.

major comments (3)
  1. [Section 4, Table 4] The node-classification experiments do not state the train/val/test split. Appendix A.2 (Table 9) lists epochs, layers, clusters, hidden units, learning rates, and weight decay, but no split and no label-rate protocol. Table 1 reports a GCN baseline of 81.59, which matches the standard 20-labels-per-class Cora split, while Table 4 reports 94.18 for the best CNA architecture on Cora. If the CNA numbers were obtained on a different split, the comparison to the PwC leaderboard entries is not apples-to-apples, and the abstract's headline 94.18% and 95.75% figures are unsupported. The authors should report the exact split and label rates for every dataset and confirm that the leaderboard baselines use the same protocol.
  2. [Section 4, Table 4] The 'Best CNA Result' column selects the best of five CNA-equipped architectures per dataset, while the PwC leaderboard entry is a single method. This is a best-of-five selection and introduces winner's bias in the claimed 8/11 win count. The paper should either report results for all CNA-equipped architectures, or fix a single architecture for the leaderboard comparison and report the selection rule explicitly.
  3. [Section 3.4] The extremal-case argument contains a technical error. With K=N clusters, Eq. (1) normalizes each node over a singleton cluster: the mean equals the feature value and the variance is zero, so every normalized feature becomes zero (up to the numerical epsilon in the denominator). This does not 'render the normalization step ineffective' and does not recover the standard MPNN; it collapses all inputs to zero before the activation. The statement that K=1 corresponds to standard MPNN is also inaccurate, since K=1 is global normalization. Consequently, the section does not provide a sound argument that CNA escapes oversmoothing; the practical claim must rest on the empirical results, or the theoretical discussion should be revised.
minor comments (5)
  1. [Section 2] 'Oversquasching' is a typo for 'oversquashing'.
  2. [Section 2] GraphCON is cited to McCallum et al. [2000]; this reference is for the Cora dataset, not for GraphCON. The GraphCON citation should be corrected.
  3. [Section 3.3 and Appendix A.2] Eq. (2) defines the rational activation with numerator degree m and denominator degree n, but Appendix A.2 states 'n=5 for the numerator and m=4 for the denominator'. The notation is inconsistent and should be fixed.
  4. [Section 4, Table 3] The NMSE values for GCN and GAT on Chameleon are identical (0.207) and nearly identical on Squirrel (0.143). This looks suspicious; please double-check the baseline results or clarify whether these are shared or rounded values.
  5. [Section 3.1] The cluster assignment from k-means is non-differentiable and changes across layers. The paper does not state how gradients flow through the clustering step (e.g., whether the assignment is detached and treated as a constant in backpropagation). Please clarify this in the method description.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: CNA's claims are empirical benchmark comparisons, not derivations from fitted parameters or self-citation chains.

full rationale

The paper's central claims are experimental: CNA is a constructively defined module (Eq. 1 for cluster-wise normalization, Eq. 2 for rational activations), and its reported gains come from benchmark evaluations, not from deriving a result from its own fitted constants. There is no 'prediction' that reduces to a fitted parameter: the k-means cluster assignment is a non-differentiable preprocessing/grouping step, and the reported accuracies are measured outcomes, not identities imposed by construction. The SOTA comparison in Table 4 is potentially weakened by selecting the best CNA architecture per dataset and by not stating the train/val/test split, but that is an evaluation-protocol concern, not circularity. Rational activations are cited from prior work, including work by the same research group, but they are used as an off-the-shelf component and are not invoked as a uniqueness theorem or as the sole justification for the method; the paper's own limitations section explicitly acknowledges that no formal link to oversmoothing theory is established. Thus, while the strength of the SOTA claim may be debatable on evidential grounds, the derivation chain is not circular.

Assumptions & free parameters 5 free parameters · 4 assumptions · 1 invented entities

The paper adds a new module with several design choices (number of clusters K, hidden sizes, learning rates, weight decays, rational degrees) that are tuned per dataset. It relies on standard assumptions about k-means clustering, rational activation expressivity, and the trainability of hard assignments. No new physical or mathematical entities are introduced; 'super nodes' is a conceptual grouping.

free parameters (5)
  • Number of clusters K per layer = varies per dataset (e.g., 12 for Cora, 10 for ogbn-arxiv, 8 for Enzymes)
    Chosen by validation; controls granularity of super-nodes. Not derived from first principles.
  • Hidden feature dimension = varies per dataset (e.g., 28 for Cora, 400 for ogbn-arxiv)
    Standard architecture search, not derived from the method.
  • Weight decay = varies per dataset (e.g., 5e-6 for Cora, 1e-4 for ogbn-arxiv)
    Tuned per dataset; listed in Appendix Table 9.
  • Learning rate for activations = 1e-5 for most datasets, 1e-8 for Chameleon and Squirrel
    Tuned per dataset for the rational activation parameters.
  • Rational activation degrees = numerator degree 5, denominator degree 4
    Fixed by the authors; a design choice, not fitted to data.
assumptions (4)
  • domain assumption k-means with Euclidean distance on node features yields stable and meaningful clusters.
    Section 3.1 states Euclidean distance 'worked reliably', but no stability analysis is given; the module's behavior depends on this.
  • standard math Rational activations are universal approximators and trainable end-to-end.
    Section 3.3 relies on properties from Leshno et al. (1993) and Telgarsky (2017), as cited.
  • ad hoc to paper Breaking the assumptions of existing oversmoothing proofs (non-Lipschitz activations) is sufficient to avoid the oversmoothing fixed point in practice.
    Section 3.4 argues this but provides no formal result; the limitations section admits no formal link to oversmoothing theory.
  • domain assumption Hard cluster assignments can be used in a differentiable pipeline without hurting gradient flow.
    Not stated explicitly; the paper uses hard k-means and does not discuss straight-through estimation or gradients through cluster assignments.
invented entities (1)
  • Super nodes
    purpose: Conceptual grouping of nodes with similar features to apply a shared per-cluster transformation.
    No falsifiable prediction attached; it is a framing device for per-cluster activations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Graph Neural Networks Need Cluster-Normalize-Activate Modules." pith.science (2026). https://pith.science/paper/WZNIAJWH

@misc{pith2026241204064,
  author       = {Pith},
  title        = {Pith review of: Graph Neural Networks Need Cluster-Normalize-Activate Modules},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WZNIAJWH}},
  note         = {Machine review of arXiv:2412.04064}
}
read the original abstract

Graph Neural Networks (GNNs) are non-Euclidean deep learning models for graph-structured data. Despite their successful and diverse applications, oversmoothing prohibits deep architectures due to node features converging to a single fixed point. This severely limits their potential to solve complex tasks. To counteract this tendency, we propose a plug-and-play module consisting of three steps: Cluster-Normalize-Activate (CNA). By applying CNA modules, GNNs search and form super nodes in each layer, which are normalized and activated individually. We demonstrate in node classification and property prediction tasks that CNA significantly improves the accuracy over the state-of-the-art. Particularly, CNA reaches 94.18% and 95.75% accuracy on Cora and CiteSeer, respectively. It further benefits GNNs in regression tasks as well, reducing the mean squared error compared to all baselines. At the same time, GNNs with CNA require substantially fewer learnable parameters than competing architectures.

Figures

Figures reproduced from arXiv: 2412.04064 by the authors.

Figure 1
Figure 1. Evolution of node embeddings for the Cora dataset. The colors indicate the membership of one of the seven target classes. Graph Neural Networks (GNNs) are a promising ap￾proach to leveraging the full extent of the geometric properties of various types of data in many different key domains [Zhou et al., 2020a, Bronstein et al., 2021, Waikhom and Patgiri, 2023]. For instance, they are used to predict the stability of … view at source ↗
Figure 2
Figure 2. CNA replaces the activation function in each iteration of any GNN architecture. When employing classical activations like ReLU to all nodes undifferentiatedly, we observe oversmoothing. With CNA, we cluster the node features and then normalize and project them with a separate learned activation function each, effectively increasing their expressiveness even in deeper networks. A natural approach to increasing expres… view at source ↗
Figure 3
Figure 3. The components of CNA modules: They cluster node features without changing the adjacency matrix, normalize them separately, and finally activate with distinct learned functions. features. Instead, we argue that simple hard clustering, for example, provided by the classic k-means algorithm, is sufficient and more desirable. Zhao and Akoglu [2020] suggest PairNorm, where layerwise normalization ensures a constant tota… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: CNA limits oversmoothing and improves the performance of deep GNNs. 4 Experiments To evaluate the effectiveness of CNA with GNNs, we aim to answer the following research questions: (Q1) Does CNA limit oversmoothing? (Q2) Does CNA improve the performance in node classif…
Figure 6
Figure 6. Figure 6: Hyperparameter sensitivity analysis. find that CNA is very robust to the choice of these hyperparameters and works best with moderate numbers of features, as the results from (Q3) would suggest. Answering (Q4), we observed that all three operations of CNA are necessary…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 51 canonical work pages

  1. [1]

    On the Bottleneck of Graph Neural Networks and its Practical Implications

    Uri Alon and Eran Yahav. On the Bottleneck of Graph Neural Networks and its Practical Implications . In The Ninth International Conference on Learning Representations , 2020

  2. [2]

    A survey on modern trainable activation functions

    Andrea Apicella, Francesco Donnarumma, Francesco Isgrò, and Roberto Prevete. A survey on modern trainable activation functions. Neural Networks, 138: 0 14--32, June 2021. doi:10.1016/j.neunet.2021.01.026

  3. [3]

    Heba Askr, Enas Elgeldawi, Heba Aboul Ella, Yaseen A. M. M. Elshaier, Mamdouh M. Gomaa, and Aboul Ella Hassanien. Deep learning in drug discovery: an integrative review and future challenges. Artificial Intelligence Review, 56: 0 5975--6037, July 2023. doi:10.1007/s10462-022-10306-1

  4. [4]

    Battaglia, Jessica B

    Peter W. Battaglia, Jessica B. Hamrick, Victor Bapst, Alvaro Sanchez-Gonzalez, Vinicius Zambaldi, Mateusz Malinowski, Andrea Tacchetti, David Raposo, Adam Santoro, Ryan Faulkner, Caglar Gulcehre, Francis Song, Andrew Ballard, Justin Gilmer, George Dahl, Ashish Vaswani, Kelsey Allen, Charles Nash, Victoria Langston, Chris Dyer, Nicolas Heess, Daan Wierstra...

  5. [5]

    Ahmed Begga, Francisco Escolano, Miguel Angel Lozano, and Edwin R. Hancock. Diffusion- Jump GNNs : Homophiliation via Learnable Metric Filters , June 2023. URL http://arxiv.org/abs/2306.16976

  6. [6]

    Stability of k- Means Clustering

    Shai Ben-David, Dávid Pál, and Hans Ulrich Simon. Stability of k- Means Clustering . In Nader H. Bshouty and Claudio Gentile, editors, Learning Theory , pages 20--34, Berlin, Heidelberg, 2007. Springer. ISBN 978-3-540-72927-3. doi:10.1007/978-3-540-72927-3_4

  7. [7]

    Christopher M. Bishop. Pattern recognition and machine learning. Springer, 2006

  8. [8]

    Learnable Extended Activation Function for Deep Neural Networks

    Yevgeniy Bodyanskiy and Serhii Kostiuk. Learnable Extended Activation Function for Deep Neural Networks . International Journal of Computing, 22 0 (3), October 2023. doi:10.47839/ijc.22.3.3225

Show all 75 references
  1. [9]

    Deep Gaussian Embedding of Graphs : Unsupervised Inductive Learning via Ranking

    Aleksandar Bojchevski and Stephan Günnemann. Deep Gaussian Embedding of Graphs : Unsupervised Inductive Learning via Ranking . In The Sixth International Conference on Learning Representations , 2018

  2. [10]

    Borgwardt, Cheng Soon Ong, Stefan Schönauer, S

    Karsten M. Borgwardt, Cheng Soon Ong, Stefan Schönauer, S. V. N. Vishwanathan, Alex J. Smola, and Hans-Peter Kriegel. Protein function prediction via graph kernels. Bioinformatics (Oxford, England), 21 Suppl 1: 0 i47--56, June 2005. ISSN 1367-4803. doi:10.1093/bioinformatics/bti1007

  3. [11]

    Rational neural networks

    Nicolas Boulle, Yuji Nakatsukasa, and Alex Townsend. Rational neural networks. In Advances in Neural Information Processing Systems , 2020

  4. [12]

    Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković

    Michael M. Bronstein, Joan Bruna, Taco Cohen, and Petar Veličković. Geometric Deep Learning : Grids , Groups , Graphs , Geodesics , and Gauges , May 2021. URL http://arxiv.org/abs/2104.13478v2

  5. [13]

    GraphNorm : A Principled Approach to Accelerating Graph Neural Network Training

    Tianle Cai, Shengjie Luo, Keyulu Xu, Di He, Tie-Yan Liu, and Liwei Wang. GraphNorm : A Principled Approach to Accelerating Graph Neural Network Training . In Proceedings of the 38th International Conference on Machine Learning , 2021. ISSN: 2640-3498

  6. [14]

    Renormalized Graph Neural Networks , June 2023

    Francesco Caso, Giovanni Trappolini, Andrea Bacciu, Pietro Liò, and Fabrizio Silvestri. Renormalized Graph Neural Networks , June 2023

  7. [15]

    Learning graph normalization for graph neural networks

    Yihao Chen, Xin Tang, Xianbiao Qi, Chun-Guang Li, and Rong Xiao. Learning graph normalization for graph neural networks. Neurocomputing, 493: 0 613--625, July 2022. ISSN 0925-2312

  8. [16]

    Cluster- GCN : An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks

    Wei-Lin Chiang, Xuanqing Liu, Si Si, Yang Li, Samy Bengio, and Cho-Jui Hsieh. Cluster- GCN : An Efficient Algorithm for Training Deep and Large Graph Convolutional Networks . In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ...

  9. [17]

    Lopez de Compadre, Gargi Debnath, Alan J

    Asim Kumar Debnath, Rosa L. Lopez de Compadre, Gargi Debnath, Alan J. Shusterman, and Corwin Hansch. Structure-activity relationship of mutagenic aromatic and heteroaromatic nitro compounds. Correlation with molecular orbital energies and hydrophobicity. Journal of Medicinal C...

  10. [18]

    Adaptive Rational Activations to Boost Deep Reinforcement Learning

    Quentin Delfosse, Patrick Schramowski, Martin Mundt, Alejandro Molina, and Kristian Kersting. Adaptive Rational Activations to Boost Deep Reinforcement Learning . In The Twelfth International Conference on Learning Representations , 2024

  11. [19]

    SimTeG : A Frustratingly Simple Approach Improves Textual Graph Learning , August 2023

    Keyu Duan, Qian Liu, Tat-Seng Chua, Shuicheng Yan, Wei Tsang Ooi, Qizhe Xie, and Junxian He. SimTeG : A Frustratingly Simple Approach Improves Textual Graph Learning , August 2023. URL http://arxiv.org/abs/2308.02565

  12. [20]

    Fast Graph Representation Learning with PyTorch Geometric

    Matthias Fey and Jan Eric Lenssen. Fast Graph Representation Learning with PyTorch Geometric . In Representation Learning on Graphs and Manifolds , 2019

  13. [21]

    Schoenholz, Patrick F

    Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for Quantum chemistry. In Proceedings of the 34th International Conference on Machine Learning , volume 70, August 2017

  14. [22]

    M. Gori, G. Monfardini, and F. Scarselli. A new model for learning in graph domains. In IEEE International Joint Conference on Neural Networks , 2005

  15. [23]

    Inductive Representation Learning on Large Graphs

    Will Hamilton, Zhitao Ying, and Jure Leskovec. Inductive Representation Learning on Large Graphs . In Advances in Neural Information Processing Systems , 2017

  16. [24]

    Harnessing Explanations : LLM -to- LM Interpreter for Enhanced Text - Attributed Graph Representation Learning

    Xiaoxin He, Xavier Bresson, Thomas Laurent, Adam Perold, Yann LeCun, and Bryan Hooi. Harnessing Explanations : LLM -to- LM Interpreter for Enhanced Text - Attributed Graph Representation Learning . In The Twelfth International Conference on Learning Representations , 2023

  17. [25]

    Mitigating Degree Biases in Message Passing Mechanism by Utilizing Community Structures , December 2023

    Van Thuy Hoang and O.-Joun Lee. Mitigating Degree Biases in Message Passing Mechanism by Utilizing Community Structures , December 2023. URL https://arxiv.org/abs/2312.16788

  18. [26]

    Open Graph Benchmark : Datasets for Machine Learning on Graphs

    Weihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong, Hongyu Ren, Bowen Liu, Michele Catasta, and Jure Leskovec. Open Graph Benchmark : Datasets for Machine Learning on Graphs . In Advances in Neural Information Processing Systems , 2020

  19. [27]

    Normalization Techniques in Training DNNs : Methodology , Analysis and Application

    Lei Huang, Jie Qin, Yi Zhou, Fan Zhu, Li Liu, and Ling Shao. Normalization Techniques in Training DNNs : Methodology , Analysis and Application . IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (8): 0 10173 -- 10196, August 2023. doi:10.1109/TPAMI.2023.3250241

  20. [28]

    Higher-order Graph Convolutional Network with Flower - Petals Laplacians on Simplicial Complexes

    Yiming Huang, Yujie Zeng, Qiang Wu, and Linyuan Lü. Higher-order Graph Convolutional Network with Flower - Petals Laplacians on Simplicial Complexes . In The Thirty - Eighth AAAI Conference on Artificial Intelligence , January 2024

  21. [29]

    Batch normalization: accelerating deep network training by reducing internal covariate shift

    Sergey Ioffe and Christian Szegedy. Batch normalization: accelerating deep network training by reducing internal covariate shift. In Proceedings of the 32nd International Conference on International Conference on Machine Learning - Volume 37 , ICML '15, pages 448--456, Lille, ...

  22. [30]

    Optimization of Graph Neural Networks with Natural Gradient Descent

    Mohammad Rasool Izadi, Yihao Fang, Robert Stevenson, and Lizhen Lin. Optimization of Graph Neural Networks with Natural Gradient Descent . In Proceedings of the IEEE International Conference on Big Data ( Big Data ) , 2020

  23. [31]

    Graph neural network for traffic forecasting: A survey

    Weiwei Jiang and Jiayun Luo. Graph neural network for traffic forecasting: A survey. Expert Systems with Applications, 207, November 2022. doi:10.1016/j.eswa.2022.117921

  24. [32]

    Reducing Oversmoothing in Graph Neural Networks by Changing the Activation Function

    Dimitrios Kelesis, Dimitrios Vogiatzis, Georgios Katsimpras, Dimitris Fotakis, and Georgios Paliouras. Reducing Oversmoothing in Graph Neural Networks by Changing the Activation Function . In Kobi Gal, Ann Nowé, Grzegorz J. Nalepa, Roy Fairstein, and Roxana Rădulescu, editors,...

  25. [33]

    On the power of graph neural networks and the role of the activation function, May 2024

    Sammy Khalife and Amitabh Basu. On the power of graph neural networks and the role of the activation function, May 2024

  26. [34]

    Kipf and Max Welling

    Thomas N. Kipf and Max Welling. Semi- Supervised Classification with Graph Convolutional Networks . In 5th International Conference on Learning Representations , 2016

  27. [35]

    Mathematical Expressiveness of Graph Neural Networks

    Guillaume Lachaud, Patricia Conde-Cespedes, and Maria Trocan. Mathematical Expressiveness of Graph Neural Networks . Mathematics. Mathematical Foundations of Deep Neural Networks, 10 0 (24): 0 4770, January 2022. doi:10.3390/math10244770

  28. [36]

    Lin, Allan Pinkus, and Shimon Schocken

    Moshe Leshno, Vladimir Ya. Lin, Allan Pinkus, and Shimon Schocken. Multilayer feedforward networks with a nonpolynomial activation function can approximate any function. Neural Networks, 6 0 (6): 0 861--867, January 1993. doi:10.1016/S0893-6080(05)80131-5

  29. [37]

    Training Graph Neural Networks with 1000 Layers

    Guohao Li, Matthias Müller, Bernard Ghanem, and Vladlen Koltun. Training Graph Neural Networks with 1000 Layers . In Proceedings of the 38th International Conference on Machine Learning , 2021

  30. [38]

    Deeper Insights Into Graph Convolutional Networks for Semi - Supervised Learning

    Qimai Li, Zhichao Han, and Xiao-ming Wu. Deeper Insights Into Graph Convolutional Networks for Semi - Supervised Learning . In Proceedings of the AAAI Conference on Artificial Intelligence , volume 32 of 1, April 2018. doi:10.1609/aaai.v32i1.11604

  31. [39]

    Revisiting Heterophily For Graph Neural Networks

    Sitao Luan, Chenqing Hua, Qincheng Lu, Jiaqi Zhu, Mingde Zhao, Shuyuan Zhang, Xiao-Wen Chang, and Doina Precup. Revisiting Heterophily For Graph Neural Networks . Advances in Neural Information Processing Systems, 35: 0 1362--1375, December 2022

  32. [40]

    Distilling Self - Knowledge From Contrastive Links to Classify Graph Nodes Without Passing Messages , June 2021

    Yi Luo, Aiguo Chen, Ke Yan, and Ling Tian. Distilling Self - Knowledge From Contrastive Links to Classify Graph Nodes Without Passing Messages , June 2021. URL http://arxiv.org/abs/2106.08541

  33. [41]

    MacQueen

    J. MacQueen. Some methods for classification and analysis of multivariate observations. Proceedings of the Fifth Berkeley Symposium on Mathematical Statistics and Probability, 1, 1967

  34. [42]

    Introduction to Information Retrieval

    Christopher Manning, Prabhakar Raghavan, and Hinrich Schuetze. Introduction to Information Retrieval . Cambridge University Press, 2009

  35. [43]

    Automating the Construction of Internet Portals with Machine Learning

    Andrew Kachites McCallum, Kamal Nigam, Jason Rennie, and Kristie Seymore. Automating the Construction of Internet Portals with Machine Learning . Information Retrieval, 3: 0 127--163, July 2000. doi:10.1023/A:1009953814988

  36. [44]

    Padé Activation Units : End -to-end Learning of Flexible Activation Functions in Deep Networks

    Alejandro Molina, Patrick Schramowski, and Kristian Kersting. Padé Activation Units : End -to-end Learning of Flexible Activation Functions in Deep Networks . In The Eighth International Conference on Learning Representations , 2019

  37. [45]

    Very Fast EM - Based Mixture Model Clustering Using Multiresolution Kd - Trees

    Andrew Moore. Very Fast EM - Based Mixture Model Clustering Using Multiresolution Kd - Trees . In Advances in Neural Information Processing Systems , 1998

  38. [46]

    TUDataset : A collection of benchmark datasets for learning with graphs

    Christopher Morris, Nils M Kriege, Franka Bause, Kristian Kersting, Petra Mutzel, and Marion Neumann. TUDataset : A collection of benchmark datasets for learning with graphs. In ICML 2020 Workshop on Graph Representation Learning and Beyond ( GRL + 2020) , 2020. URL www.graphl...

  39. [47]

    Predicting basin stability of power grids using graph neural networks

    Christian Nauck, Michael Lindner, Konstantin Schürholt, Haoming Zhang, Paul Schultz, Jürgen Kurths, Ingrid Isenhardt, and Frank Hellmann. Predicting basin stability of power grids using graph neural networks. New Journal of Physics, 24, April 2022. doi:10.1088/1367-2630/ac54c9

  40. [48]

    Revisiting over-smoothing and over-squashing using ollivier-ricci curvature

    Khang Nguyen, Hieu Nong, Vinh Nguyen, Nhat Ho, Stanley Osher, and Tan Nguyen. Revisiting over-smoothing and over-squashing using ollivier-ricci curvature. In Proceedings of the 40th International Conference on Machine Learning , volume 202 of ICML '23 , pages 25956--25979, Hon...

  41. [49]

    Revisiting Graph Neural Networks : All We Have is Low - Pass Filters , May 2019

    Hoang NT and Takanori Maehara. Revisiting Graph Neural Networks : All We Have is Low - Pass Filters , May 2019. URL http://arxiv.org/abs/1905.09550

  42. [50]

    Geom- GCN : Geometric Graph Convolutional Networks

    Hongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei, and Bo Yang. Geom- GCN : Geometric Graph Convolutional Networks . In The Eighth International Conference on Learning Representations . arXiv, February 2020

  43. [51]

    Edge Directionality Improves Learning on Heterophilic Graphs , November 2023

    Emanuele Rossi, Bertrand Charpentier, Francesco Di Giovanni, Fabrizio Frasca, Stephan Günnemann, and Michael Bronstein. Edge Directionality Improves Learning on Heterophilic Graphs , November 2023. URL http://arxiv.org/abs/2305.10498

  44. [52]

    Multi- Scale attributed node embedding

    Benedek Rozemberczki, Carl Allen, and Rik Sarkar. Multi- Scale attributed node embedding. Journal of Complex Networks, 9 0 (2), April 2021. doi:10.1093/comnet/cnab014

  45. [53]

    Graph- Coupled Oscillator Networks

    T Konstantin Rusch, Benjamin P Chamberlain, James Rowbottom, Siddhartha Mishra, and Michael M Bronstein. Graph- Coupled Oscillator Networks . In Proceedings of the 39th International Conference on Machine Learning , volume 162, Baltimore, Maryland, USA, 2022. PMLR

  46. [54]

    Konstantin Rusch, Michael M

    T. Konstantin Rusch, Michael M. Bronstein, and Siddhartha Mishra. A Survey on Oversmoothing in Graph Neural Networks , March 2023 a . URL http://arxiv.org/abs/2303.10993

  47. [55]

    Konstantin Rusch, Benjamin Paul Chamberlain, Michael W

    T. Konstantin Rusch, Benjamin Paul Chamberlain, Michael W. Mahoney, Michael M. Bronstein, and Siddhartha Mishra. Gradient Gating for Deep Multi - Rate Learning on Graphs . In The Eleventh International Conference on Learning Representations , February 2023 b

  48. [56]

    The Graph Neural Network Model

    Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. The Graph Neural Network Model . IEEE Transactions on Neural Networks, 20 0 (1): 0 61 -- 80, December 2008. doi:10.1109/TNN.2008.2005605

  49. [57]

    Collective Classification in Network Data

    Prithviraj Sen, Galileo Namata, Mustafa Bilgic, Lise Getoor, Brian Galligher, and Tina Eliassi-Rad. Collective Classification in Network Data . AI Magazine, 29 0 (3), September 2008. doi:10.1609/aimag.v29i3.2157

  50. [58]

    Pitfalls of Graph Neural Network Evaluation , June 2019

    Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, and Stephan Günnemann. Pitfalls of Graph Neural Network Evaluation , June 2019. URL http://arxiv.org/abs/1811.05868

  51. [59]

    Han Shi, Jiahui Gao, Hang Xu, Xiaodan Liang, Zhenguo Li, Lingpeng Kong, Stephen M. S. Lee, and James Kwok. Revisiting Over -smoothing in BERT from the Perspective of Graph . In The Tenth International Conference on Learning Representations , 2021 a

  52. [60]

    Masked Label Prediction : Unified Message Passing Model for Semi - Supervised Classification

    Yunsheng Shi, Zhengjie Huang, Shikun Feng, Hui Zhong, Wenjing Wang, and Yu Sun. Masked Label Prediction : Unified Message Passing Model for Semi - Supervised Classification . In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence , page 1548,...

  53. [61]

    Exphormer: Sparse Transformers for Graphs

    Hamed Shirzad, Ameya Velingker, Balaji Venkatachalam, Danica J Sutherland, and Ali Kemal Sinop. Exphormer: Sparse Transformers for Graphs . In Proceedings of the 40th International Conference on Machine Learning , 2023

  54. [62]

    ArnetMiner : extraction and mining of academic social networks

    Jie Tang, Jing Zhang, Limin Yao, Juanzi Li, Li Zhang, and Zhong Su. ArnetMiner : extraction and mining of academic social networks. In Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining , August 2008

  55. [63]

    Neural Networks and Rational Functions

    Matus Telgarsky. Neural Networks and Rational Functions . In Proceedings of the 34th International Conference on Machine Learning , pages 3387--3393. PMLR, July 2017

  56. [64]

    Bronstein

    Jake Topping, Francesco Di Giovanni, Benjamin Paul Chamberlain, Xiaowen Dong, and Michael M. Bronstein. Understanding over-squashing and bottlenecks on graphs via curvature. In The Fortieth International Conference on Machine Learning , 2021

  57. [65]

    ERA : Enhanced Rational Activations

    Martin Trimmel, Mihai Zanfir, Richard Hartley, and Cristian Sminchisescu. ERA : Enhanced Rational Activations . In Shai Avidan, Gabriel Brostow, Moustapha Cissé, Giovanni Maria Farinella, and Tal Hassner, editors, Proceedings of the European Conference on Computer Vision , 2022

  58. [66]

    Attention is All you Need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is All you Need . In Advances in Neural Information Processing Systems , 2017

  59. [67]

    Graph Attention Networks

    Petar Veličković, Guillem Cucurull, Arantxa Casanova, Adriana Romero, Pietro Liò, and Yoshua Bengio. Graph Attention Networks . In The Thirty -fifth International Conference on Machine Learning , 2018

  60. [68]

    A survey of graph neural networks in various learning paradigms: methods, applications, and challenges

    Lilapati Waikhom and Ripon Patgiri. A survey of graph neural networks in various learning paradigms: methods, applications, and challenges. Artificial Intelligence Review, 56: 0 6295--6364, July 2023. doi:10.1007/s10462-022-10321-2

  61. [69]

    Graph Neural Networks for Molecules

    Yuyang Wang, Zijie Li, and Amir Barati Farimani. Graph Neural Networks for Molecules . Machine Learning in Molecular Sciences, 36: 0 21--66, 2023

  62. [70]

    Simplifying Graph Convolutional Networks

    Felix Wu, Amauri Souza, Tianyi Zhang, Christopher Fifty, Tao Yu, and Kilian Weinberger. Simplifying Graph Convolutional Networks . In Proceedings of the 36th International Conference on Machine Learning , 2019. ISBN 2640-3498

  63. [71]

    Link Prediction Based on Graph Neural Networks

    Muhan Zhang and Yixin Chen. Link Prediction Based on Graph Neural Networks . In Advances in Neural Information Processing Systems , 2018

  64. [72]

    PairNorm : Tackling Oversmoothing in GNNs

    Lingxiao Zhao and Leman Akoglu. PairNorm : Tackling Oversmoothing in GNNs . In International Conference on Learning Representations ( ICLR ) , 2020

  65. [73]

    Graph neural networks: A review of methods and applications

    Jie Zhou, Ganqu Cui, Shengding Hu, Zhengyan Zhang, Cheng Yang, Zhiyuan Liu, Lifeng Wang, Changcheng Li, and Maosong Sun. Graph neural networks: A review of methods and applications. AI Open, 1: 0 57--81, January 2020 a . doi:10.1016/j.aiopen.2021.01.001

  66. [74]

    Towards deeper graph neural networks with differentiable group normalization

    Kaixiong Zhou, Xiao Huang, Yuening Li, Daochen Zha, Rui Chen, and Xia Hu. Towards deeper graph neural networks with differentiable group normalization. In Proceedings of the 34th International Conference on Neural Information Processing Systems , NIPS '20, pages 4917--4928, Re...

  67. [75]

    Deep Graph Contrastive Representation Learning , 2020

    Yanqiao Zhu, Yichen Xu, Feng Yu, Qiang Liu, Shu Wu, and Liang Wang. Deep Graph Contrastive Representation Learning , 2020

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.