REVIEW 4 major objections 6 minor 155 references
Asymmetric Learning for Spectral Graph Neural Networks
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Spectral GNNs train better when their two parameter blocks are updated at matched speeds; the paper proves this via a reduced block condition number and shows consistent accuracy gains on eighteen datasets.
desk verdict Useful empirical trick with a broken proof; the LARS-like preconditioner consistently helps heterophilic GNN training, but the claimed block-condition-number reduction is not established. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The engine of the argument is the block condition number of the Hessian, $\kappa'(H)$, defined as the ratio of the larger to the smaller of the two largest eigenvalues of the diagonal blocks for the filter parameter $\Theta$ and the feature parameter $W$. The method itself is the asymmetric preconditioner $R_t$: before each optimizer step, the gradient is rescaled block-wise by $s_{\Theta}=\|\Theta\|/\|\nabla_{\Theta}L\|$ and $s_W=\|W\|/\|\nabla_W L\|$, so the gradient-parameter norm ratios of both blocks become 1. This rescaling acts as a Hessian preconditioner $H'_t=R_tH_t$, and Theorem 17 shows it shrinks $\kappa'$ when the stated assumptions hold. The paper reads this as making the two parameter groups equally sensitive to updates, so neither block dominates the landscape.
What would settle it
Compute the block condition number at every iteration during ChebNet training on Texas with and without asymmetric learning. If $\kappa'(H')$ is not smaller than $\kappa'(H)$ in the early iterations while the method still improves final accuracy, the reduction in block condition number is not the mechanism explaining the gains.
Extended reading notes
Core claim
The paper's central claim is that the poor optimization of spectral GNNs is caused by unequal curvature in the two Hessian blocks for filter parameters $\Theta$ and transformation parameters $W$, and that this can be fixed by asymmetric gradient preconditioning. Concretely, it defines the block condition number $\kappa'(H)=\max(\lambda_{\max}(H_{\Theta,\Theta}),\lambda_{\max}(H_{W,W}))/\min(\lambda_{\max}(H_{\Theta,\Theta}),\lambda_{\max}(H_{W,W}))$, proves that applying the diagonal preconditioner $R_t=\mathrm{diag}(s_{\Theta} I_{d_\Theta}, s_W I_{d_W})$ with $s_{\Theta}=\|\Theta\|/\|\nabla_{\Theta}L\|$ and $s_W=\|W\|/\|\nabla_W L\|$ makes $\kappa'(H'_t)\leq \kappa'(H_t)$ under Assumptions 7, 11, and 13, and reports consistent accuracy improvements when this rescaling is inserted before the optimizer on eighteen datasets. The empirical pattern it stresses is that heterophilic graphs have larger block condition numbers and receive the largest improvements, sometimes more than ten accuracy points on small graphs.
Load-bearing premise
The load-bearing premise is that the ordering of largest Hessian eigenvalues between two nearby points is mirrored by the ordering of gradient-parameter norm ratios; the paper's Table 8 shows this premise fails on Cornell, Actor, Chameleon, and Squirrel, so the proof's guarantee does not cover those cases.
Editorial extensions
If this is right
- Asymmetric learning can be added to any optimizer, so existing spectral GNN baselines improve without architectural changes; the paper demonstrates this with Adam and five spectral models.
- Heterophilic graphs, where the block condition number is largest, stand to gain the most; average improvements are about 3.08 accuracy points on six small heterophilic datasets and about 0.46 points on homophilic ones.
- Small graphs with sparse training labels show the largest gains, with improvements above 10 points on Texas, Wisconsin, Cornell, and Chameleon, suggesting the method helps when gradient estimates are noisy.
- Orthogonal polynomial-basis GNNs benefit more than non-orthogonal ones, consistent with the paper's view that interference among filter coefficients weakens a single uniform scaling.
Reading between the lines
- The same block-condition view could extend to other architectures with heterogeneous parameter groups, such as GNNs with separate encoder and classifier blocks; the paper does not test this.
- The paper's own Table 8 shows Assumption 11 fails on Cornell, Actor, Chameleon, and Squirrel, so the proof's guarantee does not cover those datasets; the consistent accuracy gains there suggest the method may still work when the theorem does not apply, but the theoretical explanation would need a weaker assumption.
- A direct test of the causal story would be to check whether the measured block condition number actually drops at every step on a dataset where the assumption holds; the paper provides empirical support on some datasets but not a per-iteration trace for all.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies optimization of spectral GNNs whose parameters split into graph-filter parameters Θ and feature-transformation parameters W. It introduces a “block condition number” of the Hessian based on the largest eigenvalues of the two diagonal Hessian blocks, argues that spectral GNNs—especially on heterophilic graphs—are poorly conditioned in this sense, and proposes an asymmetric preconditioner that scales the Θ and W gradient blocks by the ratios of parameter norms to gradient norms. The authors claim (Theorem 17) that under Assumptions 7, 11, and 13 this preconditioning reduces the block condition number, and they report accuracy gains on 18 benchmark datasets with five spectral GNN baselines. The empirical study is broad and the method is simple to implement, but the theoretical argument as written has load-bearing gaps.
Significance. If the theoretical claim were established, the paper would provide a principled explanation for the common practice of using different effective learning rates for Θ and W in spectral GNNs, and it would offer a cheap training-time modification with clear practical payoff. The empirical contribution is substantial: 18 datasets, 5 spectral GNN architectures, code released, and consistent positive average gains that are especially large on small heterophilic graphs. However, the central theorem is not proven as stated: the proof uses an assumption in a regime it does not cover, the empirical validation tests a different quantity than the proof requires, and the preconditioned object H'=RH is not the Hessian of the actual transformed optimization problem. These issues are fixable in a revision, but they are central rather than cosmetic.
major comments (4)
- [§3, Eq. (9) and Theorem 17] The paper states that preconditioning the gradient with R_t is equivalent to preconditioning the Hessian as H' = R_t H_t. This is not the standard equivalence. A block-diagonal positive scaling of the gradient corresponds, under a linear reparameterization, to a congruent transformation of the Hessian, H' = R_t^{1/2} H_t R_t^{1/2} (with the usual caveats about symmetrization), not to the left multiplication H' = R_t H_t. The matrix R_t H_t is generally non-symmetric, so the subsequent use of eigenvalue interlacing and the interpretation of the block condition number as a property of a Hessian are not justified. The authors should either prove the reduction for the Hessian that actually governs the preconditioned dynamics, or explicitly reframe the theorem as a statement about the diagonal blocks of R_t H_t without calling it the Hessian of the loss.
- [Proof of Theorem 17, Eq. (26)] The decisive step in the proof of Theorem 17 is the implication: if λmax(H_t_{Θ,Θ}) ≥ λmax(H_t_{W,W}), then ρ_t_Θ ≥ ρ_t_W, which is then used to conclude s_t_Θ / s_t_W ≤ 1. The paper justifies this by invoking Assumption 11. But Assumption 11 is a statement about the full Hessian eigenvalues λmax(H_Ψ) and full-parameter GPNRs at two distinct nearby points Ψ and Ψ'. It does not relate the two diagonal-block eigenvalues λmax(H_Θ,Θ) and λmax(H_W,W) at a single point. Proposition 9 only provides separate upper bounds ρ_Θ ≤ λmax(H_Θ,Θ) and ρ_W ≤ λmax(H_W,W); these bounds cannot yield the required cross-parameter implication. Therefore Eq. (26) does not follow from the stated assumptions, and the reduction κ'(H'^t) ≤ κ'(H^t) is not established.
- [Table 8 and Assumption 11] The empirical validation of Assumption 11 in Table 8 is not a test of the assumption as used in Theorem 17. The table compares full-Hessian eigenvalues and full GPNRs at two noise-perturbed points, whereas the proof needs a monotonicity relation between the two diagonal block eigenvalues at the single current point. Moreover, even the stated full-Hessian version of Assumption 11 is violated on Chameleon: λmax(H_Ψ1) > λmax(H_Ψ2) but ρ_Ψ1 < ρ_Ψ2. Chameleon is one of the datasets where asymmetric learning gives the largest reported gains (Table 1), so the theoretical result does not cover a primary success case. The authors should either repair the theorem so it no longer requires this implication, or state and validate the actual block-level assumption used in Eq. (26).
- [Assumption 13 and the structure of the proof] The role of Assumption 13 in the proof is to preserve the ordering of the two block eigenvalues after preconditioning, but the actual reduction κ'(H'^t) ≤ κ'(H^t) is obtained from Eq. (25) together with the upper bound s_t_Θ / s_t_W ≤ 1. Since s_t_Θ / s_t_W = ρ_t_W / ρ_t_Θ by Eq. (6), the theorem is essentially equivalent to assuming that the parameter block with the larger Hessian eigenvalue also has the larger GPNR. This property is the real content of the theorem, and it is not derived from Assumption 11 or from Proposition 9. The authors should make this dependence explicit and provide empirical evidence for the specific block-level GPNR ordering rather than for a different assumption.
minor comments (6)
- [§3, Eq. (8)] Equation (8) as printed, "[∇ΘLS; ∇WLS] = R_t[∇ΘLS; ∇WLS]", is false unless R_t is the identity; it should introduce a new symbol for the preconditioned gradient or use an assignment arrow, e.g., "[∇̃Θ; ∇̃W] ← R_t[∇Θ; ∇W]".
- [Remark 6] In the displayed formulas for ρ_t_Θ and ρ_t_W, the denominators are written as ‖Θ‖2 and ‖W‖2 without the time index t; for consistency with Definition 5 and Eq. (5) they should be ‖Θ^t‖2 and ‖W^t‖2.
- [Proof of Theorem 17] The last line of the proof states κ′(H′^t) ≤ κ(H^t); the right-hand side should be κ′(H^t), since the theorem compares block condition numbers.
- [Table 8] Several entries in Table 8 contain stray spaces (e.g., "0 .016569"), and the table caption should state explicitly which datasets violate Assumption 11; Chameleon is a violation, not merely a case where the assumption is untested.
- [Abstract and Q1] The claim that asymmetric learning "consistently improves" performance is stronger than the tables support: Table 2 contains negative deltas for ChebNetII on Cora, GPRGNN on Cora, and BernNet on Coauthor-Physics, and Table 6 contains negative deltas for ChebNetII on Questions and BernNet on Questions. The authors should qualify "consistently" or report a paired statistical test.
- [Figure 2] The caption of Figure 2 refers to "ChbNet," which is a typo for ChebNet.
Circularity Check
Theorem 17's block-condition reduction is built into the definitions: s := 1/ρ (Eq. 6) makes the reduction factor (Eq. 25) equal to the GPNR ratio ρW/ρΘ, and the proof invokes Assumption 11 — whose needed reading is exactly the conclusion — for that ratio being ≤ 1. Only the benchmark gains are independent.
-
self definitional
[Asymmetric Learning section, 'Asymmetric Preconditioner' (Eq. 6); Theorem 17 proof, Eqs. (25)-(26)]
"stΘ = ∥Θt∥2/∥∇ΘLS(Θt, Wt)∥2; stW = ∥Wt∥2/∥∇WLS(Θt, Wt)∥2. (6) ... κ′(H′t) = stΘλmax(HtΘ,Θ)/stWλmax(HtW,W) = (stΘ/stW)κ′(Ht). (25) ... According to Assumption 11, when λmax(HtΘ,Θ) ≥ λmax(HtW,W), we have ρtΘ ≥ ρtW ... stΘ/stW = ... ≤ 1. (26) ... By substituting Eq. (26) into Eq. (25), we obtain κ′(H′t) ≤ κ(Ht)."
By Eq. (6) the preconditioner is the inverse GPNR: sΘ = 1/ρΘ and sW = 1/ρW, so sΘ/sW = ρW/ρΘ. By the paper's own Eq. (25), κ′(H′t) = (sΘ/sW)·κ′(Ht); hence the claimed reduction κ′(H′t) ≤ κ′(Ht) is equivalent, by construction, to ρΘ ≥ ρW (in the case λmax(HΘΘ) ≥ λmax(HWW)). The proof's only substantive move is to assert this inequality 'according to Assumption 11'. Since Assumption 11 is exactly the statement that GPNRs track the largest eigenvalues, the theorem's conclusion is the assumption restated: the reduction factor is defined as the inverse-GPNR ratio, and the proof assumes that ratio ≤ 1. No independent derivation of the GPNR ordering is given, so the theoretical claim is built into the definitions of s and of the assumption.
-
other
[Appendix 'Proofs': Assumption 11 and Theorem 17 proof; Appendix Table 8]
"Assumption 11 (Proportional GPNR and Maximum Eigenvalue). Let Ψ and Ψ′ be two points near a critical point Ψ∗ such that ∥Ψ − Ψ∗∥2 = ϵ1 and ∥Ψ′ − Ψ∗∥2 = ϵ2, where ϵ1 → 0 and ϵ2 → 0. Denote HΨ and HΨ′ as the Hessian matrices of the empirical loss at points Ψ and Ψ′, respectively. If λmax(HΨ) ≥ λmax(HΨ′), then ρΨ ≥ ρΨ′."
The step 'when λmax(HtΘ,Θ) ≥ λmax(HtW,W), we have ρtΘ ≥ ρtW' is not an instance of Assumption 11 as stated: the assumption relates full-Hessian largest eigenvalues at two distinct points Ψ and Ψ′, while the proof needs an implication between the two diagonal-block eigenvalues of one Hessian at a single point. Proposition 9 supplies only upper bounds ρΘ ≤ λmax(HΘΘ) and ρW ≤ λmax(HWW), from which the needed cross-parameter implication cannot be derived. The only way to make the proof go through is to read Assumption 11 as asserting the block-eigenvalue/GPNR proportionality — that is, to assume the theorem's conclusion.
full rationale
This paper has two independent strands: a theoretical claim that asymmetric learning reduces the block-condition number of the Hessian, and an empirical evaluation on eighteen benchmarks. The empirical strand is not circular: the accuracy/ROC-AUC comparisons in Tables 1, 2, and 6 are standard external benchmarks computed from the released code, and they do not rely on Theorem 17; they support the abstract's empirical claim on their own. The theoretical strand is circular in a specific, quotable way. Definition 5 defines GPNR ρ = ∥∇L∥/∥θ∥; Eq. (6) defines the preconditioner as s = 1/ρ; Eq. (25) shows the post-preconditioning block condition number is κ′(H′) = (sΘ/sW)·κ′(H) = (ρW/ρΘ)·κ′(H). Therefore the claimed reduction κ′(H′) ≤ κ′(H) is, by the paper's own equations, equivalent to the ordering ρΘ ≥ ρW. The proof of Theorem 17 obtains exactly this ordering by invoking Assumption 11 ('Proportional GPNR and Maximum Eigenvalue'). As literally stated, Assumption 11 compares full-Hessian largest eigenvalues at two different points and cannot license the needed statement about the two diagonal blocks of one Hessian (Proposition 9 gives only upper bounds); the only reading under which the proof works is that GPNR ordering follows block-eigenvalue ordering — which is the theorem's conclusion. The reduction thus reduces to its own assumption, and Theorem 17's content is the assumption restated through the definitions. The paper's own Table 8 further reports violations of the stated assumption on Cornell, Actor, and Chameleon, and the Appendix concedes it holds only 'on most datasets'. There are no load-bearing self-citations: all cited results (Taylor expansion, Weyl's inequality, the Hessian learning-rate bound from Granziol et al.) are standard external results. Because the central theoretical claim reduces by construction while the empirical results are genuinely independent, the appropriate score is 6 (partial circularity), not 8-10.
Assumptions & free parameters
free parameters (2)
- beta_pi_Theta
- beta_pi_W
assumptions (4)
- domain assumption Assumption 7 (Point Proximity): near a critical point, the distance to the critical point is at most the parameter norm.
- ad hoc to paper Assumption 11 (Proportional GPNR and Maximum Eigenvalue): if one point has a larger max Hessian eigenvalue than another nearby point, then its GPNR is also larger.
- ad hoc to paper Assumption 13 (Mild Scaling): the preconditioner ratio s_Theta/s_W is at least the inverse of the block eigenvalue ratio.
- domain assumption The Hessian matrix is positive semi-definite (implicitly assumed for the interlacing argument and for defining block condition number).
Cite this review
Pith. "Pith review of Asymmetric Learning for Spectral Graph Neural Networks." pith.science (2026). https://pith.science/paper/4MXK7S4P
@misc{pith2026241211739,
author = {Pith},
title = {Pith review of: Asymmetric Learning for Spectral Graph Neural Networks},
year = {2026},
howpublished = {\url{https://pith.science/paper/4MXK7S4P}},
note = {Machine review of arXiv:2412.11739}
}
read the original abstract
Optimizing spectral graph neural networks (GNNs) remains a critical challenge in the field, yet the underlying processes are not well understood. In this paper, we investigate the inherent differences between graph convolution parameters and feature transformation parameters in spectral GNNs and their impact on the optimization landscape. Our analysis reveals that these differences contribute to a poorly conditioned problem, resulting in suboptimal performance. To address this issue, we introduce the concept of the block condition number of the Hessian matrix, which characterizes the difficulty of poorly conditioned problems in spectral GNN optimization. We then propose an asymmetric learning approach, dynamically preconditioning gradients during training to alleviate poorly conditioned problems. Theoretically, we demonstrate that asymmetric learning can reduce block condition numbers, facilitating easier optimization. Extensive experiments on eighteen benchmark datasets show that asymmetric learning consistently improves the performance of spectral GNNs for both heterophilic and homophilic graphs. This improvement is especially notable for heterophilic graphs, where the optimization process is generally more complex than for homophilic graphs. Code is available at https://github.com/Mia-321/asym-opt.git.
Figures
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Morris, Christopher and Frasca, Fabrizio and Dym, Nadav and Maron, Haggai and Ceylan, Ismail Ilkan and Levie, Ron and Lim, Derek and Bronstein, Michael M and Grohe, Martin and Jegelka, Stefanie Position: Future Directions in the Theory of Graph Machine Learning In: ICML
-
[4]
Jin, Chi and Ge, Rong and Netrapalli, Praneeth and Kakade, Sham M and Jordan, Michael I How to escape saddle points efficiently In: ICML (2017)
2017
-
[5]
Awasthi, Pranjal and Das, Abhimanyu and Gollapudi, Sreenivas A convergence analysis of gradient descent on graph neural networks In: NeurIPS (2021)
2021
-
[6]
Kaddour, Jean and Liu, Linqing and Silva, Ricardo and Kusner, Matt J When do flat minima optimizers work? In: NeurIPS (2022)
2022
-
[7]
Zhang, Yushun and Chen, Congliang and Ding, Tian and Li, Ziniu and Sun, Ruoyu and Luo, Zhi-Quan Why transformers need adam: A hessian perspective In: arXiv preprint arXiv:2402.16788 (2024)
arXiv 2024
-
[8]
Tao, Terence Topics in random matrix theory (2012) Volume: 132
2012
Show all 155 references
-
[9]
Zeng, Hanqing and Zhou, Hongkuan and Srivastava, Ajitesh and Kannan, Rajgopal and Prasanna, Viktor GraphSAINT: Graph Sampling Based Inductive Learning Method In: ICML (2020)
2020
-
[10]
Horn, Roger A and Johnson, Charles R Matrix analysis (2012)
2012
-
[11]
Liao, Renjie and Zhao, Zhizhen and Urtasun, Raquel and Zemel, Richard LanczosNet: Multi-Scale Deep Graph Convolutional Networks In: ICLR (2019)
2019
-
[12]
Bianchi, Filippo Maria and Grattarola, Daniele and Livi, Lorenzo and Alippi, Cesare Graph neural networks with convolutional arma filters IEEE transactions on pattern analysis and machine intelligence (2021) Volume: 44 Pages: 3496--3507
2021
-
[13]
Bronstein CayleyNets: Graph Convolutional Neural Networks With Complex Rational Spectral Filters IEEE Transactions on Signal Processing (2019) Volume: 67 Pages: 97-109
Ron Levie and Federico Monti and Xavier Bresson and Michael M. Bronstein CayleyNets: Graph Convolutional Neural Networks With Complex Rational Spectral Filters IEEE Transactions on Signal Processing (2019) Volume: 67 Pages: 97-109
2019
-
[14]
Golub, Gene H and Van Loan, Charles F Matrix computations (2013)
2013
-
[17]
Shalev-Shwartz, Shai and Ben-David, Shai Understanding machine learning: From theory to algorithms (2014)
2014
-
[18]
Fike, Jeffrey and Jongsma, Sietse and Alonso, Juan and Van Der Weide, Edwin Optimization with gradient and hessian information calculated using hyper-dual numbers In: AIAA Applied Aerodynamics Conference (2011) Pages: 3807
2011
-
[19]
Jia, Xixi and Wang, Hailin and Peng, Jiangjun and Feng, Xiangchu and Meng, Deyu Preconditioning Matters: Fast Global Convergence of Non-convex Matrix Factorization via Scaled Gradient Descent In: NeurIPS (2024)
2024
-
[20]
Dong, Mingze and Kluger, Yuval Towards understanding and reducing graph structural noise for GNNs In: ICML (2023)
2023
-
[21]
Platonov, Oleg and Kuznedelev, Denis and Diskin, Michael and Babenko, Artem and Prokhorenkova, Liudmila A critical look at the evaluation of GNNs under heterophily: Are we really making progress? In: ICLR (2023)
2023
-
[22]
Li, Hao and Xu, Zheng and Taylor, Gavin and Studer, Christoph and Goldstein, Tom Visualizing the loss landscape of neural nets In: NeurIPS (2018)
2018
-
[23]
Bonnans, Joseph-Fr \'e d \'e ric and Gilbert, Jean Charles and Lemar \'e chal, Claude and Sagastiz \'a bal, Claudia A Numerical optimization: theoretical and practical aspects (2006)
2006
-
[24]
Goldfarb, Donald and Ren, Yi and Bahamou, Achraf Practical quasi-newton methods for training deep neural networks In: NeurIPS (2020)
2020
-
[25]
Lewis and Michael L
Adrian S. Lewis and Michael L. Overton Nonsmooth optimization via quasi-Newton methods Mathematical Programming (2012) Volume: 141 Pages: 135 - 163
2012
-
[26]
Vineet Gupta and Tomer Koren and Yoram Singer Shampoo: Preconditioned Stochastic Tensor Optimization In: ICML (2018)
2018
-
[27]
Yen, Jui-Nan and Duvvuri, Sai Surya and Dhillon, Inderjit and Hsieh, Cho-Jui Block Low-Rank Preconditioner with Shared Basis for Stochastic Optimization In: NeurIPS (2024)
2024
-
[28]
Guo, Yuhe and Wei, Zhewei Graph neural networks with learnable and optimal polynomial bases In: ICML (2023)
2023
-
[29]
Loukas, Andreas What graph neural networks cannot learn: depth vs width In: ICLR (2020)
2020
-
[30]
Dwivedi, Vijay Prakash and Joshi, Chaitanya K and Luu, Anh Tuan and Laurent, Thomas and Bengio, Yoshua and Bresson, Xavier Benchmarking graph neural networks Journal of Machine Learning Research (2023) Volume: 24 Pages: 1--48
2023
-
[31]
Velickovic, Petar and Cucurull, Guillem and Casanova, Arantxa and Romero, Adriana and Lio, Pietro and Bengio, Yoshua Graph attention networks In: ICLR (2017)
2017
-
[32]
Xu, Keyulu and Hu, Weihua and Leskovec, Jure and Jegelka, Stefanie How powerful are graph neural networks? In: ICLR (2019)
2019
-
[33]
Zhu, Jiong and Yan, Yujun and Zhao, Lingxiao and Heimann, Mark and Akoglu, Leman and Koutra, Danai Beyond homophily in graph neural networks: Current limitations and effective designs In: NeurIPS (2020)
2020
-
[34]
Michael Yu Wang and Xiaoming Wang and Dong-ming Guo A level set method for structural topology optimization Computer Methods in Applied Mechanics and Engineering (2003) Volume: 192 Pages: 227-246
2003
-
[35]
Yurii Nesterov Efficiency of Coordinate Descent Methods on Huge-Scale Optimization Problems SIAM J. Optim. (2012) Volume: 22 Pages: 341-362
2012
-
[36]
Amur Ghose and Yingxue Zhang and Jianye Hao and Mark Coates Spectral Augmentations for Graph Contrastive Learning In: AISTATS (2023)
2023
-
[37]
Yang, Mingqi and Shen, Yanming and Li, Rui and Qi, Heng and Zhang, Qiang and Yin, Baocai A new perspective on the effects of spectrum in graph neural networks In: ICML (2022)
2022
-
[38]
Levie, Ron and Monti, Federico and Bresson, Xavier and Bronstein, Michael M Cayleynets: Graph convolutional neural networks with complex rational spectral filters IEEE Transactions on Signal Processing (2018) Volume: 67 Pages: 97--109
2018
-
[39]
Wu, Felix and Souza, Amauri and Zhang, Tianyi and Fifty, Christopher and Yu, Tao and Weinberger, Kilian Simplifying graph convolutional networks In: ICML (2019)
2019
-
[40]
Chen, Ming and Wei, Zhewei and Huang, Zengfeng and Ding, Bolin and Li, Yaliang Simple and deep graph convolutional networks In: ICML (2020)
2020
-
[41]
Yu A Comprehensive Survey on Graph Neural Networks IEEE Transactions on Neural Networks and Learning Systems (2019) Volume: 32 Pages: 4-24
Zonghan Wu and Shirui Pan and Fengwen Chen and Guodong Long and Chengqi Zhang and Philip S. Yu A Comprehensive Survey on Graph Neural Networks IEEE Transactions on Neural Networks and Learning Systems (2019) Volume: 32 Pages: 4-24
2019
-
[42]
M. Balcilar and Guillaume Renton and Pierre H \'e roux and Benoit Ga \"u z \`e re and S \'e bastien Adam and Paul Honeine Analyzing the Expressive Power of Graph Neural Networks in a Spectral Perspective In: ICLR (2021)
2021
-
[43]
Petar Velickovic and Guillem Cucurull and Arantxa Casanova and Adriana Romero and Pietro Lio’ and Yoshua Bengio Graph Attention Networks In: ICLR (2018)
2018
-
[44]
Wang, Xiyuan and Zhang, Muhan How powerful are spectral graph neural networks In: ICML (2022) Pages: 23341--23362
2022
-
[45]
Defferrard, Micha \"e l and Bresson, Xavier and Vandergheynst, Pierre Convolutional neural networks on graphs with fast localized spectral filtering In: NeurIPS (2016)
2016
-
[46]
Kipf, Thomas N and Welling, Max Semi-Supervised Classification with Graph Convolutional Networks In: ICLR (2016)
2016
-
[47]
Gasteiger, Johannes and Bojchevski, Aleksandar and G \"u nnemann, Stephan Predict then propagate: Graph neural networks meet personalized pagerank In: ICLR (2019)
2019
-
[48]
Hu and Igor Babuschkin and Szymon Sidor and Xiaodong Liu and David Farhi and Nick Ryder and Jakub W
Ge Yang and Edward J. Hu and Igor Babuschkin and Szymon Sidor and Xiaodong Liu and David Farhi and Nick Ryder and Jakub W. Pachocki and Weizhu Chen and Jianfeng Gao Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer In: NeurIPS (2021)
2021
-
[49]
Das, Rudrajit and Agarwal, Naman and Sanghavi, Sujay and Dhillon, Inderjit S Towards Quantifying the Preconditioning Effect of Adam arXiv preprint arXiv:2402.07114 (2024)
2024 arXiv
-
[50]
Peter Richt \'a rik and Martin Tak \'a c Parallel coordinate descent methods for big data optimization Mathematical Programming (2012) Volume: 156 Pages: 433 - 484
2012
-
[51]
Wright Coordinate descent algorithms Mathematical Programming (2015) Volume: 151 Pages: 3 - 34
Stephen J. Wright Coordinate descent algorithms Mathematical Programming (2015) Volume: 151 Pages: 3 - 34
2015
-
[52]
Trends Mach
S \'e bastien Bubeck Convex Optimization: Algorithms and Complexity Found. Trends Mach. Learn. (2014) Volume: 8 Pages: 231-357
2014
-
[53]
Hammond, David K and Vandergheynst, Pierre and Gribonval, R \'e mi Wavelets on graphs via spectral graph theory Applied and Computational Harmonic Analysis (2011) Volume: 30 Pages: 129--150
2011
-
[54]
Stevenson and Lizhen Lin Optimization of Graph Neural Networks with Natural Gradient Descent In: Big Data (2020)
Mohammad Rasool Izadi and Yihao Fang and Robert L. Stevenson and Lizhen Lin Optimization of Graph Neural Networks with Natural Gradient Descent In: Big Data (2020)
2020
-
[55]
Tianlong Chen and Kaixiong Zhou and Keyu Duan and Wenqing Zheng and Peihao Wang and Xia Hu and Zhangyang Wang Bag of Tricks for Training Deeper Graph Neural Networks: A Comprehensive Benchmark Study IEEE Transactions on Pattern Analysis and Machine Intelligence (2021) Volume: ...
2021
-
[56]
Keyulu Xu and Mozhi Zhang and Stefanie Jegelka and Kenji Kawaguchi Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More Depth In: ICML (2021)
2021
-
[57]
Wright and Azalia Mirhoseini and Joseph Gonzalez and Ion Stoica Representing Long-Range Context for Graph Neural Networks with Global Attention ArXiv (2022)
Zhanghao Wu and Paras Jain and Matthew A. Wright and Azalia Mirhoseini and Joseph Gonzalez and Ion Stoica Representing Long-Range Context for Graph Neural Networks with Global Attention ArXiv (2022)
2022
-
[58]
Wu, Zhanghao and Jain, Paras and Wright, Matthew and Mirhoseini, Azalia and Gonzalez, Joseph E and Stoica, Ion Representing long-range context for graph neural networks with global attention In: NeurIPS (2021)
2021
-
[59]
Zhuang, Juntang and Tang, Tommy and Ding, Yifan and Tatikonda, Sekhar C and Dvornek, Nicha and Papademetris, Xenophon and Duncan, James Adabelief optimizer: Adapting stepsizes by the belief in observed gradients NeurIPS (2020)
2020
-
[60]
Liu, Yanli and Zhang, Kaiqing and Basar, Tamer and Yin, Wotao An improved analysis of (variance-reduced) policy gradient and natural policy gradient methods NeurIPS (2020)
2020
-
[61]
Nguyen, Lam M and Liu, Jie and Scheinberg, Katya and Tak \'a c , Martin SARAH: A novel method for machine learning problems using stochastic recursive gradient In: ICML (2017)
2017
-
[62]
Bach and Simon Lacoste-Julien SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives In: NeurIPS (2014)
Aaron Defazio and Francis R. Bach and Simon Lacoste-Julien SAGA: A Fast Incremental Gradient Method With Support for Non-Strongly Convex Composite Objectives In: NeurIPS (2014)
2014
-
[63]
Zeyuan Allen-Zhu and Elad Hazan Variance Reduction for Faster Non-Convex Optimization In: ICML (2016)
2016
-
[64]
Rie Johnson and Tong Zhang Accelerating Stochastic Gradient Descent using Predictive Variance Reduction In: NeurIPS (2013)
2013
-
[65]
Cutkosky, Ashok and Orabona, Francesco Momentum-based variance reduction in non-convex sgd NeurIPS (2019)
2019
-
[66]
Schmidt and Francis R
Robert Mansel Gower and Mark W. Schmidt and Francis R. Bach and Peter Richt \'a rik Variance-Reduced Methods for Machine Learning Proceedings of the IEEE (2020) Volume: 108 Pages: 1968-1983
2020
-
[67]
Yang, Zhilin and Cohen, William and Salakhudinov, Ruslan Revisiting semi-supervised learning with graph embeddings In: ICML (2016)
2016
-
[68]
Complex Networks (2021) Volume: 9
Benedek Rozemberczki and Carl Allen and Rik Sarkar Multi-scale Attributed Node Embedding J. Complex Networks (2021) Volume: 9
2021
-
[69]
Saeed Ghadimi and Guanghui Lan Stochastic First- and Zeroth-Order Methods for Nonconvex Stochastic Programming SIAM J. Optim. (2013) Volume: 23 Pages: 2341-2368
2013
-
[70]
Reddi, Sashank J and Hefny, Ahmed and Sra, Suvrit and Poczos, Barnabas and Smola, Alex Stochastic variance reduction for nonconvex optimization In: ICML (2016)
2016
-
[71]
Soufiane Hayou and Nikhil Ghosh and Bin Yu LoRA+: Efficient Low Rank Adaptation of Large Models arXiv preprint arXiv:2402.12354 (2024)
2024 arXiv
-
[72]
Yunwen Lei and Ke Tang Learning Rates for Stochastic Gradient Descent With Nonconvex Objectives IEEE Transactions on Pattern Analysis and Machine Intelligence (2021) Volume: 43 Pages: 4505-4511
2021
-
[73]
Chaoyue Liu and Libin Zhu and Mikhail Belkin Toward a theory of optimization for over-parameterized systems of non-linear equations: the lessons of deep learning ArXiv (2020) Volume: abs/2003.00307
2020 arXiv
-
[74]
Wainwright and Bin Yu Statistical guarantees for the EM algorithm: From population to sample-based analysis The Annals of Statistics (2017) Volume: 45 Pages: 77--120
Sivaraman Balakrishnan and Martin J. Wainwright and Bin Yu Statistical guarantees for the EM algorithm: From population to sample-based analysis The Annals of Statistics (2017) Volume: 45 Pages: 77--120
2017
-
[75]
Zheng and Yiqun Hui and Yanan Niu and Yang Song and Depeng Jin and Yong Li Sequential Recommendation with Graph Neural Networks In: SIGIR (2021)
Jianxin Chang and Chen Gao and Y. Zheng and Yiqun Hui and Yanan Niu and Yang Song and Depeng Jin and Yong Li Sequential Recommendation with Graph Neural Networks In: SIGIR (2021)
2021
-
[76]
Susheel Suresh and Vinith Budde and Jennifer Neville and Pan Li and Jianzhu Ma Breaking the Limit of Graph Neural Networks by Improving the Assortativity of Graphs with Local Mixing Patterns In: KDD (2021)
2021
-
[77]
Sunil Kumar Maurya and Xin Liu and Tsuyoshi Murata Improving Graph Neural Networks with Simple Architecture Design ArXiv (2021) Volume: abs/2105.07634
2021 arXiv
-
[78]
Sami Abu-El-Haija and Bryan Perozzi and Amol Kapoor and Hrayr Harutyunyan and Nazanin Alipourfard and Kristina Lerman and Greg Ver Steeg and A. G. Galstyan MixHop: Higher-Order Graph Convolutional Architectures via Sparsified Neighborhood Mixing In: ICML (2019)
2019
-
[79]
Song Mei and Yu Bai and Andrea Montanari The landscape of empirical risk for nonconvex losses The Annals of Statistics (2016)
2016
-
[80]
Yuanzhi Li and Yang Yuan Convergence Analysis of Two-layer Neural Networks with ReLU Activation In: NeurIPS (2017)
2017
-
[81]
Schmidt Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-Łojasiewicz Condition In: ECML/PKDD (2016)
Hamed Karimi and Julie Nutini and Mark W. Schmidt Linear Convergence of Gradient and Proximal-Gradient Methods Under the Polyak-Łojasiewicz Condition In: ECML/PKDD (2016)
2016
-
[82]
Yunwen Lei and Yiming Ying Sharper Generalization Bounds for Learning with Gradient-dominated Objective Functions In: ICLR (2021)
2021
-
[83]
Yunwen Lei and Ting Hu and Ke Tang Generalization Performance of Multi-pass Stochastic Gradient Descent with Convex Loss Functions J. Mach. Learn. Res. (2021) Volume: 22 Pages: 25:1-25:41
2021
-
[84]
Li, Xiaoyu and Orabona, Francesco A High Probability Analysis of Adaptive SGD with Momentum In: Workshop at ICML (2020)
2020
-
[85]
Shaojie Li and Yong Liu Improved Learning Rates for Stochastic Optimization: Two Theoretical Viewpoints ArXiv (2021) Volume: abs/2107.08686
2021
-
[86]
Shaojie Li and Yong Liu High Probability Guarantees for Nonconvex Stochastic Gradient Descent with Heavy Tails In: ICML (2022)
2022
-
[87]
Jiang, Yiding and Neyshabur, Behnam and Mobahi, Hossein and Krishnan, Dilip and Bengio, Samy Fantastic Generalization Measures and Where to Find Them In: ICLR (2019)
2019
-
[88]
Keskar, Nitish Shirish and Mudigere, Dheevatsa and Nocedal, Jorge and Smelyanskiy, Mikhail and Tang, Ping Tak Peter On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima In: ICLR (2017)
2017
-
[89]
Kaiyue Wen and Tengyu Ma and Zhiyuan Li How Does Sharpness-Aware Minimization Minimize Sharpness? ArXiv (2022) Volume: abs/2211.05729
2022 arXiv
-
[90]
Laurent Dinh and Razvan Pascanu and Samy Bengio and Yoshua Bengio Sharp Minima Can Generalize For Deep Nets In: ICML (2017)
2017
-
[91]
Wang, Xiyuan and Zhang, Muhan How powerful are spectral graph neural networks In: ICML (2022)
2022
-
[92]
Mingguo He and Zhewei Wei and Zengfeng Huang and Hongteng Xu BernNet: Learning Arbitrary Graph Spectral Filters via Bernstein Approximation In: NeurIPS (2021)
2021
-
[93]
Jiong Zhu and Yujun Yan and Lingxiao Zhao and Mark Heimann and Leman Akoglu and Danai Koutra Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs In: NeurIPS (2020)
2020
-
[94]
Derek Lim and Felix Hohne and Xiuyu Li and Sijia Huang and Vaishnavi Gupta and Omkar Bhalerao and Ser-Nam Lim Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and Strong Simple Methods In: NeurIPS (2021)
2021
-
[95]
Yan, Yujun and Hashemi, Milad and Swersky, Kevin and Yang, Yaoqing and Koutra, Danai Two sides of the same coin: Heterophily and oversmoothing in graph convolutional neural networks In: ICDM (2022)
2022
-
[96]
Xiang Li and Renyu Zhu and Yao Cheng and Caihua Shan and Siqiang Luo and Dongsheng Li and Wei Qian Finding Global Homophily in Graph Neural Networks When Meeting Heterophily In: ICML (2022)
2022
-
[97]
Welling Semi-Supervised Classification with Graph Convolutional Networks In: ICLR (2017)
Thomas Kipf and M. Welling Semi-Supervised Classification with Graph Convolutional Networks In: ICLR (2017)
2017
-
[98]
Pei, Hongbin and Wei, Bingzhe and Chang, Kevin Chen-Chuan and Lei, Yu and Yang, Bo Geom-GCN: Geometric Graph Convolutional Networks In: ICLR (2020)
2020
-
[100]
Lukas Balles and Philipp Hennig Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients In: ICML (2017)
2017
-
[101]
Chien, Eli and Peng, Jianhao and Li, Pan and Milenkovic, Olgica Adaptive Universal Generalized PageRank Graph Neural Network In: ICLR (2021)
2021
-
[102]
He, Mingguo and Wei, Zhewei and Wen, Ji-Rong Convolutional neural networks on graphs with chebyshev approximation, revisited NeurIPS (2022)
2022
-
[103]
Micha \"e l Defferrard and Xavier Bresson and Pierre Vandergheynst Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering In: NeurIPS (2016)
2016
-
[104]
Keyulu Xu and Weihua Hu and Jure Leskovec and Stefanie Jegelka How Powerful are Graph Neural Networks? In: ICLR (2019)
2019
-
[105]
Roberts Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training J
Diego Granziol and Stefan Zohren and Stephen J. Roberts Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training J. Mach. Learn. Res. (2020) Volume: 23 Pages: 173:1-173:65
2020
-
[106]
Simard and Barak A
Yann LeCun and Patrice Y. Simard and Barak A. Pearlmutter Automatic Learning Rate Maximization by On-Line Estimation of the Hessian's Eigenvectors In: NeurIPS (1992)
1992
-
[107]
Yimin Ding The Impact of Learning Rate Decay and Periodical Learning Rate Restart on Artificial Neural Network In: AIEE (2021)
2021
-
[108]
Shiv Ram Dubey and Soumendu Chakraborty and Swalpa Kumar Roy and Snehasis Mukherjee and Satish Kumar Singh and Bidyut Baran Chaudhuri diffGrad: An Optimization Method for Convolutional Neural Networks IEEE Transactions on Neural Networks and Learning Systems (2019) Volume: 31 ...
2019
-
[109]
Yash Deshpande and Andrea Montanari and Elchanan Mossel and Subhabrata Sen Contextual Stochastic Block Models In: NeurIPS (2018)
2018
-
[110]
Duchi and Elad Hazan and Yoram Singer Adaptive Subgradient Methods for Online Learning and Stochastic Optimization J
John C. Duchi and Elad Hazan and Yoram Singer Adaptive Subgradient Methods for Online Learning and Stochastic Optimization J. Mach. Learn. Res. (2011) Volume: 12 Pages: 2121--2159
2011
-
[111]
Tieleman, Tijmen and Hinton, Geoffrey Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude COURSERA: Neural networks for machine learning (2012) Volume: 4 Pages: 26--31
2012
-
[112]
Richtárik, Peter and Takáč, Martin On Optimal Probabilities in Stochastic Coordinate Descent Methods Optimization Letters (2013) Volume: 10 Pages:
2013
-
[113]
Saad, Yousef Iterative Methods for Sparse Linear Systems (2003)
2003
-
[114]
Benzi, Martin Preconditioning Techniques for Large Linear Systems: A Survey Journal of Computational Physics (2002) Volume: 182 Pages: 418--477
2002
-
[115]
Reddi and On the Convergence of Adam and Beyond In: ICLR (2018) main.bbl0000664000000000000000000002376414730021623011175 0ustar rootroot thebibliography 43 [1] #1
Sashank J. Reddi and On the Convergence of Adam and Beyond In: ICLR (2018) main.bbl0000664000000000000000000002376414730021623011175 0ustar rootroot thebibliography 43 [1] #1
2018
-
[116]
Balcilar, M.; Renton, G.; H \'e roux, P.; Ga \"u z \`e re, B.; Adam, S.; and Honeine, P. 2021. Analyzing the Expressive Power of Graph Neural Networks in a Spectral Perspective. In ICLR
2021
-
[117]
M.; Grattarola, D.; Livi, L.; and Alippi, C
Bianchi, F. M.; Grattarola, D.; Livi, L.; and Alippi, C. 2021. Graph neural networks with convolutional arma filters. IEEE transactions on pattern analysis and machine intelligence, 44(7): 3496--3507
2021
-
[118]
C.; Lemar \'e chal, C.; and Sagastiz \'a bal, C
Bonnans, J.-F.; Gilbert, J. C.; Lemar \'e chal, C.; and Sagastiz \'a bal, C. A. 2006. Numerical optimization: theoretical and practical aspects. Springer Science & Business Media
2006
-
[119]
Chien, E.; Peng, J.; Li, P.; and Milenkovic, O. 2021. Adaptive Universal Generalized PageRank Graph Neural Network. In ICLR
2021
-
[120]
Defferrard, M.; Bresson, X.; and Vandergheynst, P. 2016. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. In NeurIPS
2016
-
[121]
Ding, Y. 2021. The Impact of Learning Rate Decay and Periodical Learning Rate Restart on Artificial Neural Network. In AIEE
2021
-
[122]
C.; Hazan, E.; and Singer, Y
Duchi, J. C.; Hazan, E.; and Singer, Y. 2011. Adaptive Subgradient Methods for Online Learning and Stochastic Optimization. J. Mach. Learn. Res., 12: 2121--2159
2011
-
[123]
P.; Joshi, C
Dwivedi, V. P.; Joshi, C. K.; Luu, A. T.; Laurent, T.; Bengio, Y.; and Bresson, X. 2023. Benchmarking graph neural networks. Journal of Machine Learning Research, 24(43): 1--48
2023
-
[124]
H.; and Van Loan, C
Golub, G. H.; and Van Loan, C. F. 2013. Matrix computations. JHU press
2013
-
[125]
Granziol, D.; Zohren, S.; and Roberts, S. J. 2020. Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training. J. Mach. Learn. Res., 23: 173:1--173:65
2020
-
[126]
Gupta, V.; Koren, T.; and Singer, Y. 2018. Shampoo: Preconditioned Stochastic Tensor Optimization. In ICML
2018
-
[127]
K.; Vandergheynst, P.; and Gribonval, R
Hammond, D. K.; Vandergheynst, P.; and Gribonval, R. 2011. Wavelets on graphs via spectral graph theory. Applied and Computational Harmonic Analysis, 30(2): 129--150
2011
-
[128]
He, M.; Wei, Z.; Huang, Z.; and Xu, H. 2021. BernNet: Learning Arbitrary Graph Spectral Filters via Bernstein Approximation. In NeurIPS
2021
-
[129]
He, M.; Wei, Z.; and Wen, J.-R. 2022. Convolutional neural networks on graphs with chebyshev approximation, revisited. NeurIPS
2022
-
[130]
A.; and Johnson, C
Horn, R. A.; and Johnson, C. R. 2012. Matrix analysis. Cambridge university press
2012
-
[131]
R.; Fang, Y.; Stevenson, R
Izadi, M. R.; Fang, Y.; Stevenson, R. L.; and Lin, L. 2020. Optimization of Graph Neural Networks with Natural Gradient Descent. In Big Data
2020
-
[132]
Jia, X.; Wang, H.; Peng, J.; Feng, X.; and Meng, D. 2024. Preconditioning Matters: Fast Global Convergence of Non-convex Matrix Factorization via Scaled Gradient Descent. In NeurIPS
2024
-
[133]
P.; and Ba, J
Kingma, D. P.; and Ba, J. 2014. Adam: A Method for Stochastic Optimization. CoRR, abs/1412.6980
2014 arXiv
-
[134]
Kipf, T.; and Welling, M. 2017. Semi-Supervised Classification with Graph Convolutional Networks. In ICLR
2017
-
[135]
Y.; and Pearlmutter, B
LeCun, Y.; Simard, P. Y.; and Pearlmutter, B. A. 1992. Automatic Learning Rate Maximization by On-Line Estimation of the Hessian's Eigenvectors. In NeurIPS
1992
-
[136]
Levie, R.; Monti, F.; Bresson, X.; and Bronstein, M. M. 2019. CayleyNets: Graph Convolutional Neural Networks With Complex Rational Spectral Filters. IEEE Transactions on Signal Processing, 67: 97--109
2019
-
[137]
Liao, R.; Zhao, Z.; Urtasun, R.; and Zemel, R. 2019. LanczosNet: Multi-Scale Deep Graph Convolutional Networks. In ICLR
2019
-
[138]
Lim, D.; Hohne, F.; Li, X.; Huang, S.; Gupta, V.; Bhalerao, O.; and Lim, S.-N. 2021. Large Scale Learning on Non-Homophilous Graphs: New Benchmarks and Strong Simple Methods. In NeurIPS
2021
-
[139]
Loukas, A. 2020. What graph neural networks cannot learn: depth vs width. In ICLR
2020
-
[140]
C.-C.; Lei, Y.; and Yang, B
Pei, H.; Wei, B.; Chang, K. C.-C.; Lei, Y.; and Yang, B. 2020. Geom-GCN: Geometric Graph Convolutional Networks. In ICLR
2020
-
[141]
Platonov, O.; Kuznedelev, D.; Diskin, M.; Babenko, A.; and Prokhorenkova, L. 2023. A critical look at the evaluation of GNNs under heterophily: Are we really making progress? In ICLR
2023
-
[142]
Rozemberczki, B.; Allen, C.; and Sarkar, R. 2021. Multi-scale Attributed Node Embedding. J. Complex Networks, 9
2021
-
[143]
Saad, Y. 2003. Iterative Methods for Sparse Linear Systems. SIAM
2003
-
[144]
Shalev-Shwartz, S.; and Ben-David, S. 2014. Understanding machine learning: From theory to algorithms. Cambridge university press
2014
-
[145]
Shchur, O.; Mumme, M.; Bojchevski, A.; and G \"u nnemann, S. 2018. Pitfalls of graph neural network evaluation. arXiv preprint arXiv:1811.05868
2018 arXiv
-
[146]
Tao, T. 2012. Topics in random matrix theory, volume 132. American Mathematical Soc
2012
-
[147]
Tieleman, T.; and Hinton, G. 2012. Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude. COURSERA: Neural networks for machine learning, 4(2): 26--31
2012
-
[148]
Velickovic, P.; Cucurull, G.; Casanova, A.; Romero, A.; Lio’, P.; and Bengio, Y. 2018. Graph Attention Networks. In ICLR
2018
-
[149]
Wang, X.; and Zhang, M. 2022 a . How powerful are spectral graph neural networks. In ICML, 23341--23362
2022
-
[150]
Wang, X.; and Zhang, M. 2022 b . How powerful are spectral graph neural networks. In ICML
2022
-
[151]
Wu, Z.; Pan, S.; Chen, F.; Long, G.; Zhang, C.; and Yu, P. S. 2019. A Comprehensive Survey on Graph Neural Networks. IEEE Transactions on Neural Networks and Learning Systems, 32: 4--24
2019
-
[152]
Xu, K.; Hu, W.; Leskovec, J.; and Jegelka, S. 2019. How Powerful are Graph Neural Networks? In ICLR
2019
-
[153]
Xu, K.; Zhang, M.; Jegelka, S.; and Kawaguchi, K. 2021. Optimization of Graph Neural Networks: Implicit Acceleration by Skip Connections and More Depth. In ICML
2021
-
[154]
Yang, M.; Shen, Y.; Li, R.; Qi, H.; Zhang, Q.; and Yin, B. 2022. A new perspective on the effects of spectrum in graph neural networks. In ICML
2022
-
[155]
S.; Dhillon, I.; and Hsieh, C.-J
Yen, J.-N.; Duvvuri, S. S.; Dhillon, I.; and Hsieh, C.-J. 2024. Block Low-Rank Preconditioner with Shared Basis for Stochastic Optimization. In NeurIPS
2024
-
[156]
You, Y.; Gitman, I.; and Ginsburg, B. 2017. Large batch training of convolutional networks. arXiv preprint arXiv:1708.03888
2017 arXiv
-
[157]
Zeng, H.; Zhou, H.; Srivastava, A.; Kannan, R.; and Prasanna, V. 2020. GraphSAINT: Graph Sampling Based Inductive Learning Method. In ICML
2020
-
[158]
Zhu, J.; Yan, Y.; Zhao, L.; Heimann, M.; Akoglu, L.; and Koutra, D. 2020. Beyond Homophily in Graph Neural Networks: Current Limitations and Effective Designs. In NeurIPS
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.