REVIEW 3 major objections 6 minor 76 references
Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper argues that GNN gains for language model representations come from feature transformation, not from message passing, so plain MLPs suffice.
desk verdict The MLP-versus-GNN comparison is clean and useful, but the paper's claim that message passing is not the driver rests on a Message module that is neither parameter-matched nor free of feature transformation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the modular probing framework: a frozen LM produces token representations, a selectable control module transforms them, and a classifier is trained on Edge Probing tasks. The control module has three instantiations: a two-layer GCN (Eq. 2), a Message module that keeps only DAGNN-style propagation along Universal Dependency trees with a single trainable projection vector s weighting each hop (Eqs. 3-7), and an MLP that applies only the ReLU feature transformation (Eq. 8). Because Eq. 2 is algebraically the composition of Eq. 3 and Eq. 8, differences among the three configurations isolate what each operation contributes.
What would settle it
Run the same probing suite with the Message module's hop-weighting projection s frozen to uniform values and with GNN, MLP, and Message matched for parameter count; if Message-only then matches or exceeds MLP on the eight tasks, the paper's central attribution fails.
Extended reading notes
Core claim
The central discovery is that when a GCN is decoupled into its two micro-operations, the feature-transformation operation (an MLP applied to each token representation) is the part that helps probing classifiers find linguistic structure in frozen LM representations, and the message-passing operation (propagating representations along the dependency tree) is at best neutral and often detrimental. On average over eight tasks, the MLP control module improves macro F1 over the no-module baseline by 4.52 on BERT, 5.82 on T5, and 3.06 on Llama-3-8B, whereas the message-passing-only module changes scores by -1.22, +0.18, and -5.47 respectively. The paper reads this as evidence that explicit structural modeling is not the operative ingredient; the 'simpler yet stronger' MLP suffices.
Load-bearing premise
The attribution of all gains to feature transformation rests on the Message module being a pure message-passing operation, yet it contains a trainable projection vector that learns how much each node trusts each propagation hop, and the three modules are not matched in parameter count.
Editorial extensions
If this is right
- If feature transformation is the driver, then GNN modules for LM-enhancement can be replaced by MLPs with equal or better probing performance and lower computational cost.
- Message-passing-only modules should be avoided for tasks where the paper finds them harmful, such as syntactic probing with Llama-3-8B, where the message-passing module dropped average F1 by 5.47.
- The conclusion holds across encoder-only, encoder-decoder, and decoder-only LM architectures (BERT, T5, and Llama-3-8B) and across both syntactic and semantic probing tasks, so it is not an artifact of one model family.
- KAN as an alternative feature-transformation module performs close to MLP, supporting the claim that the specific transformation family matters less than having a transformation at all.
Reading between the lines
- The paper's probing setup uses frozen LMs; an untested extension is whether the same conclusion holds when the LM is fine-tuned jointly with the graph module, where message passing might have more room to help.
- Because the Message module contains a learned hop-weighting vector s, a fairer test of 'pure' message passing would freeze s to uniform weights; the paper's results could change if s itself is acting as a feature transformation.
- The span-distance analysis hints that message passing has a marginal advantage only for very long spans (greater than 15 tokens); a targeted study on long-range relation extraction could reveal a narrow regime where structure still earns its cost.
- Parameter counts are not reported as matched across GNN, MLP, and Message modules, so part of the MLP advantage could be capacity; a controlled capacity-matched comparison would sharpen the claim.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a probing framework for assessing whether explicit syntactic structure, encoded through GNN message passing, is needed to enhance language model representations. The framework adds a control module before a probing classifier, with four configurations: no control module, a GCN, a message-passing-only module, and an MLP feature-transformation module. Experiments on the Edge Probing Suite across BERT, T5, and Llama-3-8B indicate that MLPs consistently improve or match GNN performance, while the message-passing-only module often underperforms the baseline. The authors conclude that observed gains are primarily driven by feature transformation rather than message passing and that MLPs are sufficient for enhancing LM representations.
Significance. If the conclusion holds, the paper is valuable for the NLP and graph representation learning communities: it provides a reusable probing setup, a direct parameter-matched comparison between GCNs and MLPs, and several positive controls such as random-graph degradation, layer-wise probing, span-distance grouping, and a KAN comparison. The parameter-matched GNN-versus-MLP comparison in particular is a credible and useful result, since Equations (2) and (8) share the same W(l) and b(l). However, the central attribution claim that performance gains are primarily driven by feature transformation rests on the Message module, which is not a clean isolation of message passing.
major comments (3)
- [III.B-b, Eq. (5)] The Message component is not a clean isolation of message passing. Equation (5) computes S = σ(Qs) with a trainable projection vector s ∈ R^d, producing a per-node, feature-dependent gate that reweights the stacked hop representations in Equations (4)–(7). This is a learned nonlinear transformation of the input features, so the module labeled 'message-passing only' already contains a feature-transformation operation. Consequently, the negative results for the Message module in Table III do not establish that message passing itself is ineffective; they establish only that this particular gated propagation module underperforms. The authors should either use a parameter-free propagation baseline (e.g., averaging over hops, or fixing s to a constant) or explicitly justify why the learned gate should not be regarded as feature transformation.
- [V.A.3 and Table III] The Message module has only O(d) trainable parameters (the vector s), while the GNN and MLP modules contain two d×d matrices plus biases. Its lower performance relative to the baseline is therefore confounded with lower model capacity. The claim that 'message-passing operations contribute minimally' is not supported unless the comparison is capacity-matched. The authors should add a parameter-matched Message variant, for example by giving the propagation module a per-dimension or low-rank linear readout with a comparable number of parameters, and also test a low-capacity MLP (e.g., diagonal or rank-1 W) to determine whether the observed gap is attributable to capacity rather than to the absence of structure.
- [III.B-b and VII] The statement in Section VII that 'the observed performance gains are primarily driven by feature transformation operations, rather than the message-passing mechanisms' is stronger than the current evidence supports. The direct GNN-versus-MLP comparison shows that structure is not necessary for these probing tasks when transformations are matched, but it does not localize the source of GNN gains because the Message ablation is confounded as described above. The authors should soften the attribution claim to what the experiments establish, or add the additional controls needed to support the causal attribution.
minor comments (6)
- [Abstract] The abstract contains the typo 'reinforcing the the principle'; it should read 'reinforcing the principle'.
- [II.A] There is a stray comma in 'Furthermore, , Sachan et al. [59] proposed'; the extra comma should be removed.
- [III.B-b] The sentence describing s is punctuated incorrectly: '...different hops, It is the only learnable parameter...' should use a period or semicolon instead of a comma before 'It'.
- [References] Several references are duplicated: [26] and [53] are the same paper, [32], [54], and [66] are the same Kipf and Welling paper, and [36] and [39] are the same Zhang et al. paper. The reference list should be deduplicated.
- [VI.E] The opening sentence of Section VI.E is repeated verbatim twice in the same paragraph; one copy should be removed.
- [Fig. 7] Figure 7 is not readable in the provided version; the loss curves should be regenerated at adequate resolution and with legible labels.
Circularity Check
No circularity: the MLP-sufficiency claim rests on direct empirical ablations; the Eq. 2 = Eq. 3 + Eq. 8 identity is a definitional decomposition, not a hidden reuse of the result.
full rationale
The paper's central claim, that GNN gains on edge probing are primarily driven by feature transformation rather than message passing, is supported by direct empirical ablations in Table III comparing GNN, Message, and MLP control modules against a frozen-LM baseline. No constant is fitted to a subset of data and then reported as a prediction; the comparisons are held-out evaluations on standard benchmarks (OntoNotes, EWT, SPR, SemEval) with frozen LMs. The only exact identity in the framework is architectural: Eq. 2 (GCN layer) is literally Eq. 3 (normalized propagation) followed by Eq. 8 (MLP transform), so the statement that Eq. 2 is formed by directly combining Eq. 3 and Eq. 8 is a definitional decomposition, not a hidden reuse of the paper's conclusion. The mentions of prior work by co-authors ([9] as motivation and [26] for residual connections) are not load-bearing: the empirical verdict does not reduce to those citations. One substantive caveat, noted for correctness rather than circularity, is that the Message module (Eqs. 4-7) uses a learned projection vector s to gate hop contributions, so it is not a pure structural propagation; this is a construct-validity confound for the attribution to feature transformation, but it is not a case of the paper's result being equivalent to its input by construction. Hence no circular step is present.
Assumptions & free parameters
free parameters (2)
- Message hop count k =
not explicitly reported
- Control module depth and width =
two layers, width not reported
assumptions (4)
- domain assumption Stanza Universal Dependencies parses provide faithful syntactic structure for each token in every probing task (Section V.A.3).
- domain assumption The Graph Convolutional Network is a representative GNN for structure-aware NLP (Section III.B.a).
- domain assumption Probing classifiers with trained control modules measure what is encoded in LM representations (Section III).
- domain assumption LLM2Vec adaptation preserves Llama-3-8B as a representative decoder-only LM encoder (Section V.A.2).
Cite this review
Pith. "Pith review of Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations." pith.science (2026). https://pith.science/paper/ZSQHI7JV
@misc{pith2026250621682,
author = {Pith},
title = {Pith review of: Do We Really Need GNNs with Explicit Structural Modeling? MLPs Suffice for Language Model Representations},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZSQHI7JV}},
note = {Machine review of arXiv:2506.21682}
}
read the original abstract
Explicit structural information has been proven to be encoded by Graph Neural Networks (GNNs), serving as auxiliary knowledge to enhance model capabilities and improve performance in downstream NLP tasks. However, recent studies indicate that GNNs fail to fully utilize structural information, whereas Multi-Layer Perceptrons (MLPs), despite lacking the message-passing mechanisms inherent to GNNs, exhibit a surprising ability in structure-aware tasks. Motivated by these findings, this paper introduces a comprehensive probing framework from an information-theoretic perspective. The framework is designed to systematically assess the role of explicit structural modeling in enhancing language model (LM) representations and to investigate the potential of MLPs as efficient and scalable alternatives to GNNs. We extend traditional probing classifiers by incorporating a control module that allows for selective use of either the full GNN model or its decoupled components, specifically, the message-passing and feature-transformation operations.This modular approach isolates and assesses the individual contributions of these operations, avoiding confounding effects from the complete GNN architecture. Using the Edge Probing Suite, a diagnostic tool for evaluating the linguistic knowledge encoded in LMs, we find that MLPs, when used as feature-transformation modules, consistently improve the linguistic knowledge captured in LM representations across different architectures. They effectively encode both syntactic and semantic patterns. Similarly, GNNs that incorporate feature-transformation operations show beneficial effects. In contrast, models that rely solely on message-passing operations tend to underperform, often leading to negative impacts on probing task performance.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
arXiv 2023
-
[2]
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Yang, A. Fan et al. , “The llama 3 herd of models,” arXiv preprint arXiv:2407.21783 , 2024
arXiv 2024
-
[3]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi et al., “Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,” arXiv preprint arXiv:2501.12948 , 2025
arXiv 2025
-
[4]
A primer in BERTology: What we know about how BERT works,
A. Rogers, O. Kovaleva, and A. Rumshisky, “A primer in BERTology: What we know about how BERT works,” Transactions of the Association for Computational Linguistics , vol. 8, pp. 842–866, 2020. [Online]. Available: https://aclanthology.org/2020.tacl-1.54
work page 2020
-
[5]
Language models as knowledge bases?
F. Petroni, T. Rockt ¨aschel, S. Riedel, P. Lewis, A. Bakhtin, Y . Wu, and A. Miller, “Language models as knowledge bases?” in Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) , K. Inui, J. Jiang, V . Ng, and X. Wan, Eds. Hong Kon...
work page 2019
-
[6]
Do PLMs know and understand ontological knowledge?
W. Wu, C. Jiang, Y . Jiang, P. Xie, and K. Tu, “Do PLMs know and understand ontological knowledge?” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada: Association for Computational Linguistics, Jul. 2023, pp. 3080–3101. [Onlin...
work page 2023
-
[7]
Attention can reflect syntactic structure (if you let it),
V . Ravishankar, A. Kulmizev, M. Abdou, A. Søgaard, and J. Nivre, “Attention can reflect syntactic structure (if you let it),” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , P. Merlo, J. Tiedemann, and R. Tsarfaty, Eds. Online: Association for Computational Linguistics, Apr. 20...
work page 2021
-
[8]
Do transformer models show similar attention patterns to task-specific human gaze?
O. Eberle, S. Brandl, J. Pilot, and A. Søgaard, “Do transformer models show similar attention patterns to task-specific human gaze?” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , S. Muresan, P. Nakov, and A. Villavicencio, Eds. Dublin, Ireland: Association for Computational Linguistics...
work page 2022
Show all 76 references
-
[9]
Mlps compass: What is learned when mlps are combined with plms?
L. Zhou, W. Chen, Y . Cao, D. Zeng, W. Liu, and H. Qu, “Mlps compass: What is learned when mlps are combined with plms?” in ICASSP 2024- 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE, 2024, pp. 11 346–11 350
2024
-
[10]
What does BERT learn about the structure of language?
G. Jawahar, B. Sagot, and D. Seddah, “What does BERT learn about the structure of language?” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. M `arquez, Eds. Florence, Italy: Association for Computationa...
2019
-
[11]
Acceptability judgements via examining the topology of attention maps,
D. Cherniavskii, E. Tulchinskii, V . Mikhailov, I. Proskurina, L. Kushnareva, E. Artemova, S. Barannikov, I. Piontkovskaya, D. Piontkovski, and E. Burnaev, “Acceptability judgements via examining the topology of attention maps,” in Findings of the Association for Computational...
2022
-
[12]
How well do text embedding models understand syntax?
Y . Zhang, Z. Feng, Z. Teng, Z. Liu, and H. Li, “How well do text embedding models understand syntax?” in Findings of the Association for Computational Linguistics: EMNLP 2023 , H. Bouamor, J. Pino, and K. Bali, Eds. Singapore: Association for Computational Linguistics, Dec. 2...
2023
-
[13]
Do LLMs learn a true syntactic uni- versal?
J. T. Hale and M. Stanojevi ´c, “Do LLMs learn a true syntactic uni- versal?” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y . Al-Onaizan, M. Bansal, and Y .- N. Chen, Eds. Miami, Florida, USA: Association for Computational Lingui...
2024
-
[14]
Investigating entity knowledge in BERT with simple neural end-to-end entity linking,
S. Broscheit, “Investigating entity knowledge in BERT with simple neural end-to-end entity linking,” in Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL) , M. Bansal and A. Villavicencio, Eds. Hong Kong, China: Association for Computational ...
2019
-
[15]
What’s in a name? are BERT named entity representations just as good for any other name?
S. Balasubramanian, N. Jain, G. Jindal, A. Awasthi, and S. Sarawagi, “What’s in a name? are BERT named entity representations just as good for any other name?” in Proceedings of the 5th Workshop on Representation Learning for NLP , S. Gella, J. Welbl, M. Rei, F. Petroni, P. Le...
2020
-
[16]
Reassessing semantic knowledge encoded in large language models through the word-in-context task,
Y . Hayashi, “Reassessing semantic knowledge encoded in large language models through the word-in-context task,” in Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , N. Calzolari, M.-Y . ...
2024
-
[17]
Evaluating llms’ capability to identify lexical semantic equiv- alence: Probing with the word-in-context task,
——, “Evaluating llms’ capability to identify lexical semantic equiv- alence: Probing with the word-in-context task,” in Proceedings of the 31st International Conference on Computational Linguistics , O. Ram- bow, L. Wanner, M. Apidianaki, H. Al-Khalifa, B. D. Eugenio, and S. S...
2025
-
[18]
How much knowledge can you pack into the parameters of a language model?
A. Roberts, C. Raffel, and N. Shazeer, “How much knowledge can you pack into the parameters of a language model?” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y . He, and Y . Liu, Eds. Online: Associatio...
2020
-
[19]
MuLan: A study of fact mutability in language models,
C. Fierro, N. Garneau, E. Bugliarello, Y . Kementchedjhieva, and A. Søgaard, “MuLan: A study of fact mutability in language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologie...
2024
-
[20]
Does mapo tofu contain coffee? probing llms for food- related cultural knowledge,
L. Zhou, T. Karidi, W. Liu, N. Garneau, Y . Cao, W. Chen, H. Li, and D. Hershcovich, “Does mapo tofu contain coffee? probing llms for food- related cultural knowledge,” arXiv preprint arXiv:2404.06833 , 2024
2024 arXiv
-
[21]
BERT rediscovers the classical NLP pipeline,
I. Tenney, D. Das, and E. Pavlick, “BERT rediscovers the classical NLP pipeline,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. M `arquez, Eds. Florence, Italy: Association for Computational Linguisti...
2019
-
[22]
Graph convolutional encoders for syntax-aware neural machine translation,
J. Bastings, I. Titov, W. Aziz, D. Marcheggiani, and K. Sima’an, “Graph convolutional encoders for syntax-aware neural machine translation,” in Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing , M. Palmer, R. Hwa, and S. Riedel, Eds. Copen...
2017
-
[23]
Dsisa: A new neural machine translation combining dependency weight and neighbors,
L. Li, A. Zhang, and M.-X. Luo, “Dsisa: A new neural machine translation combining dependency weight and neighbors,” ACM Trans. Asian Low-Resour. Lang. Inf. Process. , vol. 23, no. 2, Feb. 2024
2024
-
[24]
Structured neural summarization,
P. Fernandes, M. Allamanis, and M. Brockschmidt, “Structured neural summarization,” in International Conference on Learning Representa- tions, 2019
2019
-
[25]
Cross-document distillation via graph-based summarization of extracted essential knowl- edge,
L. Ragazzi, G. Moro, L. Valgimigli, and R. Fiorani, “Cross-document distillation via graph-based summarization of extracted essential knowl- edge,” IEEE/ACM Transactions on Audio, Speech, and Language Pro- cessing, 2024
2024
-
[26]
A weighted gcn with logical adjacency matrix for relation extraction,
L. Zhou, T. Wang, H. Qu, L. Huang, and Y . Liu, “A weighted gcn with logical adjacency matrix for relation extraction,” in ECAI 2020 . iOS Press, 2020, pp. 2314–2321
2020
-
[27]
Document-level relation extraction with structure enhanced transformer encoder,
W. Liu, L. Zhou, D. Zeng, and H. Qu, “Document-level relation extraction with structure enhanced transformer encoder,” in Proc. of IJCNN, 2022
2022
-
[28]
Diffusyn: A diffusion- driven framework with syntactic dependency for aspect sentiment triplet extraction,
Q. Yi, X. Kong, L. Zhu, C. Zhang, and G. Shen, “Diffusyn: A diffusion- driven framework with syntactic dependency for aspect sentiment triplet extraction,” IEEE Transactions on Audio, Speech and Language Process- ing, 2025
2025
-
[29]
Dualgcn: Exploring syntactic and semantic information for aspect-based sentiment analysis,
R. Li, H. Chen, F. Feng, Z. Ma, X. Wang, and E. Hovy, “Dualgcn: Exploring syntactic and semantic information for aspect-based sentiment analysis,” IEEE Transactions on Neural Networks and Learning Systems, 2022. 12
2022
-
[30]
Integrated syntactic and semantic tree for targeted sentiment classification using dual- channel graph convolutional network,
P. Zhang, R. Zhao, B. Yang, Y . Li, and Z. Yang, “Integrated syntactic and semantic tree for targeted sentiment classification using dual- channel graph convolutional network,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 1109–1124, 2024
2024
-
[31]
Incorporating syntax and lexical knowledge to multilingual sentiment classification on large language models,
H. Kanayama, Y . Zhao, R. Iwamoto, and T. Ohko, “Incorporating syntax and lexical knowledge to multilingual sentiment classification on large language models,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V . Srikumar, Eds....
2024
-
[32]
Semi-Supervised Classification with Graph Convolutional Networks,
T. N. Kipf and M. Welling, “Semi-Supervised Classification with Graph Convolutional Networks,” in Proceedings of the 5th International Con- ference on Learning Representations , ser. ICLR ’17, 2017
2017
-
[33]
Graph attention networks,
P. Veli ˇckovi´c, G. Cucurull, A. Casanova, A. Romero, P. Li `o, and Y . Bengio, “Graph attention networks,” 6th International Conference on Learning Representations , 2017
2017
-
[34]
Dpgnn: Dual-perception graph neural network for representation learning,
L. Zhou, W. Chen, D. Zeng, S. Cheng, W. Liu, M. Zhang, and H. Qu, “Dpgnn: Dual-perception graph neural network for representation learning,” Knowledge-Based Systems, vol. 268, p. 110377, 2023
2023
-
[35]
Graph convolution over pruned dependency trees improves relation extraction,
Y . Zhang, P. Qi, and C. D. Manning, “Graph convolution over pruned dependency trees improves relation extraction,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds. Brussels, Be...
2018
-
[37]
A fair comparison of graph neural networks for graph classification,
F. Errica, M. Podda, D. Bacciu, and A. Micheli, “A fair comparison of graph neural networks for graph classification,” in International Conference on Learning Representations , 2020
2020
-
[38]
Gnn is a counter? revisiting gnn for question answering,
K. Wang, Y . Zhang, D. Yang, L. Song, and T. Qin, “Gnn is a counter? revisiting gnn for question answering,” in International Conference on Learning Representations, 2021
2021
-
[39]
Graph-less neural networks: Teaching old mlps new tricks via distillation,
S. Zhang, Y . Liu, Y . Sun, and N. Shah, “Graph-less neural networks: Teaching old mlps new tricks via distillation,” in International Confer- ence on Learning Representations , 2021
2021
-
[40]
Learning mlps on graphs: A unified view of effectiveness, robustness, and efficiency,
Y . Tian, C. Zhang, Z. Guo, X. Zhang, and N. Chawla, “Learning mlps on graphs: A unified view of effectiveness, robustness, and efficiency,” in The eleventh international conference on learning representations , 2022
2022
-
[41]
SA-MLP: Distilling graph knowledge from GNNs into structure-aware MLP,
J. Chen, M. Bai, S. Chen, J. Gao, J. Zhang, and J. Pu, “SA-MLP: Distilling graph knowledge from GNNs into structure-aware MLP,” Transactions on Machine Learning Research , 2024
2024
-
[42]
Bag-of-words vs. graph vs. sequence in text classification: Questioning the necessity of text-graphs and the surprising strength of a wide MLP,
L. Galke and A. Scherp, “Bag-of-words vs. graph vs. sequence in text classification: Questioning the necessity of text-graphs and the surprising strength of a wide MLP,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long ...
2022
-
[43]
Do we really need gcns in traffic forecasting? a graph-less pure- mlp architecture,
Y . Zhang, X. Wang, X. Yu, K. Yang, Z. Zhou, and Y . Wang, “Do we really need gcns in traffic forecasting? a graph-less pure- mlp architecture,” in Companion Proceedings of the ACM on Web Conference 2025, ser. WWW ’25. New York, NY , USA: Association for Computing Machinery, 2...
2025
-
[44]
How well apply simple MLP to incomplete utterance rewriting?
J. Li, X. Su, X. Ma, and G. Gao, “How well apply simple MLP to incomplete utterance rewriting?” in Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , A. Rogers, J. Boyd-Graber, and N. Okazaki, Eds. Toronto, Canada...
2023
-
[45]
Re-tacred: Addressing shortcomings of the tacred dataset,
G. Stoica, E. A. Platanios, and B. P ´oczos, “Re-tacred: Addressing shortcomings of the tacred dataset,” in Proc. of AAAI , 2021
2021
-
[46]
SemEval- 2010 task 8: Multi-way classification of semantic relations between pairs of nominals,
I. Hendrickx, S. N. Kim, Z. Kozareva, P. Nakov, D. ´O S ´eaghdha, S. Pad ´o, M. Pennacchiotti, L. Romano, and S. Szpakowicz, “SemEval- 2010 task 8: Multi-way classification of semantic relations between pairs of nominals,” in Proceedings of the 5th International Workshop on Se...
2010
-
[47]
Filter-enhanced mlp is all you need for sequential recommendation,
K. Zhou, H. Yu, W. X. Zhao, and J.-R. Wen, “Filter-enhanced mlp is all you need for sequential recommendation,” in Proceedings of the ACM web conference 2022 , 2022, pp. 2388–2399
2022
-
[48]
Smlp4rec: an efficient all-mlp architecture for sequential recommendations,
J. Gao, X. Zhao, M. Li, M. Zhao, R. Wu, R. Guo, Y . Liu, and D. Yin, “Smlp4rec: an efficient all-mlp architecture for sequential recommendations,” ACM Transactions on Information Systems , vol. 42, no. 3, pp. 1–23, 2024
2024
-
[49]
Predicting temporal sets with simplified fully connected networks,
L. Yu, Z. Liu, T. Zhu, L. Sun, B. Du, and W. Lv, “Predicting temporal sets with simplified fully connected networks,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 37, no. 4, 2023, pp. 4835–4844
2023
-
[50]
Mlpst: Mlp is all you need for spatio-temporal prediction,
Z. Zhang, Z. Huang, Z. Hu, X. Zhao, W. Wang, Z. Liu, J. Zhang, S. J. Qin, and H. Zhao, “Mlpst: Mlp is all you need for spatio-temporal prediction,” in Proceedings of the 32nd ACM International Conference on Information and Knowledge Management , 2023, pp. 3381–3390
2023
-
[51]
Mlp-mixer: An all-mlp architecture for vision,
I. O. Tolstikhin, N. Houlsby, A. Kolesnikov, L. Beyer, X. Zhai, T. Un- terthiner, J. Yung, A. Steiner, D. Keysers, J. Uszkoreitet al., “Mlp-mixer: An all-mlp architecture for vision,” Advances in neural information processing systems, vol. 34, pp. 24 261–24 272, 2021
2021
-
[52]
As-mlp: An axial shifted mlp architecture for vision,
D. Lian, Z. Yu, X. Sun, and S. Gao, “As-mlp: An axial shifted mlp architecture for vision,” in International Conference on Learning Representations, 2022
2022
-
[53]
A weighted gcn with logical adjacency matrix for relation extraction,
L. Zhou, T. Wang, H. Qu, L. Huang, and Y . Liu, “A weighted gcn with logical adjacency matrix for relation extraction,” in European Conference on Artificial Intelligence , 2020
2020
-
[54]
Semi-supervised classification with graph convolutional networks,
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in International Conference on Learning Rep- resentations (ICLR), 2017
2017
-
[55]
Universal Dependencies,
M.-C. de Marneffe, C. D. Manning, J. Nivre, and D. Zeman, “Universal Dependencies,” Computational Linguistics, vol. 47, no. 2, pp. 255–308, Jun. 2021. [Online]. Available: https://aclanthology.org/2021.cl-2.11
2021
-
[56]
Heterogeneous graph transformer for graph-to-sequence learning,
S. Yao, T. Wang, and X. Wan, “Heterogeneous graph transformer for graph-to-sequence learning,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computa...
2020
-
[57]
Improving neural machine translation with the Abstract Meaning Representation by combining graph and sequence transformers,
C. Li and J. Flanigan, “Improving neural machine translation with the Abstract Meaning Representation by combining graph and sequence transformers,” in Proceedings of the 2nd Workshop on Deep Learning on Graphs for Natural Language Processing (DLG4NLP 2022) , L. Wu, B. Liu, R....
2022
-
[58]
Abstract Meaning Representation for sembanking,
L. Banarescu, C. Bonial, S. Cai, M. Georgescu, K. Griffitt, U. Hermjakob, K. Knight, P. Koehn, M. Palmer, and N. Schneider, “Abstract Meaning Representation for sembanking,” in Proceedings of the 7th Linguistic Annotation Workshop and Interoperability with Discourse, A. Pareja...
2013
-
[59]
Do syntax trees help pre-trained transformers extract information?
D. Sachan, Y . Zhang, P. Qi, and W. L. Hamilton, “Do syntax trees help pre-trained transformers extract information?” in Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume , P. Merlo, J. Tiedemann, and R. Ts...
2021
-
[60]
Graph attention networks,
P. Veli ˇckovi´c, G. Cucurull, A. Casanova, A. Romero, and P. Li `o, “Graph attention networks,” International Conference on Learning Representa- tions, 2018
2018
-
[61]
Probing classifiers: Promises, shortcomings, and advances,
Y . Belinkov, “Probing classifiers: Promises, shortcomings, and advances,” Computational Linguistics , vol. 48, no. 1, pp. 207–219, Mar. 2022. [Online]. Available: https://aclanthology.org/2022.cl-1.7
2022
-
[62]
Information-theoretic probing for linguistic structure,
T. Pimentel, J. Valvoda, R. H. Maudslay, R. Zmigrod, A. Williams, and R. Cotterell, “Information-theoretic probing for linguistic structure,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai, N. Schluter, and J. ...
2020
-
[63]
An information theoretic view on selecting linguistic probes,
Z. Zhu and F. Rudzicz, “An information theoretic view on selecting linguistic probes,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , B. Webber, T. Cohn, Y . He, and Y . Liu, Eds. Online: Association for Computational Lingui...
2020
-
[64]
Probing pretrained language models for lexical semantics,
I. Vuli ´c, E. M. Ponti, R. Litschko, G. Glavaˇs, and A. Korhonen, “Probing pretrained language models for lexical semantics,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), B. Webber, T. Cohn, Y . He, and Y . Liu, Eds. Onlin...
2020
-
[65]
Beyond binary: Towards fine-grained llm-generated text detection via role recognition and involvement measurement,
Z. Cheng, L. Zhou, F. Jiang, B. Wang, and H. Li, “Beyond binary: Towards fine-grained llm-generated text detection via role recognition and involvement measurement,” arXiv preprint arXiv:2410.14259, 2024
2024 arXiv
-
[66]
Semi-supervised classification with graph convolutional networks,
T. N. Kipf and M. Welling, “Semi-supervised classification with graph convolutional networks,” in Proc. of ICLR , 2017
2017
-
[67]
Towards deeper graph neural networks,
M. Liu, H. Gao, and S. Ji, “Towards deeper graph neural networks,” in Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining , 2020, pp. 338–348
2020
-
[68]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
-
[69]
What do you learn from context? probing for sentence structure in contextualized word representations,
I. Tenney, P. Xia, B. Chen, A. Wang, A. Poliak, R. T. McCoy, N. Kim, B. V . Durme, S. Bowman, D. Das, and E. Pavlick, “What do you learn from context? probing for sentence structure in contextualized word representations,” in Proc. of ICLR , 2019
2019
-
[70]
A gold standard dependency corpus for English,
N. Silveira, T. Dozat, M.-C. de Marneffe, S. Bowman, M. Connor, J. Bauer, and C. Manning, “A gold standard dependency corpus for English,” in Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC’14) , N. Calzolari, K. Choukri, T. Declerc...
2014
-
[71]
Neural-Davidsonian semantic proto-role labeling,
R. Rudinger, A. Teichert, R. Culkin, S. Zhang, and B. Van Durme, “Neural-Davidsonian semantic proto-role labeling,” in Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, E. Riloff, D. Chiang, J. Hockenmaier, and J. Tsujii, Eds. Brussels, Be...
2018
-
[72]
BERT: Pre- training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre- training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technolog...
2019
-
[73]
Exploring the limits of transfer learning with a unified text-to-text transformer,
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research, vol. 21, no. 140, pp. 1–67, 2020
2020
-
[74]
Llama 3 model card,
AI@Meta, “Llama 3 model card,” 2024
2024
-
[75]
LLM2Vec: Large language models are secretly powerful text encoders,
P. BehnamGhader, V . Adlakha, M. Mosbach, D. Bahdanau, N. Chapados, and S. Reddy, “LLM2Vec: Large language models are secretly powerful text encoders,” in First Conference on Language Modeling , 2024
2024
-
[76]
Stanza: A python natural language processing toolkit for many human languages,
P. Qi, Y . Zhang, Y . Zhang, J. Bolton, and C. D. Manning, “Stanza: A python natural language processing toolkit for many human languages,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations , A. Celikyilmaz and T....
2020
-
[77]
Kan: Kolmogorov-arnold networks,
Z. Liu, Y . Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Solja ˇci´c, T. Y . Hou, and M. Tegmark, “Kan: Kolmogorov-arnold networks,”arXiv preprint arXiv:2404.19756, 2024
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.