{"id":"d6383387-eaa1-401d-98e9-7b871d6cc8be","arxiv_id":"2412.12456","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey proposing a Data-Model-Task taxonomy for GNN-LLM integration, but with unreliable citations and overlapping categories.","lead":"This paper is a survey that groups recent research on combining graph neural networks with large language models into categories based on data type, model architecture, and task. It is mainly useful as a starting bibliography, but it contains many citation errors and includes unverifiable anonymous submissions, so it should not be treated as a reliable reference.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The taxonomy is not a systematic classification: the same methods are placed in mutually exclusive categories (§2, §4, §6), so the central claim of a novel, foundational framework is unsupported.","rationale":"I read the paper as attempting a comprehensive, foundational survey whose organizing contribution is the Data-Model-Task taxonomy. The reader's weakest-assumption analysis correctly identifies the load-bearing condition: for the taxonomy to be a systematic classification, the categories must be distinct and each method must be determinately placeable. The paper's own category lists show repeated violations: GraphBridge, GOFA, GraphFM, GraphProp, and others appear in multiple categories that the paper presents as alternatives. This is an internal inconsistency, not a disagreement with community consensus; it can be checked directly from the manuscript, and the cited references only make the duplication harder to resolve because 15 references are anonymous under-review submissions. The survey does contain useful dataset descriptions and method summaries, but those do not rescue the central claim of a novel, systematic framework. Because my concern matches the reader's weakest assumption and supports the existing REJECT verdict, no adjustment to the reader's verdict is needed.","tokens_in":28416,"tokens_out":5335,"duration_ms":45490,"concrete_test":"Build a method-to-category incidence table by parsing every named-method list in §2, §4, and §6 with its citation number. For each name that appears in more than one category within a single perspective, resolve the assignment against the cited paper; for references [1]–[15], resolve against the anonymized submission if available. The minimal decisive case is GraphBridge: check whether its original paper supports both 'Single-task & Single-domain' and 'Multi-task & Multi-domain' in §2 and both 'independent collaborative modules' and 'GNN-only' in §4. If any method remains in two categories without an explicit multi-label designation, the 'systematically categorize' claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is a classification framework that 'systematically categorize[s]' the GNN-LLM literature (§8). That claim requires the five categories in each perspective (data, model, task) to be coherent and each method to be assigned to one category. The paper's own lists violate this. In §2, GraphBridge appears under both 'Single-task & Single-domain' and 'Multi-task & Multi-domain'; in §4 it appears under both 'GNN and LLM as independent collaborative modules' and 'GNN-only.' GOFA is cited as [43] in §2 and as [78] in §4; GraphFM appears in 'Single-task & Multi-domain' (§2) and 'GNN-only' (§4); GraphProp appears in 'Single-task & Multi-domain' (§2) and 'LLM-enhanced GNN' (§4); GraphPrompt appears twice in §2. These are not subtle borderline cases: the same named method is placed in categories that the paper defines as distinct, or cited to different references. Since the categories are meant to organize the field, these inconsistencies undermine the 'novel classification framework' claim. The issue is compounded by 15 anonymous 'under review' references ([1]–[15]) that cannot be checked, making it impossible to verify whether duplicated entries are the same method or different ones. A reader cannot use the survey as a foundational reference if category assignment is not reproducible.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a survey of methods that integrate Graph Neural Networks (GNNs) and Large Language Models (LLMs) for learning on text-attributed graphs. The authors propose a classification framework built on three dimensions—Data, Model, and Task—and they assign the surveyed methods to five categories in each dimension (e.g., Section 2 for data, Section 4 for model architectures, Section 6 for training/application scenarios). The paper also tabulates datasets (Section 3), gives short method descriptions (Section 5), and discusses pre-training, fine-tuning, and inference phases (Section 7). The central claim, stated in the abstract and Section 8, is that this is a novel, systematic classification framework that can serve as a foundational reference for the field and that text can act as a medium for cross-domain generalization of graph learning models.","tokens_in":28706,"tokens_out":3920,"duration_ms":34224,"significance":"A reliable survey organizing the rapidly growing GNN-LLM literature would be valuable, and the choice of the Data-Model-Task triad is a reasonable organizing principle. The paper also provides useful dataset tables and maintains an open-source repository, which are helpful community resources. However, the significance of the contribution depends entirely on the validity and reproducibility of the proposed taxonomy. The manuscript's own lists violate the exclusivity of its categories and contain citation inconsistencies, so the claimed 'novel classification framework' is not currently supported. Because the core contribution is the taxonomy, these problems are load-bearing rather than cosmetic.","major_comments":[{"comment":"The proposed categories are not mutually exclusive as applied. GraphBridge is listed under both 'Single-task & Single-domain' and 'Multi-task & Multi-domain' in Section 2, and also under both 'GNN and LLM as independent collaborative modules' and 'GNN-only' in Section 4. GraphFM is listed under both 'LLM-enhanced GNN' and 'GNN-only' in Section 4, and GraphProp appears under both 'Single-task & Multi-domain' in Section 2 and 'LLM-enhanced GNN' in Section 4. Since Section 8 claims that the framework 'systematically categorize[s]' the field, categories within a single perspective must be exclusive; these duplicate placements invalidate the classification claim as stated.","section":"§2 vs §4"},{"comment":"The same method is cited inconsistently across sections. GOFA is cited as [43] in Section 2 and in the Section 5 method description, but appears as 'GOFA[78]' in Sections 4 and 6, where reference [78] is in fact AnyGraph, not GOFA. GraphProp is cited as [53] in Section 2, but reference [53] is GraphPrompt; GraphProp is later cited as [8] in Sections 4 and 6. These inconsistencies make it impossible for a reader to verify which method is being classified in each category.","section":"§2, §4, §5, §6"},{"comment":"Fifteen references, [1] through [15], are anonymous 'under review' ICLR 2024 submissions. These works are used as substantive entries in Sections 2, 4, and 6 and are described in detail in Section 5. Because their content cannot be checked, any category placement involving them is unverifiable, and duplicated entries such as GraphBridge[6] appearing in two categories cannot be resolved by reading the cited source. The survey should either remove these entries or clearly mark them as unverified and exclude them from the central taxonomy.","section":"References [1]–[15]"},{"comment":"The survey claims to be comprehensive and systematic, but it does not state its literature search strategy, inclusion criteria, or method-selection protocol. The lists in Sections 2 and 4 are labeled 'main works' and 'representative research papers,' which is not the same as a systematic categorization. Without explicit selection criteria, the 'comprehensive' and 'foundational' claims in the abstract and Section 8 are not supportable.","section":"§1, §2, §8"}],"minor_comments":[{"comment":"The method description for WalkLM is headed 'WalkFM [68]' in the text, while the reference list and Section 6 use 'WalkLM'; the heading should be consistent.","section":"§5"},{"comment":"The heading 'SimTEG [26]' should read 'SimTeG' to match the reference and the rest of the manuscript.","section":"§5"},{"comment":"There are typos in the tables: 'Tokoler' should be 'Tolokers' and 'conncetivity' should be 'connectivity.'","section":"Tables 3 and 4"},{"comment":"In Section 6, category (2), 'TAGA[57]' is inconsistent with Section 2 and the reference list, where TAGA is [87]; reference [57] is CIKM-KD.","section":"§6"},{"comment":"The dataset citations in the tables appear unreliable; for example, Cora is cited as [13] and [32] in Table 1, but those reference numbers correspond to anonymous OMOG and TAPE, not to the original Cora dataset sources.","section":"Tables 1–5"}],"recommendation":"major_revision","confidential_remarks":"The paper's central contribution is a classification framework, and the framework's internal inconsistencies are extensive enough that the current version cannot serve as a reference. I chose major_revision rather than reject because the taxonomy could in principle be repaired by reclassifying entries with a clear protocol, removing or bracketing the anonymous under-review references, and correcting the citation errors. However, the revision would be substantial: it would require re-checking every entry in Sections 2, 4, and 6 against the cited papers and against each other. The anonymous references are a particular concern for a survey, since they make a large fraction of the taxonomy unverifiable by the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a handy pointer into a fast-moving field, but it is not yet a reliable systematic survey. The taxonomy that is supposed to be the paper's main contribution is self-contradictory. The same methods land in mutually exclusive categories: GraphBridge appears in both Single-task & Single-domain and Multi-task & Multi-domain in Section 2, and in both independent collaborative modules and GNN-only in Section 4; GraphFM sits in both GNN-only and LLM-enhanced GNN in Section 4; GOFA is cited as [43] in Section 2 and [78] in Section 4. These are not borderline cases. The paper claims a novel, systematic classification framework in Section 8, and that claim is unsupported by its own lists.\n\nCredit where it is due: the paper gathers a broad slice of the recent GNN-LLM literature and provides dataset statistics plus short method summaries. The open-source repo is a genuinely useful index. For a newcomer, skimming this survey gives a quick sense of the landscape. That utility is real, but it is the utility of a blog post, not a foundational reference.\n\nThe soft spots are serious. The 15 anonymous under-review references ([1]–[15]) are a load-bearing problem: some of the most-cited methods in the survey (GraphBridge, GraphFM, GraphProp) are exactly these unverifiable entries. When a survey cites work that cannot be checked, the reader cannot trust the categories, the counts, or even whether two entries are the same method. There are also plain citation errors: SKETCH is cited as [40] when [40] is GraphAdapter, and TAGA appears as [87] in Section 2 but [57] in Section 6. These may seem like small slips, but in a survey the citations are the data, and this data is noisy.\n\nThe Data-Model-Task trichotomy is standard machine learning vocabulary, and similar LLM-as-enhancer/collaborator/predictor splits already circulate in the field. The paper does not compare its taxonomy with existing surveys, so the novelty claim is not established. If the authors are willing to do that positioning, fix the category assignments, and replace the anonymous references with published versions, a useful survey could emerge from this draft. As it stands, I would not cite it in my own work and I would not hand it to a student as a reliable map.\n\nVerdict: desk reject in current form, but invite a resubmission after major revision. The topic is timely, the material is there, and the flaws are fixable in principle. A serious referee could help the authors see where the taxonomy needs tightening, so in that sense it deserves referee time.","headline":"A useful but unreliable map of a hot field: the taxonomy that is supposed to be the paper's contribution is contradicted by the paper's own lists, and the heavy reliance on anonymous under-review papers makes it unusable as a reference in its current form.","tokens_in":29150,"tokens_out":3657,"would_cite":false,"duration_ms":31833,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey proposes a Data-Model-Task framework for GNN-LLM integration and argues that text-attributed graphs make text the medium for cross-domain graph generalization.","keywords":["graph neural networks","large language models","text-attributed graphs","graph foundation models","cross-domain generalization","graph reasoning","taxonomy","machine learning survey"],"falsifier":"Take a Multi-task & Multi-domain model from the survey and transfer it to a held-out text-attributed graph whose text is rich but whose structure follows an unusual distribution, such as mostly heterophilic edges or very long-range dependencies. If performance collapses to near-random while a domain-specific GNN trained on that graph keeps high accuracy, the claim that text alone can carry cross-domain generalization would be refuted; if the model transfers, the text-as-medium thesis survives.","tokens_in":28251,"feed_emoji":"🕸️","tokens_out":7088,"duration_ms":60338,"temperature":0.7,"pith_summary":"This paper surveys the growing body of work that combines graph neural networks (GNNs) with large language models (LLMs), and offers a way to organize it. Its central claim is that the three classic pillars of machine learning—data, model, and task—are the right lenses for understanding GNN-LLM integration. The paper argues that text-attributed graphs make text a common medium: a model that reads node and edge descriptions in natural language can transfer across domains that would otherwise require separate models. If the framework holds, it gives researchers a shared map of the field and a concrete route toward a graph foundation model.","feed_headline":"Text is the medium that lets one graph model cross domains","feed_subtitle":"A Data-Model-Task taxonomy sorts the GNN-LLM field and points toward graph foundation models.","key_machinery":"The central object is the Data-Model-Task classification framework itself. It organizes every included method into one of five data categories (single-task & single-domain, single-task & multi-domain, multi-task & single-domain, multi-task & multi-domain, and graph reasoning), one of five model categories (independent modules, GNN-enhanced LLM, LLM-enhanced GNN, GNN-only, and LLM-only), and one of five training and application categories (single-domain supervised, single-domain unsupervised, multi-domain supervised, multi-domain unsupervised, and few-shot and zero-shot inference). The framework does the argument's load-bearing work: it is what lets the survey claim that text can serve as a universal medium, because the same textual descriptions can be passed through any of the five model designs and evaluated across all three axes.","core_discovery":"On the paper's own terms, the discovery is that the GNN-LLM literature is not a scattering of ad hoc hybrids but a coherent space that can be classified along three dimensions. From the data side, LLMs supply high-quality semantic features for text-attributed graphs, improving data quality and enabling cross-domain generalization. From the model side, the paper distinguishes five architectures: GNN and LLM as independent collaborative modules, GNN-enhanced LLM, LLM-enhanced GNN, GNN-only, and LLM-only, with learnable integration seen as the path to a graph foundation model. From the task side, it identifies five training-and-application scenarios, from single-domain supervised fine-tuning to multi-domain unsupervised learning and few-shot and zero-shot inference. The unifying thesis is that text is the medium that lets a single graph model handle diverse tasks across different data domains.","pith_inferences":["Beyond the paper: if text is truly a universal medium, a graph foundation model should be stress-tested on graphs with sparse or noisy text, such as user-generated reviews or low-resource languages, where LLM features degrade; the survey does not single out this failure mode.","Beyond the paper: the taxonomy could be operationalized as a machine-readable registry where each method is tagged with coordinates on the three axes; such a registry would automatically expose overlaps like GraphBridge's dual placement, turning the taxonomy from a static survey into a living map.","Beyond the paper: the model-axis distinctions suggest a concrete experiment—hold the training scenario fixed and compare GNN-enhanced LLM versus LLM-enhanced GNN on the same text-attributed graph benchmark; the survey lists both as promising but does not specify when one should be preferred.","Beyond the paper: the graph-reasoning category implies that evaluation should move beyond node classification to question-answering and link-prediction settings; a next step would be a cross-domain graph-reasoning benchmark that combines the multi-domain datasets in the paper's tables with natural-language graph queries."],"forward_implications":["If text is a workable common medium, then graph models should be designed and evaluated for cross-domain transfer from the start, rather than trained per domain, and multi-task & multi-domain methods become the main line of development.","The five model categories give practitioners a design menu: choose independent modules for simplicity, GNN-enhanced LLM when reasoning is primary, LLM-enhanced GNN when structure is primary, and the learnable options when aiming at a graph foundation model.","The five training categories imply a matching rule: supervised fine-tuning suits single-domain deployments, while generalizable systems need unsupervised multi-domain pre-training followed by few-shot or zero-shot inference.","Graph reasoning is presented as the frontier where integrated models move beyond classification and prediction toward inference and question answering over graph structure."],"supporting_citations":[{"why":"Shows how a language model can be fine-tuned with neighborhood prediction to produce node features, grounding the data-perspective claim.","marker":"[23]"},{"why":"Supplies the LLM-to-LM interpreter pipeline used to convert LLM explanations into GNN features, supporting the LLM-enhanced GNN category.","marker":"[32]"},{"why":"Demonstrates joint optimization of a language model and GNN via variational expectation-maximization, a central example of model collaboration.","marker":"[89]"},{"why":"Introduces GOFA, a generative one-for-all graph language model that integrates GNN layers into an LLM, used as evidence for the graph foundation model direction.","marker":"[43]"},{"why":"GraphText translates graphs into text space for training-free reasoning, supporting the LLM-only and graph-reasoning categories.","marker":"[90]"},{"why":"Provides the NLGraph benchmark of graph reasoning problems, the basis for the task-perspective discussion of graph reasoning.","marker":"[72]"},{"why":"OFA unifies graph tasks via text-attributed graphs and nodes-of-interest prompting, a key example of single-model cross-domain generalization.","marker":"[52]"},{"why":"AnyGraph is used as a multi-domain unsupervised foundation model with mixture-of-experts, anchoring that training category.","marker":"[78]"}],"fun_headline_variants":["Text bridges GNNs and LLMs for cross-domain graphs","Unifying GNN-LLM models: text as the cross-domain glue","Survey maps GNN-LLM space, finds text is the key","From data to tasks: text powers graph foundation models","LLM+GNN synergy: text enables one graph model for all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The framework assumes that every method fits exactly one of the five data categories, one of the five model categories, and one of the five training categories—an assumption the survey's own lists strain, since methods such as GraphBridge are placed in several cells.","fun_headline_variants_meta":{"raw":{"variants":["Text bridges GNNs and LLMs for cross-domain graphs","Unifying GNN-LLM models: text as the cross-domain glue","Survey maps GNN-LLM space, finds text is the key","From data to tasks: text powers graph foundation models","LLM+GNN synergy: text enables one graph model for all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1366,"prompt_tokens":1016,"completion_tokens":350,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":632,"completion_tokens_details":{"reasoning_tokens":260}},"tokens_in":632,"tokens_out":350,"duration_ms":3767,"temperature":1.0,"reasoning_tokens":260,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:03:04.309611+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a Multi-task & Multi-domain model from the survey and transfer it to a held-out text-attributed graph whose text is rich but whose structure follows an unusual distribution, such as mostly heterophilic edges or very long-range dependencies. If performance collapses to near-random while a domain-specific GNN trained on that graph keeps high accuracy, the claim that text alone can carry cross-domain generalization would be refuted; if the model transfers, the text-as-medium thesis survives.","supporting_citations":[{"cited_title":"N ode feature extraction by self-supervised multi-scale neighborhood prediction","cited_arxiv_id":null,"evidence_quote":"Shows how a language model can be fine-tuned with neighborhood prediction to produce node features, grounding the data-perspective claim."},{"cited_title":"Harnessing explanations: Llm -to-lm interpreter for enhanced text-attributed graph representation learning","cited_arxiv_id":null,"evidence_quote":"Supplies the LLM-to-LM interpreter pipeline used to convert LLM explanations into GNN features, supporting the LLM-enhanced GNN category."},{"cited_title":"Learning on large-scale text-at tributed graphs via variational inference","cited_arxiv_id":null,"evidence_quote":"Demonstrates joint optimization of a language model and GNN via variational expectation-maximization, a central example of model collaboration."},{"cited_title":"Gofa: A generative o ne-for-all model for joint graph language modeling","cited_arxiv_id":null,"evidence_quote":"Introduces GOFA, a generative one-for-all graph language model that integrates GNN layers into an LLM, used as evidence for the graph foundation model direction."},{"cited_title":"Graphtext: Gra ph reasoning in text space","cited_arxiv_id":null,"evidence_quote":"GraphText translates graphs into text space for training-free reasoning, supporting the LLM-only and graph-reasoning categories."},{"cited_title":"Can language models solve graph problems in natural language? 2024","cited_arxiv_id":null,"evidence_quote":"Provides the NLGraph benchmark of graph reasoning problems, the basis for the task-perspective discussion of graph reasoning."},{"cited_title":"One for all: Towards tra ining one graph model for all classiﬁcation tasks","cited_arxiv_id":null,"evidence_quote":"OFA unifies graph tasks via text-attributed graphs and nodes-of-interest prompting, a key example of single-model cross-domain generalization."},{"cited_title":"Anygraph: Graph foundatio n model in the wild","cited_arxiv_id":null,"evidence_quote":"AnyGraph is used as a multi-domain unsupervised foundation model with mixture-of-experts, anchoring that training category."}],"review_version":1}