Pith. sign in

REVIEW 3 major objections 5 minor 181 references

Graph Data Management and Graph Machine Learning: Synergies and Opportunities

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Graph data management and graph machine learning reinforce each other at every stage of the data pipeline, and this survey is the first to map the relationship end to end.

desk verdict A useful survey map of the GDM-GML intersection, but the 'first survey' claim needs either a search protocol or a softer wording. read the letter →

arxiv 2502.00529 v1 pith:E5R63KUB submitted 2025-02-01 cs.DB

classification cs.DB
keywords graphdatamanagementmachinelearningneuralnetworksembeddingsgraph-basedvectorindexesGNNexplainabilityknowledgequeryansweringretrieval-augmentedgeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that graph data management (GDM) and graph machine learning (GML) form a two-way street: database-style techniques such as cleaning, indexing, and scalable systems improve graph learning, while graph learning and large language models improve database tasks such as query answering and knowledge-graph retrieval. It organizes this relationship through a five-phase graph data pipeline—cleaning and augmentation, embedding, GNN training, downstream tasks, and explainability—and labels where each side helps the other. The authors claim this is the first survey to cover the synergy over the full pipeline, since earlier surveys treated relational data management and ML, or graph ML alone. If the map is right, researchers and system builders can see concrete openings where database methods and graph ML should be co-designed rather than developed separately.

What carries the argument

The load-bearing device is the five-phase graph data pipeline of Figure 1: graph data extraction, integration, cleaning, and augmentation; graph embedding; GNN training; downstream tasks; and explainability, with end-to-end learning possible across phases. Each phase is tagged as a GDM task, a GML task, or both, and the survey's argument proceeds phase by phase, showing where the other discipline intervenes. The second structural device is the three-scenario classification: GDM benefits GML, GML benefits GDM, and GDM plus GML together serve downstream tasks. This classification is what turns a collection of examples into a map of the full two-way relationship.

What would settle it

A systematic literature search that finds a prior peer-reviewed survey or tutorial, published before February 2025, that explicitly covers both directions of the graph data management and graph machine learning synergy across the full graph pipeline would refute the paper's firstness claim.

Watch

Extended reading notes

Core claim

The central claim is that graph data management and graph machine learning are mutually reinforcing across the entire graph data science lifecycle, and that this interdependence has not been surveyed before. On one side, GDM contributes graph data cleaning and augmentation that improve GNN accuracy, distributed and parallel systems that make embedding and training scale to billion-edge graphs, graph-based vector indexes that make high-dimensional embeddings searchable, and view-based or queryable explanation structures that make GNN outputs understandable. On the other side, GML contributes inference over incomplete knowledge graphs, natural-language query translation, cardinality estimation and query optimization, and graph-based retrieval-augmented generation that grounds large language models with structured facts. The paper presents these as synergies already visible in existing systems, and concludes that the integration should be treated as a deliberate design goal.

Load-bearing premise

The survey's usefulness depends on the five-phase pipeline in Figure 1 being a complete and representative model of the graph data science lifecycle, and on the chosen examples being the most relevant illustrations of each synergy.

Editorial extensions

If this is right

  • Graph data cleaning and augmentation should be treated as a first-class step in GNN pipelines, since dirty or noisy graphs directly limit model accuracy.
  • Distributed training systems, graph partitioning, and graph databases can carry GNN training and embedding past billion-edge scale, making scalability a data-management problem as much as an ML problem.
  • Graph-based vector indexes make GNN-produced embeddings queryable, so vector data management becomes part of the graph ML workflow.
  • Knowledge graph query answering can absorb ML-based inference and natural-language interfaces, allowing answers on incomplete, schema-flexible graphs.
  • Graph retrieval-augmented generation can ground LLM outputs in structured facts, and graph databases become plausible semantic caches for LLM question-answer pairs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension: build a citation-driven map of the same five phases using explicit, reproducible inclusion criteria; if the selected examples shift substantially, the paper's illustrative choices may reflect the authors' research agenda more than the field's full landscape.
  • If the synergy claim holds, graph database vendors will likely fold embedding generation, vector search, and LLM grounding into core query engines rather than shipping them as add-on libraries.
  • The paper's future-direction point about cleaning graphs for robustness rather than only correctness suggests a concrete benchmark: evaluate GNN accuracy under label noise before and after cleaning that targets robustness metrics.
  • Graph RAG as a semantic cache implies a cost model: indexing question-answer pairs in a graph or vector space could reduce LLM API calls for repeated or similar queries; measuring hit-rate versus latency would test that promise.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey by Khan, Ke, and Wu reviews the two-way interaction between graph data management (GDM) and graph machine learning (GML) across an end-to-end graph data pipeline. The pipeline, shown in Figure 1, spans data cleaning and augmentation, graph embedding and GNN training, vector data management, explainability, knowledge-graph query answering, and graph-based retrieval-augmented generation for LLMs. The paper is organized around three scenarios: GDM benefits GML, GML benefits GDM, and GDM+GML jointly support downstream tasks. It surveys representative systems and algorithms (e.g., DistDGL, HNSW, GVEX, and graph RAG) and closes with future research directions. The central claim, stated in Section 5, is that this is the first survey exploring GDM-GML synergies over the end-to-end graph data pipeline.

Significance. If the claimed novelty and coverage hold, the survey fills a genuine gap: prior work has addressed ML-for-DM and DM-for-ML mostly for relational data, while graph-specific surveys have focused on one side (e.g., graph representation learning, GNNs, or data-centric GML) without organizing the two-way interaction around a pipeline. The paper's strengths are its broad and current topic coverage, its explicit 'Synergy' bullets that connect each area to the central GDM-GML theme, and its inclusion of emerging topics such as graph RAG and vector indexes. It also gives a practical overview of existing systems, which will help newcomers identify entry points. The survey is not technically derivational—there are no proofs or experiments—so its value rests on accurate characterization and comprehensive organization of prior work; both are generally reasonable, but the 'first' and 'comprehensive' status is asserted rather than demonstrated.

major comments (3)
  1. [Section 5, Related Work] The load-bearing claim that 'ours is the first survey exploring the synergies between graph data management and graph ML over the end-to-end graph data pipeline' is not supported by a systematic literature search. The paper gives no search protocol, no inclusion/exclusion criteria, no time window or venue list, and no comparison table mapping earlier surveys to the pipeline phases in Figure 1. The closest prior survey, [166] on data-centric graph ML, is set aside with a one-line distinction despite substantial overlap with Sections 3.1 and 3.2. To keep the 'first' and 'comprehensive' claims, the revision should supply this evidence; alternatively, the wording should be downgraded to something like 'a survey organized around the GDM-GML pipeline.'
  2. [Section 3.3, Graph-based Vector Data Indexes] The statement that graph-based approaches present 'unparalleled effectiveness' is an unsupported superlative and the citations given for it, [26] and [123], do not support the claim: [26] is a paper on graph dependencies and [123] is a materials-science KG exploration work, not an approximate nearest neighbor benchmark. Graph-based ANNS methods have known trade-offs against IVF, HNSW variants, and learned indexes, and the survey itself later discusses hardware-aware optimizations and hybrid methods, which suggests a more qualified phrasing is appropriate.
  3. [Sections 3.2-3.4 and 4.1-4.2, selection of examples] A disproportionate share of the systems highlighted as signature examples are the authors' own prior works: GVEX [16] and RoboGExp [91] for explainability, DistGER [27] for distributed embedding, GraphLingo [60] for KG-LLM exploration, MUST [121] and Starling [125] for vector indexes. No selection criteria are given for choosing these examples over alternatives, so a reader cannot determine whether the coverage is comprehensive or tailored to the authors' research agenda. The revision should either state the selection methodology or deliberately diversify the examples, especially for claims of 'comprehensive' coverage.
minor comments (5)
  1. [Section 3.1] The text says 'editing-based GP A' and 'editing-based GPA'; this appears to be a typo for 'GDA' (graph data augmentation).
  2. [Section 3.3] The phrase 'an order of magnitude increase in efficiency' is vague; specifying the comparison baseline and workload would make the claim more informative.
  3. [Section 3.4] The term 'F orwardexplainability' should be 'Forward explainability' (with a space), and the same formatting issue appears elsewhere in the paper.
  4. [Section 4.2] The sentence 'the later retrieves the most relevant paths' should read 'the latter retrieves...' since two categories are being contrasted.
  5. [Section 6] The phrase 'how to create a holistic embedding across multiple modalities' should use 'holistic embeddings' or a singular noun consistently; the current phrasing is slightly awkward.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's synergy thesis is supported by external systems, and the authors' own prior works are illustrative rather than load-bearing.

full rationale

This is a survey with no derived equations, no fitted parameters, and no empirical predictions that could be forced by construction. The central claim that graph data management and graph machine learning mutually reinforce each other across a five-phase pipeline is supported by a broad set of external systems and results, including DistDGL, HNSW, DiskANN, PyTorch BigGraph, Query2box, and graph RAG works. The authors' own prior systems (GVEX, RoboGExp, DistGER, GraphLingo, MUST, Starling) appear as illustrative examples of each synergy area, not as the evidence establishing that the synergy exists; removing them would not collapse the survey's argument. The §5 assertion that this is 'the first survey exploring the synergies between graph data management and graph ML' is a novelty and coverage claim that lacks a stated search protocol, but an unsupported novelty claim is a verification gap, not circularity: it does not reduce to its own input by definition, and no equation or fitted value is reused as a prediction. Accordingly, no significant circularity is present.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

The survey introduces no free parameters and no new entities. Its central claim rests on two domain assumptions: the completeness of the pipeline model and the representativeness of the selected literature, the latter weakened by self-citation.

assumptions (2)
  • domain assumption The five-phase pipeline (cleaning/augmentation, embedding, GNN training, downstream tasks, explainability) in Figure 1 is a representative model of the graph data science pipeline.
    The entire survey (sections 3.1-3.4 and 4) is organized around this pipeline; if it omits important phases such as storage engines, transaction processing, or data acquisition, the synergy map is incomplete. Introduced in §1 with Figure 1.
  • domain assumption The papers chosen to illustrate each synergy are the most representative recent works in the field.
    No systematic search or inclusion/exclusion criteria are provided; the selection relies on the authors' judgment, and many illustration examples are the authors' own publications (e.g., GVEX, RoboGExp, DistGER, GraphLingo, MUST, Starling). Stated implicitly throughout §3 and §4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Graph Data Management and Graph Machine Learning: Synergies and Opportunities." pith.science (2026). https://pith.science/paper/E5R63KUB

@misc{pith2026250200529,
  author       = {Pith},
  title        = {Pith review of: Graph Data Management and Graph Machine Learning: Synergies and Opportunities},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E5R63KUB}},
  note         = {Machine review of arXiv:2502.00529}
}
read the original abstract

The ubiquity of machine learning, particularly deep learning, applied to graphs is evident in applications ranging from cheminformatics (drug discovery) and bioinformatics (protein interaction prediction) to knowledge graph-based query answering, fraud detection, and social network analysis. Concurrently, graph data management deals with the research and development of effective, efficient, scalable, robust, and user-friendly systems and algorithms for storing, processing, and analyzing vast quantities of heterogeneous and complex graph data. Our survey provides a comprehensive overview of the synergies between graph data management and graph machine learning, illustrating how they intertwine and mutually reinforce each other across the entire spectrum of the graph data science and machine learning pipeline. Specifically, the survey highlights two crucial aspects: (1) How graph data management enhances graph machine learning, including contributions such as improved graph neural network performance through graph data cleaning, scalable graph embedding, efficient graph-based vector data management, robust graph neural networks, user-friendly explainability methods; and (2) how graph machine learning, in turn, aids in graph data management, with a focus on applications like query answering over knowledge graphs and various data science tasks. We discuss pertinent open problems and delineate crucial research directions.

Figures

Figures reproduced from arXiv: 2502.00529 by the authors.

Figure 1
Figure 1. Graph data pipeline in data science and machine learning applications. Graph embedding can be task-specific or task-agnostic. Graph neural network (GNN) training can be end-to-end based on downstream tasks. We show which phases belong to GDM and which belong to GML, and can bene￾fit from each other. sential, as data is foundational to both. First, effec￾tive collaboration between data management and ML is necessary … view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

181 extracted references · 74 canonical work pages

  1. [166]

    Zhang, Q

    Z. Zhang, Q. Liu, H. Wang, C. Lu, and C. Lee. ProtGNN: Towards Self-Explaining Graph Neural Networks. In AAAI, 2022

  2. [26]

    M. Chen, Z. Wei, B. Ding, Y . Li, Y . Y uan, X. Du, and J. Wen. Scalable Graph Neural Networks via Bidirectional Propagation. In NeurIPS, 2020

  3. [123]

    Y . Tian. The world of graph databases from an industry perspective. SIGMOD Rec., 51(4):60–67, 2022

  4. [16]

    Antelmi, G

    A. Antelmi, G. Cordasco, M. Polato, V . Scarano, C. Spagnuolo, and D. Y ang. A Survey on Hypergraph Representation Learning. ACM Computing Surveys, 56(1):1–38, 2023

  5. [91]

    disturbance

    introduces a new class of explanation structures to provide robust, both counterfactual and factual ex- planations for graph neural networks. Given a GNN, a robust explanation refers to the fraction of a graph that are counterfactual and factual explanation of the re- sults of the GNN over the graph, but also remains so for any “disturbance” by flipping up...

  6. [27]

    T. Chen, D. Qiu, Y . Wu, A. Khan, X. Ke, and Y . Gao. View-based Explanations for Graph Neural Networks. Proc. ACM Manag. Data, 2(1):40:1–40:27, 2024

  7. [60]

    Juillard, A

    P . Juillard, A. Bonifati, and A. Maur. Interactive Grap h Repairs for Neighborhood Constraints. In EDBT, 2024

  8. [121]

    D. Tang, J. Wang, R. Chen, L. Wang, W. Y u, J. Zhou, and K. Li. XGNN: Boosting Multi-GPU GNN Training via Global GNN MemoryStore. PVLDB, 17(5):1105–1118, 2024

  9. [125]

    V asimuddin, S

    M. V asimuddin, S. Misra, G. Ma, R. Mohanty, E. Georgana s, A. Heinecke, D. D. Kalamkar, N. K. Ahmed, and S. Avancha. Distgnn: Scalable distributed training for large-scale gr aph neural networks. In SC, 2021

Show all 181 references
  1. [1]

    Abdallah and E

    H. Abdallah and E. Mansour. Towards A GML-enabled Knowledge Graph Platform. In ICDE, 2023

  2. [2]

    Graph neural networks (GNNs) are deep learning models to tackle graph-related tasks in an end-to-end manner [135]

    BACKGROUND We introduce background materials on graph neural networks and graph embeddings. Graph neural networks (GNNs) are deep learning models to tackle graph-related tasks in an end-to-end manner [135]. GNNs have many variants, e.g., graph convolutional network (GCN) [57],...

  3. [3]

    editing-based

    GRAPH DATA MANAGEMENT FOR GRAPH ML We discuss applications of graph data management such as data cleaning and augmentation in improving the GNN performance, graph algorithms, databases, and systems for scalable embedding learning, graph indexes for vector data management, and ...

  4. [4]

    4.1 Knowledge Graphs Query Answering Query answering over datasets is an important data management task

    GRAPH ML FOR GRAPH DATA MANAGEMENT We illustrate applications of graph machine learning and graph-based LLMs in knowledge graphs query an- swering and other data science tasks. 4.1 Knowledge Graphs Query Answering Query answering over datasets is an important data management t...

  5. [5]

    However, graph data result in unique challenges to both data management and ML (§1), justifying the importance of our survey

    RELATED WORK The closest to our work are surveys and tutorials on ML for data management and data management for ML, emphasizing on relational data and RDBMS [14, 59, 88, 43]. However, graph data result in unique challenges to both data management and ML (§1), justifying the i...

  6. [6]

    Real-time Graph Learning and Inference

    FUTURE DIRECTIONS Future work can be in several directions. Real-time Graph Learning and Inference . The in- tegration of spatiotemporal GNNs and dynamic graphs would enable the real-time decision that can rapidly ex- plore evolving nodes and links. This calls for adaptive gra...

  7. [7]

    Ke is supported by Zhejiang Province’s “Lingyan” R&D Project under Grant No

    ACKNOWLEDGMENT Khan acknowledges support from the Novo Nordisk Foundation grant NNF 22OC0072415. Ke is supported by Zhejiang Province’s “Lingyan” R&D Project under Grant No. 2024C01259 and Y ongjiang Talent Intro- duction Programme (2022A-237-G). Wu is supported by NSF under C...

  8. [8]

    M. Bawa, T. Condie, and P . Ganesan. LSH Forest: Self-tuni ng Indexes for Similarity Search. In WWW, 2005

  9. [9]

    Borgeaud, A

    S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherfor d, K. Millican, G. v. d. Driessche, J.-B. Lespiau, B. Damoc, A. Clark, D. d. L. Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals,...

  10. [10]

    Bruna, W

    J. Bruna, W. Zaremba, A. Szlam, and Y . LeCun. Spectral Networks and Locally Connected Networks on Graphs. In ICLR, 2014

  11. [11]

    H. Cai, V . W. Zheng, and K. C. Chang. A Comprehensive Survey of Graph Embedding: Problems, Techniques, and Applications. IEEE Trans. Knowl. Data Eng. , 30(9):1616–1637, 2018

  12. [12]

    L. Cai, J. Li, J. Wang, and S. Ji. Line Graph Neural Networ ks for Link Prediction. IEEE Trans. Pattern Anal. Mach. Intell., 44(9):5103–5113, 2022

  13. [13]

    M. Ali, M. Berrendorf, C. T. Hoyt, L. V ermue, M. Galkin, S. Sharifzadeh, A. Fischer, V . Tresp, and J. Lehmann. Bringing Light Into the Dark: A Large-Scale Evaluation of Knowledge Graph Embedding Models Under a Unified Framework. IEEE Trans. Pattern Anal. Mach. Intell., 44(12)...

  14. [14]

    Andoni, P

    A. Andoni, P . Indyk, T. Laarhoven, I. Razenshteyn, and L. Schmidt. Practical and Optimal LSH for Angular Distance. In NeurIPS, 2015

  15. [15]

    Andoni, P

    A. Andoni, P . Indyk, and I. Razenshteyn. Approximate Nea rest Neighbor Search in High Dimensions. In International Congress of Mathematicians: Rio de Janeiro , 2018

  16. [17]

    Babenko and V

    A. Babenko and V . Lempitsky. The Inverted Multi-index. IEEE Trans. on Pattern Anal. Mach. Intell., 37(6):1247–1260, 2014

  17. [18]

    C. D. T. Barros, M. R. F. Mendonça, A. B. Vieira, and A. Ziviani. A Survey on Embedding Dynamic Graphs. ACM Comput. Surv., 55(1), 2021

  18. [19]

    Datta, M

    A. Datta, M. Fredrikson, K. Leino, K. Lu, S. Sen, and Z. Wang. Machine Learning Explainability and Robustness: Connected at the Hip. In KDD, 2021

  19. [20]

    W. Dong, C. Moses, and K. Li. Efficient k-nearest Neighbo r Graph Construction for Generic Similarity Measures. In WWW, 2011

  20. [21]

    In spectral methods, filters are ap- plied on a graph’s frequency modes computed via graph Fourier transform

    approaches. In spectral methods, filters are ap- plied on a graph’s frequency modes computed via graph Fourier transform. Spectral formulations rely on the fixed spectrum of the graph Laplacian, and are suitable only for graphs with a single structure (and varying fea- tures on ...

  21. [22]

    Echihabi, T

    K. Echihabi, T. Palpanas, and K. Zoumpatianos. New Tren ds in High-D V ector Similarity Search: AI-driven, Progressive, and Distributed. PVLDB, 14(12):3198–3201, 2021

  22. [23]

    Faber, A

    L. Faber, A. K. Moghaddam, and R. Wattenhofer. When Comparing to Ground Truth is Wrong: On Evaluating GNN Explanation Methods. In KDD, 2021

  23. [24]

    C. Chai, N. Tang, J. Fan, and Y . Luo. Demystifying Artific ial Intelligence for Data Preparation. In SIGMOD, 2023

  24. [25]

    C. Chai, J. Wang, Y . Luo, Z. Niu, and G. Li. Data Managemen t for Machine Learning: A Survey. IEEE Trans. Knowl. Data Eng., 35(5):4646–4667, 2023

  25. [28]

    P . Cui, X. Wang, J. Pei, and W. Zhu. A Survey on Network Embedding. IEEE Trans. Knowl. Data Eng. , 31(5):833–852, 2019

  26. [29]

    Dai and S

    E. Dai and S. Wang. Towards Self-Explainable Graph Neur al Network. In CIKM, 2021

  27. [30]

    Gao and C

    J. Gao and C. Long. High-Dimensional Approximate Neare st Neighbor Search: with Reliable and Efficient Distance Comparison Operations. PACMMOD, 1(2):1–27, 2023

  28. [31]

    Gollapudi, N

    S. Gollapudi, N. Karia, V . Sivashankar, R. Krishnaswam y, N. Begwani, S. Raz, Y . Lin, Y . Zhang, N. Mahapatro, P . Srinivasan, et al. Filtered-DiskANN: Graph Algorithms for Approximate Nearest Neighbor Search with Filters. In WWW, 2023

  29. [32]

    Duvenaud, D

    D. Duvenaud, D. Maclaurin, J. Aguilera-Iparraguirre, R. Gómez-Bombarelli, T. Hirzel, A. Aspuru-Guzik, and R. P . Adams. Convolutional Networks on Graphs for Learning Molecular Fingerprints. In NeurIPS, 2015

  30. [33]

    S. Guan, H. Ma, P . Lin, and Y . Wu. GEDet: Adversarially Learned Few-shot Detection of Erroneous Nodes in Graphs. In IEEE BigData, 2020

  31. [34]

    S. Guan, H. Ma, M. Wang, and Y . Wu. GALE: Active Adversarial Learning for Erroneous Node Detection in Graphs. In ICDE, 2023

  32. [35]

    W. Fan, Z. Fan, C. Tian, and X. L. Dong. Keys for Graphs. PVLDB, 8(12):1590–1601, 2015

  33. [36]

    W. Fan, R. Jin, M. Liu, P . Lu, C. Tian, and J. Zhou. Capturi ng Associations in Graphs. PVLDB, 13(11):1863–1876, 2020

  34. [37]

    Fan and P

    W. Fan and P . Lu. Dependencies for Graphs. TODS, 44(2):5:1–5:40, 2019

  35. [38]

    P . Fang, A. Khan, S. Luo, F. Wang, D. Feng, Z. Li, W. Yin, an d Y . Cao. Distributed Graph Embedding with Information-Oriented Random Walks. PVLDB, 16(7):1643–1656, 2023

  36. [39]

    C. Fu, C. Xiang, C. Wang, and D. Cai. Fast Approximate Nearest Neighbor Search with the Navigating Spreading-out Graph. PVLDB, 12:461––474, 2019

  37. [40]

    Funke, M

    T. Funke, M. Khosla, M. Rathee, and A. Anand. Zorro: V ali d, Sparse, and Stable Explanations in Graph Neural Networks. IEEE Trans. Knowl. Data Eng. , 35(8):8687–8698, 2023

  38. [41]

    Huang, X

    P .-S. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck. Learning Deep Structured Semantic Models for Web Search Using Clickthrough Data. In CIKM, 2013

  39. [42]

    (5) KG embedding methods can be useful as well

    for simple NLQs and EmbedKGQA [99] for multi- hop NLQs. (5) KG embedding methods can be useful as well. Wang et al. [128, 127] decompose multi-hop and complex queries into smaller subqueries, answer each subquery via single-hop reasoning with KG embedding, and then assemble th...

  40. [43]

    Grover and J

    A. Grover and J. Leskovec. Node2vec: Scalable Feature Learning for Networks. In KDD, 2016

  41. [44]

    J. Jang, H. Choi, H. Bae, S. Lee, M. Kwon, and M. Jung. CXL-ANNS: Software-Hardware Collaborative Memory Disaggregation and Computation for Billion-Scale Approximate Nearest Neighbor Search. In USENIX ATC, 2023

  42. [45]

    Jayaram Subramanya, F

    S. Jayaram Subramanya, F. Devvrit, H. V . Simhadri, R. Krishnawamy, and R. Kadekodi. Diskann: Fast Accurate Billion-point Nearest Neighbor Search on a Single Node. NeurIPS, 32, 2019

  43. [46]

    Gupta, A

    P . Gupta, A. Goel, J. Lin, A. Sharma, D. Wang, and R. Zadeh . WTF: The Who to Follow Service at Twitter. In WWW, 2013

  44. [47]

    W. L. Hamilton, Z. Ying, and J. Leskovec. Inductive Representation Learning on Large Graphs. In NeurIPS, 2017

  45. [48]

    X. Han, Z. Jiang, N. Liu, and X. Hu. G-Mixup: Graph Data Augmentation for Graph Classification. In ICML, 2022

  46. [49]

    X. He, Y . Tian, Y . Sun, N. V . Chawla, T. Laurent, Y . LeCun, X. Bresson, and B. Hooi. G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering. arXiv preprint arXiv:2402.07630, 2024

  47. [50]

    Horchidan and P

    S. Horchidan and P . Carbone. ORB: Empowering Graph Queries through Inference. In Joint Proceedings of the ESWC 2023 W orkshops and Tutorials, 2023

  48. [51]

    J. Hu, R. Cheng, Z. Huang, Y . Fang, and S. Luo. On Embedding Uncertain Graphs. In CIKM, 2017

  49. [52]

    A. Khan. Knowledge Graphs Querying. SIGMOD Rec., 52(2):18–29, 2023

  50. [53]

    Huang, J

    X. Huang, J. Zhang, D. Li, and P . Li. Knowledge Graph Embedding Based Question Answering. In WSDM, 2019

  51. [54]

    Hulsebos, X

    M. Hulsebos, X. Deng, H. Sun, and P . Papotti. Models and Practice of Neural Table Representations. In SIGMOD, 2023

  52. [55]

    Kim and F

    B. Kim and F. Doshi-V elez. Machine Learning Techniques for Accountability. AI Mag., 42(1):47–52, 2021

  53. [56]

    T. N. Kipf and M. Welling. V ariational Graph Auto-Encod ers. In NeurIPS W orkshop on Bayesian Deep Learning, 2016

  54. [57]

    Jegou, M

    H. Jegou, M. Douze, and C. Schmid. Product Quantization for Nearest Neighbor Search. IEEE Trans. Pattern Anal. Mach. Intell., 33(1):117–128, 2010

  55. [58]

    B. Jin, G. Liu, C. Han, M. Jiang, H. Ji, and J. Han. Large Language Models on Graphs: A Comprehensive Survey. IEEE Trans. Knowl. Data Eng., 36(12):8622–8642, 2024

  56. [59]

    W. Jin, L. Zhao, S. Zhang, Y . Liu, J. Tang, and N. Shah. Gra ph Condensation for Graph Neural Networks. In ICLR, 2022

  57. [61]

    Kakkad, J

    J. Kakkad, J. Jannu, K. Sharma, C. C. Aggarwal, and S. Medya. A Survey on Explainability of Graph Neural Networks. IEEE Data Eng. Bull. , 46(2):35–63, 2023

  58. [62]

    Keivani and K

    O. Keivani and K. Sinha. Improved Nearest Neighbor Sear ch Using Auxiliary Information and Priority Functions. In ICML, 2018

  59. [63]

    C. Li, M. Zhang, D. G. Andersen, and Y . He. Improving Approximate Nearest Neighbor Search through Learned Adaptive Early Termination. In SIGMOD, 2020

  60. [64]

    Khan and E

    A. Khan and E. B. Mobaraki. Interpretability Methods fo r Graph Neural Networks. In DSAA, 2023

  61. [65]

    Khandelwal, O

    U. Khandelwal, O. Levy, D. Jurafsky, L. Zettlemoyer, an d M. Lewis. Generalization through Memorization: Nearest Neighbor Language Models. In ICLR, 2020

  62. [66]

    Y . Li, V . Zakhozhyi, D. Zhu, and L. J. Salazar. Domain Specific Knowledge Graphs as a Service to the Public. In KDD, 2020

  63. [67]

    H. Lin, M. Y an, X. Y e, D. Fan, S. Pan, W. Chen, and Y . Xie. A Comprehensive Survey on Distributed Training of Graph Neural Networks. Proc. IEEE, 111(12):1572–1606, 2023

  64. [68]

    T. N. Kipf and M. Welling. Semi-Supervised Classificati on with Graph Convolutional Networks. In ICLR, 2017

  65. [69]

    Klicpera, A

    J. Klicpera, A. Bojchevski, and S. Günnemann. Predict t hen Propagate: Graph Neural Networks meet Personalized PageRank. In ICLR, 2019

  66. [70]

    Kumar, M

    A. Kumar, M. Boehm, and J. Y ang. Data Management in Machine Learning: Challenges, Techniques, and Systems. In SIGMOD, 2017

  67. [71]

    D. Le, K. Zhao, M. Wang, and Y . Wu. GraphLingo: Domain Knowledge Exploration by Synchronizing Knowledge Graphs and Large Language Models. In ICDE, 2024

  68. [72]

    Lerer, L

    A. Lerer, L. Wu, J. Shen, T. Lacroix, L. Wehrstedt, A. Bos e, and A. Peysakhovich. Pytorch-BigGraph: A Large Scale Graph Embedding System. In MLSys, 2019

  69. [73]

    B. Li, W. Wang, Y . Sun, L. Zhang, M. A. Ali, and Y . Wang. GraphER: Token-Centric Entity Resolution with Graph Convolutional Neural Networks. In AAAI, 2020

  70. [74]

    K. Lu, C. Xiao, and Y . Ishikawa. Probabilistic Routing f or Graph-Based Approximate Nearest Neighbor Search. In ICML, 2024

  71. [75]

    W. Li, Y . Zhang, Y . Sun, W. Wang, M. Li, W. Zhang, and X. Lin. Approximate Nearest Neighbor Search on High Dimensional Data—Experiments, Analyses, and Improvement. IEEE Trans. Knowl. Data Eng. , 32(8):1475–1488, 2019

  72. [76]

    Y . Li, Z. Li, P . Wang, J. Li, X. Sun, H. Cheng, and J. X. Y u. A Survey of Graph Meets Large Language Model: Progress and Future Directions. In IJCAI, 2024

  73. [77]

    Y . A. Malkov and D. A. Y ashunin. Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs. IEEE Trans. Pattern Anal. Mach. Intell., 42(4):824–836, 2018

  74. [78]

    M. D. Manohar, Z. Shen, G. Blelloch, L. Dhulipala, Y . Gu, H. V . Simhadri, and Y . Sun. ParlayANN: Scalable and Deterministic Parallel Graph-Based Approximate Nearest Neighbor Search Algorithms. In PPoPP24, 2024

  75. [79]

    P . Lin, Q. Song, Y . Wu, and J. Pi. Repairing Entities usin g Star Constraints in Multirelational Graphs. In ICDE, 2020

  76. [80]

    X. Lin, Y . Peng, B. Choi, and J. Xu. Human-Powered Data Cleaning for Probabilistic Reachability Queries on Uncert ain Graphs. IEEE Trans. Knowl. Data Eng. , 29(7):1452–1465, 2017

  77. [81]

    Z. C. Lipton. The Mythos of Model Interpretability. Commun. ACM, 61(10), 2018

  78. [82]

    C. Liu, H. Liu, H. Jin, X. Liao, Y . Zhang, Z. Duan, J. Xu, an d H. Li. ReGNN: A ReRAM-Based Heterogeneous Architecture for General Graph Neural Networks. In Proceedings of the 59th ACM/IEEE Design Automation Conference, 2022

  79. [83]

    J. Liu, C. Y ang, Z. Lu, J. Chen, Y . Li, M. Zhang, T. Bai, Y . Fang, L. Sun, P . S. Y u, and C. Shi. Towards Graph Foundation Models: A Survey and Beyond. CoRR, abs/2310.11829, 2023

  80. [84]

    K. Lu, M. Kudo, C. Xiao, and Y . Ishikawa. HVS: Hierarchic al Graph Structure Based on V oronoi Diagrams for Solving Approximate Nearest Neighbor Search. PVLDB, 15(2):246–258, 2021

  81. [85]

    R. D. Pasquale and S. Represa. Empowering Domain-Speci fic Language Models with Graph-Oriented Databases: A Paradigm Shift in Performance and Model Maintenance. CoRR, abs/2410.03867, 2024

  82. [86]

    D. Luo, W. Cheng, D. Xu, W. Y u, B. Zong, H. Chen, and X. Zhang. Parameterized Explainer for Graph Neural Network. In NeurIPS, 2020

  83. [87]

    Ma and J

    Y . Ma and J. Tang. Deep Learning on Graphs . Cambridge University Press, 2021

  84. [88]

    Polyzotis, S

    N. Polyzotis, S. Roy, S. E. Whang, and M. Zinkevich. Data Management Challenges in Production Machine Learning. In SIGMOD, 2017

  85. [89]

    P . E. Pope, S. Kolouri, M. Rostami, C. E. Martin, and H. Hoffmann. Explainability Methods for Graph Convolutional Neural Networks. In CVPR, 2019

  86. [90]

    Mavromatis and G

    C. Mavromatis and G. Karypis. GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning. CoRR, 2024

  87. [92]

    E. Min, R. Chen, Y . Bian, T. Xu, K. Zhao, W. Huang, P . Zhao, J. Huang, S. Ananiadou, and Y . Rong. Transformer for Graphs: An Overview from Architecture Perspective. CoRR, abs/2202.08455, 2022

  88. [93]

    Mohoney, R

    J. Mohoney, R. Waleffe, H. Xu, T. Rekatsinas, and S. V enkataraman. Marius: Learning Massive Graph Embeddings on a Single Machine. In OSDI, 2021

  89. [94]

    Ootomo, A

    H. Ootomo, A. Naruse, C. Nolet, R. Wang, T. Feher, and Y . Wang. CAGRA: Highly Parallel Graph Construction and Approximate Nearest Neighbor Search for GPUs. In ICDE, 2024

  90. [95]

    J. J. Pan, J. Wang, and G. Li. Survey of V ector Database Management Systems. VLDB J., 33(5):1591–1615, 2024

  91. [96]

    S. Pan, L. Luo, Y . Wang, C. Chen, J. Wang, and X. Wu. Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Trans. Knowl. Data Eng. , 36(7):3580–3599, 2024

  92. [97]

    and follow-up works train on multi-hop queries – they embed multi-hop logic queries and their answers (i.e., entities from a KG) in the same embedding space to reduce the query processing cost via inference. Domain-specific knowledge graphs (KGs) [123, 66] have been curated to ...

  93. [98]

    Patel, P

    L. Patel, P . Kraft, C. Guestrin, and M. Zaharia. ACORN: Performant and Predicate-Agnostic Search Over V ector Embeddings and Structured Data. In SIGMOD, 2024

  94. [99]

    Y . Peng, B. Choi, T. N. Chan, J. Y ang, and J. Xu. Efficient Approximate Nearest Neighbor Search in Multi-dimensional Databases. PACMMOD, 1(1):1–27, 2023

  95. [100]

    M. S. Schlichtkrull, N. D. Cao, and I. Titov. Interpret ing Graph Neural Networks for NLP with Differentiable Edge Masking. In ICLR, 2021

  96. [101]

    Pradhan, A

    R. Pradhan, A. Lahiri, S. Galhotra, and B. Salimi. Expla inable AI: Foundations, Applications, Opportunities for Data Management Research. In ICDE, 2022

  97. [102]

    D. Qiu, M. Wang, A. Khan, and Y . Wu. Generating Robust Counterfactual Witnesses for Graph Neural Networks. In ICDE, 2024

  98. [103]

    J. Qiu, Y . Dong, H. Ma, J. Li, K. Wang, and J. Tang. Network Embedding as Matrix Factorization: Unifying DeepWalk, LINE, PTE, and Node2vec. In WSDM, 2018

  99. [104]

    Quamar, V

    A. Quamar, V . Efthymiou, C. Lei, and F. Özcan. Natural Language Interfaces to Data. F ound. Trends Databases, 11(4):319–414, 2022

  100. [105]

    Rabbani, M

    K. Rabbani, M. Lissandrini, and K. Hose. Shactor: impro ving the quality of large-scale knowledge graphs with validatin g shapes. In Companion of the International Conference on Management of Data (SIGMOD) , pages 151–154, 2023

  101. [106]

    Rathee, T

    M. Rathee, T. Funke, A. Anand, and M. Khosla. Bagel: A Benchmark for Assessing Graph Neural Network Explanations. CoRR, abs/2206.13983, 2022

  102. [107]

    H. Ren, M. Galkin, M. Cochez, Z. Zhu, and J. Leskovec. Neural Graph Reasoning: Complex Logical Query Answering Meets Graph Databases. CoRR, abs/2303.14617, 2023

  103. [108]

    H. Ren, W. Hu, and J. Leskovec. Query2box: Reasoning ove r Knowledge Graphs in V ector Space Using Box Embeddings. In ICLR, 2020

  104. [109]

    J. Ren, M. Zhang, and D. Li. HM-ANN: efficient billion-po int nearest neighbor search on heterogeneous memory. In NeurIPS, 2020

  105. [110]

    Saxena, A

    A. Saxena, A. Tripathi, and P . P . Talukdar. Improving Multi-hop Question Answering over Knowledge Graphs using Knowledge Base Embeddings. In ACL, 2020

  106. [111]

    J. Tang, Y . Y ang, W. Wei, L. Shi, L. Su, S. Cheng, D. Yin, a nd C. Huang. GraphGPT: Graph Instruction Tuning for Large Language Models. In SIGIR, pages 491–500, 2024

  107. [112]

    M. S. Schlichtkrull, T. N. Kipf, P . Bloem, R. van den Ber g, I. Titov, and M. Welling. Modeling Relational Data with Graph Convolutional Networks. In ESWC, 2018

  108. [113]

    Schnake, O

    T. Schnake, O. Eberle, J. Lederer, S. Nakajima, K. T. Sc hütt, K. Müller, and G. Montavon. Higher-Order Explanations of Graph Neural Networks via Relevant Walks. IEEE Trans. Pattern Anal. Mach. Intell., 44(11):7581–7596, 2022

  109. [114]

    Y . Shao, H. Li, X. Gu, H. Yin, Y . Li, X. Miao, W. Zhang, B. Cui, and L. Chen. Distributed Graph Neural Network Training: A Survey. ACM Comput. Surv., 56(8):191:1–191:39, 2024

  110. [115]

    J. Si, X. Gan, T. Xiao, B. Y ang, D. Dong, and Z. Pang. STEGNN: Spatial-Temporal Embedding Graph Neural Networks for Road Network Forecasting. In ICPADS, 2022

  111. [116]

    J. Sun, C. Xu, L. Tang, S. Wang, C. Lin, Y . Gong, H. Shum, and J. Guo. Think-on-Graph: Deep and Responsible Reasoning of Large Language Model with Knowledge Graph. In ICLR, 2024

  112. [117]

    L. Sun, L. He, Z. Huang, B. Cao, C. Xia, X. Wei, and P . S. Y u . Joint Embedding of Meta-Path and Meta-Graph for Heterogeneous Information Networks. In ICBK, 2018

  113. [118]

    L. Sun, Z. Tao, Y . Li, and H. Arakawa. ODA: Observation-Driven Agent for integrating LLMs and Knowledge Graphs. In ACL, 2024

  114. [119]

    Suresh, P

    S. Suresh, P . Li, C. Hao, and J. Neville. Adversarial Gr aph Augmentation to Improve Graph Contrastive Learning. In NeurIPS, 2021

  115. [120]

    J. Tan, S. Geng, Z. Fu, Y . Ge, S. Xu, Y . Li, and Y . Zhang. Learning and Evaluating Graph Neural Network Explanations Based on Counterfactual and Factual Reasoning. In W eb Conference, 2022

  116. [122]

    M. Wang, L. Lv, X. Xu, Y . Wang, Q. Y ue, and J. Ni. An Efficient and Robust Framework for Approximate Nearest Neighbor Search with Attribute Constraint. In NeurIPS, 2023

  117. [124]

    Y . Tian, X. Zhao, and X. Zhou. DB-LSH 2.0: Locality-Sensitive Hashing With Query-Based Dynamic Bucketing. IEEE Trans. Knowl. Data Eng. , 2023

  118. [126]

    V elickovic, G

    P . V elickovic, G. Cucurull, A. Casanova, A. Romero, P .Liò, and Y . Bengio. Graph Attention Networks. In ICLR, 2018

  119. [127]

    V elickovic, R

    P . V elickovic, R. Ying, M. Padovano, R. Hadsell, and C. Blundell. Neural Execution of Graph Algorithms. In ICLR, 2020

  120. [128]

    M. N. Vu and M. T. Thai. PGM-Explainer: Probabilistic Graphical Model Explanations for Graph Neural Networks. In NeurIPS, 2020

  121. [129]

    H. Wang, J. Wang, J. Wang, M. Zhao, W. Zhang, F. Zhang, X. Xie, and M. Guo. GraphGAN: Graph Representation Learning With Generative Adversarial Nets. In AAAI, 2018

  122. [130]

    J. Wang, P . Huang, H. Zhao, Z. Zhang, B. Zhao, and D. L. Le e. Billion-scale Commodity Embedding for E-commerce Recommendation in Alibaba. In KDD, 2018

  123. [131]

    J. Wang, X. Yi, R. Guo, H. Jin, P . Xu, S. Li, X. Wang, X. Guo , C. Li, X. Xu, et al. Milvus: A Purpose-built V ector Data Management System. In SIGMOD, 2021

  124. [132]

    M. Wang, X. Ke, X. Xu, L. Chen, Y . Gao, et al. MUST: An Effective and Scalable Framework for Multimodal Search of Target Modality. In ICDE, 2024

  125. [133]

    Y . Wu, K. Ma, Z. Cai, T. Jin, B. Li, C. Zheng, J. Cheng, and F. Y u. Seastar: V ertex-centric Programming for Graph Neural Networks. In EuroSys, 2021

  126. [134]

    M. Wang, H. Ma, A. Daundkar, S. Guan, Y . Bian, A. Sehirlioglu, and Y . Wu. CRUX: Crowdsourced Materials Science Resource and Workflow Exploration. In CIKM, 2022

  127. [135]

    M. Wang, H. Wu, X. Ke, Y . Gao, X. Xu, and L. Chen. An Interactive Multi-modal Query Answering System with Retrieval-Augmented Large Language Models. PVLDB, 17(12):1643–1656, 2024

  128. [136]

    M. Wang, W. Xu, X. Yi, S. Wu, Z. Peng, X. Ke, Y . Gao, X. Xu, R. Guo, and C. Xie. Starling: An I/O-Efficient Disk-Resident Graph Index Framework for High-Dimensional V ector Similarity Search on Data Segment. In SIGMOD, 2024

  129. [137]

    M. Wang, X. Xu, Q. Y ue, and Y . Wang. A Comprehensive Survey and Experimental Comparison of Graph-based Approximate Nearest Neighbor Search. PVLDB, 14(11):1964—-1978, 2021

  130. [138]

    Y . Wang, A. Khan, T. Wu, J. Jin, and H. Y an. Semantic Guided and Response Times Bounded Top-k Similarity Search over Knowledge Graphs. In ICDE, 2020

  131. [139]

    Y . Wang, A. Khan, X. Xu, J. Jin, Q. Hong, and T. Fu. Aggregate Queries on Knowledge Graphs: Fast Approximation with Semantic-aware Sampling. In ICDE, 2022

  132. [140]

    Y . Wang, N. Lipka, R. A. Rossi, A. F. Siu, R. Zhang, and T. Derr. Knowledge Graph Prompting for Multi-Document Question Answering. In AAAI, 2024

  133. [141]

    C. Wei, B. Wu, S. Wang, R. Lou, C. Zhan, F. Li, and Y . Cai. Analyticdb-v: A Hybrid Analytical Engine Towards Query Fusion for Structured and Unstructured Data. PVLDB, 13(12):3152–3165, 2020

  134. [142]

    se- mantic caches

    and GreaseLM [152] fine-tune a vanilla LM with a KG on downstream tasks, whereas DRAGON [141] and JAKET [144] perform self-supervised pre-training from both text and KGs at scale. Synergy. Query processing is the bread-and-butter for the data management community. We highlight ...

  135. [143]

    Weikum, X

    G. Weikum, X. L. Dong, S. Razniewski, and F. M. Suchanek . Machine Knowledge: Creation and Curation of Comprehensive Knowledge Bases. F ound. Trends Databases, 10(2-4):108–490, 2021

  136. [144]

    Y . Wu, N. Hu, S. Bi, G. Qi, J. Ren, A. Xie, and W. Song. Retrieve-Rewrite-Answer: A KG-to-Text Enhanced LLMs Framework for Knowledge Graph Question Answering. CoRR, abs/2309.11206, 2023

  137. [145]

    Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and S. Y . Philip . A Comprehensive Survey on Graph Neural Networks. IEEE Trans. Neural Networks Learn. Syst. , 2020

  138. [146]

    Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P . S. Y u. A Comprehensive Survey on Graph Neural Networks. IEEE Trans. Neural Networks Learn. Syst. , 32(1):4–24, 2021

  139. [147]

    F. Xia, K. Sun, S. Y u, A. Aziz, L. Wan, S. Pan, and H. Liu. Graph Learning: A Survey. IEEE Trans. Artif. Intell., 2(2):109–127, 2021

  140. [148]

    J. Xia, H. Lin, Y . Xu, C. Tan, L. Wu, S. Li, and S. Z. Li. GNN Cleaner: Label Cleaner for Graph Structured Data. IEEE Trans. Knowl. Data Eng., 2023

  141. [149]

    K. Xu, W. Hu, J. Leskovec, and S. Jegelka. How Powerful a re Graph Neural Networks? In ICLR, 2019

  142. [150]

    Z. Xu, M. J. Cruz, M. Guevara, T. Wang, M. Deshpande, X. Wang, and Z. Li. Retrieval-Augmented Generation with Knowledge Graphs for Customer Service Question Answering. In SIGIR, 2024

  143. [151]

    Y ang, J

    R. Y ang, J. Shi, X. Xiao, Y . Y ang, J. Liu, and S. S. Bhowmick. Scaling Attributed Network Embedding to Massive Graphs. PVLDB, 14(1):37–49, 2020

  144. [152]

    Y asunaga, A

    M. Y asunaga, A. Bosselut, H. Ren, X. Zhang, C. D. Mannin g, P . Liang, and J. Leskovec. Deep Bidirectional Language-Knowledge Graph Pretraining. In NeurIPS, 2022

  145. [153]

    Y asunaga, H

    M. Y asunaga, H. Ren, A. Bosselut, P . Liang, and J. Lesko vec. QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering. In NAACL-HLT, 2021

  146. [154]

    Z. Ying, D. Bourgeois, J. Y ou, M. Zitnik, and J. Leskove c. GNNExplainer: Generating Explanations for Graph Neural Networks. In NeurIPS, 2019

  147. [155]

    D. Y u, C. Zhu, Y . Y ang, and M. Zeng. JAKET: Joint Pre-training of Knowledge Graph and Language Understanding. In AAAI, 2022

  148. [156]

    Y . Y u, D. Wen, Y . Zhang, L. Qin, W. Zhang, and X. Lin. GPU-accelerated Proximity Graph Approximate Nearest Neighbor Search and Construction. In ICDE, 2022

  149. [157]

    Y uan, J

    H. Y uan, J. Tang, X. Hu, and S. Ji. XGNN: Towards Model-level Explanations of Graph Neural Networks. In KDD, 2020

  150. [158]

    Y uan, H

    H. Y uan, H. Y u, S. Gui, and S. Ji. Explainability in Grap h Neural Networks: A Taxonomic Survey. TPAMI, 45(5):5782–5799, 2023

  151. [159]

    Y uan, H

    H. Y uan, H. Y u, J. Wang, K. Li, and S. Ji. On Explainabili ty of Graph Neural Networks via Subgraph Explorations. In ICML, 2021

  152. [160]

    Y uan, X

    S. Y uan, X. Wu, and Y . Xiang. SNE: Signed Network Embedding. In PAKDD, 2017

  153. [161]

    Q. Y ue, X. Xu, Y . Wang, Y . Tao, and X. Luo. Routing-Guided Learned Product Quantization for Graph-Based Approximate Nearest Neighbor Search. In ICDE, 2024

  154. [162]

    Zhang, S

    Q. Zhang, S. Xu, Q. Chen, G. Sui, J. Xie, Z. Cai, Y . Chen, Y . He, Y . Y ang, F. Y ang, et al. VBASE: Unifying Online V ector Similarity Search and Relational Queries via Relaxed Monotonicity. In OSDI, 2023

  155. [163]

    Zhang, A

    X. Zhang, A. Bosselut, M. Y asunaga, H. Ren, P . Liang, C. D. Manning, and J. Leskovec. GreaseLM: Graph REASoning Enhanced Language Models for Question Answering. In ICLR, 2022

  156. [164]

    Zhang, H

    Y . Zhang, H. Zhu, Z. Song, P . Koniusz, and I. King. COSTA : Covariance-Preserving Feature Augmentation for Graph Contrastive Learning. In KDD, 2022

  157. [165]

    Zhang, P

    Z. Zhang, P . Cui, and W. Zhu. Deep Learning on Graphs: A Survey. IEEE Trans. Knowl. Data Eng. , 34(1):249–270, 2022

  158. [167]

    J. Zhao, Y . Dong, M. Ding, E. Kharlamov, and J. Tang. Adaptive Diffusion in Graph Neural Networks. In NeurIPS, 2021

  159. [168]

    K. Zhao, J. X. Y u, H. Zhang, Q. Li, and Y . Rong. A Learned Sketch for Subgraph Counting. In SIGMOD, 2021

  160. [169]

    T. Zhao, G. Liu, S. Günnemann, and M. Jiang. Graph Data Augmentation for Graph Machine Learning: A Survey. IEEE Data Eng. Bull., 46(2):140–165, 2023

  161. [170]

    T. Zhao, Y . Liu, L. Neves, O. J. Woodford, M. Jiang, and N. Shah. Data Augmentation for Graph Neural Networks. In AAAI, 2021

  162. [171]

    T. Zhao, X. Tang, D. Zhang, H. Jiang, N. Rao, Y . Song, P . Agrawal, K. Subbian, B. Yin, and M. Jiang. AutoGDA: Automated Graph Data Augmentation for Node Classification. In Learning on Graphs Conference, 2022

  163. [172]

    T. Zhao, X. Zhang, and S. Wang. GraphSMOTE: Imbalanced Node Classification on Graphs with Graph Neural Networks. In WSDM, 2021

  164. [173]

    W. Zhao, S. Tan, and P . Li. Song: Approximate Nearest Neighbor Search on GPU. In ICDE, 2020

  165. [174]

    X. Zhao, Y . Tian, K. Huang, B. Zheng, and X. Zhou. Toward s Efficient Index Construction and Approximate Nearest Neighbor Search in High-Dimensional Spaces. PVLDB, 16(8):1979–1991, 2023

  166. [175]

    Zheng, B

    C. Zheng, B. Zong, W. Cheng, D. Song, J. Ni, W. Y u, H. Chen , and W. Wang. Robust Graph Representation Learning via Neural Sparsification. In ICML, 2020

  167. [176]

    Zheng, C

    D. Zheng, C. Ma, M. Wang, J. Zhou, Q. Su, X. Song, Q. Gan, Z. Zhang, and G. Karypis. DistDGL: Distributed Graph Neural Network Training for Billion-Scale Graphs. In IA3, 2020

  168. [177]

    Zheng, Y

    X. Zheng, Y . Liu, Z. Bao, M. Fang, X. Hu, A. W. Liew, and S. Pan. Towards Data-centric Graph Machine Learning: Review and Outlook. CoRR, abs/2309.10979, 2023

  169. [178]

    R. Zhu, K. Zhao, H. Y ang, W. Lin, C. Zhou, B. Ai, Y . Li, and J. Zhou. Aligraph: A comprehensive graph neural network platform. Proc. VLDB Endow., 12(12):2094–2105, 2019

  170. [179]

    Y . Zhu, Y . Xu, F. Y u, Q. Liu, S. Wu, and L. Wang. Graph Contrastive Learning with Adaptive Augmentation. In WWW, 2021

  171. [180]

    Z. Zhu, S. Xu, J. Tang, and M. Qu. GraphVite: A High-Performance CPU-GPU Hybrid System for Node Embedding. In WWW, 2019

  172. [181]

    C. Zuo, M. Qiao, W. Zhou, F. Li, and D. Deng. SeRF: Segment Graph for Range-Filtering Approximate Nearest Neighbor Search. In SIGMOD, 2024

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.