Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

APEX$^2$: Adaptive and Extreme Summarization for Personalized Knowledge Graphs

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Personalized knowledge graphs can be summarized to 0.1% of their original size and still track a user's shifting interests, because APEX2 models interest as decaying heat and updates only the local neighborhood touched by each new query.

desk verdict A genuinely new adaptive PKG summarization framework with shipped code, but the best variant can exceed its own triple budget and the main comparison gives it an unfair update-frequency advantage. read the letter →

arxiv 2412.17336 v2 pith:C6EEHSOL submitted 2024-12-23 cs.LG cs.AIcs.DBcs.SC

classification cs.LGcs.AIcs.DBcs.SC
keywords adaptivesummarizationpersonalizedknowledgegraphheatdiffusionincrementalsortingextremecompressionevolvinguserinterestsqueryanswering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that personalized knowledge graphs can be kept extremely small—down to 0.1% or less of the full graph—while still answering a user's next query accurately, even as the user's interests drift from one topic to another. It proposes APEX2, a framework that models user interest as heat on the knowledge graph, lets that heat decay and diffuse after each query, and incrementally updates a ranked summary so re-summarization does not require scanning the whole graph. The reported experiments on graphs with up to 12 million triples show APEX2 and its entity-focused variant APEX2-N beating the existing baselines on next-query F1 while running tens of times faster than the strongest baseline. If the claims hold, storing a personalized copy of a knowledge graph on a device becomes practical for very large graphs and for users whose interests change during the day.

What carries the argument

The load-bearing object is the sparse heat tensor $\boldsymbol{H}$, defined entrywise as $\boldsymbol{H}^{(T)}[i][j][k] = \boldsymbol{e}^{(T)}[i]\,\boldsymbol{r}^{(T)}[j]\,\boldsymbol{e}^{(T)}[k]$, where $\boldsymbol{e}$ is the diffused entity-interest vector and $\boldsymbol{r}$ is the relation-frequency vector. The framework applies a decay-inject-diffuse cycle: at each timestamp all nonzero entries are multiplied by $\gamma^3$, new query heat is injected, and only entries whose entities or relation changed are recalculated. A second mechanism, incremental binary insertion sort, reuses the previous sorted order and inserts the few changed entries in $O(k \log n)$ comparisons, which Theorem 3.1 proves optimal. Together these make the per-timestamp update cost $O(c \cdot |\mathcal{Q}|^2 \log(c|\mathcal{Q}|))$, with $c$ the average number of neighbors within $d$ hops, independent of the number of entities in the full KG; an elimination threshold can reduce this to $O(c \log c)$.

What would settle it

Run the adaptation experiment with two topics deliberately chosen to have very different average connectivity or very different sizes (e.g., a dense hub topic and a sparse peripheral topic) and record the number of queries until the summary's F1 on the new topic surpasses that on the old topic. If that measured count systematically falls outside the bound given by Theorem 4.1, or if replacing the synthetic query logs with a real anonymized SPARQL log changes the F1 ranking of APEX2 versus the baselines, the central claim would be weakened.

Watch

Extended reading notes

Core claim

APEX2 is presented as the first adaptive personalized knowledge-graph summarization framework that keeps the summary useful under extreme compression (budgets at or below 0.1% of the full graph). The core mechanism is a heat-based model of user interest: each query injects heat at the queried entity and its neighbors, heat decays by a factor $\gamma$ each timestamp, and the summary is simply the $K$ triples with highest heat. Because decay only scales all scores and new queries affect only a small neighborhood, the ordering of triples can be maintained by incremental binary insertion sort, making the per-query update cost depend on the local connectivity of the queried area rather than on the size of the whole knowledge graph. The paper also proves an adaptation bound: after a user switches from topic $U$ to topic $V$, APEX2 needs at most about $\log_\gamma \frac{1}{\frac{A}{B}(1-\gamma^a)}+1$ queries to re-adapt, where $A$ and $B$ are derived from the average connectivity of the two topics.

Load-bearing premise

The adaptation bound assumes the two topics have similar size and a well-defined average connectivity, and that products of user preferences can be replaced by products of their expectations; these approximations are not validated on the benchmark graphs, and the experiments further assume that synthetic topic-block query logs mimic real user behavior.

Editorial extensions

If this is right

  • Personalized knowledge graphs can be stored at compression ratios below 0.1% (as small as one triple per million) while still answering next queries with competitive F1, making on-device PKGs feasible for graphs with hundreds of millions of facts.
  • Users' shifting interests can be tracked without re-summarizing the whole graph: each new query triggers a local heat update and an incremental re-sort, so the cost per adapting phase is independent of the KG's total size.
  • APEX2-N, which ignores relations and tracks only entity heat, achieves higher next-query F1 than the full APEX2 in the experiments, suggesting that for short-horizon interest tracking entities matter more than relations.
  • The decay factor $\gamma$ directly controls the trade-off between adapting to new interests and retaining useful old facts; setting $\gamma$ near 1 (no forgetting) causes F1 to collapse under extreme compression.
  • Baseline methods that re-summarize periodically every 9 timestamps are dominated in both accuracy and speed, so existing static PKG summarizers are not a viable fallback for evolving interest.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the heat model is right, the same decay-inject-diffuse mechanism could be applied to other evolving personalization tasks, such as adaptive retrieval-augmented generation contexts or personalized recommendation, where the 'graph' is a user-specific interaction graph rather than a knowledge graph.
  • The paper leaves the entity-relation weight trade-off unresolved; a natural testable extension is to learn the weight per user or per query type from the query log instead of fixing it to 0 or 1.
  • The synthetic query logs are constructed as blocks of 10 same-topic queries; real interest shifts are likely more gradual and interleaved, so a stress test with probabilistic topic mixtures would show whether the adaptation bound degrades gracefully.
  • Because the update cost is independent of KG size once a threshold zeroes out decayed heat, the framework could in principle scale to graphs far beyond the 12-million-triple benchmarks tested, provided the local connectivity $c$ stays bounded.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes APEX2 and a variant APEX2-N for adaptive, extremely compressed personalized knowledge graph (PKG) summarization under evolving user interests. The method maintains a heat-based interest model with decay, incrementally updates entity/relation/triple preferences, and uses incremental sorting to select top-heat triples under a storage budget. The authors provide theoretical claims on adaptation speed and incremental time complexity, and report experiments on YAGO, DBPedia, MetaQA, and Freebase showing that APEX2 and APEX2-N outperform GLIMPSE, PEGASUS, iSummary, and personalized PageRank in next-query F1 at compression ratios below 0.1%.

Significance. If the claimed results hold, the paper would make an important step toward practical on-device PKGs: it targets the regime of extreme compression (≤0.1%) and continuous interest shift, which prior work (e.g., GLIMPSE, PEGASUS) was not designed for. The paper ships code and provides a clean problem formulation with a heat-decay mechanism that is intuitive and likely extensible. The theoretical results, though conditional on several assumptions, offer a starting point for reasoning about adaptation speed. The main empirical claim hinges on the correctness of the experiments, and two load-bearing issues currently prevent acceptance: APEX2-N may violate the stated storage budget, and the baselines are evaluated at a disadvantageous update frequency. The central idea is promising, but the current evidence is not sufficient to substantiate the headline claims.

major comments (3)
  1. [§5.2, Algorithm 5 (Appendix A.5), Eq. (2)] APEX2-N does not enforce the size budget K on triples. The algorithm selects top-K entities in line 6 and then in lines 7 and 15 constructs T_p as all triples induced by those entities, with no final trimming. Because an induced subgraph on K entities can contain substantially more than K triples, the resulting PKG can violate the constraint |P| ≤ K from the problem definition in Eq. (2). The compression ratios reported in Section 5.2 are triple-based (e.g., 0.01% of MetaQA’s 231,103 triples corresponds to about 23 triples), so APEX2-N’s summaries may be much larger than the claimed budget. The F1 results for APEX2-N in Figure 2 and Table 3 are therefore not valid evidence for effectiveness at the stated compression ratios unless the code enforces a triple-level cap that is missing from Algorithm 5. Please revise the algorithm to guarantee |T_p| ≤ K (for example, by keeping only the top-K triples from the induced set) and rerun the experiments.
  2. [§5.1.3 and §5.3] The main experimental comparison is not at equal update opportunity. APEX2 and APEX2-N update the summary every timestamp (R_APEX = 1), while the baselines GLIMPSE, PEGASUS, iSummary, and PageRank re-summarize only every R = 9 timestamps, as described in Section 5.1.3. Thus APEX receives nine times more update chances than the baselines. The additional experiments in Section 5.6 show that APEX2-N’s F1 drops from 0.858 to 0.680 when R_APEX increases from 1 to 6 on MetaQA, which suggests that a significant part of the advantage may be due to the asymmetric update schedule. To support the claim that APEX2 and APEX2-N outperform baselines, please report results at matched update frequencies (e.g., R_APEX = 9 for APEX methods or R = 1 for baselines, if computationally feasible).
  3. [§4.1, Theorem 4.1 (proof in Appendix E.6) and Theorem 4.3 (proof in Appendix E.7)] The query-count adaptation bounds in Theorems 4.1 and 4.3 rely on assumptions that are not stated in the theorem and are not validated empirically. The proofs in Appendices E.6 and E.7 replace products of entity/relation preferences by products of their expectations and assume the topics U and V have similar sizes, i.e., |E_u| ≈ |E_v| and |R_u| ≈ |R_v|, in order to obtain the closed-form bound b > log_gamma(...). Without these assumptions, the derivation does not produce the stated bound. Please state these assumptions explicitly in the theorems and either validate them on the benchmark KGs or discuss the sensitivity of the bound when they are violated.
minor comments (5)
  1. [Appendix E: theorem numbering] The appendix proof labels do not match the main-text theorem numbers: E.2 proves the APEX2 time complexity (Theorem 4.2) but is titled Theorem 4.3; E.3 proves APEX2-N time complexity (Theorem 4.4) but is titled Theorem 4.5; E.6 proves APEX2 effectiveness (Theorem 4.1) but is titled Theorem 4.2; E.7 proves APEX2-N effectiveness (Theorem 4.3) but is titled Theorem 4.4. Please renumber for consistency.
  2. [Algorithm 5, line 7] The condition in line 7, "v in e", is unclear and appears to contain a typo; it should presumably refer to entities being endpoints in E_p^(0) (or to the relation k). Please correct.
  3. [Eq. (5) and Eq. (6)] Equation (5) writes Pr(e|Q) with an unweighted indicator 1(e_o in q), whereas the vector q_total in Eq. (3) weights answer entities by 1/|A_i|. If Eq. (6) is the actual computation, please clarify that Eq. (5) is an informal illustration and point to Eq. (3)–(4) for the exact weighting.
  4. [Table 3] The PageRank row for YAGO reports a mean of 22.81 with standard deviation 259.7, which is implausibly large relative to the mean; please check whether this is a typo or an artifact of a small number of outlier runs.
  5. [Section 5.2 and Table 3] The notation "APEX-N2" in Table 3 is inconsistent with the rest of the paper, which uses "APEX2-N"; please unify.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the APEX2 objective and experiments are grounded in external baselines and held-out next-query evaluation, and the effectiveness theorems are self-consistency statements rather than input-equivalent reductions.

full rationale

The paper's derivation chain is self-contained. The triple-preference objective in Eq. 14-15 is explicitly inherited from the external GLIMPSE baseline ('we stick to GLIMPSE's choice for triple preference'), not from a fitted constant or self-citation. The experimental F1 comparisons are measured against GLIMPSE, PEGASUS, iSummary, and personalized PageRank on 'the very next query' after adaptation, and no parameter is fitted to the F1 numbers, so the reported comparisons are not forced by construction. The effectiveness theorems (4.1 and 4.3) analyze the heat-decay model's own scoring function: 'adaptation' is formalized as the point where the model's preference for the new topic exceeds that for the old topic, and the bounds in Appendix E.6/E.7 follow from the model equations under explicit assumptions (|E_u|≈|E_v|, |R_u|≈|R_v|, and expectation-of-products approximations). This is an internal consistency analysis rather than a circular prediction, because the theorems do not assume their conclusions and are not used to define the input quantities. Self-citations such as [34] for the Neumann-series closed form and the PPR baseline are not load-bearing: the closed form is a standard identity and PPR is an external benchmark. Section 8's limitation statement ('there may be no significant advantage to summarizing a PKG rather than querying the KG directly') is candid and does not conceal a circular step. Two non-circular concerns fall outside this pass: Algorithm 5 can select an induced triple set exceeding the budget K, and the Appendix E proofs rely on unvalidated homogeneity approximations; these affect correctness and rigor, not circularity.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claims rest on several modeling choices: conditional independence of triple preference, heat diffusion as a proxy for user interest, topic-as-entity, plus unquantified approximations in the adaptation proofs. The three hyperparameters gamma, alpha, and d are hand set, and the heat-elimination threshold is left unspecified. No new entities are introduced.

free parameters (5)
  • gamma (decay factor) = 0.5 (default; ablated 0.1-1.0 in Fig. 3)
    Hand-set hyperparameter controlling how quickly old interests fade; the adaptation bound in Theorem 4.1 depends on it.
  • alpha (damping factor of neighbor) = 0.3
    Hand-set neighbor diffusion weight; used in Eq. 6 and in the theoretical adaptation analysis.
  • d (diffusing diameter) = 1
    Number of diffusion hops; with d>1 the step from Eq. 5 to Eq. 6 is not exact because A^l counts walks rather than distance-l neighbors.
  • epsilon threshold for heat elimination = unspecified ("small enough value")
    Used in Section 4.1 to zero out decayed heat and make incremental complexity independent of graph size; never defined or reported in experiments.
  • entity/relation weight ratio in APEX2-N = entities: 1, relations: 0
    Deliberate modeling choice that makes APEX2-N ignore relation interest; the paper leaves tuning of this trade-off to future work.
assumptions (5)
  • standard math Convergence and invertibility of (I - alpha A)^-1 in Eq. 7
    Used to write the infinite-diffusion closed form; with unnormalized adjacency A and alpha=0.3 it need not hold for high-degree KG nodes, though experiments use finite d=1.
  • domain assumption Triple preference factorizes as Pr(e_i) Pr(r_k) Pr(e_j)
    Eq. 14; inherited from GLIMPSE (Eq. 22), not derived from query logs.
  • domain assumption User interests follow heat diffusion and topics correspond to entities/nodes
    Section 3.1 and Appendix D; used to justify query entity as topic label and the entire adaptation analysis.
  • ad hoc to paper Topic areas have similar sizes and expectation products can be replaced by products of expectations
    Appendix E.4, E.6, and E.7 derive the adaptation bounds only after assuming |E_u| approximately equals |E_v|, |R_u| approximately equals |R_v|, and approximating products by expectations; no error bound is given.
  • domain assumption Queries are simple one-hop queries with known answers
    Section 2 defines the query log this way and Section 5.2 generates only 1-hop simple queries, so claims do not extend to complex multi-hop queries.

how reviews work

0 comments
Cite this review

Pith. "Pith review of APEX$^2$: Adaptive and Extreme Summarization for Personalized Knowledge Graphs." pith.science (2026). https://pith.science/paper/C6EEHSOL

@misc{pith2026241217336,
  author       = {Pith},
  title        = {Pith review of: APEX$^2$: Adaptive and Extreme Summarization for Personalized Knowledge Graphs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/C6EEHSOL}},
  note         = {Machine review of arXiv:2412.17336}
}
abstract

Knowledge graphs (KGs), which store an extensive number of relational facts, serve various applications. Recently, personalized knowledge graphs (PKGs) have emerged as a solution to optimize storage costs by customizing their content to align with users' specific interests within particular domains. In the real world, on one hand, user queries and their underlying interests are inherently evolving, requiring PKGs to adapt continuously; on the other hand, the summarization is constantly expected to be as small as possible in terms of storage cost. However, the existing PKG summarization methods implicitly assume that the user's interests are constant and do not shift. Furthermore, when the size constraint of PKG is extremely small, the existing methods cannot distinguish which facts are more of immediate interest and guarantee the utility of the summarized PKG. To address these limitations, we propose APEX$^2$, a highly scalable PKG summarization framework designed with robust theoretical guarantees to excel in adaptive summarization tasks with extremely small size constraints. To be specific, after constructing an initial PKG, APEX$^2$ continuously tracks the interest shift and adjusts the previous summary. We evaluate APEX$^2$ under an evolving query setting on benchmark KGs containing up to 12 million triples, summarizing with compression ratios $\leq 0.1\%$. The experiments show that APEX outperforms state-of-the-art baselines in terms of both query-answering accuracy and efficiency. Code is available at https://github.com/iDEA-iSAIL-Lab-UIUC/APEX.

Figures

Figures reproduced from arXiv: 2412.17336 by the authors.

Figure 1
Figure 1. Example of Adaptive PKGs. The entire KG is stored [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Effectiveness Comparison Under Querying Scenario [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Querying Effectiveness in Multiple Decay Levels [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Parameter Study. From left to right: compression ratio [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Cosine similarity between consecutive query topics in different datasets for one user. The cosine similarity is computed [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: The distribution of cosine similarities between consecutive query topics in our synthetic queries. From left to right: [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Our pre-experiment to determine that 𝑅 = 9 is a good trade-off for accuracy and efficiency. C.4.5 Freebase. https://paperswithcode.com/dataset/f b15k￾237. The entire FB15K237 dataset can be downloaded by clicking "Homepage". C.4.6 Environments. We run all our experimen…
Figure 8
Figure 8. Figure 8: Querying Effectiveness in Multiple Decay Levels for different [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]
Figure 9
Figure 9. Figure 9: Parameter Study adjusting different 𝑅𝐴𝑃𝐸𝑋 . From left to right: compression ratio 𝜅, damping factor of neighbor 𝛼, diffusing diameter 𝑑. we adopt a heat decay-inject-diffuse framework for heat diffusion, as shown in [PITH_FULL_IMAGE:figures/full_fig_p018_9.png]
Figure 10
Figure 10. Figure 10: We adopt a heat decay-inject-diffuse framework for heat diffusion. For each timestamp, first, all the heat (i.e., interest) [PITH_FULL_IMAGE:figures/full_fig_p019_10.png]
Figure 11
Figure 11. Figure 11: The specific part of MetaQA Knowledge Graph within the interest of our case study. The summarization operates on [PITH_FULL_IMAGE:figures/full_fig_p021_11.png]
Figure 12
Figure 12. Figure 12: Case Study on APEX2 [PITH_FULL_IMAGE:figures/full_fig_p022_12.png]
Figure 13
Figure 13. Figure 13: Case Study on APEX2 -N [PITH_FULL_IMAGE:figures/full_fig_p023_13.png]
Figure 14
Figure 14. Figure 14: Case Study on GLIMPSE [PITH_FULL_IMAGE:figures/full_fig_p024_14.png]
Figure 15
Figure 15. Figure 15: Case Study on PageRank [PITH_FULL_IMAGE:figures/full_fig_p025_15.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PyG-SSL: A Graph Self-Supervised Learning Toolkit

    cs.LG 2024-12 conditional novelty 4.0 of 10

    PyG-SSL is a new open-source library that packages ten graph self-supervised learning methods with unified training, evaluation, and hyperparameter configs, and reports baseline reproductions on six datasets.

Reference graph

Works this paper leans on

85 extracted references · 60 canonical work pages · cited by 1 Pith paper

  1. [1]

    Amiri, Aditya Bharadwaj, and B

    Bijaya Adhikari, Yao Zhang, Sorour E. Amiri, Aditya Bharadwaj, and B. Aditya Prakash. 2018. Propagation-Based Temporal Network Summarization. IEEE Trans. Knowl. Data Eng. (2018)

  2. [2]

    Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary G. Ives. 2007. DBpedia: A Nucleus for a Web of Open Data. In The Semantic Web, 6th International Semantic Web Conference, 2nd Asian Semantic Web Conference, ISWC 2007 + ASWC 2007, Busan, Korea, November 11-15, 2007

  3. [3]

    Yikun Ban, Jiaru Zou, Zihao Li, Yunzhe Qi, Dongqi Fu, Jian Kang, Hanghang Tong, and Jingrui He. 2024. PageRank Bandits for Link Prediction.CoRR abs/2411.01410 (2024). https://doi.org/10.48550/ARXIV.2411.01410 arXiv:2411.01410

  4. [4]

    Caleb Belth, Xinyi Zheng, Jilles Vreeken, and Danai Koutra. 2020. What is Normal, What is Strange, and What is Missing in a Knowledge Graph: Unified Characterization via Inductive Summarization. In WWW 2020

  5. [5]

    Bollacker, Colin Evans, Praveen K

    Kurt D. Bollacker, Colin Evans, Praveen K. Paritosh, Tim Sturge, and Jamie Taylor. 2008. Freebase: a collaboratively created graph database for structuring human knowledge. In Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2008, Vancouver, BC, Canada, June 10-12, 2008 , Jason Tsong-Li Wang (Ed.)

  6. [6]

    Angela Bonifati, Wim Martens, and Thomas Timm. 2020. An analytical study of large SPARQL query logs. VLDB J. 29, 2-3 (2020), 655–679. https://doi.org/10.1 007/s00778-019-00558-9

  7. [7]

    Marco Braga. 2024. Personalized Large Language Models through Parameter Efficient Fine-Tuning Techniques. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July 14-18, 2024 , Grace Hui Yang, Hongning Wang, Sam Han, Claudia Hauff, Guido Zuccon, and Yi Zhang (Ed...

  8. [8]

    Qingyu Chen, Alexis Allot, and Zhiyong Lu. 2021. LitCovid: an open database of COVID-19 literature. Nucleic Acids Res. (2021)

Show all 85 references
  1. [9]

    George Cybenko. 1989. Dynamic Load Balancing for Distributed Memory Multi- processors. J. Parallel Distributed Comput. (1989)

  2. [10]

    Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. CoRR abs/2404.16130 (2024). https://doi.org/10.48550/ARXIV.2404.16130 arXiv:2404.16130

  3. [11]

    Patrick Ernst, Cynthia Meng, Amy Siu, and Gerhard Weikum. 2014. Knowlife: a knowledge graph for health and life sciences. In 2014 IEEE 30th International Conference on Data Engineering

  4. [12]

    Lukas Faber, Tara Safavi, Davide Mottin, Emmanuel Müller, and Danai Koutra

  5. [13]

    Michael Färber, Frederic Bartscherer, Carsten Menne, and Achim Rettinger. 2018. Linked data quality of DBpedia, Freebase, OpenCyc, Wikidata, and YAGO. Se- mantic Web (2018)

  6. [14]

    William Feller. 1991. An introduction to probability theory and its applications, Volume 2. Vol. 81. John Wiley & Sons

  7. [15]

    Torvik, and Jingrui He

    Dongqi Fu, Liri Fang, Zihao Li, Hanghang Tong, Vetle I. Torvik, and Jingrui He

  8. [16]

    Torvik, and Jingrui He

    Dongqi Fu, Liri Fang, Ross Maciejewski, Vetle I. Torvik, and Jingrui He. 2022. Meta-Learned Metrics over Multi-Evolution Temporal Graphs. In KDD

  9. [17]

    Dongqi Fu, Dawei Zhou, Ross Maciejewski, Arie Croitoru, Marcus Boyd, and Jin- grui He. 2023. Fairness-Aware Clique-Preserving Spectral Clustering of Temporal Graphs. In WWW

  10. [18]

    Dongqi Fu, Yada Zhu, Hanghang Tong, Kommy Weldemariam, Onkar Bhardwaj, and Jingrui He. 2024. Generating Fine-Grained Causality in Climate Time Series Data for Forecasting and Anomaly Detection. CoRR abs/2408.04254 (2024). https: //doi.org/10.48550/ARXIV.2408.04254 arXiv:2408.04254

  11. [19]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Qianyu Guo, Meng Wang, and Haofen Wang. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. CoRR abs/2312.10997 (2023). https://doi.org/10.48550/ARXIV.2312.10997 arX...

  12. [20]

    Cyril Goutte and Éric Gaussier. 2005. A Probabilistic Interpretation of Precision, Recall and F-Score, with Implication for Evaluation. In Advances in Informa- tion Retrieval, 27th European Conference on IR Research, ECIR 2005, Santiago de Compostela, Spain, March 21-23, 2005,...

  13. [21]

    Aditya Grover and Jure Leskovec. 2016. node2vec: Scalable Feature Learning for Networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, August 13-17, 2016

  14. [22]

    Szymanski

    Jingrui He, Hanghang Tong, Qiaozhu Mei, and Boleslaw K. Szymanski. 2012. GenDeR: A Generic Diversified Ranking Algorithm. In Advances in Neural In- formation Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting h...

  15. [23]

    Qi He, Jaewon Yang, and Baoxu Shi. 2020. Constructing knowledge graph for social networks in a deep and holistic way. In Companion Proceedings of the Web Conference 2020

  16. [24]

    Xinrui He, Yikun Ban, Jiaru Zou, Tianxin Wei, Curtiss B Cook, and Jingrui He. 2024. LLM-Forest for Health Tabular Data Imputation. arXiv preprint arXiv:2410.21520 (2024)

  17. [25]

    C. A. R. Hoare. 1961. Algorithm 64: Quicksort. Commun. ACM 4, 7 (1961), 321. https://doi.org/10.1145/366622.366644

  18. [26]

    Bowen Jin, Jinsung Yoon, Jiawei Han, and Sercan Ö. Arik. 2024. Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAG. CoRR abs/2410.05983 (2024). https://doi.org/10.48550/ARXIV.2410.05983 arXiv:2410.05983

  19. [27]

    Shinhwan Kang, Kyuhan Lee, and Kijung Shin. 2022. Personalized Graph Sum- marization: Formulation, Scalable Algorithms, and Applications. In 38th IEEE International Conference on Data Engineering, ICDE 2022, Kuala Lumpur, Malaysia, May 9-12, 2022

  20. [28]

    Matthäus Kleindessner, Pranjal Awasthi, and Jamie Morgenstern. 2019. Fair k- Center Clustering for Data Summarization. InProceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA

  21. [29]

    Jihoon Ko, Yunbum Kook, and Kijung Shin. 2020. Incremental Lossless Graph Summarization. In KDD 2020

  22. [30]

    Danai Koutra, U Kang, Jilles Vreeken, and Christos Faloutsos. 2014. VOG: Sum- marizing and Understanding Large Graphs. In Proceedings of the 2014 SIAM International Conference on Data Mining, Philadelphia, Pennsylvania, USA, April 24-26, 2014

  23. [31]

    Ravuri, Timo Ewalds, Ferran Alet, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Hol- land, Jacklynn Stott, Oriol Vinyals, Shakir Mohamed, and Peter W

    Rémi Lam, Alvaro Sanchez-Gonzalez, Matthew Willson, Peter Wirnsberger, Meire Fortunato, Alexander Pritzel, Suman V. Ravuri, Timo Ewalds, Ferran Alet, Zach Eaton-Rosen, Weihua Hu, Alexander Merose, Stephan Hoyer, George Hol- land, Jacklynn Stott, Oriol Vinyals, Shakir Mohamed, ...

  24. [32]

    Kyuhan Lee, Hyeonsoo Jo, Jihoon Ko, Sungsu Lim, and Kijung Shin. 2020. SSumM: Sparse Summarization of Massive Graphs. In KDD 2020

  25. [33]

    Zihao Li, Yuyi Ao, and Jingrui He. 2024. SpherE: Expressive and Interpretable Knowledge Graph Embedding for Set Retrieval. In Proceedings of the 47th In- ternational ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2024, Washington DC, USA, July...

  26. [34]

    Zihao Li, Dongqi Fu, and Jingrui He. 2023. Everything Evolves in Personalized PageRank. In Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023

  27. [35]

    Zihao Li, Dongqi Fu, Hengyu Liu, and Jingrui He. 2024. Hypergraphs as Weighted Directed Self-Looped Graphs: Spectral Properties, Clustering, Cheeger Inequality. arXiv preprint arXiv:2411.03331 (2024)

  28. [36]

    Zihao Li, Dongqi Fu, Hengyu Liu, and Jingrui He. 2024. Provably Extending PageRank-based Local Clustering Algorithm to Weighted Directed Graphs with Self-Loops and to Hypergraphs. arXiv preprint arXiv:2412.03008 (2024)

  29. [37]

    Zihao Li, Lecheng Zheng, Bowen Jin, Dongqi Fu, Baoyu Jing, Yikun Ban, Jingrui He, and Jiawei Han. 2024. Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision? arXiv preprint arXiv:2412.08174 (2024)

  30. [38]

    Jue Liu, Zhuocheng Lu, and Wei Du. 2019. Combining enterprise knowledge graph and news sentiment analysis for stock price prediction. (2019)

  31. [39]

    Lihui Liu, Yuzhong Chen, Mahashweta Das, Hao Yang, and Hanghang Tong. 2023. Knowledge Graph Question Answering with Ambiguous Query. In Proceedings of the ACM Web Conference 2023

  32. [40]

    Lihui Liu, Boxin Du, Yi Ren Fung, Heng Ji, Jiejun Xu, and Hanghang Tong. 2021. KompaRe: A Knowledge Graph Comparative Reasoning System. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Virtual Event, Singapore) (KDD ’21). Association for...

  33. [41]

    Lihui Liu, Boxin Du, Jiejun Xu, Yinglong Xia, and Hanghang Tong. 2022. Joint Knowledge Graph Completion and Question Answering. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (Washington DC, USA) (KDD ’22). Association for Computing Mach...

  34. [42]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. CoRR abs/1907.11692 (2019). arXiv:1907.11692 http://arxiv.org/abs/1907.11692

  35. [43]

    Yike Liu, Tara Safavi, Abhilash Dighe, and Danai Koutra. 2018. Graph Summa- rization Methods and Applications: A Survey. ACM Comput. Surv. (2018). APEX2: Adaptive and Extreme Summarization for Personalized Knowledge Graphs KDD ’25, August 3–7, 2025, Toronto, ON, Canada

  36. [44]

    Zhining Liu, Ruizhong Qiu, Zhichen Zeng, Hyunsik Yoo, David Zhou, Zhe Xu, Yada Zhu, Kommy Weldemariam, Jingrui He, and Hanghang Tong. 2024. Class- Imbalanced Graph Learning without Class Rebalancing. InForty-first International Conference on Machine Learning, ICML 2024, Vienna...

  37. [46]

    Lyu, and Irwin King

    Hao Ma, Haixuan Yang, Michael R. Lyu, and Irwin King. 2008. Mining social networks using heat diffusion processes for marketing candidates selection. In Proceedings of the 17th ACM Conference on Information and Knowledge Manage- ment, CIKM 2008, Napa Valley, California, USA, O...

  38. [47]

    Tharun Medini, Beidi Chen, and Anshumali Shrivastava. 2021. SOLAR: Sparse Orthogonal Learned and Random Embeddings. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021

  39. [48]

    Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston

    Alexander H. Miller, Adam Fisch, Jesse Dodge, Amir-Hossein Karimi, Antoine Bordes, and Jason Weston. 2016. Key-Value Memory Networks for Directly Reading Documents. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, EMNLP 2016, Austin, T...

  40. [49]

    Maximilian Nickel, Kevin Murphy, Volker Tresp, and Evgeniy Gabrilovich. 2016. A Review of Relational Machine Learning for Knowledge Graphs. Proc. IEEE (2016)

  41. [50]

    Sancheng Peng, Yongmei Zhou, Lihong Cao, Shui Yu, Jianwei Niu, and Weijia Jia. 2018. Influence analysis in social networks: A survey. J. Netw. Comput. Appl. (2018)

  42. [51]

    Yunzhe Qi, Yikun Ban, and Jingrui He. 2023. Graph neural bandits. InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1920–1931

  43. [52]

    Hongyu Ren, Mikhail Galkin, Michael Cochez, Zhaocheng Zhu, and Jure Leskovec

  44. [53]

    Mariia Rizun et al. 2019. Knowledge graph application in education: a literature review. Acta Universitatis Lodziensis. Folia Oeconomica (2019)

  45. [54]

    Rossi and Rong Zhou

    Ryan A. Rossi and Rong Zhou. 2018. GraphZIP: a clique-based sparse graph compression method. J. Big Data (2018)

  46. [55]

    Tara Safavi, Caleb Belth, Lukas Faber, Davide Mottin, Emmanuel Müller, and Danai Koutra. 2019. Personalized Knowledge Graph Summarization: From the Cloud to Your Pocket. In 2019 IEEE International Conference on Data Mining, ICDM 2019, Beijing, China, November 8-11, 2019

  47. [56]

    Neil Shah, Danai Koutra, Tianmin Zou, Brian Gallagher, and Christos Faloutsos

  48. [57]

    Suchanek, Gjergji Kasneci, and Gerhard Weikum

    Fabian M. Suchanek, Gjergji Kasneci, and Gerhard Weikum. 2007. Yago: a core of semantic knowledge. In Proceedings of the 16th International Conference on World Wide Web, WWW 2007, Banff, Alberta, Canada, May 8-12, 2007

  49. [58]

    Tadesse, Hongfei Lin, Bo Xu, and Liang Yang

    Michael M. Tadesse, Hongfei Lin, Bo Xu, and Liang Yang. 2019. Detection of Depression-Related Posts in Reddit Social Media Forum. IEEE Access (2019)

  50. [59]

    Nan Tang, Qing Chen, and Prasenjit Mitra. 2016. Graph Stream Summarization: From Big Bang to Big Crunch. In SIGMOD 2016

  51. [60]

    James Thorne, Majid Yazdani, Marzieh Saeidi, Fabrizio Silvestri, Sebastian Riedel, and Alon Y. Levy. 2021. From Natural Language Processing to Neural Databases. Proc. VLDB Endow. (2021)

  52. [61]

    Giannis Vassiliou, Fanouris Alevizakis, Nikolaos Papadakis, and Haridimos Kondylakis. 2023. iSummary: Workload-Based, Personalized Summaries for Knowledge Graphs. In The Semantic Web - 20th International Conference, ESWC 2023, Hersonissos, Crete, Greece, May 28 - June 1, 2023,...

  53. [62]

    Denny Vrandecic and Markus Krötzsch. 2014. Wikidata: a free collaborative knowledgebase. Commun. ACM (2014)

  54. [63]

    Stanislaw Wozniak, Bartlomiej Koptyra, Arkadiusz Janz, Przemyslaw Kazienko, and Jan Kocon. 2024. Personalized Large Language Models. CoRR abs/2402.09269 (2024). https://doi.org/10.48550/ARXIV.2402.09269 arXiv:2402.09269

  55. [64]

    Birge, and Jingrui He

    Ziwei Wu, Lecheng Zheng, Yuancheng Yu, Ruizhong Qiu, John R. Birge, and Jingrui He. 2024. Fair Anomaly Detection For Imbalanced Groups. CoRR abs/2409.10951 (2024). https://doi.org/10.48550/ARXIV.2409.10951 arXiv:2409.10951

  56. [65]

    Yuchen Yan, Lihui Liu, Yikun Ban, Baoyu Jing, and Hanghang Tong. 2021. Dy- namic knowledge graph alignment. In Proceedings of the AAAI conference on artificial intelligence, Vol. 35. 4564–4572

  57. [66]

    Ban Yikun, Liu Xin, Huang Ling, Duan Yitao, Liu Xue, and Xu Wei. 2019. No place to hide: Catching fraudulent entities in tensors. In The World Wide Web Conference. 83–93

  58. [67]

    Quinton Yong, Mahdi Hajiabadi, Venkatesh Srinivasan, and Alex Thomo. 2021. Efficient Graph Summarization using Weighted LSH at Billion-Scale. In SIGMOD 2021

  59. [68]

    Zhichen Zeng, Boxin Du, Si Zhang, Yinglong Xia, Zhining Liu, and Hang- hang Tong. 2024. Hierarchical Multi-Marginal Optimal Transport for Net- work Alignment. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applicat...

  60. [69]

    Smola, and Le Song

    Yuyu Zhang, Hanjun Dai, Zornitsa Kozareva, Alexander J. Smola, and Le Song

  61. [70]

    Birge, Yifang Zhang, and Jingrui He

    Lecheng Zheng, John R. Birge, Yifang Zhang, and Jingrui He. 2024. Towards Multi- view Graph Anomaly Detection with Similarity-Guided Contrastive Clustering. CoRR abs/2409.09770 (2024). https://doi.org/10.48550/ARXIV.2409.09770 arXiv:2409.09770

  62. [71]

    Lecheng Zheng, Baoyu Jing, Zihao Li, Hanghang Tong, and Jingrui He. 2024. Heterogeneous Contrastive Learning for Foundation Models and Beyond. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2024, Barcelona, Spain, August 25-29, 202...

  63. [72]

    Ming Zhong, Chenxin An, Weizhu Chen, Jiawei Han, and Pengcheng He. 2024. Seeking Neural Nuggets: Knowledge Transfer in Large Language Models from a Parametric Perspective. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11...

  64. [73]

    Dawei Zhou, Kangyang Wang, Nan Cao, and Jingrui He. 2015. Rare Category Detection on Time-Evolving Graphs. In 2015 IEEE International Conference on Data Mining, ICDM 2015, Atlantic City, NJ, USA, November 14-17, 2015 , Charu C. Aggarwal, Zhi-Hua Zhou, Alexander Tuzhilin, Hui X...

  65. [74]

    Dawei Zhou, Lecheng Zheng, Dongqi Fu, Jiawei Han, and Jingrui He. 2022. MentorGNN: Deriving Curriculum for Pre-Training GNNs. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, October 17-21, 2022 , Mohammad Al Hasa...

  66. [75]

    Variational Reasoning for Question Answering With Knowledge Graph. In Proceedings of the Thirty-Second AAAI Conference on Artificial Intelligence, (AAAI- 18), the 30th innovative Applications of Artificial Intelligence (IAAI-18), and the 8th AAAI Symposium on Educational Advan...

  67. [76]

    Xinwen Zhu, Zihao Li, Yuxuan Jiang, Jiazhen Xu, Jie Wang, and Xuyang Bai

  68. [77]

    Jiaru Zou, Mengyu Zhou, Tao Li, Shi Han, and Dongmei Zhang. 2024. PromptIn- tern: Saving Inference Costs by Internalizing Recurrent Prompt during Large Lan- guage Model Fine-tuning. InFindings of the Association for Computational Linguis- tics: EMNLP 2024, Miami, Florida, USA,...

  69. [78]

    ≜” means “defined as

    Xiaohan Zou. 2020. A survey on application of knowledge graph. In Journal of Physics: Conference Series. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Li et al. A Pseudo-code of Algorithms A.1 GLIMPSE Algorithm 1 The GLIMPSE framework Require: knowledge graphG; query logQ; si...

  70. [81]

    Dawei Zhou, Lecheng Zheng, Jiawei Han, and Jingrui He. 2020. A Data-Driven Graph Generative Model for Temporal Interaction Networks. InKDD ’20: The 26th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Virtual Event, CA, USA, August 23-27, 2020 , Rajesh Gupta, Yan...

  71. [83]

    CoRR abs/2410.17576 (2024)

    Real-time Vehicle-to-Vehicle Communication Based Network Cooper- ative Control System through Distributed Database and Multimodal Percep- tion: Demonstrated in Crossroads. CoRR abs/2410.17576 (2024). https: //doi.org/10.48550/ARXIV.2410.17576 arXiv:2410.17576

  72. [237]

    Homepage

    The entire FB15K237 dataset can be downloaded by clicking "Homepage". C.4.6 Environments. We run all our experiment on a Windows 10 machine with Intel(R) Core(TM) i7-10750H CPU @ 2.60GHz and 32GB RAM. For other different platforms such as Linux, you may need to reset the path ...

  73. [2015]

    In KDD 2015

    TimeCrunch: Interpretable Dynamic Graph Summarization. In KDD 2015

  74. [2018]

    In MLG Workshop (with KDD)

    Adaptive personalized knowledge graph summarization. In MLG Workshop (with KDD)

  75. [2022]

    CoRR abs/2212.12794 (2022)

    GraphCast: Learning skillful medium-range global weather forecasting. CoRR abs/2212.12794 (2022). https://doi.org/10.48550/ARXIV.2212.12794 arXiv:2212.12794

  76. [2023]

    CoRR abs/2303.14617 (2023)

    Neural Graph Reasoning: Complex Logical Query Answering Meets Graph Databases. CoRR abs/2303.14617 (2023)

  77. [2024]

    CoRR abs/2410.12126 (2024)

    Parametric Graph Representations in the Era of Foundation Models: A Survey and Position. CoRR abs/2410.12126 (2024). https://doi.org/10.48550/ARX IV.2410.12126 arXiv:2410.12126

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.