Pith. sign in

REVIEW 53 references

Enhancing LLMs via High-Knowledge Data Selection

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A knowledge-element density and coverage scorer selects pre-training data that improves LLM performance on knowledge-intensive and general understanding benchmarks by 2 to 3 points.

arxiv 2505.14070 v2 pith:7T3KL53Y submitted 2025-05-20 cs.CL

classification cs.CL
keywords knowledgedatahigh-knowledgeproposescorerselectioncomprehensivedomain-specific
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper addresses a practical problem: when training a large language model, which web texts should you keep? Most filters score text by fluency, such as perplexity, or by similarity to a reference corpus. The authors instead score text by how many 'knowledge elements' it contains. A knowledge element is a term like a concept name, pulled from Wikipedia titles, academic paper keywords, and GPT-4 extractions, organized into five domains. For every text, they count how many of the 5M knowledge elements appear per token (density) and how many different ones appear (coverage). The product of density and ln(1+coverage) becomes the text's score. They test this High-Knowledge Scorer by training a 1.1B-parameter bilingual model on the top-scoring portions of the Pile and Wudao corpora (20B tokens) and by continuing to train Llama-3-8B on 100B tokens. HKS-selected data outperform random selection by about 2.4 percentage points on average across knowledge-intensive and general understanding benchmarks, and outperform perplexity, error-L2-norm, and DSIR baselines in most comparisons. Restricting the knowledge elements to science gives a domain-specific version that improves science questions. The method is attractive because scoring only requires string matching on a CPU, so it is far cheaper than running a model to score every document. The main weaknesses are that the exact knowledge-element pool is not released, the coverage definition in the text is inconsistent with the example numbers, and the scoring formula was chosen by fitting to a small set of human preferences rather than derived from theory.
Extended reading notes

Core claim

The paper states that HKS 'improves the model's performance in knowledge-intensive and general comprehension tasks,' with an average improvement of 2.37 pp over random selection for a 1.1B model trained on 20B tokens, and an average increase of 2.4 pp in continual pretraining of Llama-3-8B on 100B tokens. If correct, HKS is a cheap, effective data-selection method that outperforms PPL, EL2N, and DSIR on the tested benchmarks.

Load-bearing premise

The method equates knowledge with exact matches to a pre-built pool of surface n-grams, as defined in Definition 1 and implemented with Aho-Corasick matching in Section 2.3. If a text conveys the same fact with different wording, it scores zero; the scorer therefore measures lexical overlap with Wikipedia- and OAG-derived terms, not knowledge as such. This assumption is load-bearing because the entire causal story, that knowledge-rich data improve models, depends on the pool and exact matching being a valid proxy for knowledge.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 2 free parameters · 5 assumptions · 1 invented entities

The central claim rests on the knowledge-element operationalization, exact-match tagging, the multiplicative scorer form, and the human preference labels used to choose the formula. These are not derived from first principles, and the knowledge pool itself is not released, so independent verification currently requires reconstructing the entire pipeline.

free parameters (2)
  • Scorer functional form f(d)*g(c) = d * ln(c+1)
    Chosen from nine candidate functions by maximizing Spearman correlation with 2K human pairwise annotations (Appendix D, Table 5); the paper itself notes there is no theoretical validation.
  • Softmax temperature tau for sampling = 2
    Hand-set in Eq. 6 for the Gumbel softmax sampling strategy; no tuning procedure is described.
assumptions (5)
  • ad hoc to paper Knowledge can be represented as n-gram noun phrases (knowledge elements) that encapsulate concepts, facts, theories, and definitions.
    Definition 1 in Section 2.2 is the core operationalization of knowledge for the entire method.
  • domain assumption Exact substring matching with Aho-Corasick identifies all knowledge in a text; paraphrased or reworded knowledge is not counted.
    Section 2.3 labels all knowledge elements that appear in each text using Aho-Corasick, assuming surface-form matches capture knowledge.
  • ad hoc to paper Knowledge density and coverage are independent and combine multiplicatively with concave f and g.
    Section 2.4 Eq. 2-3 asserts independence and concavity as empirical assumptions without derivation.
  • domain assumption Human pairwise judgments of 'more informative signal' are a valid proxy for pretraining data quality.
    Appendix D uses three annotators on 2K text pairs to select the scoring formula; this assumes human informativeness ratings transfer to downstream model performance.
  • domain assumption The 20B-token bilingual sample from Pile and Wudao is representative enough to test data-selection methods.
    Section 3.1 fixes the dataset size and source corpora; results may not transfer to other corpora, languages, or scales.
invented entities (1)
  • Knowledge element
    purpose: A surface n-gram from Wikipedia titles, OAG keywords, or GPT-4 extractions, used as the unit of knowledge for computing density and coverage scores.
    The paper defines knowledge elements internally and builds a 5M pool, but there is no external benchmark validating that these n-grams are the correct decomposition of knowledge; validation is limited to internal human annotation and downstream task improvement.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Enhancing LLMs via High-Knowledge Data Selection." pith.science (2026). https://pith.science/paper/7T3KL53Y

@misc{pith2026250514070,
  author       = {Pith},
  title        = {Pith review of: Enhancing LLMs via High-Knowledge Data Selection},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7T3KL53Y}},
  note         = {Machine review of arXiv:2505.14070}
}
read the original abstract

The performance of Large Language Models (LLMs) is intrinsically linked to the quality of its training data. Although several studies have proposed methods for high-quality data selection, they do not consider the importance of knowledge richness in text corpora. In this paper, we propose a novel and gradient-free High-Knowledge Scorer (HKS) to select high-quality data from the dimension of knowledge, to alleviate the problem of knowledge scarcity in the pre-trained corpus. We propose a comprehensive multi-domain knowledge element pool and introduce knowledge density and coverage as metrics to assess the knowledge content of the text. Based on this, we propose a comprehensive knowledge scorer to select data with intensive knowledge, which can also be utilized for domain-specific high-knowledge data selection by restricting knowledge elements to the specific domain. We train models on a high-knowledge bilingual dataset, and experimental results demonstrate that our scorer improves the model's performance in knowledge-intensive and general comprehension tasks, and is effective in enhancing both the generic and domain-specific capabilities of the model.

Figures

Figures reproduced from arXiv: 2505.14070 by the authors.

Figure 1
Figure 1. The overall framework of HKS. Our methodology begins with sourcing knowledge from Wikipedia articles and [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Score density in Pile subsets. There is a noticeable difference in the distribution of knowledge density ( [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Perplexity evaluation on Wudao validation dataset. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Perplexity evaluation on Pile validation dataset. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: We compare the costs of the various ap￾proaches based on cloud server rental fees. PPLEL2NDSIR d cHKS PPL EL2N DSIR d c HKS 0.75 0.50 0.25 0.00 0.25 0.50 0.75 1.00 [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 7
Figure 7. Figure 7: Knowledge elements distribution over different [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: Knowledge element cases over different categories. [PITH_FULL_IMAGE:figures/full_fig_p012_8.png]
Figure 9
Figure 9. Figure 9: Data selection cases on Pile [PITH_FULL_IMAGE:figures/full_fig_p014_9.png]
Figure 10
Figure 10. Figure 10: Data selection cases on Wudao [PITH_FULL_IMAGE:figures/full_fig_p015_10.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

53 extracted references · 28 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al

    Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F. L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  4. [4]

    Allen-Zhu, Z.; and Li, Y. 2024. Physics of language models: Part 3.3, knowledge capacity scaling laws. arXiv preprint arXiv:2404.05405

  5. [5]

    Bhakthavatsalam, S.; Khashabi, D.; Khot, T.; Dalvi Mishra, B.; Richardson, K.; Sabharwal, A.; Schoenick, C.; Tafjord, O.; and Clark, P. 2021. Think you have Solved Direct-Answer Question Answering? Try ARC-DA, the Direct-Answer AI2 Reasoning Challenge. arXiv e-prints, arXiv--2102

  6. [6]

    D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al

    Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J. D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. 2020. Language models are few-shot learners. Advances in neural information processing systems, 33: 1877--1901

  7. [7]

    Cai, Z.; Cao, M.; Chen, H.; Chen, K.; Chen, K.; Chen, X.; Chen, X.; Chen, Z.; Chen, Z.; Chu, P.; et al. 2024. Internlm2 technical report. arXiv preprint arXiv:2403.17297

  8. [8]

    W.; Sutton, C.; Gehrmann, S.; et al

    Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H. W.; Sutton, C.; Gehrmann, S.; et al. 2023. Palm: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240): 1--113

Show all 53 references
  1. [9]

    Clark, C.; Lee, K.; Chang, M.-W.; Kwiatkowski, T.; Collins, M.; and Toutanova, K. 2019. BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:...

  2. [10]

    Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding. In Burstein, J.; Doran, C.; and Solorio, T., eds., Proceedings of the 2019 Conference of the North A merican Chapter of the Association...

  3. [11]

    Dodge, J.; Sap, M.; Marasovi \'c , A.; Agnew, W.; Ilharco, G.; Groeneveld, D.; Mitchell, M.; and Gardner, M. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Langua...

  4. [12]

    Dong, Q.; Li, L.; Dai, D.; Zheng, C.; Wu, Z.; Chang, B.; Sun, X.; Xu, J.; and Sui, Z. 2022. A survey on in-context learning. arXiv preprint arXiv:2301.00234

  5. [13]

    Dubey, A.; Jauhri, A.; Pandey, A.; Kadian, A.; Al-Dahle, A.; Letman, A.; Mathur, A.; Schelten, A.; Yang, A.; Fan, A.; et al. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783

  6. [14]

    Engstrom, L.; Feldmann, A.; and Madry, A. 2024. DsDm: Model-Aware Dataset Selection with Datamodels. arXiv preprint arXiv:2401.12926

  7. [15]

    Gao, L.; Biderman, S.; Black, S.; Golding, L.; Hoppe, T.; Foster, C.; Phang, J.; He, H.; Thite, A.; Nabeshima, N.; et al. 2020. The pile: An 800gb dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027

  8. [16]

    Gauthier, T. D. 2001. Detecting trends using Spearman's rank correlation coefficient. Environmental forensics, 2(4): 359--362

  9. [17]

    Gururangan, S.; Card, D.; Dreier, S.; Gade, E.; Wang, L.; Wang, Z.; Zettlemoyer, L.; and Smith, N. A. 2022. Whose Language Counts as High Quality? Measuring Language Ideologies in Text Data Selection. In Proceedings of the 2022 Conference on Empirical Methods in Natural Langua...

  10. [18]

    Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300

  11. [19]

    Hoffmann, J.; Borgeaud, S.; Mensch, A.; Buchatskaya, E.; Cai, T.; Rutherford, E.; Casas, D. d. L.; Hendricks, L. A.; Welbl, J.; Clark, A.; et al. 2022. Training compute-optimal large language models. arXiv preprint arXiv:2203.15556

  12. [20]

    Huang, L.; Yu, W.; Ma, W.; Zhong, W.; Feng, Z.; Wang, H.; Chen, Q.; Peng, W.; Feng, X.; Qin, B.; and Liu, T. 2023 a . A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions. ArXiv, abs/2311.05232

  13. [21]

    Huang, Y.; Bai, Y.; Zhu, Z.; Zhang, J.; Zhang, J.; Su, T.; Liu, J.; Lv, C.; Zhang, Y.; Lei, J.; Fu, Y.; Sun, M.; and He, J. 2023 b . C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models. arXiv:2305.08322

  14. [22]

    Ilyas, A.; Park, S.; Engstrom, L.; Leclerc, G.; and Madry, A. 2022. Datamodels: Predicting Predictions from Training Data, Baltimore. In Proceedings of the 39 th International Conference on Machine Learning, volume 162

  15. [23]

    Kandpal, N.; Deng, H.; Roberts, A.; Wallace, E.; and Raffel, C. 2023. Large language models struggle to learn long-tail knowledge. In International Conference on Machine Learning, 15696--15707. PMLR

  16. [24]

    Kool, W.; Van Hoof, H.; and Welling, M. 2019. Stochastic beams and where to find them: The gumbel-top-k trick for sampling sequences without replacement. In International Conference on Machine Learning, 3499--3508. PMLR

  17. [25]

    Kreutzer, J.; Caswell, I.; Wang, L.; Wahab, A.; van Esch, D.; Ulzii-Orshikh, N.; Tapo, A.; Subramani, N.; Sokolov, A.; Sikasote, C.; et al. 2022. Quality at a glance: An audit of web-crawled multilingual datasets. Transactions of the Association for Computational Linguistics, ...

  18. [26]

    Lee, K.; Ippolito, D.; Nystrom, A.; Zhang, C.; Eck, D.; Callison-Burch, C.; and Carlini, N. 2022. Deduplicating Training Data Makes Language Models Better. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8424--8445

  19. [27]

    Li, H.; Zhang, Y.; Koto, F.; Yang, Y.; Zhao, H.; Gong, Y.; Duan, N.; and Baldwin, T. 2024. CMMLU: Measuring massive multitask language understanding in Chinese. arXiv:2306.09212

  20. [28]

    Li, Y.; Bubeck, S.; Eldan, R.; Del Giorno, A.; Gunasekar, S.; and Lee, Y. T. 2023. Textbooks are all you need ii: phi-1.5 technical report. arXiv preprint arXiv:2309.05463

  21. [29]

    Lu, K.; Yuan, H.; Yuan, Z.; Lin, R.; Lin, J.; Tan, C.; Zhou, C.; and Zhou, J. 2023. \# InsTag: Instruction Tagging for Analyzing Supervised Fine-tuning of Large Language Models. In The Twelfth International Conference on Learning Representations

  22. [30]

    Lucy, L.; Gururangan, S.; Soldaini, L.; Strubell, E.; Bamman, D.; Klein, L.; and Dodge, J. 2024. AboutMe: Using Self-Descriptions in Webpages to Document the Effects of English Pretraining Data Filters. arXiv preprint arXiv:2401.06408

  23. [31]

    Marion, M.; \"U st \"u n, A.; Pozzobon, L.; Wang, A.; Fadaee, M.; and Hooker, S. 2023. When less is more: Investigating data pruning for pretraining llms at scale. arXiv preprint arXiv:2309.04564

  24. [32]

    Mihaylov, T.; Clark, P.; Khot, T.; and Sabharwal, A. 2018. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In Conference on Empirical Methods in Natural Language Processing

  25. [33]

    Muennighoff, N.; Rush, A.; Barak, B.; Le Scao, T.; Tazi, N.; Piktus, A.; Pyysalo, S.; Wolf, T.; and Raffel, C. A. 2024. Scaling data-constrained language models. Advances in Neural Information Processing Systems, 36

  26. [34]

    OpenAI, R. 2023. GPT-4 technical report. ArXiv, 2303

  27. [35]

    Pan, S.; Luo, L.; Wang, Y.; Chen, C.; Wang, J.; and Wu, X. 2024. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering

  28. [36]

    Pao, D.; Lin, W.; and Liu, B. 2010. A memory-efficient pipelined implementation of the aho-corasick string-matching algorithm. ACM Transactions on Architecture and Code Optimization (TACO), 7(2): 1--27

  29. [37]

    Patel, J. M. 2020. Introduction to Common Crawl Datasets, 277--324. Berkeley, CA: Apress. ISBN 978-1-4842-6576-5

  30. [38]

    T.; and Camacho-Collados, J

    Pilehvar, M. T.; and Camacho-Collados, J. 2019. WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techn...

  31. [39]

    W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; et al

    Rae, J. W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; et al. 2021. Scaling language models: Methods, analysis & insights from training gopher. arXiv preprint arXiv:2112.11446

  32. [40]

    A.; and Gordon, A

    Roemmele, M.; Bejan, C. A.; and Gordon, A. S. 2011. Choice of plausible alternatives: An evaluation of commonsense causal reasoning. In 2011 AAAI Spring Symposium Series

  33. [41]

    H.; Caverlee, J.; McAuley, J.; and Cheng, D

    Sachdeva, N.; Coleman, B.; Kang, W.-C.; Ni, J.; Hong, L.; Chi, E. H.; Caverlee, J.; McAuley, J.; and Cheng, D. Z. 2024. How to Train Data-Efficient LLMs. arXiv preprint arXiv:2402.09668

  34. [42]

    W.; Chowdhery, A.; Le, Q.; Chi, E.; Zhou, D.; et al

    Suzgun, M.; Scales, N.; Sch \"a rli, N.; Gehrmann, S.; Tay, Y.; Chung, H. W.; Chowdhery, A.; Le, Q.; Chi, E.; Zhou, D.; et al. 2023. Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them. In Findings of the Association for Computational Linguistics: ACL 2023,...

  35. [43]

    Thakkar, M.; Bolukbasi, T.; Ganapathy, S.; Vashishth, S.; Chandar, S.; and Talukdar, P. 2023. Self-Influence Guided Data Reweighting for Language Model Pre-training. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2033--2045

  36. [44]

    Touvron, H.; Martin, L.; Stone, K.; Albert, P.; Almahairi, A.; Babaei, Y.; Bashlykov, N.; Batra, S.; Bhargava, P.; Bhosale, S.; et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  37. [45]

    Wang, A.; Singh, A.; Michael, J.; Hill, F.; Levy, O.; and Bowman, S. 2018. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 353--355

  38. [46]

    Wettig, A.; Gupta, A.; Malik, S.; and Chen, D. 2024. QuRating: Selecting High-Quality Data for Training Language Models. arXiv preprint arXiv:2402.09739

  39. [47]

    L.; Fan, A.; Akiki, C.; Pavlick, E.; Ili \'c , S.; Hesslow, D.; Castagn \'e , R.; Luccioni, A

    Workshop, B.; Scao, T. L.; Fan, A.; Akiki, C.; Pavlick, E.; Ili \'c , S.; Hesslow, D.; Castagn \'e , R.; Luccioni, A. S.; Yvon, F.; et al. 2022. Bloom: A 176b-parameter open-access multilingual language model. arXiv preprint arXiv:2211.05100

  40. [48]

    M.; Pham, H.; Dong, X.; Du, N.; Liu, H.; Lu, Y.; Liang, P

    Xie, S. M.; Pham, H.; Dong, X.; Du, N.; Liu, H.; Lu, Y.; Liang, P. S.; Le, Q. V.; Ma, T.; and Yu, A. W. 2024 a . Doremi: Optimizing data mixtures speeds up language model pretraining. Advances in Neural Information Processing Systems, 36

  41. [49]

    M.; Santurkar, S.; Ma, T.; and Liang, P

    Xie, S. M.; Santurkar, S.; Ma, T.; and Liang, P. S. 2024 b . Data selection for language models via importance resampling. Advances in Neural Information Processing Systems, 36

  42. [50]

    Xu, L.; Hu, H.; Zhang, X.; Li, L.; Cao, C.; Li, Y.; Xu, Y.; Sun, K.; Yu, D.; Yu, C.; et al. 2020. CLUE: A Chinese Language Understanding Evaluation Benchmark. In Proceedings of the 28th International Conference on Computational Linguistics, 4762--4772

  43. [51]

    Xu, L.; Lu, X.; Yuan, C.; Zhang, X.; Xu, H.; Yuan, H.; Wei, G.; Pan, X.; Tian, X.; Qin, L.; et al. 2021. Fewclue: A chinese few-shot learning evaluation benchmark. arXiv preprint arXiv:2107.07498

  44. [52]

    Yuan, S.; Zhao, H.; Du, Z.; Ding, M.; Liu, X.; Cen, Y.; Zou, X.; Yang, Z.; and Tang, J. 2021. Wudaocorpora: A super large-scale chinese corpora for pre-training language models. AI Open, 2: 65--68

  45. [53]

    Zhang, F.; Liu, X.; Tang, J.; Dong, Y.; Yao, P.; Zhang, J.; Gu, X.; Wang, Y.; Kharlamov, E.; Shao, B.; et al. 2022. Oag: Linking entities across large-scale heterogeneous knowledge graphs. IEEE Transactions on Knowledge and Data Engineering

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.