Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Improving Biomedical Knowledge Graph Quality: A Community Approach

T0 review · 4 major / 5 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A 28-point audit of 16 biomedical knowledge graphs finds that only one, RTX-KG2, meets every reusability and transparency check, while most fail on versioning and public request tracking.

desk verdict A useful, honest survey of 16 biomedical KGs with real descriptive value, but the 28-item scorecard and the 'higher score equals greater trustworthiness' leap need a serious haircut before this should be treated as a ranking. read the letter →

arxiv 2508.21774 v1 pith:PD47LAX2 submitted 2025-08-29 q-bio.OT

classification q-bio.OT
keywords biomedicalknowledgegraphsgraphqualityreusabilityFAIRprinciplesversioningBiolinkModelmetadatastandardsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that biomedical knowledge graphs, unlike ontologies, lack shared conventions for construction, documentation, and release, and that this is why many graphs are hard for outsiders to reuse. To demonstrate the gap, the authors derive a 28-point evaluation rubric from widely accepted data-governance and ontology principles and apply it to 16 publicly accessible biomedical KGs. The result: only one graph, RTX-KG2, satisfies all 28 binary checks; the weakest practices across the field are having a public tracker for requests, identifying stable versions, and making prior versions accessible with documented changes. The paper reads this as evidence that even graphs that appear to follow best practices often hide the information needed for external reuse, and it calls for community-wide adoption of shared criteria, machine-readable metadata, and standardized exchange formats such as Biolink and KGX.

What carries the argument

The load-bearing instrument is the 28-criteria scoring rubric, organized into six principles, with each criterion scored strictly yes/no using only publicly accessible information. It is supplemented by mapping each KG's node-type labels onto the Biolink Model—a standardized schema for biological entity types and relationships—so that graphs with different native vocabularies can be compared on content. The rubric's countable yes/no output is what carries the comparative claim: it converts undocumented practice into a transparent score across the 16 graphs.

What would settle it

Re-run the 28-criteria review with several independent curator teams on the same 16 KGs and test inter-rater agreement; or check whether the scores correlate with independently observed reuse such as download counts, citations, and successful external integrations. Low agreement or no correlation with reuse would falsify the claim that the rubric measures trustworthiness.

Watch

Extended reading notes

Core claim

The central claim is that a deliberately small set of transparency and reusability criteria can expose real differences among biomedical KGs, and that current practice is far from adequate. The authors built a rubric with six principles—access level and type; provenance of nodes and edges; documented standards, schema, and construction; update frequency and versioning; evaluation and fitness for purpose; and licensing—totaling 28 yes/no items, and scored 16 KGs by manual review of websites, code repositories, and publications. RTX-KG2 was the only graph with all 'yes' answers. Even graphs that look aligned with best practices often fail to provide, in one findable place, the versions of sour

Load-bearing premise

The rubric is assumed to be a valid measure of trustworthiness and reusability; the paper scores graphs against it but never validates the scores against actual reuse, user experience, or downstream performance, and partial compliance was counted as a full yes.

Editorial extensions

If this is right

  • If the criteria are adopted, KG builders get a concrete checklist for what to document before release, and users get a basis for comparing graphs on access, provenance, versioning, evaluation, and licensing.
  • The weakest observed practices—public request trackers, clearly identified stable versions, and archived prior versions with documented changes—identify versioning and community feedback as the first targets for improvement.
  • Adoption of machine-readable metadata files and shared exchange formats such as Biolink and KGX would make cross-graph comparison routine rather than a manual, ad hoc exercise.
  • Standardized documentation would make it easier to detect when two graphs ingest the same source differently, which the paper shows is common and currently hard to see.
  • If evaluation and fitness-for-purpose criteria are taken seriously, KG authors would need to provide case studies, comparisons, and confidence measures as part of the resource, not as optional extras.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural extension is to turn the 28 criteria into a machine-checkable metadata schema so scoring becomes automatic and continuous instead of a one-time curator review.
  • The paper's node-type mapping suggests a testable claim: after mapping to a common schema, many 'different' graphs are more similar in content than their native vocabularies suggest; this could be quantified by measuring overlap in entities and edges.
  • The assumption that higher scores equal greater trustworthiness could be tested against behavioral data—for example, whether scored graphs are downloaded, cited, or reused more often, or whether users report fewer integration errors.
  • A public registry publishing these scores, updated as graphs change, would create ongoing pressure for improvement, much as ontology dashboards have done for ontologies.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a 28-criterion evaluation rubric for biomedical knowledge graphs (KGs), grounded in FAIR, TRUST, O3, and OBO Foundry principles, and applies it manually to 16 publicly available KGs. It reports that only RTX-KG2 scored 'yes' on every criterion, that common deficiencies include lack of public issue trackers (7/16), clearly identified stable versions (9/16), and accessible prior versions with documented changes (9/16), and that KGs vary widely in the number of node types and ingested sources. The authors argue that the community should adopt shared criteria, standards such as Biolink and KGX, and machine-readable metadata to improve transparency, comparability, and reusability.

Significance. If the scorecard and descriptive findings hold, the paper provides a practical checklist and a baseline survey that could catalyze community-wide documentation standards for biomedical KGs. Strengths include the explicit, publicly available criteria; the detailed per-KG evidence in the supplemental analysis; falsifiable claims about under-documentation; and candid statements of scope limitations. The paper also names its own restrictions, such as the exclusion of deep code/file review and the lenient 'partial compliance counts as a full yes' rule, which increases the robustness of the finding that documentation is often missing. However, the paper's stronger inference—that higher scores reflect greater 'trustworthiness'—is not empirically validated, and the comparative ranking depends on a manual, non-blinded scoring protocol with no inter-rater reliability assessment.

major comments (4)
  1. [Methods (KG Evaluation) and Results (KG Evaluation)] The central comparative result—'RTX-KG2 was the only knowledge graph that scored yes on every criteria'—rests entirely on a manual scoring protocol with no inter-rater reliability statistic and no blinding. With six individual reviewers, only secondary review, and the rule that 'partial compliance counted as a full yes' while 'could not be found' counted as a no, each binary call is a subjective judgment about documentation discoverability and completeness. A second set of independent raters could plausibly flip items, altering the unique all-yes status and the reported deficiency rates (e.g., D.2 = 7/16, D.1 and D.5 = 9/16). Please report inter-rater agreement on a subset of KGs, publish the full evidence trail per item, or temper the ranking to a descriptive documentation survey.
  2. [Methods (KG Evaluation)] The paper explicitly states that 'A critical review of KG code and files is outside the scope of this work and was not performed.' The evaluation was limited to websites, software repositories, and publications. Consequently, for criteria that describe properties of the KG itself (e.g., B.4 'Nodes and edges have source information', B.5 'Duplicate edge management', C.4 'Documented data transforms'), a 'no' may reflect absence from the public documentation rather than absence from the KG. The abstract's claim that KGs 'obscure essential information' is well supported by the documentation-based findings, but the scorecard conflates documentation completeness with KG-internal properties. The authors should either reframe the criteria as 'documentation transparency' or calibrate the distinction by inspecting at least a sample of KG files.
  3. [Discussion] The assertion that 'higher scores across these criteria reflect greater trustworthiness' is load-bearing for the paper's recommendations but is not validated. The criteria are derived from FAIR, TRUST, and O3 principles, which is a reasonable normative basis, but no evidence links scores to actual reuse outcomes, downstream task performance, or user trust. A concrete test would be to correlate scores with independent usage indicators (e.g., downloads, citations, successful third-party integrations) or to have external users rate the same KGs. Without such evidence, the claim should be weakened to 'documentation completeness and transparency,' which is what the rubric directly measures.
  4. [Results (KG Evaluation)] Only the Monarch Initiative KG was reviewed by external experts; the remaining 15 KGs were reviewed by members of the same ecosystem that maintains or co-authors several of them, including RTX-KG2, ROBOKOP, Clinical KG, NCATS GARD, HRA-KG, and Monarch. Since RTX-KG2—a KG developed by co-authors—is the unique perfect scorer, the comparative ranking is vulnerable to privileged knowledge and conflicts of interest. Please disclose which reviewer scored each KG, use external reviewers for all KGs (or at least for the highest-ranked ones), or provide an independent audit of the scores.
minor comments (5)
  1. [Results (KG Evaluation)] The sentence 'Of note, KGs such as EmBiology and SPOKE are largely accessible only through a paywall' appears inconsistent with the supplemental table, where SPOKE is scored as openly accessible (A.3/A.4/A.5 = yes, license CC BY 4.0). Please correct or clarify.
  2. [Results (KG Evaluation)] 'scored yes on every criteria' should be 'every criterion' (criteria is plural).
  3. [Discussion] 'data providence' should be 'data provenance' (the intended concept).
  4. [Methods (KG Node Types)] The ChatGPT-based mapping of node labels to Biolink types lacks reproducibility details: model version, prompt, date, temperature, and any manual validation. This mapping drives Figure 3B and the clustering comparison; please provide the exact mapping procedure and validation results, or replace it with a deterministic manual mapping.
  5. [Throughout] Capitalization is inconsistent: 'Biolink' and 'BioLink' are both used. Please standardize.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an explicit rubric-based audit, not a derivation that reduces to its own inputs.

full rationale

The paper contains no equations, fitted parameters, or predictive model. Its central claims are descriptive summaries of a manual 28-criterion audit of 16 biomedical KGs. The criteria are reported in full (Methods, Principles A–F) and are grounded in external frameworks (FAIR, TRUST, O3, OBO Foundry) rather than derived from the outcome being claimed. The finding that “RTX-KG2 was the only knowledge graph that scored yes on every criteria” is a direct tally of the published binary scoring matrix, not the result of fitting a parameter or of a self-citation chain. The statement that higher scores reflect greater trustworthiness is explicitly proposed rather than derived; it is a definitional framing of the scorecard, but the underlying criteria are transparent, and the observed variation in node types and source integration is independently descriptive. The main limitations — manual scoring, non-blinded review except for Monarch, partial compliance counted as “yes,” and author overlap with some scored ecosystems — are reliability and conflict-of-interest concerns rather than circularity under the strict definition used here. No load-bearing conclusion reduces to its own input by construction.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

Nothing is fitted numerically and nothing is invented. The load is carried by three domain assumptions (criteria validity, scoring reliability, sample representativeness) plus a fourth specific to this author community: Biolink as the canonical target vocabulary for node-type harmonization. These assumptions bound what the survey can establish: they support a documentation-hygiene comparison, not a correctness or trustworthiness certificate.

free parameters (3)
  • Binary leniency rule = 'partial compliance counted as a full yes'
    Hand-chosen scoring rule in Methods that inflates every KG's score; it makes the main finding conservative but shapes the exact ranking and the uniqueness of RTX-KG2's perfect score.
  • Equal item weighting = 28 items weighted equally
    No justification is given for treating each subcriterion as equally important to reusability; licensing is one item while access has five.
  • ChatGPT node-label mapping = unvalidated label-to-Biolink mapping
    The node-type comparison (Figure 3B) depends on an LLM mapping of heterogeneous labels to Biolink classes; no manual accuracy check or confidence report is provided.
assumptions (4)
  • domain assumption The 28 binary criteria, grounded in FAIR, TRUST, O3, and OBO Foundry principles, measure KG reusability and trustworthiness
    Principles and Discussion: the paper states its goal is to determine 'if one can easily access and gain specific information about a given KG so as to determine its utility for reuse' and later asserts higher scores reflect greater trustworthiness. This linkage is asserted, not validated against measured reuse or accuracy outcomes.
  • domain assumption Manual yes/no scoring by six curators with secondary review yields scores reliable enough to rank the 16 KGs and identify RTX-KG2 as the unique perfect scorer
    Methods: one primary reviewer per KG plus a secondary review; no inter-rater reliability statistics are reported, and 'partial compliance counted as a full yes' compresses variability.
  • domain assumption The 16 KGs selected are representative of the biomedical KG landscape
    Methods and Results: selection by 'varying purposes, sizes, and creators' plus inclusion of KG-Registry members; the set skews toward consortium-affiliated graphs (Monarch, ROBOKOP, RTX-KG2, SPOKE, Clinical KG).
  • ad hoc to paper Biolink Model node types are a valid common vocabulary for comparing KG node labels
    Results: ChatGPT maps each KG's node labels to Biolink classes used for the Figure 3 clustering; no validation rate for the mapping is reported, and several authors are Biolink maintainers.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Biomedical Knowledge Graph Quality: A Community Approach." pith.science (2026). https://pith.science/paper/PD47LAX2

@misc{pith2026250821774,
  author       = {Pith},
  title        = {Pith review of: Improving Biomedical Knowledge Graph Quality: A Community Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PD47LAX2}},
  note         = {Machine review of arXiv:2508.21774}
}
read the original abstract

Biomedical knowledge graphs (KGs) are widely used across research and translational settings, yet their design decisions and implementation are often opaque. Unlike ontologies that more frequently adhere to established creation principles, biomedical KGs lack consistent practices for construction, documentation, and dissemination. To address this gap, we introduce a set of evaluation criteria grounded in widely accepted data standards and principles from related fields. We apply these criteria to 16 biomedical KGs, revealing that even those that appear to align with best practices often obscure essential information required for external reuse. Moreover, biomedical KGs, despite pursuing similar goals and ingesting the same sources in some cases, display substantial variation in models, source integration, and terminology for node types. Reaping the potential benefits of knowledge graphs for biomedical research while reducing wasted effort requires community-wide adoption of shared criteria and maturation of standards such as BioLink and KGX. Such improvements in transparency and standardization are essential for creating long-term reusability, improving comparability across resources, and enhancing the overall utility of KGs within biomedicine.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Koza and Koza-Hub for born-interoperable knowledge graph generation using KGX

    cs.DB 2025-09 conditional novelty 6.0 of 10

    Koza and Koza-Hub provide a modular, YAML-configured Python pipeline that converts biomedical source data into standardized KGX knowledge graph artifacts, with ingest recipes for 14 repositories.

Reference graph

Works this paper leans on

66 extracted references · 53 canonical work pages · cited by 1 Pith paper

  1. [1]

    ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering

    Morton K, Wang P, Bizon C, Cox S, Balhoff J, Kebede Y, et al. ROBOKOP: an abstraction layer and user interface for knowledge graphs to support question answering. Bioinformatics. 2019;35: 5382–5384

  2. [2]

    ROBOKOP KG and KGB: Integrated Knowledge Graphs from Federated Sources

    Bizon C, Cox S, Balhoff J, Kebede Y, Wang P, Morton K, et al. ROBOKOP KG and KGB: Integrated Knowledge Graphs from Federated Sources. J Chem Inf Model. 2019;59: 4968–4973

  3. [3]

    Comparative Toxicogenomics Database (CTD): update 2023

    Davis AP, Wiegers TC, Johnson RJ, Sciaky D, Wiegers J, Mattingly CJ. Comparative Toxicogenomics Database (CTD): update 2023. Nucleic Acids Res. 2023;51: D1257–D1262

  4. [4]

    The Gene Ontology knowledgebase in 2023

    Gene Ontology Consortium, Aleksander SA, Balhoff J, Carbon S, Cherry JM, Drabkin HJ, et al. The Gene Ontology knowledgebase in 2023. Genetics. 2023;224. doi:10.1093/genetics/iyad031

  5. [5]

    DrugBank 6.0: The DrugBank knowledgebase for 2024

    Knox C, Wilson M, Klinger CM, Franklin M, Oler E, Wilson A, et al. DrugBank 6.0: The DrugBank knowledgebase for 2024. Nucleic Acids Res. 2024;52: D1265–D1275

  6. [6]

    Knowledge graphs: Opportunities and challenges

    Peng C, Xia F, Naseriparsa M, Osborne F. Knowledge graphs: Opportunities and challenges. Artif Intell Rev. 2023; 1–32

  7. [7]

    A Literature-Based Knowledge Graph Embedding Method for Identifying Drug Repurposing Opportunities in Rare Diseases

    Sosa DN, Derry A, Guo M, Wei E, Brinton C, Altman RB. A Literature-Based Knowledge Graph Embedding Method for Identifying Drug Repurposing Opportunities in Rare Diseases

  8. [8]

    Clustering rare diseases within an ontology-enriched knowledge graph

    Sanjak J, Zhu Q, Mathé EA. Clustering rare diseases within an ontology-enriched knowledge graph. doi:10.1101/2023.02.15.528673

Show all 66 references
  1. [9]

    Integrating and formatting biomedical data as pre-calculated knowledge graph embeddings in the Bioteque

    Fernández-Torras A, Duran-Frigola M, Bertoni M, Locatelli M, Aloy P. Integrating and formatting biomedical data as pre-calculated knowledge graph embeddings in the Bioteque. Nat Commun. 2022;13: 5304

  2. [10]

    The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species

    Putman TE, Schaper K, Matentzoglu N, Rubinetti VP, Alquaddoomi FS, Cox C, et al. The Monarch Initiative in 2024: an analytic platform integrating phenotypes, genes and diseases across species. Nucleic Acids Res. 2024;52: D938–D949

  3. [11]

    Efficient reinterpretation of rare disease cases using Exomiser

    Vestito L, Jacobsen JOB, Walker S, Cipriani V, Harris NL, Haendel MA, et al. Efficient reinterpretation of rare disease cases using Exomiser. NPJ Genom Med. 2024;9: 65

  4. [12]

    Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery

    Buniello A, Suveges D, Cruz-Castillo C, Llinares MB, Cornu H, Lopez I, et al. Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery. Nucleic Acids Res. 2025;53: D1467–D1475

  5. [13]

    Biomedical knowledge graph-optimized prompt generation for large language models

    Soman K, Rose PW, Morris JH, Akbas RE, Smith B, Peetoom B, et al. Biomedical knowledge graph-optimized prompt generation for large language models. Bioinformatics. 2024;40. doi:10.1093/bioinformatics/btae560

  6. [14]

    Foundation model for biomedical graphs: Integrating knowledge graphs and protein structures to large language models

    Kim Y. Foundation model for biomedical graphs: Integrating knowledge graphs and protein structures to large language models. Fu X, Fleisig E, editors. Annu Meet Assoc Comput Linguistics. 2024; 346–355

  7. [15]

    Knowledge graph-based thought: a knowledge graph-enhanced LLM framework for pan-cancer question answering

    Feng Y, Zhou L, Ma C, Zheng Y, He R, Li Y. Knowledge graph-based thought: a knowledge graph-enhanced LLM framework for pan-cancer question answering. Gigascience. 2025;14. doi:10.1093/gigascience/giae082

  8. [16]

    BioGPT: generative pre-trained transformer for biomedical text generation and mining

    Luo R, Sun L, Xia Y, Qin T, Zhang S, Poon H, et al. BioGPT: generative pre-trained transformer for biomedical text generation and mining. Brief Bioinform. 2022;23. doi:10.1093/bib/bbac409

  9. [17]

    ScispaCy: Fast and robust models for biomedical natural language processing

    Neumann M, King D, Beltagy I, Ammar W. ScispaCy: Fast and robust models for biomedical natural language processing. arXiv [cs.CL]. 2019. Available: https://www.aclweb.org/anthology/W19-5034.pdf

  10. [18]

    Leveraging medical knowledge graphs into Large Language Models for diagnosis prediction: Design and application study

    Gao Y, Li R, Croxford E, Caskey J, Patterson BW, Churpek M, et al. Leveraging medical knowledge graphs into Large Language Models for diagnosis prediction: Design and application study. arXiv [cs.CL]. 2023. Available: http://arxiv.org/abs/2308.14321

  11. [19]

    KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA

    Su X, Wang Y, Gao S, Liu X, Giunchiglia V, Clevert D-A, et al. KGARevion: An AI Agent for Knowledge-Intensive Biomedical QA. arXiv [cs.AI]. 2024. Available: http://arxiv.org/abs/2410.04660

  12. [20]

    Adaptive knowledge graphs enhance medical question answering: Bridging the gap between LLMs and evolving medical knowledge. arXiv. Available: https://arxiv.org/html/2502.13010v1

  13. [21]

    Deep Bidirectional Language-Knowledge Graph Pretraining

    Yasunaga M, Bosselut A, Ren H, Zhang X, Manning CD, Liang P, et al. Deep Bidirectional Language-Knowledge Graph Pretraining. arXiv [cs.CL]. 2022. doi:10.48550/ARXIV.2210.09338

  14. [22]

    Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning

    Caufield JH, Hegde H, Emonet V, Harris NL, Joachimiak MP, Matentzoglu N, et al. Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES): a method for populating knowledge bases using zero-shot learning. Bioinformatics. 2024;40. doi:10.1093/bioinformatics/btae104

  15. [23]

    Building a knowledge graph to enable precision medicine

    Chandak P, Huang K, Zitnik M. Building a knowledge graph to enable precision medicine. Sci Data. 2023;10: 67

  16. [24]

    KG4SL: Knowledge graph neural network for synthetic lethality prediction in human cancers

    Wang S, Xu F, Li Y, Wang J, Zhang K, Liu Y, et al. KG4SL: Knowledge graph neural network for synthetic lethality prediction in human cancers. Bioinformatics. 7 2021;37: I418–I425

  17. [25]

    Rare disease-based scientific annotation knowledge graph

    Hofmann-Apitius M, Romacker M, Wolstencroft SK, Zhu Q. Rare disease-based scientific annotation knowledge graph. Available: https://neooj.com/

  18. [26]

    An integrative knowledge graph for rare diseases, derived from the Genetic and Rare Diseases Information Center (GARD)

    Zhu Q, Nguyen DT, Grishagin I, Southall N, Sid E, Pariser A. An integrative knowledge graph for rare diseases, derived from the Genetic and Rare Diseases Information Center (GARD). J Biomed Semantics. 12 2020;11. doi:10.1186/s13326-020-00232-y

  19. [27]

    ATOM: Construction of Anti-tumor Biomaterial Knowledge Graph by Biomedicine Literature

    Wang T, Duan L, He C, Deng G, Qin R, Zhang Y. ATOM: Construction of Anti-tumor Biomaterial Knowledge Graph by Biomedicine Literature. 2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM). IEEE; 2019. pp. 1256–1258

  20. [28]

    Knowledge Graph-Enabled Cancer Data Analytics

    Hasan SMS, Rivera D, Wu X-C, Durbin EB, Christian JB, Tourassi G. Knowledge Graph-Enabled Cancer Data Analytics. IEEE J Biomed Health Inform. 2020;24: 1952–1967

  21. [29]

    RDBridge: a knowledge graph of rare diseases based on large-scale text mining

    Xing H, Zhang D, Cai P, Zhang R, Hu QN. RDBridge: a knowledge graph of rare diseases based on large-scale text mining. Bioinformatics. 7 2023;39. doi:10.1093/bioinformatics/btad440

  22. [30]

    Taming the Generative AI Wild West: Integrating Knowledge Graphs in Digital Library Systems

    D’Souza J. Taming the Generative AI Wild West: Integrating Knowledge Graphs in Digital Library Systems. Available: https://journal.code4lib.org/articles/18277

  23. [31]

    [cited 15 Aug 2025]

    Knowledge graphs and exploration: How to find your way in the data wilderness. [cited 15 Aug 2025]. Available: https://projet.liris.cnrs.fr/bd/?q=node/358

  24. [32]

    The Ontology of Biological Attributes (OBA)-computational traits for the life sciences

    Stefancsik R, Balhoff JP, Balk MA, Ball RL, Bello SM, Caron AR, et al. The Ontology of Biological Attributes (OBA)-computational traits for the life sciences. Mamm Genome. 2023;34: 364–378

  25. [33]

    OWL web ontology language reference

    Dean M, McGuinness DL. OWL web ontology language reference. [cited 21 Aug 2025]. Available: https://www.w3.org/TR/owl-ref/

  26. [34]

    Dead simple OWL design patterns

    Osumi-Sutherland D, Courtot M, Balhoff JP, Mungall C. Dead simple OWL design patterns. J Biomed Semantics. 2017;8: 18

  27. [35]

    ROBOT: A tool for automating ontology workflows

    Jackson RC, Balhoff JP, Douglass E, Harris NL, Mungall CJ, Overton JA. ROBOT: A tool for automating ontology workflows. BMC Bioinformatics. 2019;20: 407

  28. [36]

    OBO Foundry in 2021: operationalizing open data principles to evaluate ontologies

    Jackson R, Matentzoglu N, Overton JA, Vita R, Balhoff JP, Buttigieg PL, et al. OBO Foundry in 2021: operationalizing open data principles to evaluate ontologies. Database (Oxford). 2021;2021. doi:10.1093/database/baab069

  29. [37]

    Unifying the identification of biomedical entities with the Bioregistry

    Hoyt CT, Balk M, Callahan TJ, Domingo-Fernández D, Haendel MA, Hegde HB, et al. Unifying the identification of biomedical entities with the Bioregistry. Sci Data. 2022;9: 714

  30. [38]

    BioPortal: an open community resource for sharing, searching, and utilizing biomedical ontologies

    Vendetti J, Harris NL, Dorf MV, Skrenchuk A, Caufield JH, Gonçalves RS, et al. BioPortal: an open community resource for sharing, searching, and utilizing biomedical ontologies. Nucleic Acids Res. 2025;53: W84–W94

  31. [39]

    Industry-scale knowledge graphs: Lessons and challenges

    Noy N, Gao Y, Jain A, Narayanan A, Patterson A, Taylor J. Industry-scale knowledge graphs: Lessons and challenges. In: Communications of the ACM [Internet]. 1 Aug 2019 [cited 26 Jun 2025]. Available: https://cacm.acm.org/practice/industry-scale-knowledge-graphs/

  32. [40]

    Adoption of knowledge-graph best development practices for scalable and optimized manufacturing processes

    Jawad MS, Dhawale C, Ramli AAB, Mahdin H. Adoption of knowledge-graph best development practices for scalable and optimized manufacturing processes. MethodsX. 2023;10: 102124

  33. [41]

    Scholarly knowledge graphs through structuring scholarly communication: a review

    Verma S, Bhatia R, Harit S, Batish S. Scholarly knowledge graphs through structuring scholarly communication: a review. Complex Intell Syst. 2023;9: 1059–1095

  34. [42]

    [cited 27 Aug 2025]

    KG-registry. [cited 27 Aug 2025]. Available: https://kghub.org/kg-registry/

  35. [43]

    An analysis and metric of reusable data licensing practices for biomedical resources

    Carbon S, Champieux R, McMurry JA, Winfree L, Wyatt LR, Haendel MA. An analysis and metric of reusable data licensing practices for biomedical resources. PLoS One. 2019;14: e0213090

  36. [44]

    Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science

    Unni DR, Moxon SAT, Bada M, Brush M, Bruskiewich R, Caufield JH, et al. Biolink Model: A universal schema for knowledge graphs in clinical, biomedical, and translational science. Clin Transl Sci. 2022;15: 1848–1855

  37. [45]

    The Biomedical Data Translator Program: Conception, Culture, and Community

    Biomedical Data Translator Consortium. The Biomedical Data Translator Program: Conception, Culture, and Community. Clin Transl Sci. 2019;12: 91–94

  38. [46]

    RTX-KG2: a system for building a semantically standardized knowledge graph for translational biomedicine

    Wood EC, Glen AK, Kvarfordt LG, Womack F, Acevedo L, Yoon TS, et al. RTX-KG2: a system for building a semantically standardized knowledge graph for translational biomedicine. BMC Bioinformatics. 2022;23: 400

  39. [47]

    Clinical Knowledge Graph Integrates Proteomics Data into Clinical Decision-Making

    Santos A, Colaço AR, Nielsen AB, Niu L, Geyer PE, Coscia F, et al. Clinical Knowledge Graph Integrates Proteomics Data into Clinical Decision-Making. bioRxiv. 2020. p. 2020.05.09.084897. doi:10.1101/2020.05.09.084897

  40. [48]

    The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information

    Morris JH, Soman K, Akbas RE, Zhou X, Smith B, Meng EC, et al. The scalable precision medicine open knowledge engine (SPOKE): a massive knowledge graph of biomedical information. Bioinformatics. 2023;39. doi:10.1093/bioinformatics/btad080

  41. [49]

    Interpretation of biological experiments changes with evolution of the Gene Ontology and its annotations

    Tomczak A, Mortensen JM, Winnenburg R, Liu C, Alessi DT, Swamy V, et al. Interpretation of biological experiments changes with evolution of the Gene Ontology and its annotations. Sci Rep. 2018;8: 5115

  42. [50]

    Semantic units: organizing knowledge graphs into semantically meaningful units of representation

    Vogt L, Kuhn T, Hoehndorf R. Semantic units: organizing knowledge graphs into semantically meaningful units of representation. J Biomed Semantics. 2024;15: 7

  43. [51]

    The TRUST Principles for digital repositories

    Lin D, Crabtree J, Dillo I, Downs RR, Edmunds R, Giaretta D, et al. The TRUST Principles for digital repositories. Sci Data. 2020;7: 144

  44. [52]

    The O3 guidelines: open data, open code, and open infrastructure for sustainable curated scientific resources

    Hoyt CT, Gyori BM. The O3 guidelines: open data, open code, and open infrastructure for sustainable curated scientific resources. Sci Data. 2024;11: 547

  45. [53]

    Biomedical knowledge graph: A survey of domains, tasks, and real-world applications

    Lu Y, Goi SY, Zhao X, Wang J. Biomedical knowledge graph: A survey of domains, tasks, and real-world applications. arXiv [cs.CL]. 2025. Available: http://arxiv.org/abs/2501.11632

  46. [54]

    A review of biomedical datasets relating to drug discovery: a knowledge graph perspective

    Bonner S, Barrett IP, Ye C, Swiers R, Engkvist O, Bender A, et al. A review of biomedical datasets relating to drug discovery: a knowledge graph perspective. Brief Bioinform. 2022;23. doi:10.1093/bib/bbac404

  47. [55]

    Knowledge Graphs for drug repurposing: a review of databases and methods

    Perdomo-Quinteiro P, Belmonte-Hernández A. Knowledge Graphs for drug repurposing: a review of databases and methods. Brief Bioinform. 2024;25. doi:10.1093/bib/bbae461

  48. [56]

    PharmKG: A dedicated knowledge graph benchmark for bomedical data mining

    Zheng S, Rao J, Song Y, Zhang J, Xiao X, Fang EF, et al. PharmKG: A dedicated knowledge graph benchmark for bomedical data mining. Brief Bioinform. 7 2021;22. doi:10.1093/bib/bbaa344

  49. [57]

    Systematic integration of biomedical knowledge prioritizes drugs for repurposing

    Himmelstein DS, Lizee A, Hessler C, Brueggeman L, Chen SL, Hadley D, et al. Systematic integration of biomedical knowledge prioritizes drugs for repurposing. 2017. doi:10.7554/eLife.26726.001

  50. [58]

    Construction, deployment, and usage of the Human Reference Atlas Knowledge Graph

    Bueckle A, Herr BW 2nd, Hardi J, Quardokus EM, Musen MA, Börner K. Construction, deployment, and usage of the Human Reference Atlas Knowledge Graph. Sci Data. 2025;12: 1100

  51. [59]

    Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data

    Stear BJ, Mohseni Ahooyi T, Simmons JA, Kollar C, Hartman L, Beigel K, et al. Petagraph: A large-scale unifying knowledge graph framework for integrating biomolecular and biomedical data. Sci Data. 2024;11: 1338

  52. [60]

    [cited 18 Aug 2025]

    GenomicKB. [cited 18 Aug 2025]. Available: https://gkb.dcmb.med.umich.edu/

  53. [61]

    EmBiology explainer video

    Us PW. EmBiology explainer video. Elsevier; 2023. Available: https://www.elsevier.com/products/embiology

  54. [62]

    DrugMechDB: A curated database of drug mechanisms

    Gonzalez-Cavazos AC, Tanska A, Mayers M, Carvalho-Silva D, Sridharan B, Rewers PA, et al. DrugMechDB: A curated database of drug mechanisms. Sci Data. 2023;10: 632

  55. [63]

    Using predicate and provenance information from a knowledge graph for drug efficacy screening

    Vlietstra WJ, Vos R, Sijbers AM, van Mulligen EM, Kors JA. Using predicate and provenance information from a knowledge graph for drug efficacy screening. J Biomed Semantics. 2018;9: 23

  56. [64]

    Ensembl 2025

    Dyer SC, Austine-Orimoloye O, Azov AG, Barba M, Barnes I, Barrera-Enriquez VP, et al. Ensembl 2025. Nucleic Acids Res. 2025;53: D948–D957

  57. [65]

    Piloting a model-to-data approach to enable predictive analytics in health care through patient mortality prediction

    Bergquist T, Yan Y, Schaffter T, Yu T, Pejaver V, Hammarlund N, et al. Piloting a model-to-data approach to enable predictive analytics in health care through patient mortality prediction. J Am Med Inform Assoc. 2020;27: 1393–1400

  58. [66]

    recent versions

    Guinney J, Saez-Rodriguez J. Alternative models for sharing confidential biomedical data. Nat Biotechnol. 2018;36: 391–392. Detailed Principle AnalysisA. Access Level and TypesB. Provenance of Nodes and EdgesC. Documented standards, schema, constructionD. Update frequency and ...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.