Pith. sign in

REVIEW 3 cited by

BIOS: An Algorithmically Generated Biomedical Knowledge Graph

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2203.09975 v2 pith:RNL4VDGS submitted 2022-03-18 cs.CL cs.LG

classification cs.CLcs.LG
keywords biomedicalbiostermscurationdevelopmentgeneratedknowledgemachine
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Biomedical knowledge graphs (BioMedKGs) are essential infrastructures for biomedical and healthcare big data and artificial intelligence (AI), facilitating natural language processing, model development, and data exchange. For decades, these knowledge graphs have been developed via expert curation; however, this method can no longer keep up with today's AI development, and a transition to algorithmically generated BioMedKGs is necessary. In this work, we introduce the Biomedical Informatics Ontology System (BIOS), the first large-scale publicly available BioMedKG generated completely by machine learning algorithms. BIOS currently contains 4.1 million concepts, 7.4 million terms in two languages, and 7.3 million relation triplets. We present the methodology for developing BIOS, including the curation of raw biomedical terms, computational identification of synonymous terms and aggregation of these terms to create concept nodes, semantic type classification of the concepts, relation identification, and biomedical machine translation. We provide statistics on the current BIOS content and perform preliminary assessments of term quality, synonym grouping, and relation extraction. The results suggest that machine learning-based BioMedKG development is a viable alternative to traditional expert curation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CliniQ: A Multi-faceted Benchmark for Electronic Health Record Retrieval with Semantic Match Assessment

    cs.IR 2025-02 conditional novelty 6.0 of 10

    CliniQ is a public EHR retrieval benchmark with 77,206 LLM-annotated relevance judgments, showing that BM25 is a strong baseline and that semantic matches drive dense-retriever gains.

  2. GENIE: Generative Note Information Extraction model for structuring EHR data

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A single fine-tuned 8B-parameter LLM extracts medical terms plus six clinical attributes from EHR notes in one pass, outperforming cTAKES and MetaMap on a small human-annotated test set.

  3. MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

    cs.CL 2025-08 reject novelty 4.0 of 10

    A submission whose abstract describes a large temporal medical knowledge graph built by LLM agents, but whose full text is an unrelated paper on histogram regression, leaving the announced claims unsupported.

Pith tools