Pith. sign in

REVIEW 2 cited by

CEAR: Automatic construction of a knowledge graph of chemical entities and roles from scientific literature

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21708 v1 pith:JJIOZZOW submitted 2024-07-31 cs.AI

classification cs.AI
keywords knowledgechemicalentitiesrolesscientificchebiliteraturecear
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Ontologies are formal representations of knowledge in specific domains that provide a structured framework for organizing and understanding complex information. Creating ontologies, however, is a complex and time-consuming endeavor. ChEBI is a well-known ontology in the field of chemistry, which provides a comprehensive resource for defining chemical entities and their properties. However, it covers only a small fraction of the rapidly growing knowledge in chemistry and does not provide references to the scientific literature. To address this, we propose a methodology that involves augmenting existing annotated text corpora with knowledge from Chebi and fine-tuning a large language model (LLM) to recognize chemical entities and their roles in scientific text. Our experiments demonstrate the effectiveness of our approach. By combining ontological knowledge and the language understanding capabilities of LLMs, we achieve high precision and recall rates in identifying both the chemical entities and roles in scientific literature. Furthermore, we extract them from a set of 8,000 ChemRxiv articles, and apply a second LLM to create a knowledge graph (KG) of chemical entities and roles (CEAR), which provides complementary information to ChEBI, and can help to extend it.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Multi-Hop Reasoning in Large Language Models: A Chemistry-Centric Case Study

    cs.CL 2025-04 conditional novelty 6.0 of 10

    A new 971-question chemistry benchmark shows that even the best large language models, given full context, still fail on many multi-step reasoning questions.

  2. MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

    cs.CL 2025-08 reject novelty 4.0 of 10

    A submission whose abstract describes a large temporal medical knowledge graph built by LLM agents, but whose full text is an unrelated paper on histogram regression, leaving the announced claims unsupported.

Pith tools