Pith. sign in

REVIEW 1 cited by

Can Language Models be Biomedical Knowledge Bases?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.07154 v1 pith:7XJ3RS5J submitted 2021-09-15 cs.CL

classification cs.CL
keywords biomedicalknowledgeprobingbeenlanguagetherebasesbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained language models (LMs) have become ubiquitous in solving various natural language processing (NLP) tasks. There has been increasing interest in what knowledge these LMs contain and how we can extract that knowledge, treating LMs as knowledge bases (KBs). While there has been much work on probing LMs in the general domain, there has been little attention to whether these powerful LMs can be used as domain-specific KBs. To this end, we create the BioLAMA benchmark, which is comprised of 49K biomedical factual knowledge triples for probing biomedical LMs. We find that biomedical LMs with recently proposed probing methods can achieve up to 18.51% Acc@5 on retrieving biomedical knowledge. Although this seems promising given the task difficulty, our detailed analyses reveal that most predictions are highly correlated with prompt templates without any subjects, hence producing similar results on each relation and hindering their capabilities to be used as domain-specific KBs. We hope that BioLAMA can serve as a challenging benchmark for biomedical factual probing.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BioHopR: A Benchmark for Multi-Hop, Multi-Answer Reasoning in Biomedical Domain

    cs.CL 2025-05 conditional novelty 6.0 of 10

    BioHopR introduces 1-hop and 2-hop question-answer benchmarks over PrimeKG with multiple correct answers, and shows LLMs achieve low precision, dropping sharply from 1-hop (best 37.93%) to 2-hop (14.57%).

Pith tools