Pith. sign in

REVIEW 5 cited by

Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge Bases

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.09231 v1 pith:LJV4I2B7 submitted 2021-06-17 cs.CL cs.AI

classification cs.CLcs.AI
keywords mlmsknowledgepreviousbasesextractionfactuallanguagemainly
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Previous literatures show that pre-trained masked language models (MLMs) such as BERT can achieve competitive factual knowledge extraction performance on some datasets, indicating that MLMs can potentially be a reliable knowledge source. In this paper, we conduct a rigorous study to explore the underlying predicting mechanisms of MLMs over different extraction paradigms. By investigating the behaviors of MLMs, we find that previous decent performance mainly owes to the biased prompts which overfit dataset artifacts. Furthermore, incorporating illustrative cases and external contexts improve knowledge prediction mainly due to entity type guidance and golden answer leakage. Our findings shed light on the underlying predicting mechanisms of MLMs, and strongly question the previous conclusion that current MLMs can potentially serve as reliable factual knowledge bases.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MALAMUTE: A Multilingual, Highly-granular, Template-free, Education-based Probing Dataset

    cs.CL 2024-12 conditional novelty 7.0 of 10

    MALAMUTE is a 116k-prompt cloze-style dataset derived from 71 university textbooks in three languages, used to probe language models' fine-grained subject knowledge.

  2. How Do Programming Students Use Generative AI?

    cs.HC 2025-01 conditional novelty 6.0 of 10

    A controlled study of 37 programming students finds that most who used a chatbot requested full code solutions and repeatedly relayed error messages back to it, evidence for over-reliance on generative AI in learning.

  3. ComprehendEdit: A Comprehensive Dataset and Evaluation Framework for Multimodal Knowledge Editing

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A new eight-task benchmark with in-domain metrics KGI and KPI reveals that existing multimodal editing methods degrade on related samples, and the proposed HICE method achieves a better balance.

  4. Never Come Up Empty: Adaptive HyDE Retrieval for Improving LLM Developer Support

    cs.SE 2025-07 conditional novelty 4.0 of 10

    A HyDE retrieval pipeline with full-answer context and adaptive similarity thresholding improves LLM answers to Stack Overflow questions over zero-shot prompting for three of four open-source models.

  5. WisdomBot: Tuning Large Language Models with Artificial Intelligence Knowledge

    cs.CL 2025-01 conditional novelty 4.0 of 10

    Fine-tuning Chinese LLMs on Bloom's Taxonomy-guided AI textbook instruction data plus retrieval augmentation improves performance on the authors' tests and on C-Eval, though evidence is weakened by circular GPT-4 eval...

Pith tools