Pith. sign in

REVIEW 6 cited by

Elements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09605 v2 pith:6ZQJ2IEL submitted 2024-05-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords worldknowledgemodelslanguagemodelingdomainsevaluatingewok
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to build and reason about models of the world is essential for situated language understanding. But evaluating world modeling capabilities in modern AI systems -- especially those based on language models -- has proven challenging, in large part because of the difficulty of disentangling conceptual knowledge about the world from knowledge of surface co-occurrence statistics. This paper presents Elements of World Knowledge (EWoK), a framework for evaluating language models' understanding of the conceptual knowledge underlying world modeling. EWoK targets specific concepts from multiple knowledge domains known to be important for world modeling in humans, from social interactions (help, deceive) to spatial relations (left, right). Objects, agents, and locations in the items can be flexibly filled in, enabling easy generation of multiple controlled datasets. We then introduce EWoK-core-1.0, a dataset of 4,374 items covering 11 world knowledge domains. We evaluate 20 open-weights large language models (1.3B--70B parameters) and compare them with human performance. All tested models perform worse than humans, with results varying drastically across domains. Performance on social interactions and social properties was highest and performance on physical relations and spatial relations was lowest. Overall, this dataset highlights simple cases where even large models struggle and presents rich avenues for targeted research on LLM world modeling capabilities.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Kolmogorov--Arnold Networks for Small Language Models

    cs.LG 2026-07 conditional novelty 7.0 of 10

    In small language models, KAN feed-forward blocks are auditable and pruneable, but on standardized benchmarks and scale tests they show no consistent accuracy, quality, or latency advantage over MLP baselines.

  2. Influence-driven Curriculum Learning for Pre-training on Limited Data

    cs.CL 2025-08 unverdicted novelty 6.0 of 10

    Sorting pre-training examples by training-data influence instead of human-judged difficulty reportedly gives over 10 percentage point benchmark gains over random order in limited-data language model pre-training.

  3. Potemkin Understanding in Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLMs frequently pass definition questions yet fail to use the same concepts in classification, generation, and editing tasks, a gap the authors call potemkin understanding.

  4. Trick or Neat: Adversarial Ambiguity and Language Model Evaluation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Language models answer prompts about sentence ambiguity poorly, but linear probes on their hidden states classify ambiguous versus unambiguous sentences with high accuracy on the new AmbAdv dataset.

  5. Mitigating Geospatial Knowledge Hallucination in Large Language Models: Benchmarking and Dynamic Factuality Aligning

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A new benchmark called GEOHALUBENCH measures how often LLMs invent, omit, or confuse real-world places and relations, and a dynamic-beta KTO method reduces these errors on the benchmark.

  6. Masked Diffusion Language Models with Frequency-Informed Training

    cs.CL 2025-09 conditional novelty 4.0 of 10

    Masked diffusion language models trained on 100M words match a hybrid GPT-BERT baseline on BabyLM tests, with a rare-word-focused masking variant.

Pith tools