Pith. sign in

REVIEW 2 cited by

HLB: Benchmarking LLMs' Humanlikeness in Language Use

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15890 v1 pith:6CY7NWUK submitted 2024-09-24 cs.CL

classification cs.CL
keywords languagehumanlikenessllmshumanmodelsbenchmarkdistributionsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As synthetic data becomes increasingly prevalent in training language models, particularly through generated dialogue, concerns have emerged that these models may deviate from authentic human language patterns, potentially losing the richness and creativity inherent in human communication. This highlights the critical need to assess the humanlikeness of language models in real-world language use. In this paper, we present a comprehensive humanlikeness benchmark (HLB) evaluating 20 large language models (LLMs) using 10 psycholinguistic experiments designed to probe core linguistic aspects, including sound, word, syntax, semantics, and discourse (see https://huggingface.co/spaces/XufengDuan/HumanLikeness). To anchor these comparisons, we collected responses from over 2,000 human participants and compared them to outputs from the LLMs in these experiments. For rigorous evaluation, we developed a coding algorithm that accurately identified language use patterns, enabling the extraction of response distributions for each task. By comparing the response distributions between human participants and LLMs, we quantified humanlikeness through distributional similarity. Our results reveal fine-grained differences in how well LLMs replicate human responses across various linguistic levels. Importantly, we found that improvements in other performance metrics did not necessarily lead to greater humanlikeness, and in some cases, even resulted in a decline. By introducing psycholinguistic methods to model evaluation, this benchmark offers the first framework for systematically assessing the humanlikeness of LLMs in language use.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On the use of foundation models in cognitive science

    cs.CL 2026-08 accept novelty 5.0 of 10

    Behavioral fit alone never justifies treating a foundation model as an explanatory cognitive model; explicit theoretical commitments, diagnostic tasks, and contrastive evaluation are required.

  2. The potential -- and the pitfalls -- of using pre-trained language models as cognitive science theories

    cs.CL 2025-01 accept novelty 4.0 of 10

    Pretrained language models can serve as credible cognitive science theories only if researchers validate linking hypotheses and avoid pitfalls of commission and omission.

Pith tools