Pith. sign in

REVIEW 2 cited by

Challenging the Validity of Personality Tests for Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.05297 v2 pith:5Q4O7EJE submitted 2023-11-09 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords llmspersonalitytestshumanresultsacrossevaluatelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With large language models (LLMs) like GPT-4 appearing to behave increasingly human-like in text-based interactions, it has become popular to attempt to evaluate personality traits of LLMs using questionnaires originally developed for humans. While reusing measures is a resource-efficient way to evaluate LLMs, careful adaptations are usually required to ensure that assessment results are valid even across human subpopulations. In this work, we provide evidence that LLMs' responses to personality tests systematically deviate from human responses, implying that the results of these tests cannot be interpreted in the same way. Concretely, reverse-coded items ("I am introverted" vs. "I am extraverted") are often both answered affirmatively. Furthermore, variation across prompts designed to "steer" LLMs to simulate particular personality types does not follow the clear separation into five independent personality factors from human samples. In light of these results, we believe that it is important to investigate tests' validity for LLMs before drawing strong conclusions about potentially ill-defined concepts like LLMs' "personality".

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Two-Process Theory of Machine Self-Report

    cs.CL 2026-07 conditional novelty 8.0 of 10

    The single 'Pinocchio Axis' of LLM self-report splits into two independent, training-dependent dimensions—persona installation (B) and attribution gating (A)—measurable with a reproducible 48-item inventory.

  2. From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology

    cs.CY 2025-06 conditional novelty 5.0 of 10

    This Perspective paper proposes that LLM research in psychology must combine psychometric validity and causal inference standards, mapping evidence requirements to the type of claim being made.

Pith tools