Pith. sign in

REVIEW 3 major objections 5 minor 5 references

From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Psychological claims about large language models fail without dual validation—psychometric and causal—this paper argues.

desk verdict A useful dual-validity synthesis for LLM psychology, but the central 'measurement phantom' claim is broader than the paper's own adapted-instrument evidence supports. read the letter →

arxiv 2506.16697 v1 pith:EVWIAKTB submitted 2025-06-20 cs.CY cs.AIcs.CLcs.HC

classification cs.CYcs.AIcs.CLcs.HC
keywords largelanguagemodelspsychometricsconstructvaliditycausalinferencereliabilitymeasurementphantomsAIpsychologypsychological
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This Perspective argues that using human psychological instruments on large language models can produce 'measurement phantoms'—statistical artifacts mistaken for genuine psychological phenomena—because current practice skips the two validation traditions psychology built for human research: psychometric validity (does the instrument measure the construct?) and causal inference validity (does the design support the conclusion?). The paper tries to establish that evidence requirements must scale with scientific ambition: classifying text asks for accuracy checks, while claiming a model 'is anxious' or 'has theory of mind' asks for much more. It proposes a dual-validity framework that maps four uses of LLMs in psychology (research tool, behavioral characterization, human simulator, cognitive model) onto the validity evidence each requires. If right, much of the current literature on machine personality, moral reasoning, and theory of mind is built on unreliable measures and would need revalidation before its claims can stand.

What carries the argument

The framework's load-bearing object is the pairing of the psychometric and causal-inference validity traditions into a single pipeline. From psychometrics it takes the reliability ceiling—no measure can be more valid than it is reliable—and construct validity as an accumulating evidence argument, sharpened by the causal theory of validity, which requires both that the attribute exists and that variations in it causally produce observed scores. From experimental methodology it takes the four parallel threats to causal inference (internal, external, construct, and statistical conclusion validity). The framework maps the four uses of LLMs in psychology onto these standards and classifies failure modes: training artifact contamination, prompt hypersensitivity, and stochastic degradation violate psychometric assumptions, while temperature confounds, version drift, dynamic scenario reconstruction, and non-independence violate causal assumptions. The unifying move is the observation that in LLM research the same output serves simultaneously as a measurement indicator and as experimental data, so a failure on either side corrupts the other.

What would settle it

Concretely: administer a personality or moral-decision inventory to a single model across many semantically equivalent prompt variants (changed option labels, order, punctuation, phrasing) at fixed temperature and version; if scores stay stable, factorially coherent, and predictive of external outcomes across all variants—rather than shifting more than 70% as the paper reports—the reliability crisis is refuted for that model. A second refutation would be a model whose 'anxiety' responses are lawfully modulated by threat manipulations, remain stable across sessions, and predict downstream outputs within a nomological network, which would challenge the claim that no such attribute exists.

Watch

Extended reading notes

Core claim

The paper's central claim is that LLM responses can look psychologically meaningful while being generated by statistical pattern matching, and that the field currently has no way to tell the difference because it applies human measurement tools without their validation scaffolding. The same output—a model endorsing 'I am anxious'—requires entirely different validation strategies depending on whether the claim is to measure anxiety, characterize model behavior, simulate human responses, or model a cognitive mechanism. A measurement claim demands that the attribute exist in the model and causally produce the response; the paper argues that psychological constructs such as anxiety presuppose temporal experience, a persistent self, and embodied consequences that LLMs lack, so without such evidence 'any resulting output is a pattern of words masquerading as a psychological phenomenon.' The paper proposes that reliability and validity be established before causal experimentation, that construct validation draw on five sources of evidence (content, response processes, internal structure, relations with other variables, consequences), and that causal claims address four parallel validity types (internal, external, construct, statistical conclusion). Its constructive proposal is to study computational analogues—'anxiety-analogous patterns' rather than anxiety—so AI psychology proceeds on mechanistic, not biological, terms.

Load-bearing premise

The paper's strongest conclusion rests on the assumption that psychological constructs such as anxiety presuppose embodiment, temporal continuity, and a persistent self, so that a language model lacking these cannot possess the attribute being measured; if a functionalist account of mental states—where internal states are defined by causal roles rather than physical substrate—is correct, that premise fails.

Editorial extensions

If this is right

  • Publications claiming to measure personality, theory of mind, or moral reasoning in LLMs would need to document reliability (test–retest, parallel forms, internal consistency) and evidence from all five construct-validity sources before their claims could be credited.
  • Studies that manipulate prompts to test causal hypotheses would need to rule out computational confounds—temperature, prompt formatting, model version, non-independence—through factorial designs, ablations, or unblinding procedures.
  • Simulation claims (LLM responses stand in for human responses) would be restricted to demonstrated behavioral correspondence and would not license claims about human-like underlying mechanisms.
  • Many reported LLM psychological effects are predicted to be unstable: they should shift or vanish under trivial prompt variations, and re-analysis with factorial designs should shrink or eliminate them.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A natural next step, left implicit in the paper, is a public audit instrument that scores LLM psychological studies on reliability and validity evidence; the paper calls for infrastructure but does not specify one.
  • The ontological boundary is testable: if future models with persistent memory and embodiment-like training show stable, lawfully connected anxiety-analogous responses, the 'attribute does not exist' premise blurs and the framework would need to decide when an analogue becomes a construct.
  • The framework could also be applied retroactively as a taxonomy to meta-analyze existing LLM findings, sorting which reported effects survive prompt perturbation—an extension of the reliability discussion the paper leaves implicit.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The manuscript is a Perspective proposing a dual-validity framework for the use of large language models in psychological research. It argues that reliable measurement and sound causal inference must be integrated when LLM outputs are used to measure, characterize, simulate, or model psychological constructs, and that the required evidence should scale with the scientific ambition of the claim. The paper reviews three reliability threats (training artifact contamination, prompt hypersensitivity, and stochastic degradation), five sources of psychometric construct validity evidence, and four types of causal-inference validity, illustrating each with recent empirical work. It concludes that current practice systematically fails these requirements and that much current LLM psychological research produces 'measurement phantoms' or 'patterns of words masquerading as a psychological phenomenon.'

Significance. The manuscript is a timely and broadly useful synthesis. Its strengths are the extensive and current literature integration; the concrete catalog of reliability threats; the careful separation of four application categories (research tool, characterization, human simulation, cognitive model); and the empirically grounded acknowledgment that some LLM-adapted instruments, such as Ye et al.'s generative psychometrics, Lee et al.'s scenario-based TRAIT, and Ma et al.'s implicit sentiment measure, can achieve structural or predictive validity. The paper does not present new data, code, or machine-checked proofs, but its framework offers testable expectations about which measurement approaches are likely to succeed and which are likely to fail. If the scope of the central claim is appropriately calibrated, the paper could provide a useful methodological reference for researchers and reviewers.

major comments (3)
  1. [Internal Structure; Relations with Other Variables; Conclusions and Future Directions] The manuscript's categorical conclusion that LLM psychological research 'systematically fails' and that when the measured attribute does not exist 'any resulting output is a pattern of words masquerading as a psychological phenomenon' is undercut by the manuscript's own evidence. In the Internal Structure section, Ye et al.'s generative psychometrics reproduced the Schwartz value circumplex and Lee et al.'s scenario-based TRAIT produced theoretically coherent inter-trait correlations; in Relations with Other Variables, Ma et al.'s implicit Core Sentiment Inventory achieved predictive correlations above 0.85. The paper even concedes that 'Structural validity failures may thus indicate methodological mismatch rather than construct absence.' Because these successes involve adapted or LLM-specific instruments, the categorical phantom diagnosis is overbroad. The conclusion should be explicitly restricted to unadapted human measures, or the paper should explain why these successes do not count as evidence relevant to psychological constructs.
  2. [Construct Validity from the Psychometric Foundation] The claim that 'anxiety presupposes temporal experience, a persistent self, and embodied consequences—ontological properties the model lacks' is asserted rather than defended. This ontological premise is load-bearing: it is the basis for saying the attribute does not exist and hence that LLM outputs are 'patterns of words masquerading as a psychological phenomenon.' A functionalist account of psychological attributes, in which internal states are defined by causal roles rather than by embodiment, would block this inference. The manuscript should either provide an argument against such functionalism or present the conclusion conditionally, for example by stating that under a constitution-based, non-functionalist account the attribute does not exist in LLMs.
  3. [Why LLM Research Requires Both; Table 1] The paper's central organizing claim is that 'validity requirements scale with psychological ambition,' but the framework does not specify how this scaling works. The text assigns requirements to four application categories—research tools, characterization, human simulation, cognitive modeling—but the basis for these assignments is not given; for example, why human simulation requires only behavioral correspondence while characterization requires construct validation is stated but not argued. Table 1 lists validity types and threats but provides no decision rule connecting a claim's ambition to the evidence required. Without such a rule, the framework is a useful checklist rather than a framework for determining validation demands. The authors should either provide explicit criteria or clearly frame the contribution as a checklist.
minor comments (5)
  1. [Abstract] The abstract promises that the same model output requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model a construct, but no section provides a worked illustration of this point; adding a brief example would make the framework concrete.
  2. [References] The reference 'Stanley, J. C., & Campbell, D. T. (1963)' is conventionally cited as 'Campbell, D. T., & Stanley, J. C.'; please correct or justify the ordering.
  3. [Table 1] Table 1 labels the psychometric and causal-inference 'Construct' rows identically; renaming them, for example 'Construct validity (psychometric)' and 'Construct validity (causal inference),' would prevent confusion between construct validity in measurement and construct validity of causal claims.
  4. [References] The arXiv identifier for Guan et al. (2025) duplicates the identifier for Sclar et al. (2023), 2310.11324; the correct identifier for Guan et al. appears to be 2502.04134.
  5. [Formatting] The headers 'LLM VALIDITY 1' and similar appear to be page headers from the submission and should be removed before publication.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is an evidence-synthesizing Perspective whose central claims rest on external empirical studies, not on self-definition or fitted inputs.

full rationale

The paper does not derive predictions from its own definitions or fit parameters and then relabel them as findings. Its central conclusion—that current LLM psychological research often mistakes statistical pattern matching for psychological phenomena—is supported by an extensive body of external empirical work (e.g., Oh & Demberg 2025; Gao et al. 2024; Peereboom et al. 2025; Sühr et al. 2023), not by an equation set equal to its inputs. The dual-validity framework is explicitly a synthesis of established psychometric validity (Cronbach & Meehl 1955; Borsboom et al. 2004) and causal-inference validity (Cook & Campbell 1979), presented as an organizing framework rather than as a novel first-principles derivation. The few self-citations (Lin 2025a, 2025b) are methodological guides cited for the classification of LLM uses (research tools, simulators, cognitive models) and are not load-bearing for the central validity claims; citing them does not make the argument circular. The paper also contains an internal concession that structural validity failures 'may thus indicate methodological mismatch rather than construct absence,' which narrows the scope of its phantom diagnosis rather than revealing a definitional circularity. The strongest challenges to the paper are substantive: the embodiment presupposition behind construct absence is contestable under functionalist accounts, and the paper's own adapted-instrument successes (Ye et al.; Lee et al.; Ma et al.) complicate the categorical version of the claim. But these are correctness and scope concerns, not circularity. No step in the paper's reasoning reduces by construction to its own assumptions or to a self-citation chain.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The framework introduces no new empirical entities or fitted parameters. Its load-bearing assumptions are domain-specific philosophical stances about psychological constructs and LLM ontology, plus the applicability of human psychometric standards to machines.

assumptions (4)
  • domain assumption Psychological constructs such as anxiety presuppose temporal experience, a persistent self, and embodied consequences.
    Invoked in the Construct Validity section to argue that LLMs cannot possess the attribute, making outputs 'a pattern of words masquerading as a psychological phenomenon'. This is a philosophical claim, not empirically established.
  • domain assumption The Borsboom causal theory of validity is the correct account: an instrument is valid only if the attribute exists and causally produces the observed scores.
    Adopted to sharpen the ontological challenge; not universally accepted, and it is load-bearing for the conclusion that LLM responses cannot be valid measures of human-like constructs.
  • domain assumption Human-oriented psychometric standards (reliability, validity) are appropriate starting points for evaluating LLM measurement, with adaptation.
    The whole framework presupposes that concepts like test-retest reliability and construct validity apply meaningfully to LLM responses. Some might argue that LLMs are not psychological entities, so the framework might not apply.
  • domain assumption The cited empirical studies accurately represent the state of LLM research.
    The paper's broad claim that current practice 'systematically fails' rests on a convenience sample of studies, not a systematic review or meta-analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology." pith.science (2026). https://pith.science/paper/EVWIAKTB

@misc{pith2026250616697,
  author       = {Pith},
  title        = {Pith review of: From Prompts to Constructs: A Dual-Validity Framework for LLM Research in Psychology},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EVWIAKTB}},
  note         = {Machine review of arXiv:2506.16697}
}
read the original abstract

Large language models (LLMs) are rapidly being adopted across psychology, serving as research tools, experimental subjects, human simulators, and computational models of cognition. However, the application of human measurement tools to these systems can produce contradictory results, raising concerns that many findings are measurement phantoms--statistical artifacts rather than genuine psychological phenomena. In this Perspective, we argue that building a robust science of AI psychology requires integrating two of our field's foundational pillars: the principles of reliable measurement and the standards for sound causal inference. We present a dual-validity framework to guide this integration, which clarifies how the evidence needed to support a claim scales with its scientific ambition. Using an LLM to classify text may require only basic accuracy checks, whereas claiming it can simulate anxiety demands a far more rigorous validation process. Current practice systematically fails to meet these requirements, often treating statistical pattern matching as evidence of psychological phenomena. The same model output--endorsing "I am anxious"--requires different validation strategies depending on whether researchers claim to measure, characterize, simulate, or model psychological constructs. Moving forward requires developing computational analogues of psychological constructs and establishing clear, scalable standards of evidence rather than the uncritical application of human measurement tools.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

5 extracted references · 4 linked inside Pith

  1. [1]

    J., Trager, J., Park, P

    Abdurahman, S., Atari, M., Karimi-Malekabadi, F., Xue, M. J., Trager, J., Park, P. S., . . . Dehghani, M. (2024). Perils and opportunities in using large language models in psychological research. PNAS Nexus, 3(7), pgae245. https://doi.org/10.1093/pnasnexus/pgae245 Abdurahman, S., Salkhordeh Ziabari, A., Moore, A. K., Bartels, D. M., & Dehghani, M. (2025)...

  2. [80]

    https://doi.org/10.1038/s44271-025-00258-x Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2023). Quantifying language models’ sensitivity to spurious features in prompt design or: How I learned to start worrying about prompt formatting. arXiv:2310.11324. https://arxiv.org/abs/2310.11324 LLM VALIDITY 23 Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askel...

  3. [519]

    self-report

    https://doi.org/10.1038/s41598-024-84109-5 Webb, T., Holyoak, K. J., & Lu, H. (2023). Emergent analogical reasoning in large language models. Nature Human Behaviour, 7(9), 1526-1541. https://doi.org/10.1038/s41562- 023-01659-w Xu, R., Sun, Y., Ren, M., Guo, S., Pan, R., Lin, H., . . . Han, X. (2024). AI for social science and social science of AI: A surve...

  4. [1095]

    https://doi.org/10.1057/s41599-024-03609-x Riemer, M., Ashktorab, Z., Bouneffouf, D., Das, P., Liu, M., Weisz, J., & Campbell, M. (2025). Position: Theory of mind benchmarks are broken for large language models. International Conference on Machine Learning, Vancouver, Canada. Salecha, A., Ireland, M. E., Subrahmanya, S., Sedoc, J., Ungar, L. H., & Eichsta...

  5. [2311]

    https://doi.org/10.48550/arXiv.2311.05297 Takemoto, K. (2024). The moral machine experiment on large language models. Royal Society Open Science, 11(2), 231393. https://doi.org/10.1098/rsos.231393 Taylor, J. E. T., & Taylor, G. W. (2021). Artificial cognition: How experimental psychology can help generate explainable artificial intelligence. Psychonomic B...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.