Pith. sign in

REVIEW 1 cited by

Estimating the Personality of White-Box Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.12000 v2 pith:RA7U47MB submitted 2022-04-25 cs.CL cs.AI

classification cs.CLcs.AI
keywords modelslanguagepersonalitytexttraitsbiasesdatasetsdesigned
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Technology for open-ended language generation, a key application of artificial intelligence, has advanced to a great extent in recent years. Large-scale language models, which are trained on large corpora of text, are being used in a wide range of applications everywhere, from virtual assistants to conversational bots. While these language models output fluent text, existing research shows that these models can and do capture human biases. Many of these biases, especially those that could potentially cause harm, are being well-investigated. On the other hand, studies that infer and change human personality traits inherited by these models have been scarce or non-existent. Our work seeks to address this gap by exploring the personality traits of several large-scale language models designed for open-ended text generation and the datasets used for training them. We build on the popular Big Five factors and develop robust methods that quantify the personality traits of these models and their underlying datasets. In particular, we trigger the models with a questionnaire designed for personality assessment and subsequently classify the text responses into quantifiable traits using a Zero-shot classifier. Our estimation scheme sheds light on an important anthropomorphic element found in such AI models and can help stakeholders decide how they should be applied as well as how society could perceive them. Additionally, we examined approaches to alter these personalities, adding to our understanding of how AI models can be adapted to specific contexts.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. How Personality Traits Shape LLM Risk-Taking Behaviour

    cs.CY 2025-02 conditional novelty 6.0 of 10

    Using direct certainty-equivalent questions, the authors find GPT-4o behaves close to risk-neutral and that Openness-related personality prompts shift its risk parameters in a human-like direction, while GPT-4-Turbo d...

Pith tools