Pith. sign in

Paper Citation Record · LEDGER

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

As of 15 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2412.00207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00207 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:41:03.002125Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:01:27.883371Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:09:45.977108Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 36787a03-ef3a-44ab-909e-4b1b1c06b2c6 · outbound

This paper cites Evaluating Large Language Models in Theory of Mind Tasks.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Evaluating Large Language Models in Theory of Mind Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.904636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.904636Z digest=sha256:416de7568532a18eb5383c843718eb0fd97f99114245a554f364403391df8ca5

Observation fe614e7a-2fcb-4573-911e-ad99908723ca · outbound

This paper cites UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots UPLex: Fine-Grained Personality Control in Large Language Models via Unsupervised Lexical Modulation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:41:03.320070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.931581Z digest=sha256:3f0316b54f7c1aac579aa45247c379431578d33caa1d85a394bb79fe7b6adcd7

Observation 6f410a0d-28a9-46fb-b33e-89c72df2270b · outbound

This paper cites Rethinking Model Evaluation as Narrowing the Socio-Technical Gap.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.936689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.936689Z digest=sha256:3eecd3474a6e62caa08dfe3d181c9245f28cdc3b14ee1a16ed28ebc5eb06673a

Observation 3756c180-8dff-4676-9723-cefb7ac9ec5a · outbound

This paper cites Personality-adapted multimodal dialogue system.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Personality-adapted multimodal dialogue system

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:41:03.282707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.942932Z digest=sha256:c2fff3f68459837d9d8192e055cfe4374ff7e0fe39972bd76d76d8cdc7e7973e

Observation f7139515-c948-4f04-846f-9a29eec06e4c · outbound

This paper cites Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Jeongeon Park, Bryan Min, Xiaojuan Ma, and Juho Kim

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.948359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.948359Z digest=sha256:73970d25f88345aca80733e118bb332e6106577d4f07b4d32ad806a8eea7698b

Observation 3d8c7d7c-a47c-453f-8bbb-6e0bde2641bd · outbound

This paper cites Personality Traits in Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Personality Traits in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.962924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.962924Z digest=sha256:8c8ef05004a5de8bafc05501fc833c6c42407d6c58ee1a4d4946238583a67918

Observation b80a8dae-7c40-47d5-a487-93b688367bd2 · outbound

This paper cites The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots The next big five inventory (bfi-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.505038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.967851Z digest=sha256:2526b29f7067f9663843adbc3593d506e0efbb2e4068267e1d82cfbcd3cb86af

Observation 123dfebd-e3ba-417f-80ed-5aba894cfef8 · outbound

This paper cites Will the Real Linda Please Stand up...to Large Language Models? Examining the Representativeness Heuristic in LLMs.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Will the Real Linda Please Stand up...to Large Language Models? Examining the Representativeness Heuristic in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.977982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.977982Z digest=sha256:5a57b2e4c191a7a99f14612cf09b71d8f6982a0d4352a7337acb5b4f23ef72f5

Observation 068dec37-78a4-4941-911a-32bd60fbea95 · outbound

This paper cites doi: 10.18653/v1/2023.emnlp-main.676.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots doi: 10.18653/v1/2023.emnlp-main.676

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.987172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.987172Z digest=sha256:e713260b55cca060863330e329360f05c479d9b53cfee2db44bc762d7a31c2ca

Observation 97c02d10-830d-4325-9bfe-c0ac901ff5d8 · outbound

This paper cites {personality description}.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots {personality description}

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T05:41:03.479806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.991910Z digest=sha256:b5d73e0894f2f34a5a5ec49079b659904991c1f8388ad6c4550e0e52520cb67a

Observation 0856b444-a09e-4266-a009-1c7e3bccac63 · outbound

This paper cites disagree strongly.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots disagree strongly

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.464981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.996740Z digest=sha256:a4c215e236fe14a42dc176511aeab449ca73426213e53ca51698ed73c35440fb

Observation 824a227b-6615-4f42-a0ea-f21fcd49f22f · outbound

This paper cites Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Table J details the specific instructions used, where transcript refers to the human- chatbot conversational scripts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.450407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:03.002125Z digest=sha256:ea9b399026ba4c33e471ebb2270ecc88303d4afb3215c6ed4fc63ca6ca6237f1

Observation ef3c1e2b-5243-4376-95c2-8e4db1206ddf · outbound

This paper cites Talebrush: Sketching stories with generative pretrained language models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Talebrush: Sketching stories with generative pretrained language models

Reference 1959

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.586869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.871741Z digest=sha256:32e6343d0214b5b4fb9959d6a7235c33894e2881e6c923c8f92c1c67afcee74b

Observation 726dfc18-df73-44fb-8ec0-6a9437665e9d · outbound

This paper cites Vera Liao.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Vera Liao

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.982597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.982597Z digest=sha256:30a9f00520eb43e4a89ad9575790ab6f744324a70e953bd7105701a79380eefd

Observation c6c5a87a-51e0-48b1-9a76-8663a61615e0 · outbound

This paper cites Lewis R Goldberg et al.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Lewis R Goldberg et al

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.877406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.877406Z digest=sha256:45b43e66e4cb153d58a7842735fe0e77c47ce30732e7fbc726393ce62c689168

Observation 556b9a44-2a19-40b0-bb2f-b91b5d54c2c5 · outbound

This paper cites Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Who will go the extra mile? selecting organizational citizens with a personality-based structured job interview

Reference 1999

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.570617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.882534Z digest=sha256:803aaa6fea009d032a4a28e9f4b341763458697174badfa07e2de74bf87a26bc

Observation cfb047bb-545f-4c18-8857-196d6351039b · outbound

This paper cites Behavioral change and consistency across contexts.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Behavioral change and consistency across contexts

Reference 2001

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.532629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.958572Z digest=sha256:32b11135cc88c1c85b87a07735edcdf95a7585b1b90d79d080b2b0649047dde8

Observation 3d05eea0-ff1f-47f1-acc5-a3a534dadc71 · outbound

This paper cites Revisiting the Reliability of Psychological Scales on Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Revisiting the Reliability of Psychological Scales on Large Language Models

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.893879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.893879Z digest=sha256:91cc3fa2e06ad740fdbf59d56c79584fda10ccc0575667c450c5d60b2acbe3cc

Observation 333533d2-29a9-4bb5-8291-ef8b93cd3023 · outbound

This paper cites Evaluating Human-Language Model Interaction.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Evaluating Human-Language Model Interaction

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.920337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.920337Z digest=sha256:f7be6c80e0241f9e357208dc93ed91f6fc0c07e3c72a03445e92428c3f1d9a93

Observation 1dad5904-dd25-4010-a710-ed5c86e2af45 · outbound

This paper cites Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Capturing Minds, Not Just Words: Enhancing Role-Playing Language Models with Personality-Indicative Data

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.953254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.953254Z digest=sha256:8081b8defeb12906c6791881d734637f3b65c8369679d7e9b522e7fa005443f5

Observation c8e103a3-09d0-442e-856a-27b66376ee4a · outbound

This paper cites CharacterChat: Learning towards Conversational AI with Personalized Social Support.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots CharacterChat: Learning towards Conversational AI with Personalized Social Support

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.972979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.972979Z digest=sha256:4804b1d0cf64b146104b610367223b53f791cfafaadf541772756c5b5a77e94d

Observation 9795cd0e-bb90-4f2b-b1b9-124eb638a0da · outbound

This paper cites AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots AI-TA: Towards an Intelligent Question-Answer Teaching Assistant using Open-Source LLMs

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.887989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.887989Z digest=sha256:c0bc035fb949af817db838c75096123265569caf98dba020f05544c9f31e2223

Observation 705000ef-f3fb-4196-8a70-f7aa53ac8687 · outbound

This paper cites Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.925722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.925722Z digest=sha256:476dc7fe47164b91ddac7cab03cb0da61803a8ea715574aabfe40d3daa62a878

Observation 8b744194-7571-4dfe-b54d-3ffc60e6f067 · outbound

This paper cites Psy-LLM: Scaling up Global Mental Health Psychological Services with AI-based Large Language Models.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Psy-LLM: Scaling up Global Mental Health Psychological Services with AI-based Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.910069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.910069Z digest=sha256:d3f4bd92e5e0347c8ede161f32a4a65e3031c78d66cc40f076782e6e51072c21

Observation fb89bc62-8fc3-4c31-9bb0-6e213ab858e1 · outbound

This paper cites Construction and evaluation of a user experience questionnaire.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots Construction and evaluation of a user experience questionnaire

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:41:03.550364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:41:02.915354Z digest=sha256:b6a0c80cc23d1e0b11af96c1b862ac3cbffe741f63d103a1ebdb4749624793c7

Observation d3aa69a5-462c-448c-9694-4529bc229783 · outbound

This paper cites PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits.

Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots PersonaLLM: Investigating the Ability of Large Language Models to Express Personality Traits

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T05:41:02.899394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:41:02.899394Z digest=sha256:3b5b44d0eaee995b1d070d97325207080080bf670f27a42acedee0c54e470b3a

Pith citing papers

Observation 06b9d794-77d8-46e7-a462-11c3cedf38c1 · inbound

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs cites this paper.

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:26:29.599671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T10:01:27.883371Z digest=sha256:be4af642bd867147bda41e5af18a8d1711c4115cb2dd261a992b4d680c9f3562

Observation e9a963d0-dc86-49a6-bf57-422a0464f753 · inbound

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior cites this paper.

Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:18:03.736197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T09:36:36.821058Z digest=sha256:c36448b8fe8a8b07b0214418e018f8da2fb82a100e4f8a1d2e4a2ba7b5c5d86e

Observation bc28a1b3-6bd4-4e66-8b59-1a1c5909be0d · inbound

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure cites this paper.

When Robots Rate Their Own Interactions: Engagement Validity and the Strangeness Failure Can LLM "Self-report"?: Evaluating the Validity of Self-report Scales in Measuring Personality Design in LLM-based Chatbots

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:45.978691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T08:10:46.067374Z digest=sha256:f2079d5bdc9d16c967b422d39c023ae846131c0a865e4d54042cfff43e83afc9