Pith. sign in

REVIEW 9 cited by

LLM Generated Persona is a Promise with a Catch

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16527 v1 pith:CSMOOQQR submitted 2025-03-18 cs.CL cs.AIcs.CYcs.SI

classification cs.CLcs.AIcs.CYcs.SI
keywords personagenerationpersonassignificantanalysisbiasesgeneratedincluding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Our findings underscore the need to develop a rigorous science of persona generation and outline the methodological innovations, organizational and institutional support, and empirical foundations required to enhance the reliability and scalability of LLM-driven persona simulations. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis at https://huggingface.co/datasets/Tianyi-Lab/Personas.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions

    cs.CY 2025-05 conditional novelty 7.0 of 10

    A new public 2,058-person, 500-question, four-wave dataset with a test-retest benchmark, plus initial LLM digital twin evaluations at 71.7% individual-level accuracy.

  2. Agentic Economic Modeling

    econ.EM 2025-10 conditional novelty 6.0 of 10

    Bias-corrected LLM choices, calibrated on 10% of human data, reproduce a national field experiment's treatment effect (-65 bps vs -60 bps) and reduce conjoint demand-estimation error.

  3. Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations

    cs.HC 2025-10 conditional novelty 6.0 of 10

    Four major LLMs generate occupational personas whose race and gender distributions deviate from U.S. workforce data in shared, patterned ways: White and Black workers are underrepresented while Hispanic and Asian work...

  4. Aligning LLM with human travel choices: a persona-based embedding learning approach

    cs.AI 2025-05 conditional novelty 6.0 of 10

    A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.

  5. LLM-Based Community Surveys for Operational Decision Making in Interconnected Utility Infrastructures

    cs.SI 2025-07 conditional novelty 5.0 of 10

    Simulated LLM personas can rank disaster repair priorities, and partial preference data recovers most of the full ranking.

  6. Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity

    cs.HC 2025-07 conditional novelty 5.0 of 10

    When LLMs are given more context about a real social media user, they become more ideologically consistent but also more extreme, toxic, and stereotyped than the user actually is.

  7. Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1

    cs.CL 2026-07 reject novelty 4.0 of 10

    GPT-4.1 persona simulation predicted 8/9 state winners and up to 0.94 vaccine-opinion accuracy, but the vaccine figure is weakened by post-hoc feature selection and the state results show large distribution errors.

  8. Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Anamnesis packages backstory-conditioned LLM personas into an interactive open-source survey platform that better matches real human opinion distributions than demographic-list prompting on ATP and New Yorker tasks.

  9. Leveraging LLMs for Persona-Based Visualization of Election Data

    cs.HC 2025-07 reject novelty 3.0 of 10

    LLM-generated voter personas are used to derive design criteria and prototypes for UK election visualizations, which are then evaluated by another LLM.

Pith tools