REVIEW 9 cited by
LLM Generated Persona is a Promise with a Catch
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The use of large language models (LLMs) to simulate human behavior has gained significant attention, particularly through personas that approximate individual characteristics. Persona-based simulations hold promise for transforming disciplines that rely on population-level feedback, including social science, economic analysis, marketing research, and business operations. Traditional methods to collect realistic persona data face significant challenges. They are prohibitively expensive and logistically challenging due to privacy constraints, and often fail to capture multi-dimensional attributes, particularly subjective qualities. Consequently, synthetic persona generation with LLMs offers a scalable, cost-effective alternative. However, current approaches rely on ad hoc and heuristic generation techniques that do not guarantee methodological rigor or simulation precision, resulting in systematic biases in downstream tasks. Through extensive large-scale experiments including presidential election forecasts and general opinion surveys of the U.S. population, we reveal that these biases can lead to significant deviations from real-world outcomes. Our findings underscore the need to develop a rigorous science of persona generation and outline the methodological innovations, organizational and institutional support, and empirical foundations required to enhance the reliability and scalability of LLM-driven persona simulations. To support further research and development in this area, we have open-sourced approximately one million generated personas, available for public access and analysis at https://huggingface.co/datasets/Tianyi-Lab/Personas.
Forward citations
Cited by 9 Pith papers
-
Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions
A new public 2,058-person, 500-question, four-wave dataset with a test-retest benchmark, plus initial LLM digital twin evaluations at 71.7% individual-level accuracy.
-
Agentic Economic Modeling
Bias-corrected LLM choices, calibrated on 10% of human data, reproduce a national field experiment's treatment effect (-65 bps vs -60 bps) and reduce conjoint demand-estimation error.
-
Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 Occupations
Four major LLMs generate occupational personas whose race and gender distributions deviate from U.S. workforce data in shared, patterned ways: White and Black workers are underrepresented while Hispanic and Asian work...
-
Aligning LLM with human travel choices: a persona-based embedding learning approach
A persona-based embedding learning framework aligns LLM predictions with human travel mode choices, outperforming MNL and few-shot LLM baselines on the Swissmetro dataset.
-
LLM-Based Community Surveys for Operational Decision Making in Interconnected Utility Infrastructures
Simulated LLM personas can rank disaster repair priorities, and partial preference data recovers most of the full ranking.
-
Generative Exaggeration in LLM Social Agents: Consistency, Bias, and Toxicity
When LLMs are given more context about a real social media user, they become more ideologically consistent but also more extreme, toxic, and stereotyped than the user actually is.
-
Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1
GPT-4.1 persona simulation predicted 8/9 state winners and up to 0.94 vaccine-opinion accuracy, but the vaccine figure is weakened by post-hoc feature selection and the state results show large distribution errors.
-
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
Anamnesis packages backstory-conditioned LLM personas into an interactive open-source survey platform that better matches real human opinion distributions than demographic-list prompting on ATP and New Yorker tasks.
-
Leveraging LLMs for Persona-Based Visualization of Election Data
LLM-generated voter personas are used to derive design criteria and prototypes for UK election visualizations, which are then evaluated by another LLM.
Discussion (0). Continue with ORCID to comment.