Pith. sign in

REVIEW 10 cited by

Whose Opinions Do Language Models Reflect?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.17548 v1 pith:S47FYXTW submitted 2023-03-30 cs.CL cs.AIcs.CYcs.LG

classification cs.CLcs.AIcs.CYcs.LG
keywords opinionsgroupsreflecteddemographiccurrentframeworkhumanlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Language models (LMs) are increasingly being used in open-ended contexts, where the opinions reflected by LMs in response to subjective queries can have a profound impact, both on user satisfaction, as well as shaping the views of society at large. In this work, we put forth a quantitative framework to investigate the opinions reflected by LMs -- by leveraging high-quality public opinion polls and their associated human responses. Using this framework, we create OpinionsQA, a new dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation. Across topics, we find substantial misalignment between the views reflected by current LMs and those of US demographic groups: on par with the Democrat-Republican divide on climate change. Notably, this misalignment persists even after explicitly steering the LMs towards particular demographic groups. Our analysis not only confirms prior observations about the left-leaning tendencies of some human feedback-tuned LMs, but also surfaces groups whose opinions are poorly reflected by current LMs (e.g., 65+ and widowed individuals). Our code and data are available at https://github.com/tatsu-lab/opinions_qa.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 97 citations worldwide. Full citation record

  1. Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex

    cs.LG 2026-07 conditional novelty 7.0 of 10

    Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.

  2. OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment

    cs.CL 2026-07 conditional novelty 6.5 of 10

    Inverted-pair probability elicitation shows 14 of 16 LLMs are optimistic; matched base-vs-chat pairs show family-specific post-training sets the sign of the bias.

  3. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  4. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  5. Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis

    cs.AI 2026-03 conditional novelty 6.0 of 10

    Engineered personas plus IS/OOS evidence retrieval make cheap multi-model panels produce tested claim maps and expose RLHF-induced blind spots, including asymmetric AI-risk challenge.

  6. Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility

    cs.CL 2026-03 conditional novelty 6.0 of 10

    Conditioning LLMs on survey-derived demographic belief profiles improves their ability to predict individuals' misinformation judgments, especially when belief modeling is decoupled from susceptibility prediction.

  7. Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A new VAAR metric finds that only a subset of LLMs show the human pattern where privacy concerns lower data-sharing acceptance and prosocial attitudes raise it.

  8. Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

    cs.AI 2025-10 conditional novelty 6.0 of 10

    In multi-agent debates over everyday moral dilemmas, GPT-4.1 almost never revises in simultaneous settings but conforms strongly in sequential settings, while Claude 3.7 and Gemini 2.0 Flash revise far more often.

  9. Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?

    cs.CL 2025-07 conditional novelty 6.0 of 10

    Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.

  10. Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities

    cs.CR 2025-07 reject novelty 5.0 of 10

    Threat-based prompts change LLM output length, style, and certainty, but the claimed performance gains are not backed by accuracy measures.

Pith tools