REVIEW 10 cited by
Whose Opinions Do Language Models Reflect?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models (LMs) are increasingly being used in open-ended contexts, where the opinions reflected by LMs in response to subjective queries can have a profound impact, both on user satisfaction, as well as shaping the views of society at large. In this work, we put forth a quantitative framework to investigate the opinions reflected by LMs -- by leveraging high-quality public opinion polls and their associated human responses. Using this framework, we create OpinionsQA, a new dataset for evaluating the alignment of LM opinions with those of 60 US demographic groups over topics ranging from abortion to automation. Across topics, we find substantial misalignment between the views reflected by current LMs and those of US demographic groups: on par with the Democrat-Republican divide on climate change. Notably, this misalignment persists even after explicitly steering the LMs towards particular demographic groups. Our analysis not only confirms prior observations about the left-leaning tendencies of some human feedback-tuned LMs, but also surfaces groups whose opinions are poorly reflected by current LMs (e.g., 65+ and widowed individuals). Our code and data are available at https://github.com/tatsu-lab/opinions_qa.
Forward citations
Cited by 10 Pith papers
-
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.
-
OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment
Inverted-pair probability elicitation shows 14 of 16 LLMs are optimistic; matched base-vs-chat pairs show family-specific post-training sets the sign of the bias.
-
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.
-
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.
-
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis
Engineered personas plus IS/OOS evidence retrieval make cheap multi-model panels produce tested claim maps and expose RLHF-induced blind spots, including asymmetric AI-risk challenge.
-
Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility
Conditioning LLMs on survey-derived demographic belief profiles improves their ability to predict individuals' misinformation judgments, especially when belief modeling is decoupled from susceptibility prediction.
-
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
A new VAAR metric finds that only a subset of LLMs show the human pattern where privacy concerns lower data-sharing acceptance and prosocial attitudes raise it.
-
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
In multi-agent debates over everyday moral dilemmas, GPT-4.1 almost never revises in simultaneous settings but conforms strongly in sequential settings, while Claude 3.7 and Gemini 2.0 Flash revise far more often.
-
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.
-
Analysis of Threat-Based Manipulation in Large Language Models: A Dual Perspective on Vulnerabilities and Performance Enhancement Opportunities
Threat-based prompts change LLM output length, style, and certainty, but the claimed performance gains are not backed by accuracy measures.
Discussion (0). Sign in to comment.