REVIEW 5 cited by
The Ghost in the Machine has an American accent: value conflict in GPT-3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The alignment problem in the context of large language models must consider the plurality of human values in our world. Whilst there are many resonant and overlapping values amongst the world's cultures, there are also many conflicting, yet equally valid, values. It is important to observe which cultural values a model exhibits, particularly when there is a value conflict between input prompts and generated outputs. We discuss how the co-creation of language and cultural value impacts large language models (LLMs). We explore the constitution of the training data for GPT-3 and compare that to the world's language and internet access demographics, as well as to reported statistical profiles of dominant values in some Nation-states. We stress tested GPT-3 with a range of value-rich texts representing several languages and nations; including some with values orthogonal to dominant US public opinion as reported by the World Values Survey. We observed when values embedded in the input text were mutated in the generated outputs and noted when these conflicting values were more aligned with reported dominant US values. Our discussion of these results uses a moral value pluralism (MVP) lens to better understand these value mutations. Finally, we provide recommendations for how our work may contribute to other current work in the field.
Forward citations
Cited by 5 Pith papers
-
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?
Value representations from token logits, sequence perplexity, and text generation are all sensitive to prompt and option changes, and their correlation with model behavior in value scenarios is weak.
-
Is It Bad to Work All the Time? Cross-Cultural Evaluation of Social Norm Biases in GPT-4
GPT-4 writes accurate but generic social norms for non-US cultures, defaults to US-style judgments, and still holds recoverable stereotypes about China, India, and Iran.
-
Enhancing Deliberativeness: Evaluating the Impact of Multimodal Reflection Nudges
Video-based reflection nudges improved several deliberation-quality measures over text, image, and audio formats, despite text being the subjectively preferred modality for persona-style prompts.
-
Prompt Programming for Cultural Bias and Alignment of Large Language Models
Automatically optimized prompts (DSPy) reduce survey-measured cultural distance for open-weight LLMs more often than manual cultural prompting, with MIPROv2 and a large proposer model giving the most consistent gains.
-
Do Large Language Models Understand Morality Across Cultures?
Small language models compress cross-cultural moral differences, producing more uniformly permissive and less varied judgments than international survey data.
Discussion (0). Continue with ORCID to comment.