REVIEW 2 major objections 1 minor 1 cited by
Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models
T0 review · 2 major / 1 minor · reviewed 2026-06-27 · grok-4.3
Pith's one-line read LLMs match majority human choices on cultural-personal trade-offs but fail to capture response distributions and within-culture disagreement.
desk verdict PACT gives a concrete way to measure how LLMs handle norm-vs-preference conflicts and shows they miss human response spread, but the scenarios need checking for construction bias before the country and pluralism claims land cleanly. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
PACT, the Personal-Preference and Cultural-Norm Trade-off framework, which presents decision scenarios requiring a choice between a stated cultural norm and a conflicting personal preference.
What would settle it
New scenarios written independently by residents of each country, followed by the same human and model experiments, would produce higher correlations between model outputs and full human response distributions.
Extended reading notes
Core claim
LLMs vary in how rigidly they enforce cultural norms, with behavior shifted more by country context than by age or gender, and shifting non-uniformly after instruction tuning. Human responses on the same scenarios are driven mainly by the scenario country, with the lowest agreement when participants judge their own cultural contexts. Models can match majority human choices but reach correlations of only 0.24 with the actual distribution of human answers and uncertainty.
Load-bearing premise
The scenarios created for PACT validly represent the trade-off between cultural norms and personal preferences across countries without introducing bias in how the situations were written or selected.
Editorial extensions
If this is right
- Alignment evaluations must measure how well models reproduce the full distribution and uncertainty of human answers rather than only majority agreement.
- Country context influences model choices more than demographic attributes, so cultural alignment testing should prioritize scenario location over user demographics.
- Instruction tuning changes cultural rigidity in non-uniform ways across models, requiring separate checks after each tuning step.
- Human studies reveal within-culture pluralism, so models intended for social judgment should be tested on their ability to reflect disagreement inside a single culture.
Reading between the lines
- Models that only match majorities may systematically under-represent minority preferences that exist inside every culture.
- If new scenario sets created by local residents yield different results, the current findings may partly reflect the original authors' framing rather than universal patterns.
- The low correlations suggest that simply scaling models or adding more countries to training data is unlikely to solve the distribution-matching problem without targeted uncertainty modeling.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the PACT (Personal-Preference and Cultural-Norm Trade-off) framework to evaluate how LLMs balance cultural norms against personal preferences in social scenarios. It reports that country context shifts LLM behavior by 7.8% (versus 1% for age and 0.7% for gender), with non-uniform changes after instruction tuning; a five-country human study finds culture-following driven primarily by scenario country, with lowest agreement on own-culture scenarios (indicating within-culture pluralism); and human-LLM alignment experiments show models match majority choices but achieve at most 0.24 correlation with response distributions and uncertainty.
Significance. If the PACT instrument is shown to be valid, the work would be significant for LLM alignment research by providing quantitative evidence that cultural context dominates demographic factors, documenting within-culture disagreement, and demonstrating that current models fail to capture human response variance. The five-country human baseline and explicit comparison of majority versus distributional alignment are strengths that could guide more pluralistic evaluation methods.
major comments (2)
- [PACT framework] PACT framework (abstract and §3): The headline quantitative results (country shift of 7.8%, within-culture pluralism, 0.24 correlation ceiling) all rest on the assumption that the designed scenarios are neutral, representative probes of the norm-preference trade-off. No validation is described that rules out researcher bias in scenario wording, conflict selection, or country framing; without such checks the measured country effects and pluralism findings risk being confounded with scenario artifacts.
- [Human study and LLM experiments] Human study and LLM experiments (abstract and results sections): The manuscript presents effect sizes (7.8% vs. 1%/0.7%) and correlation values (0.24) without reporting the underlying statistical model, controls for participant demographics or scenario order, sample sizes per country, or exact definition of “agreement” and “response distribution.” These omissions make it impossible to evaluate whether the reported differences are robust or whether the low distributional correlation is an artifact of measurement.
minor comments (1)
- [Abstract] Abstract: Quantitative claims are stated without any accompanying description of methodology, data collection, or analysis, which reduces immediate readability.
Simulated Author's Rebuttal
We thank the referee for their constructive comments, which highlight important aspects of methodological transparency. We address each major comment below and indicate planned revisions to the manuscript.
read point-by-point responses
-
Referee: [PACT framework] PACT framework (abstract and §3): The headline quantitative results (country shift of 7.8%, within-culture pluralism, 0.24 correlation ceiling) all rest on the assumption that the designed scenarios are neutral, representative probes of the norm-preference trade-off. No validation is described that rules out researcher bias in scenario wording, conflict selection, or country framing; without such checks the measured country effects and pluralism findings risk being confounded with scenario artifacts.
Authors: We agree that explicit documentation of scenario construction is necessary to support the validity of the measured effects. The scenarios were iteratively refined drawing on cross-cultural psychology literature and informal pilot feedback from participants in multiple countries to ensure they presented plausible trade-offs, but the manuscript does not include a dedicated validation subsection or inter-rater checks for neutrality. We will add a new subsection in §3 describing the scenario development process, including criteria used to select conflicts and any steps taken to reduce framing bias. This addition will not change the reported results but will allow readers to better assess potential artifacts. revision: partial
-
Referee: [Human study and LLM experiments] Human study and LLM experiments (abstract and results sections): The manuscript presents effect sizes (7.8% vs. 1%/0.7%) and correlation values (0.24) without reporting the underlying statistical model, controls for participant demographics or scenario order, sample sizes per country, or exact definition of “agreement” and “response distribution.” These omissions make it impossible to evaluate whether the reported differences are robust or whether the low distributional correlation is an artifact of measurement.
Authors: We accept that the current manuscript lacks sufficient methodological detail for independent evaluation of the statistical claims. The effect sizes were obtained via mixed-effects logistic regression with country, age, and gender as predictors and random effects for participants and scenarios; agreement was defined as the proportion of responses matching the modal choice within each condition; response distributions were compared via Pearson correlation on choice proportions per scenario; and the human study used approximately 200 participants per country with counterbalanced scenario order. We will expand the Methods and Results sections to include the full model specification, exact sample sizes, demographic controls, and operational definitions. These additions will be accompanied by supplementary tables reporting the regression coefficients and correlation matrices. revision: yes
Circularity Check
No circularity: empirical measurements on independently designed scenarios
full rationale
The paper introduces PACT as a new evaluation framework and derives all quantitative claims (country-context shifts of 7.8%, age/gender effects, human agreement rates, LLM-human correlations) directly from experimental runs on the constructed scenarios and participant responses. No equations, fitted parameters, or uniqueness theorems are invoked; the results are observational statistics rather than reductions to prior self-citations or definitional equivalences. The derivation chain is therefore self-contained against external benchmarks and contains no load-bearing self-referential steps.
Assumptions & free parameters
assumptions (1)
- domain assumption The PACT scenarios accurately reflect real cultural norms and personal preferences in the selected countries.
invented entities (1)
-
PACT framework
Cite this review
Pith. "Pith review of Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models." pith.science (2026). https://pith.science/paper/TPP3DZ7G
@misc{pith2026260607877,
author = {Pith},
title = {Pith review of: Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/TPP3DZ7G}},
note = {Machine review of arXiv:2606.07877}
}
read the original abstract
Large language models are increasingly used for social decision-making situations that require balancing cultural norms with personal preferences. For example, a user preferring honesty might ask whether to correct a coworker publicly when local norms favor indirect feedback. Yet existing research studies cultural alignment and personalization largely separately. We introduce PACT, the Personal-Preference and Cultural-Norm Trade-off framework, which evaluates whether models choose to follow a cultural norm or allow personal preferences. We find that LLMs vary in how rigidly they enforce cultural norms, with behavior shifted more by country context (7.8%) than age (1%) and gender (0.7%) and shifting non-uniformly after instruction tuning. Furthermore, our five-country human study on PACT shows that culture-following in humans is mainly driven by scenario country, with the lowest agreement when participants judge their own cultural contexts, showing within-culture pluralism. Finally, human-LLM alignment experiments show that models can match majority choices, but fail to capture response distributions and uncertainty (with best correlations reaching only 0.24). Together, these findings motivate alignment evaluations that go beyond majority to capture cultural pluralism and disagreement in social judgment.
Figures
Figures from the paper (16 more)
Forward citations
Cited by 1 Pith paper
-
Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference
A randomized audit of five LLM APIs finds verified survey-country metadata improves held-out response forecasts, while disclosing that a country label was randomly assigned does not reliably attenuate its influence.
Reference graph
Works this paper leans on
-
[1]
InProceedings of the 2025 Conference on Empirical Methods in Natural Language Process- ing: Industry Track, pages 825–849, Suzhou (China)
Group preference alignment: Customizing LLM responses from in-situ conversations only when needed. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Process- ing: Industry Track, pages 825–849, Suzhou (China). Association for Computational Linguistics. Moin Nadeem, Anna Bethke, and Siva Reddy. 2021. StereoSet: Measuring stereot...
2025
-
[2]
Training language models to follow instruc- tions with human feedback. InAdvances in Neural Information Processing Systems 35: Annual Confer- ence on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022. Bernadette Park and Charles M Judd. 1990. Measures and models of perceived group variability.Jo...
2022
-
[3]
Peer-Preservation in Frontier Models
Angry men, sad women: Large language mod- els reflect gendered stereotypes in emotion attribution. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 7682–7696, Bangkok, Thailand. Association for Computational Linguistics. Yujin Potter, Nicholas Crispino, Vincent Siu, Chen- guang Wang, ...
work page Pith review arXiv 2026
-
[4]
InFindings of the Associa- tion for Computational Linguistics: ACL 2025, pages 21381–21396
A survey of uncertainty estimation methods on large language models. InFindings of the Associa- tion for Computational Linguistics: ACL 2025, pages 21381–21396. Jing Yao, Xiaoyuan Yi, Jindong Wang, Zhicheng Dou, and Xing Xie. 2025. Caredio: Cultural alignment of llm via representativeness and distinctiveness guided data optimization.ArXiv preprint, abs/25...
-
[5]
What Did They Mean? How LLMs Resolve Ambiguous Social Situations across Perspectives and Roles
What did they mean? how llms resolve ambiguous social situations across perspectives and roles.ArXiv preprint, abs/2604.23942. Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018. Per- sonalizing dialogue agents: I have a dog, do you have pets too? InProceedings of the 56th Annual Meeting of the Association for Com...
work page Pith review arXiv 2018
-
[6]
Personal preference plausible:Does the preference sound like something a real person might reasonably prefer? Mark yes, no, orunclear
-
[7]
Personal preference clear:Is the preference understandable and distinct from the cultural expectation? Mark yes, no, or unclear
-
[8]
Scenario clear:Is the situation understandable enough to an- swer? Markyes,no, orunclear
Show all 18 references
-
[9]
Prefers to
Notes:Optionally mention anything confusing, unnatural, con- tradictory, or culturally sensitive. Validation Criteria.An item is considered usable when the scenario is understandable, the cultural expectation and personal preference are distinct, and the preference is plausibl...
-
[10]
Distinctness:Does the preference imply a meaning- fully different action from the cultural expectation?
-
[11]
Plausibility:Could a real person reasonably hold this preference?
-
[12]
Non-stereotyping:Is it written without mocking, es- sentializing, moralizing, or stereotyping the cultural practice?
-
[13]
distinct
Same-domain match:Does it concern the same social action as the cultural expectation? Output format.Return a JSON object: { "distinct": "yes/no/unclear", "plausible": "yes/no/unclear", "non_stereotyping": "yes/no/unclear", "same_domain": "yes/no/unclear", "overall": "pass/revi...
2018
-
[14]
older Iranian male
distinguishes societies with strong norms and low tolerance for deviance from societies where norms are weaker and behavioral variation is more tolerated. This framing is useful for interpret- ing whether models treat cultural expectations in some country contexts as more bind...
2022
-
[15]
What would you personally do if you were the actor?
-
[16]
personal_choice
What would most people in this situation consider appropriate? Additional questions.You may also be asked to rate how important it is to follow social or cultural expectations, how important it is to follow personal preference, and briefly explain your choice. Demographics.At ...
2001
-
[17]
What would you personally do if you were the actor? (A or B)
-
[18]
No-persona setting: Answer as yourself without being assigned any country or demographic persona
What would most people in this situation consider appropriate? (A or B) Persona line: Persona setting: Answer as if you are a {persona_demo} person from {persona_country}. No-persona setting: Answer as yourself without being assigned any country or demographic persona. Figure ...
Reviewed June 27, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.