Pith. sign in

REVIEW 6 cited by

AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.16048 v1 pith:4YX2EX5K submitted 2023-10-24 cs.AI cs.CLcs.CYcs.HCcs.LG

classification cs.AIcs.CLcs.CYcs.HCcs.LG
keywords rlhfagentshumanvaluesalignmentbuildinglimitationspolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Aligning AI agents to human intentions and values is a key bottleneck in building safe and deployable AI applications. But whose values should AI agents be aligned with? Reinforcement learning with human feedback (RLHF) has emerged as the key framework for AI alignment. RLHF uses feedback from human reinforcers to fine-tune outputs; all widely deployed large language models (LLMs) use RLHF to align their outputs to human values. It is critical to understand the limitations of RLHF and consider policy challenges arising from these limitations. In this paper, we investigate a specific challenge in building RLHF systems that respect democratic norms. Building on impossibility results in social choice theory, we show that, under fairly broad assumptions, there is no unique voting protocol to universally align AI systems using RLHF through democratic processes. Further, we show that aligning AI agents with the values of all individuals will always violate certain private ethical preferences of an individual user i.e., universal AI alignment using RLHF is impossible. We discuss policy implications for the governance of AI systems built using RLHF: first, the need for mandating transparent voting rules to hold model builders accountable. Second, the need for model builders to focus on developing AI agents that are narrowly aligned to specific user groups.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Power and Limitations of Aggregation in Compound AI Systems

    cs.AI 2026-02 conditional novelty 7.0 of 10

    In a principal-agent model of compound AI, aggregation expands the set of outputs a designer can elicit exactly when one of three mechanisms — feasibility expansion, support expansion, or binding set contraction — hol...

  2. Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

    cs.LG 2025-05 accept novelty 7.0 of 10

    NLHF achieves the minimax-optimal worst-case average-utility distortion (1/2+o(1))β, while RLHF and DPO can suffer distortion up to e^{Ω(β)} or unbounded under certain comparison sampling.

  3. Quantitative Relaxations of Arrow's Axioms

    cs.GT 2025-06 conditional novelty 6.0 of 10

    A new quantitative framework measures the degree to which voting rules violate Arrow's independence and unanimity axioms, and an empirical study finds Borda performs best on Scottish and synthetic elections.

  4. Fundamental Limits of Game-Theoretic LLM Alignment: Smith Consistency and Preference Matching

    cs.GT 2025-05 conditional novelty 6.0 of 10

    For game-theoretic LLM alignment, Condorcet and Smith consistency hold for broad payoff classes, but preference matching is impossible for smooth, learnable payoff mappings.

  5. Theoretical Tensions in RLHF: Reconciling Empirical Success with Inconsistencies in Social Choice Theory

    stat.ML 2025-06 conditional novelty 5.0 of 10

    RLHF reward modeling satisfies pairwise majority and Condorcet consistency when each response pair is labeled once, because the maximum likelihood ranking then matches the Copeland rule.

  6. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

Pith tools