Pith. sign in

REVIEW 2 major objections 6 minor 243 references

Toward a Theory of Value in AI Alignment

T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read AI alignment research mostly skips defining values and defaults to utility maximization, this review finds.

desk verdict A useful critical meta-analysis whose qualitative core holds up, but the headline percentages should not be read as field-level measurements because the snowball sample is seeded entirely inside the RLHF preference paradigm. read the letter →

arxiv 2608.10327 v1 pith:6MM6XRNQ submitted 2026-08-10 cs.AI

classification cs.AI
keywords AIalignmentvaluehumanvaluespreferencesutilitymaximizationreinforcementlearningfromfeedbackpluralismanthropologicaltheoryof
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that AI value alignment research has built an entire technical program without first answering what human values are. The authors annotated 94 highly cited alignment papers, drawn by snowballing from four canonical preference-based papers, and found that most never define 'values': 79% give no definition or description, 82% treat preferences as a stand-in for values, and 87% implicitly or explicitly embed utility maximization in their alignment framework. The paper claims this is not neutral engineering but the adoption of a specific economic worldview—values as measurable, individual, static, utility-maximizing preferences—while the field presents itself as aligning with human values generally. If the claim is right, the dominant alignment paradigm is silently foreclosing other conceptions of value, and 'value-aligned' claims made to governments and the public rest on unexamined philosophical commitments.

What carries the argument

The load-bearing instrument is the annotation rubric: thirteen forced-choice questions, each with binary options or 'not applicable,' built from distinctions in philosophy and anthropology—monism vs. pluralism, measurable vs. immeasurable, revealed preferences vs. prescribed principles, abstract vs. concrete, individual vs. collective, dynamic vs. static, thin vs. thick, whether values can emerge autonomously in AI, whether the paper names which humans' values are targeted, whether rater pool size is reported, and whether utility maximization is used. The rubric translates an implicit philosophical position into a countable label, so that the prevalence of each commitment across the 94 papers can be estimated and reported as field-level percentages. The snowball-sampled corpus, seeded by four preference-based papers, supplies the cases to which the rubric is applied.

What would settle it

Apply the same rubric to a new 94-paper sample built from seed papers that are not preference-based—for example, Constitutional AI, human-rights-based alignment, or normative-ethics work—and count how many define values and use utility maximization; if most of those papers define values and substantially fewer than 87% use utility maximization, the field-level claim is falsified and the finding stands only for the preference-based paradigm.

Watch

Extended reading notes

Core claim

The paper reports a systematic annotation of 94 highly cited AI value alignment papers, selected by snowball sampling from four canonical preference-based papers. Its central claim is that this literature has no explicit theory of human values: 79% of the sampled papers neither define nor describe what they mean by 'values,' 82% treat preferences as a stand-in for values, 90% treat values as measurable, 53% treat values as static, 66% give 'thin' characterizations, and 91% never specify which humans' values are being aligned. The paper further finds that 87% of the papers implicitly or explicitly use utility maximization as part of the technical framework, which it reads as an unacknowledged commitment to a specific economic worldview: values are individual, rational, monist, static, utility-maximizing preferences. The conclusion is that value alignment is not philosophically neutral; it is enacting a particular theory of value while presenting itself as aligning with human values in general.

Load-bearing premise

The entire statistical picture depends on the assumption that the four seed papers, all from the same preference-learning school, open a fair window onto the whole value alignment field, so the percentages computed from the 94 papers that cite them are treated as properties of the field rather than of that school.

Editorial extensions

If this is right

  • If the field's implicit theory of value is economic preference satisfaction, then claims that models are 'aligned with human values' overstate what reinforcement learning from human feedback (RLHF) and direct preference optimization actually encode: at best they encode the preferences of small, usually undocumented rater pools.
  • The shift from human annotators to LLM-as-a-judge and synthetic preference data means values are increasingly enacted without any human input in the loop, which the paper argues closes off alternative methods for contesting and enacting values in foundation models.
  • Because 91% of the sampled papers never identify which humans' values they are aligning to, the field's default is a universalist, culture-free picture of value that hides the situated nature of the values actually being encoded.
  • Making these commitments explicit turns 'whose values?' from a background assumption into a design question that researchers, deployers, and regulators can ask of any alignment system.
  • The paper's reading of the 'pluralist turn' implies that adding diverse preferences to a reward model does not by itself escape the preference-based theory of value; pluralism is still being expressed inside the utility-maximizing framework.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because all four starting papers belong to the same 'learn from human preferences' school, the headline percentages may describe that school's citation ecology more than the whole alignment field; a snowball seeded from constitutional or human-rights alignment work could produce different rates.
  • A direct test: apply the same rubric to a sample seeded from non-preference papers and compare rates; a large drop in utility-maximization use would reframe the finding from a property of 'the field' to a property of the dominant paradigm.
  • The observed 'alignment without humans' trend implies a governance consequence the paper leaves implicit: when models rate models, accountability for whose values are encoded becomes diffuse, and audits of training data may need to treat simulated raters as a distinct category from human raters.
  • The rubric could be repurposed as a disclosure tool: developers could state, for each deployed system, whether its theory of value is monist or pluralist, static or dynamic, preference-based or principle-based, before claiming 'value alignment.'
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. The paper annotates 94 AI value alignment papers, sampled by snowballing citations from four canonical preference/RLHF seed papers, using a 13-question rubric drawn from philosophy and anthropology. It reports that most sampled papers do not define values, commonly equate values with preferences, treat values as measurable, static, and individual, and rely on utility maximization; it argues that the field has implicitly adopted an economic theory of value, and it discusses alternative anthropological and psychological conceptions. The central qualitative claim is that the value alignment literature rarely engages critically with what human values are.

Significance. If the quantitative claims were properly scoped to the sample, the paper would be a useful critical intervention. The qualitative observation—that many influential alignment papers use 'values' and 'preferences' interchangeably without definition—is well supported by quoted examples and is robust to the authors' own criteria. The rubric is a transparent instrument for interrogating the philosophical commitments of alignment research, and the paper usefully brings anthropology and social philosophy into a debate that is usually conducted in technical terms. However, the field-level percentages are not established by the sampling design, and several rubric categories have near-chance inter-annotator agreement.

major comments (2)
  1. [Methods, Snowball sampling; Discussion, Utility maximization; Abstract] The field-level claim that 87% of papers implicitly or explicitly use utility maximization, and the closely related 82% figure for preferences-as-values, rest on a sample created by snowballing citations from four seed papers that all belong to the preference/RLHF paradigm (Christiano et al. 2017; Askell et al. 2021; Bai et al. 2022; Ouyang et al. 2022). Snowball sampling returns a citation neighborhood centered on those seeds, not a random or representative sample of AI value alignment research; principle-based, constitutional, or human-rights-based alignment work enters the sample only if it happens to cite one of the four preference seeds. The annotation filter described in Methods ('discarded if the term value alignment was used only in a cursory sense') is applied after sampling and cannot repair this frame. The Limitations section explicitly disclaims completeness, yet the Abstract and Conclusion present the percentages as properties of the field. Please either re-run the annotation on an expanded seed set that includes non-preference paradigms and report sensitivity, or reframe all quantitative findings as descriptive of the preference-RLHF citation neighborhood only.
  2. [Methods, Annotation process; Table 1] The sentence in the Annotation process section, 'we achieved a high degree of inter-annotator agreement, which averages overall around 85%,' is not supported by Table 1. The raw percentage agreement in Table 1 ranges from 51.0% to 88.8%, and several rubric categories that feed the headline claims are at or near chance: 'Thin vs Thick' at 52.8% (Krippendorff's alpha 0.073), 'Monism vs Pluralism' at 54.0%, 'Values Static or Dynamic' at 51.0%, and 'Should AI follow human values?' at 56.9% (alpha -0.015). This is a load-bearing problem for the quantitative prevalence claims that depend on these categories. Please report chance-corrected agreement for each rubric item instead of an aggregate raw-agreement figure, and temper the claims that rely on the least reliable categories.
minor comments (6)
  1. [Statement of Contributions] The second contribution bullet says 'snowballed sample of 100 commonly cited AI alignment papers,' but Methods and Findings report 94 annotated papers after discarding cursory mentions; make the numbers consistent.
  2. [Methods, Annotation rubric; Table 1] The numbered rubric lists 13 questions, but Table 1 contains 15 rows (for example, 'Sources for determining values' and 'Which humans' values?'); align the table with the rubric or explain the additional items.
  3. [Findings, Figures 1 and 2] Figures 1 and 2 are referenced in the text but are not included in the submission; ensure the final version contains the figures and their captions.
  4. [Throughout] There are several typos and wording slips: 'we conducted a analysis' (Introduction), 'a rise harms' (Abstract), 'The question ofwhat moral values are' (Related Work), and 'Think vs Thick' in the Table 1 header should be 'Thin vs Thick.'
  5. [Methods, Snowball sampling] The sentence 'We sorted these from highest to lowest Google Scholar citations. We sorted the papers by Google Scholar citation count' repeats the same information; keep one formulation.
  6. [References] Several citations to web resources (Anthropic 2025, OpenAI 2025, Institute 2025) lack URLs and access dates; please complete these entries for reproducibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the central claims are grounded in an annotated sample rather than in a self-referential derivation, and the sampling limitation is disclosed by the authors.

full rationale

This paper is a qualitative meta-analysis, not a formal derivation, so the usual circularity modes (fitted parameters renamed as predictions, uniqueness theorems imported from the authors' own prior work, ansatz smuggled in via citation) do not apply. The main quantitative claims—82% of papers treat preferences as a stand-in for values and 87% use utility maximization—are outcomes of applying an externally motivated rubric to 94 papers, with reported inter-annotator agreement around 85%. The rubric is drawn from philosophy, anthropology, and economics, not from the target result, so the measurement is not definitionally equivalent to the conclusion. The closest concern is sample construction: the four seed papers are all canonical preference/RLHF papers, so a snowball sample of papers citing them tilts toward the preference paradigm. However, the paper explicitly limits its own claim: 'we make no claims to the completeness or comprehensiveness of the sampled papers' (Limitations). Citing a preference-RLHF seed does not by construction force a paper to treat preferences as values or to use utility maximization, and indeed 18% and 13% of the sample respectively did not. The self-citations present (e.g., Ahmed et al. 2024, Birhane et al. 2022) are motivational or contextual and are not load-bearing for the quantitative findings, which rest on the authors' own annotation data. Thus no step reduces to its own input by construction, and the paper's central critical thesis has independent empirical content.

Assumptions & free parameters 3 free parameters · 4 assumptions · 2 invented entities

The paper's central claims rest on a hand-selected corpus and a rubric built from the authors' own philosophical commitments. It introduces no physical entities, but it does coin interpretive labels ('alignment without humans', 'the pluralist turn') that are not independently measurable. The main epistemic load is carried by the sampling assumption that four RLHF seed papers adequately represent the entire value alignment field, and by the annotation assumption that binary rubric categories capture the presence or absence of philosophical positions in technical papers.

free parameters (3)
  • Sample of 94 papers
    Kept 94 of 576 citing papers based on annotation consensus and resource constraints; all headline percentages depend on this hand-chosen subset.
  • Four seed papers
    Selected by expert judgment and high citation counts; all four are RLHF/preference-based, biasing the citation snowball toward the preference paradigm.
  • Rubric binary categories
    Forced nuanced evaluations of values into binary/ternary options, with the rubric itself revised iteratively after initial annotations; affects all classification counts.
assumptions (4)
  • domain assumption Google Scholar citation count is a valid weak proxy for quality and impact of AI alignment papers.
    Stated in Methods; underlies the ranking and selection of seeds and sample papers.
  • domain assumption The 4 seed papers define the core of AI value alignment research.
    Stated in Methods; different seeds would produce a different corpus and potentially different percentages.
  • domain assumption Philosophical constructs (monism/pluralism, thin/thick value, revealed vs. prescribed principles) can be reliably detected in technical ML papers via close reading.
    The entire rubric rests on this; low inter-rater kappas on several items (e.g., 0.133 for static/dynamic, 0.09 for whether AI should follow human values) suggest this assumption is fragile.
  • ad hoc to paper The authors' own socio-anthropological framework is the correct lens for evaluating value alignment.
    The rubric was built from this framework and the paper uses it to normatively judge the field, without considering whether alternative operationalizations of 'value' in ML are legitimate for engineering purposes.
invented entities (2)
  • 'Alignment without humans'
    purpose: Label for the observed trend of replacing human annotators/raters with AI feedback or simulated human data in alignment research.
    A conceptual framing category introduced by the authors; it is not independently measurable outside the paper's own annotations.
  • 'The pluralist turn'
    purpose: Term describing recent alignment research that attempts to accommodate diverse preferences or values.
    Coined by the authors to periodize the literature; it is an interpretive label rather than an independently observable entity.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Toward a Theory of Value in AI Alignment." pith.science (2026). https://pith.science/paper/6MM6XRNQ

@misc{pith2026260810327,
  author       = {Pith},
  title        = {Pith review of: Toward a Theory of Value in AI Alignment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6MM6XRNQ}},
  note         = {Machine review of arXiv:2608.10327}
}
read the original abstract

Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms spanning from toxic speech and hallucinations to AI agents executing unauthorized actions. Within the field of AI safety, these harmful instances are often framed as the alignment problem, or of models being misaligned with human values. Researchers have responded by pursuing applied and theoretical AI value alignment efforts, often without specifying what they mean by human values. How does the field of AI value alignment conceive of human values? How are these conceptions of values technically operationalized and evaluated? What does the emergent theory of value from this field signify for the future of AI? We annotated 94 value alignment research papers to discern their implicit theory of values in AI. The majority do not define values, relying heavily on preferences as a stand in that runs the risk of reducing complex culturally situated concepts down to binary choices. As researchers dispense with using human annotators for model training and evaluation, turning instead to synthetic data and autorater approaches to aligning and evaluating models, we identify the potential to close off alternative methods for contesting and enacting values in foundation models. In making AI value alignments philosophical commitments explicit, we seek to bring great specificity and under explored perspectives in the debate on whether and how AI can address human values.

Figures

Figures reproduced from arXiv: 2608.10327 by the authors.

Figure 1
Figure 1. Summary of answers to our rubric aggregated across all raters for binary yes/no answers. [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Summary of answers to our rubric aggregated across all raters for categorical questions. INS = ”information not [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

243 extracted references · 29 canonical work pages

  1. [1]

    Prabhakaran, Vinodkumar and Mitchell, Margaret and Gebru, Timnit and Gabriel, Iason , month = oct, year =. A. doi:10.48550/arXiv.2210.02667 , abstract =

  2. [2]

    Sorensen, Taylor and Moore, Jared and Fisher, Jillian and Gordon, Mitchell and Mireshghallah, Niloofar and Rytting, Christopher Michael and Ye, Andre and Jiang, Liwei and Lu, Ximing and Dziri, Nouha and Althoff, Tim and Choi, Yejin , month = aug, year =. A. doi:10.48550/arXiv.2402.05070 , abstract =

  3. [3]

    2023 , note =

    Science , author =. 2023 , note =. doi:10.1126/science.adi8982 , number =

  4. [4]

    Minds and Machines , author =

    Artificial. Minds and Machines , author =. 2020 , keywords =. doi:10.1007/s11023-020-09539-2 , abstract =

  5. [5]

    Philosophical Studies , author =

    Beyond. Philosophical Studies , author =. 2024 , note =. doi:10.1007/s11098-024-02249-w , abstract =

  6. [6]

    and Boyd, Danah and Friedler, Sorelle A

    Selbst, Andrew D. and Boyd, Danah and Friedler, Sorelle A. and Venkatasubramanian, Suresh and Vertesi, Janet , month = jan, year =. Fairness and. Proceedings of the. doi:10.1145/3287560.3287598 , abstract =

  7. [7]

    Anwar, Usman and Saparov, Abulhair and Rando, Javier and Paleka, Daniel and Turpin, Miles and Hase, Peter and Lubana, Ekdeep Singh and Jenner, Erik and Casper, Stephen and Sourbut, Oliver and Edelman, Benjamin L. and Zhang, Zhaowei and Günther, Mario and Korinek, Anton and Hernandez-Orallo, Jose and Hammond, Lewis and Bigelow, Eric and Pan, Alexander and ...

  8. [8]

    Impossibility and

    Eckersley, Peter , month = mar, year =. Impossibility and. doi:10.48550/arXiv.1901.00064 , abstract =

Show all 243 references
  1. [9]

    Interdisciplinary

    Steinert, Steffen , year =. Interdisciplinary

  2. [10]

    Buyl, Maarten and Rogiers, Alexander and Noels, Sander and Dominguez-Catena, Iris and Heiter, Edith and Romero, Raphael and Johary, Iman and Mara, Alexandru-Cristian and Lijffijt, Jefrey and Bie, Tijl De , month = oct, year =. Large. doi:10.48550/arXiv.2410.18417 , abstract =

  3. [11]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author =

    Learning. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , author =. 2024 , pages =. doi:10.1609/aies.v7i1.31716 , abstract =

  4. [12]

    Contemporary Political Theory , author =

    Mute compulsion:. Contemporary Political Theory , author =. 2023 , file =. doi:10.1057/s41296-023-00638-0 , language =

  5. [13]

    El-Mhamdi, El-Mahdi and Hoang, Lê-Nguyên , month = oct, year =. On. doi:10.48550/arXiv.2410.09638 , abstract =

  6. [14]

    Big Data & Society , author =

    On the genealogy of machine learning datasets:. Big Data & Society , author =. 2021 , note =. doi:10.1177/20539517211035955 , abstract =

  7. [15]

    Performative

    Hardt, Moritz and Mendler-Dünner, Celestine , month = oct, year =. Performative. doi:10.48550/arXiv.2310.16608 , abstract =

  8. [16]

    Arora, Arnav and Kaffee, Lucie-aimée and Augenstein, Isabelle , editor =. Probing. Proceedings of the. 2023 , pages =. doi:10.18653/v1/2023.c3nlp-1.12 , abstract =

  9. [17]

    Kliemt, Hartmut , month = jan, year =. Rawls’s. doi:10.1163/9789047431060_006 , note =

  10. [18]

    and Smith, Logan and Thibodeau, Jacques and McDonell, Kyle and Reynolds, Laria , month = jun, year =

    Kirchner, Jan H. and Smith, Logan and Thibodeau, Jacques and McDonell, Kyle and Reynolds, Laria , month = jun, year =. Researching. doi:10.48550/arXiv.2206.02841 , abstract =

  11. [19]

    Situated

    Haslanger, Sally , year =. Situated

  12. [20]

    Theory, Culture & Society , author =

    Technology as. Theory, Culture & Society , author =. 2014 , note =. doi:10.1177/0263276413488960 , abstract =

  13. [21]

    Andrews, Mel , year =. The

  14. [22]

    AI and Ethics , author =

    The obscure politics of artificial intelligence: a. AI and Ethics , author =. 2024 , keywords =. doi:10.1007/s43681-024-00476-9 , abstract =

  15. [23]

    Patterns , author =

    The reanimation of pseudoscience in machine learning and its ethical repercussions , volume =. Patterns , author =. 2024 , note =. doi:10.1016/j.patter.2024.101027 , language =

  16. [24]

    Toward an

    Graeber, David , year =. Toward an

  17. [25]

    Adilazuarda, Muhammad Farid and Mukherjee, Sagnik and Lavania, Pradhyumna and Singh, Siddhant and Aji, Alham Fikri and O'Neill, Jacki and Modi, Ashutosh and Choudhury, Monojit , month = sep, year =. Towards. doi:10.48550/arXiv.2403.15412 , abstract =

  18. [26]

    What are human values, and how do we align

    Klingefjord, Oliver and Lowe, Ryan and Edelman, Joe , month = apr, year =. What are human values, and how do we align. doi:10.48550/arXiv.2404.10636 , abstract =

  19. [27]

    Trends in Cognitive Sciences , author =

    Where. Trends in Cognitive Sciences , author =. 2019 , keywords =. doi:10.1016/j.tics.2019.07.012 , abstract =

  20. [28]

    The audit society: rituals of verification , isbn =

    Power, Michael , year =. The audit society: rituals of verification , isbn =

  21. [29]

    Jiang, Liwei and Levine, Sydney and Choi, Yejin , month = nov, year =. Can

  22. [30]

    and Alamdari, Parand A

    Klassen, Toryn Q. and Alamdari, Parand A. and McIlraith, Sheila A. , month = nov, year =. Pluralistic

  23. [31]

    and Agadakos, Nikolaos and Sasulski, Synthia and Farajzadeh, Ali and Choudhury, Sanjiban and Ziebart, Brian D

    Shah, Rushit N. and Agadakos, Nikolaos and Sasulski, Synthia and Farajzadeh, Ali and Choudhury, Sanjiban and Ziebart, Brian D. , month = nov, year =. Value-

  24. [32]

    Ovadya, Aviv and Thorburn, Luke and Redman, Kyle and Devine, Flynn and Milli, Smitha and Revel, Manon and Konya, Andrew and Kasirzadeh, Atoosa , month = nov, year =. Toward

  25. [33]

    Plurality of value pluralism and

    Kasirzadeh, Atoosa , month = nov, year =. Plurality of value pluralism and

  26. [34]

    Swaminathan, Nandhini and Danks, David , month = nov, year =

  27. [35]

    Lu, Christina and Kleek, Max Van , month = nov, year =. Model

  28. [36]

    Boldi, Ryan and Ding, Li and Spector, Lee and Niekum, Scott , month = nov, year =. Pareto-

  29. [37]

    Adaptive

    Harland, Hadassah and Dazeley, Richard and Vamplew, Peter and Senaratne, Hashini and Nakisa, Bahareh and Cruz, Francisco , month = nov, year =. Adaptive

  30. [38]

    and Foale, Cameron and Dazeley, Richard and Harland, Hadassah , month = nov, year =

    Vamplew, Peter and Hayes, Conor F. and Foale, Cameron and Dazeley, Richard and Harland, Hadassah , month = nov, year =. Multi-objective

  31. [39]

    Personalizing

    Poddar, Sriyash and Wan, Yanming and Ivison, Hamish and Gupta, Abhishek and Jaques, Natasha , month = nov, year =. Personalizing

  32. [40]

    and Benkler, Noam and Mosaphir, Drisana and Rye, Jeffrey and Schmer-Galunder, Sonja M

    Friedman, Scott E. and Benkler, Noam and Mosaphir, Drisana and Rye, Jeffrey and Schmer-Galunder, Sonja M. and Goldwater, Micah and McLure, Matthew and Wheelock, Ruta and Gottlieb, Jeremy and Goldman, Robert P. and Miller, Christopher , month = nov, year =. Bottom-

  33. [41]

    Birhane, Abeba and Kalluri, Pratyusha and Card, Dallas and Agnew, William and Dotan, Ravit and Bao, Michelle , month = jun, year =. The. 2022. doi:10.1145/3531146.3533083 , language =

  34. [42]

    The comprehensibility of the universe: a new conception of science , isbn =

    Maxwell, Nicholas , year =. The comprehensibility of the universe: a new conception of science , isbn =

  35. [43]

    and Pistilli, Giada and Menédez-González, Natalia and Duran, Leslye Denisse Dias and Panai, Enrico and Kalpokiene, Julija and Bertulfo, Donald Jay , month = mar, year =

    Johnson, Rebecca L. and Pistilli, Giada and Menédez-González, Natalia and Duran, Leslye Denisse Dias and Panai, Enrico and Kalpokiene, Julija and Bertulfo, Donald Jay , month = mar, year =. The. doi:10.48550/arXiv.2203.07785 , abstract =

  36. [44]

    Journal Moral Philosophy , author =

    Are. Journal Moral Philosophy , author =. 2023 , pages =. doi:10.1163/17455243-20234372 , number =

  37. [45]

    and Rajagopal, Dheeraj and Bolukbasi, Tolga and Dixon, Lucas and Tenney, Ian , month = dec, year =

    Chang, Tyler A. and Rajagopal, Dheeraj and Bolukbasi, Tolga and Dixon, Lucas and Tenney, Ian , month = dec, year =. Scalable. doi:10.48550/arXiv.2410.17413 , abstract =

  38. [46]

    2024 , pages =

    Science , author =. 2024 , pages =. doi:10.1126/science.adq2852 , abstract =

  39. [47]

    Cooperative inverse reinforcement learning , isbn =

    Hadfield-Menell, Dylan and Dragan, Anca and Abbeel, Pieter and Russell, Stuart , month = dec, year =. Cooperative inverse reinforcement learning , isbn =. Proceedings of the 30th

  40. [48]

    Minds and Machines , author =

    Two. Minds and Machines , author =. 2022 , keywords =. doi:10.1007/s11023-021-09569-4 , abstract =

  41. [49]

    AI & SOCIETY , author =

    Beyond model interpretability: socio-structural explanations in machine learning , issn =. AI & SOCIETY , author =. 2024 , keywords =. doi:10.1007/s00146-024-02056-1 , abstract =

  42. [50]

    and Murphy, Liam Donat , year =

    Erickson, Paul A. and Murphy, Liam Donat , year =. A history of anthropological theory , isbn =

  43. [51]

    Narayanan, Arvind and Kapoor, Sayash , year =

  44. [52]

    Elizabeth and Horowitz, Aaron and Selbst, Andrew , month = jun, year =

    Raji, Inioluwa Deborah and Kumar, I. Elizabeth and Horowitz, Aaron and Selbst, Andrew , month = jun, year =. The. 2022. doi:10.1145/3531146.3533158 , language =

  45. [53]

    Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , author =

    The fallacy of inscrutability , volume =. Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences , author =. 2018 , pages =. doi:10.1098/rsta.2018.0084 , abstract =

  46. [54]

    Internet Governance Project , author =

    The. Internet Governance Project , author =

  47. [55]

    2024 , keywords =

    Nature , author =. 2024 , keywords =. doi:10.1038/s41586-024-07856-5 , abstract =

  48. [56]

    Training a

    Bai, Yuntao and Jones, Andy and Ndousse, Kamal and Askell, Amanda and Chen, Anna and DasSarma, Nova and Drain, Dawn and Fort, Stanislav and Ganguli, Deep and Henighan, Tom and Joseph, Nicholas and Kadavath, Saurav and Kernion, Jackson and Conerly, Tom and El-Showk, Sheer and E...

  49. [57]

    Big Data & Society , author =

    How the machine ‘thinks’:. Big Data & Society , author =. 2016 , pages =. doi:10.1177/2053951715622512 , abstract =

  50. [58]

    Essays on the nature and state of modern economics , isbn =

    Lawson, Tony , year =. Essays on the nature and state of modern economics , isbn =

  51. [59]

    The master's tools will never dismantle the master's house , isbn =

    Lorde, Audre and Lorde, Audre , year =. The master's tools will never dismantle the master's house , isbn =

  52. [60]

    2021 , note =

    The tyranny of algorithms: a conversation with. 2021 , note =

  53. [61]

    Real knowing: new versions of the coherence theory , isbn =

    Alcoff, Linda , year =. Real knowing: new versions of the coherence theory , isbn =

  54. [62]

    African appropriations: cultural difference, mimesis, and media , isbn =

    Krings, Matthias , year =. African appropriations: cultural difference, mimesis, and media , isbn =

  55. [63]

    On extinction: beginning again at the end , isbn =

    Ware, Ben , year =. On extinction: beginning again at the end , isbn =

  56. [64]

    Gitelman, Lisa , year =. "

  57. [65]

    Causation in science , isbn =

    Ben-Menahem, Yemima , year =. Causation in science , isbn =

  58. [66]

    Philosophy of mind: classical and contemporary readings , isbn =

    Chalmers, David John , year =. Philosophy of mind: classical and contemporary readings , isbn =

  59. [67]

    , year =

    Shrader-Frechette, Kristin S. , year =. Risk and rationality: philosophical foundations for populist reforms , isbn =

  60. [68]

    Empire of

    Chapman, Robert , year =. Empire of

  61. [69]

    The epistemology of resistance: gender and racial oppression, epistemic injustice, and resistant imaginations , isbn =

    Medina, José , year =. The epistemology of resistance: gender and racial oppression, epistemic injustice, and resistant imaginations , isbn =

  62. [70]

    Russell , year =

    Bernard, H. Russell , year =. Research methods in anthropology: qualitative and quantitative approaches , isbn =

  63. [71]

    Measurement

    Mari, Luca and Wilson, Mark and Maul, Andrew , year =. Measurement

  64. [72]

    Debt: the first 5,000 years , isbn =

    Graeber, David , year =. Debt: the first 5,000 years , isbn =

  65. [73]

    and Whitcomb, Dennis , year =

    Goldman, Alvin I. and Whitcomb, Dennis , year =. Social epistemology: essential readings , isbn =

  66. [74]

    2012 , keywords =

    Metaphysics: an anthology , isbn =. 2012 , keywords =

  67. [75]

    Genealogy as critique:

    Koopman, Colin , year =. Genealogy as critique:

  68. [76]

    Responsibility for justice , isbn =

    Young, Iris Marion , year =. Responsibility for justice , isbn =

  69. [77]

    The idea of a critical theory:

    Geuss, Raymond , year =. The idea of a critical theory:

  70. [78]

    Postcapitalist desire: the final lectures , isbn =

    Fisher, Mark and Colquhoun, Matt , year =. Postcapitalist desire: the final lectures , isbn =

  71. [79]

    Dear science and other stories , isbn =

    McKittrick, Katherine , year =. Dear science and other stories , isbn =

  72. [80]

    and Kelley, Robin D

    Robinson, Cedric J. and Kelley, Robin D. G. and Willoughby-Herard, Tiffany and Sojoyner, Damien M. , year =. Black marxism: the making of the

  73. [81]

    , year =

    Friedman, Rachel Z. , year =. Probable justice: risk, insurance, and the welfare state , isbn =

  74. [82]

    The blind spot: why science cannot ignore human experience , isbn =

    Frank, Adam and Gleiser, Marcelo and Thompson, Evan , year =. The blind spot: why science cannot ignore human experience , isbn =

  75. [83]

    Possibilities: essays on hierarchy, rebellion, and desire , isbn =

    Graeber, David , year =. Possibilities: essays on hierarchy, rebellion, and desire , isbn =

  76. [84]

    , month = nov, year =

    Hanna, Alex and Park, Tina M. , month = nov, year =. Against. doi:10.48550/arXiv.2010.08850 , abstract =

  77. [85]

    Brown, Barry and Weilenmann, Alexandra and McMillan, Donald and Lampinen, Airi , month = may, year =. Five. Proceedings of the 2016. doi:10.1145/2858036.2858313 , abstract =

  78. [86]

    The alignment problem: machine learning and human values , isbn =

    Christian, Brian , year =. The alignment problem: machine learning and human values , isbn =

  79. [87]

    Human compatible: artificial intelligence and the problem of control , isbn =

    Russell, Stuart Jonathan , year =. Human compatible: artificial intelligence and the problem of control , isbn =

  80. [88]

    Cappelen, Herman and Dever, Josh , year =. Making

  81. [89]

    Imagining

    Cave, Stephen and Dihal, Kanta Sarasvati Monique , year =. Imagining

  82. [90]

    The ant trap: rebuilding the foundations of the social sciences , isbn =

    Epstein, Brian , year =. The ant trap: rebuilding the foundations of the social sciences , isbn =

  83. [91]

    , year =

    Mauss, Marcel and Guyer, Jane I. , year =. The gift , isbn =

  84. [92]

    Hacking, Ian , year =. The

  85. [93]

    Anthropology,

    Chibnik, Michael , year =. Anthropology,

  86. [94]

    Value alignment: a formal approach , shorttitle =

    Sierra, Carles and Osman, Nardine and Noriega, Pablo and Sabater-Mir, Jordi and Perelló, Antoni , month = oct, year =. Value alignment: a formal approach , shorttitle =. doi:10.48550/arXiv.2110.09240 , abstract =

  87. [95]

    and Gates, Monica A

    Fisac, Jaime F. and Gates, Monica A. and Hamrick, Jessica B. and Liu, Chang and Hadfield-Menell, Dylan and Palaniappan, Malayandi and Malik, Dhruv and Sastry, S. Shankar and Griffiths, Thomas L. and Dragan, Anca D. , month = feb, year =. Pragmatic-. doi:10.48550/arXiv.1707.063...

  88. [96]

    Human-centered mechanism design with

    Koster, Raphael and Balaguer, Jan and Tacchetti, Andrea and Weinstein, Ari and Zhu, Tina and Hauser, Oliver and Williams, Duncan and Campbell-Gillingham, Lucy and Thacker, Phoebe and Botvinick, Matthew and Summerfield, Christopher , month = jan, year =. Human-centered mechanis...

  89. [97]

    An introduction to social anthropology: sharing our worlds , isbn =

    Hendry, Joy , year =. An introduction to social anthropology: sharing our worlds , isbn =

  90. [98]

    The political philosophy of

    Coeckelbergh, Mark , year =. The political philosophy of

  91. [99]

    Birhane, Abeba and Dehdashtian, Sepehr and Prabhu, Vinay and Boddeti, Vishnu , month = jun, year =. The. Proceedings of the 2024. doi:10.1145/3630106.3658968 , abstract =

  92. [100]

    Superintelligence: paths, dangers, strategies , isbn =

    Bostrom, Nick , year =. Superintelligence: paths, dangers, strategies , isbn =

  93. [101]

    Christiano, Paul F and Leike, Jan and Brown, Tom and Martic, Miljan and Legg, Shane and Amodei, Dario , year =. Deep. Advances in

  94. [102]

    Aligning

    Hendrycks, Dan and Burns, Collin and Basart, Steven and Critch, Andrew and Li, Jerry and Song, Dawn and Steinhardt, Jacob , month = feb, year =. Aligning. doi:10.48550/arXiv.2008.02275 , abstract =

  95. [103]

    Scalable agent alignment via reward modeling: a research direction , shorttitle =

    Leike, Jan and Krueger, David and Everitt, Tom and Martic, Miljan and Maini, Vishal and Legg, Shane , month = nov, year =. Scalable agent alignment via reward modeling: a research direction , shorttitle =. doi:10.48550/arXiv.1811.07871 , abstract =

  96. [104]

    Online Readings in Psychology and Culture , author =

    An. Online Readings in Psychology and Culture , author =. 2012 , file =. doi:10.9707/2307-0919.1116 , number =

  97. [105]

    Geertz, Clifford , year =. “. The

  98. [106]

    arXiv preprint arXiv:2310.19852 , year=

    Ai alignment: A comprehensive survey , author=. arXiv preprint arXiv:2310.19852 , year=

  99. [107]

    arXiv preprint arXiv:2112.00861 , year=

    A general language assistant as a laboratory for alignment , author=. arXiv preprint arXiv:2112.00861 , year=

  100. [108]

    arXiv preprint arXiv:2406.18346 , year=

    AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations , author=. arXiv preprint arXiv:2406.18346 , year=

  101. [109]

    arXiv preprint arXiv:2410.11385 , year=

    Do LLMs Have the Generalization Ability in Conducting Causal Inference? , author=. arXiv preprint arXiv:2410.11385 , year=

  102. [110]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

    Unsocial Intelligence: An Investigation of the Assumptions of AGI Discourse , author=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

  103. [111]

    Scientific Reports , volume=

    STELA: a community-centred approach to norm elicitation for AI alignment , author=. Scientific Reports , volume=. 2024 , publisher=

  104. [112]

    Advances in neural information processing systems , volume=

    Training language models to follow instructions with human feedback , author=. Advances in neural information processing systems , volume=

  105. [113]

    arXiv preprint arXiv:2402.12366 , year=

    A critical evaluation of ai feedback for aligning large language models , author=. arXiv preprint arXiv:2402.12366 , year=

  106. [114]

    2007 , publisher=

    Philosophy of science AZ , author=. 2007 , publisher=

  107. [115]

    2023 , publisher=

    Mute compulsion: A Marxist theory of the economic power of capital , author=. 2023 , publisher=

  108. [116]

    Value as theory

    Introduction:“Value as theory” Comparison, cultural critique, and guerilla ethnographic theory , author=. HAU: Journal of Ethnographic Theory , volume=. 2013 , publisher=

  109. [117]

    HAU: Journal of Ethnographic Theory , volume=

    Monism, pluralism, and the structure of value relations: A Dumontian contribution to the contemporary study of value , author=. HAU: Journal of Ethnographic Theory , volume=. 2013 , publisher=

  110. [118]

    1971 , publisher=

    The elementary structures of kinship , author=. 1971 , publisher=

  111. [119]

    HAU: Journal of Ethnographic Theory , volume=

    Foreword: The return of ethnographic theory , author=. HAU: Journal of Ethnographic Theory , volume=. 2011 , publisher=

  112. [120]

    1986 , publisher=

    Essays on individualism: Modern ideology in anthropological perspective , author=. 1986 , publisher=

  113. [121]

    Swiss Journal of Economics and Statistics , volume=

    Bounded rationality: Models of fast and frugal inference , author=. Swiss Journal of Economics and Statistics , volume=

  114. [122]

    arXiv preprint arXiv:2411.04991 , year=

    Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives , author=. arXiv preprint arXiv:2411.04991 , year=

  115. [123]

    The method of paired comparisons , author=

    Rank analysis of incomplete block designs: I. The method of paired comparisons , author=. Biometrika , volume=. 1952 , publisher=

  116. [124]

    Advances in neural information processing systems , volume=

    Deep reinforcement learning from human preferences , author=. Advances in neural information processing systems , volume=

  117. [125]

    John Rawls, A Theory of Justice , pages=

    Rawls’s Critique of Utilitarianism , author=. John Rawls, A Theory of Justice , pages=. 2013 , publisher=

  118. [126]

    arXiv preprint arXiv:2404.08555 , year=

    RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs , author=. arXiv preprint arXiv:2404.08555 , year=

  119. [127]

    Advances in experimental social psychology/Academic Press , year=

    Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries , author=. Advances in experimental social psychology/Academic Press , year=

  120. [128]

    arXiv preprint arXiv:1606.06565 , year=

    Concrete problems in AI safety , author=. arXiv preprint arXiv:1606.06565 , year=

  121. [129]

    arXiv preprint arXiv:2401.10899 , year=

    Concrete problems in AI safety, revisited , author=. arXiv preprint arXiv:2401.10899 , year=

  122. [130]

    Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

    System safety and artificial intelligence , author=. Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , pages=

  123. [131]

    ASPLP 2024 Conference , year=

    Social Structures as a Site for Injustice: Why Social Theory Matters , author=. ASPLP 2024 Conference , year=

  124. [132]

    , author=

    Failures of Methodological Individualism: The Materiality of Social Systems. , author=. Journal of Social Philosophy , volume=

  125. [133]

    AI magazine , volume=

    Research priorities for robust and beneficial artificial intelligence , author=. AI magazine , volume=

  126. [134]

    Artificial intelligence safety and security , pages=

    The value learning problem , author=. Artificial intelligence safety and security , pages=. 2018 , publisher=

  127. [135]

    Kropotkin, Peter , collaborator =. Mutual

  128. [136]

    Argonauts of the

    Malinowski, Bronislaw , year =. Argonauts of the

  129. [137]

    2023 , publisher=

    Which humans? , author=. 2023 , publisher=

  130. [138]

    arXiv preprint arXiv:2307.01370 , year=

    Multilingual language models are not multicultural: A case study in emotion , author=. arXiv preprint arXiv:2307.01370 , year=

  131. [139]

    arXiv preprint arXiv:2303.17466 , year=

    Assessing cross-cultural alignment between ChatGPT and human societies: An empirical study , author=. arXiv preprint arXiv:2303.17466 , year=

  132. [140]

    International Conference on Machine Learning , pages=

    Whose opinions do language models reflect? , author=. International Conference on Machine Learning , pages=. 2023 , organization=

  133. [141]

    arXiv preprint arXiv:2203.13722 , year=

    Probing pre-trained language models for cross-cultural differences in values , author=. arXiv preprint arXiv:2203.13722 , year=

  134. [142]

    Cosmopolitanism:

    Appiah, Kwame Anthony , year =. Cosmopolitanism:

  135. [143]

    and Balkin, J

    Balkin, Jack M. and Balkin, J. M. , year =. Cultural software: a theory of ideology , isbn =

  136. [144]

    Value in ethics and economics , isbn =

    Anderson, Elizabeth , year =. Value in ethics and economics , isbn =

  137. [145]

    2013 , publisher=

    Readings for a history of anthropological theory , author=. 2013 , publisher=

  138. [146]

    arXiv preprint arXiv:2201.08239 , year=

    Lamda: Language models for dialog applications , author=. arXiv preprint arXiv:2201.08239 , year=

  139. [147]

    AI & SOCIETY , pages=

    ‘Interpretability’and ‘alignment’are fool’s errands: a proof that controlling misaligned large language models is the best anyone can hope for , author=. AI & SOCIETY , pages=. 2024 , publisher=

  140. [148]

    alignment

    The empty signifier problem: Towards clearer paradigms for operationalising" alignment" in large language models , author=. arXiv preprint arXiv:2310.02457 , year=

  141. [149]

    Trail of Bits , volume=

    Toward comprehensive risk assessments and assurance of ai-based systems , author=. Trail of Bits , volume=

  142. [150]

    First Monday , year=

    Field-building and the epistemic culture of AI safety , author=. First Monday , year=

  143. [151]

    First Monday , year=

    The TESCREAL bundle: Eugenics and the promise of utopia through artificial general intelligence , author=. First Monday , year=

  144. [152]

    The Journal of philosophy , volume=

    Rational choice and social theory , author=. The Journal of philosophy , volume=. 1994 , publisher=

  145. [153]

    Philosophy and Phenomenological Research , volume=

    Oppressive things , author=. Philosophy and Phenomenological Research , volume=. 2021 , publisher=

  146. [154]

    Advances in Neural Information Processing Systems , volume=

    Fine-tuning language models to find agreement among humans with diverse preferences , author=. Advances in Neural Information Processing Systems , volume=

  147. [155]

    arXiv preprint arXiv:2410.18417 , year=

    Large language models reflect the ideology of their creators , author=. arXiv preprint arXiv:2410.18417 , year=

  148. [156]

    Daedalus , volume=

    Artificial intelligence, humanistic ethics , author=. Daedalus , volume=. 2022 , publisher=

  149. [157]

    Now What? , author=

    AI Red-Teaming is a Sociotechnical System. Now What? , author=. arXiv preprint arXiv:2412.09751 , year=

  150. [158]

    Computer ethics , pages=

    Do artifacts have politics? , author=. Computer ethics , pages=. 2017 , publisher=

  151. [159]

    Computer , volume=

    How computer systems embody values , author=. Computer , volume=. 2001 , publisher=

  152. [160]

    2019 , publisher=

    Race after Technology: Abolitionist Tools for the New Jim Code , author=. 2019 , publisher=

  153. [161]

    Shrishak, Kris , month = apr, year =

  154. [162]

    UK AI Security Institute , month = dec, year =

  155. [163]

    Dai, Jessica , month = aug, year =

  156. [164]

    OpenAI , month = dec, year =

  157. [165]

    Anthropic , month = dec, year =

  158. [166]

    , author =

    Personalisation within bounds: A risk taxonomy and policy framework for the alignment of large language models with personalised feedback , url =. , author =

  159. [167]

    , author =

    Align on the Fly: Adapting Chatbot Behavior to Established Norms , url =. , author =

  160. [168]

    , author =

    ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time , url =. , author =

  161. [169]

    , author =

    Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF , url =. , author =

  162. [170]

    , author =

    HonestLLM: Toward an Honest and Helpful Large Language Model , url =. , author =

  163. [171]

    , author =

    Scalable agent alignment via reward modeling: a research direction , url =. , author =

  164. [172]

    , author =

    The Road to Artificial SuperIntelligence: A Comprehensive Survey of Superalignment , url =. , author =

  165. [173]

    , author =

    Constructive Large Language Models Alignment with Diverse Feedback , url =. , author =

  166. [174]

    , author =

    Online Merging Optimizers for Boosting Rewards and Mitigating Tax in Alignment , url =. , author =

  167. [175]

    , author =

    Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision , url =. , author =

  168. [176]

    , author =

    Large Language Model Alignment: A Survey , url =. , author =

  169. [177]

    , author =

    MaxMin-RLHF: Alignment with Diverse Human Preferences , url =. , author =

  170. [178]

    , author =

    SPO: Multi-Dimensional Preference Sequential Alignment With Implicit Reward Modeling , url =. , author =

  171. [179]

    , author =

    Towards Measuring the Representation of Subjective Global Opinions in Language Models , url =. , author =

  172. [180]

    , author =

    Panacea: Pareto Alignment via Preference Adaptation for LLMs , url =. , author =

  173. [181]

    Advances in Neural Information Processing Systems , volume=

    Adaptive preference scaling for reinforcement learning with human feedback , author=. Advances in Neural Information Processing Systems , volume=

  174. [182]

    arXiv preprint arXiv:2401.07836 , year=

    Two types of AI existential risk: decisive and accumulative , author=. arXiv preprint arXiv:2401.07836 , year=

  175. [183]

    EXISTENTIAL RISK FROM AI , author=

  176. [184]

    1995 , publisher=

    Artificial intelligence: A modern approach;[the intelligent agent book] , author=. 1995 , publisher=

  177. [185]

    2019 , publisher=

    Snowball sampling , author=. 2019 , publisher=

  178. [186]

    Available at SSRN 5877662 , year=

    On the Slow Death of Scaling , author=. Available at SSRN 5877662 , year=

  179. [187]

    AI Magazine , volume=

    Truth is a lie: Crowd truth and the seven myths of human annotation , author=. AI Magazine , volume=

  180. [188]

    , author =

    Distributional Preference Learning: Understanding and Accounting for Hidden Context in RLHF , url =. , author =

  181. [189]

    , author =

    Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models , url =. , author =

  182. [190]

    , author =

    Training Socially Aligned Language Models on Simulated Social Interactions , url =. , author =

  183. [191]

    , author =

    Of Models and Tin Men: A Behavioural Economics Study of Principal-Agent Problems in AI Alignment using Large-Language Models , url =. , author =

  184. [192]

    2024 , publisher=

    The blind spot: Why science cannot ignore human experience , author=. 2024 , publisher=

  185. [193]

    1983 , publisher=

    How the laws of physics lie , author=. 1983 , publisher=

  186. [194]

    Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency , pages=

    Randomness, not representation: The unreliability of evaluating cultural alignment in llms , author=. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency , pages=

  187. [195]

    arXiv preprint arXiv:2307.15217 , year=

    Open problems and fundamental limitations of reinforcement learning from human feedback , author=. arXiv preprint arXiv:2307.15217 , year=

  188. [196]

    arXiv preprint arXiv:2407.02477 , year=

    Understanding alignment in multimodal llms: A comprehensive study , author=. arXiv preprint arXiv:2407.02477 , year=

  189. [197]

    Advances in neural information processing systems , volume=

    Direct preference optimization: Your language model is secretly a reward model , author=. Advances in neural information processing systems , volume=

  190. [198]

    , author =

    Self-Alignment of Large Language Models via Monopolylogue-based Social Scene Simulation , url =. , author =

  191. [199]

    , author =

    The Alignment Problem from a Deep Learning Perspective , url =. , author =

  192. [200]

    AI Feedback , url =

    Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback , url =. , author =

  193. [201]

    , author =

    HiddenGuard: Fine-Grained Safe Generation with Specialized Representation Router , url =. , author =

  194. [202]

    , author =

    Low-Redundant Optimization for Large Language Model Alignment , url =. , author =

  195. [203]

    , author =

    REBEL: Reward Regularization-Based Approach for Robotic Reinforcement Learning from Human Feedback , url =. , author =

  196. [204]

    , author =

    Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation , url =. , author =

  197. [205]

    , author =

    Assessment of Multimodal Large Language Models in Alignment with Human Values , url =. , author =

  198. [206]

    , author =

    CREAM: CONSISTENCY REGULARIZED SELF-REWARDING LANGUAGE MODELS , url =. , author =

  199. [207]

    , author =

    Human-Instruction-Free LLM Self-Alignment with Limited Samples , url =. , author =

  200. [208]

    , author =

    Safe RLHF: Safe Reinforcement Learning from Human Feedback , url =. , author =

  201. [209]

    , author =

    Group Preference Optimization: Few-Shot Alignment of Large Language Models , url =. , author =

  202. [210]

    , author =

    Offline Regularised Reinforcement Learning for Large Language Models Alignment , url =. , author =

  203. [211]

    , author =

    Reward Learning From Preference With Ties , url =. , author =

  204. [212]

    , author =

    Secrets of RLHF in Large Language Models Part II: Reward Modeling , url =. , author =

  205. [213]

    , author =

    Self-Alignment with Instruction Backtranslation , url =. , author =

  206. [214]

    , author =

    Token-level Direct Preference Optimization , url =. , author =

  207. [215]

    , author =

    UltraFeedback: Boosting Language Models with High-quality Feedback , url =. , author =

  208. [216]

    , author =

    What Matters to You? Towards Visual Representation Alignment for Robot Learning , url =. , author =

  209. [217]

    Computational Brain & Behavior , volume=

    What makes a good theory, and how do we make a theory good? , author=. Computational Brain & Behavior , volume=. 2024 , publisher=

  210. [218]

    ACM Computing Surveys , volume=

    AI alignment: A contemporary survey , author=. ACM Computing Surveys , volume=. 2025 , publisher=

  211. [219]

    arXiv preprint arXiv:2502.08640 , year=

    Utility engineering: Analyzing and controlling emergent value systems in ais , author=. arXiv preprint arXiv:2502.08640 , year=

  212. [220]

    Nature Machine Intelligence , pages=

    AI safety for everyone , author=. Nature Machine Intelligence , pages=. 2025 , publisher=

  213. [221]

    Economic Anthropology , volume=

    Value as ethics: Climate change, crisis, and the struggle for the future , author=. Economic Anthropology , volume=. 2023 , publisher=

  214. [222]

    HAU: Journal of Ethnographic Theory , volume=

    The value of (performative) acts , author=. HAU: Journal of Ethnographic Theory , volume=. 2013 , publisher=

  215. [223]

    Advances in neural information processing systems , volume=

    Openassistant conversations-democratizing large language model alignment , author=. Advances in neural information processing systems , volume=

  216. [224]

    arXiv preprint arXiv:2312.10075 , year=

    Assessing llms for moral value pluralism , author=. arXiv preprint arXiv:2312.10075 , year=

  217. [225]

    Advances in Neural Information Processing Systems , volume=

    Evaluating the moral beliefs encoded in llms , author=. Advances in Neural Information Processing Systems , volume=

  218. [226]

    arXiv preprint arXiv:2307.11137 , year=

    Of models and tin men: a behavioural economics study of principal-agent problems in AI alignment using large-language models , author=. arXiv preprint arXiv:2307.11137 , year=

  219. [227]

    ICLR 2025 Workshop on Bidirectional Human-AI Alignment , year =

    Societal Alignment Frameworks Can Improve LLM Alignment , author=. ICLR 2025 Workshop on Bidirectional Human-AI Alignment , year =

  220. [228]

    Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

    Nothing comes without its world--practical challenges of aligning llms to situated human values through RLHF , author=. Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume=

  221. [229]

    Social cognitive and affective neuroscience , volume=

    The theory of constructed emotion: an active inference account of interoception and categorization , author=. Social cognitive and affective neuroscience , volume=. 2017 , publisher=

  222. [230]

    Normative conflicts and shallow AI alignment: R

    Milli. Normative conflicts and shallow AI alignment: R. Milli. Philosophical Studies , pages=. 2025 , publisher=

  223. [231]

    The handbook of social psychology , year=

    What’s Real? A Philosophy of Science for Social Psychology , author=. The handbook of social psychology , year=

  224. [232]

    arXiv preprint arXiv:2406.04391 , year=

    Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive? , author=. arXiv preprint arXiv:2406.04391 , year=

  225. [233]

    Advances in Neural Information Processing Systems , volume=

    The PRISM alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models , author=. Advances in Neural Information Processing Systems , volume=

  226. [234]

    Alondra Nelson , month = jul, year =

  227. [235]

    Seven masterpieces of philosophy , pages=

    Utilitarianism , author=. Seven masterpieces of philosophy , pages=. 2016 , publisher=

  228. [236]

    Ethics , volume=

    Aristotle and Kant on the Source of Value , author=. Ethics , volume=. 1986 , publisher=

  229. [237]

    Science and Engineering Ethics , volume=

    AI as an epistemic technology , author=. Science and Engineering Ethics , volume=. 2023 , publisher=

  230. [238]

    Anthropology today , volume=

    Ecological embeddedness and personhood: Have we always been capitalists? , author=. Anthropology today , volume=. 1998 , publisher=

  231. [239]

    AI and Ethics , volume=

    Culture in the code: Anthropological Concepts Decoding AI’s Hidden Assumptions , author=. AI and Ethics , volume=. 2026 , publisher=

  232. [240]

    Behavioral Sciences , volume=

    What Does ‘Human-Centred AI’Mean? , author=. Behavioral Sciences , volume=

  233. [241]

    Trends in cognitive sciences , volume=

    Efficiently irrational: deciphering the riddle of human choice , author=. Trends in cognitive sciences , volume=. 2022 , publisher=

  234. [242]

    1976 , publisher=

    The economic approach to human behavior , author=. 1976 , publisher=

  235. [243]

    2018 , publisher=

    Representation in cognitive science , author=. 2018 , publisher=

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.