Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction

T0 review · 3 major / 5 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read AI alignment should govern how systems shape evolving human preferences, not merely match fixed wants.

desk verdict Useful agenda-setting reframing of alignment as control over preference trajectories; the five meta-preferences are a clear sketch, not yet a well-posed solution. read the letter →

arxiv 2607.00001 v1 pith:TPVLVF3C submitted 2026-04-01 cs.AI cs.CY

classification cs.AIcs.CY
keywords AIalignmentpreferencedynamicsformationhuman–AIinteractioninfluencecontroltheoryDR-MDPmeta-preferences
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Standard AI alignment treats human preferences as fixed targets to recover and optimize. This paper argues that preferences are layered (immediate wants, goals, identity, values), change across time and context, and are actively built through interaction with tools and systems. Once AI systems persistently shape attention and evaluation, matching today's stated preferences is not enough and can even reward manipulation. Constructive Alignment reframes the problem as control over preference trajectories: the system must jointly influence world states and human evaluative states while staying inside explicit meta-preferences that keep trajectories coherent across layers, reflectively endorsed later, epistemically grounded, bounded in how far and how fast they may be shifted, and empowering when estimates are uncertain. The practical upshot is that alignment becomes governing long-term value formation rather than static preference satisfaction.

What carries the argument

A discrete-time control model whose state includes world state x_t, layered preference state θ_t, and belief state b_t, jointly evolving under AI actions a_t and interaction structure m_t, optimized for estimated well-being subject to five meta-preference constraints (inner coherence, reflective endorsement, bounded influence, epistemic integrity, empowerment under uncertainty).

What would settle it

Build a long-horizon personalization or recommender setting in which preference change is measurable; compare a policy that optimizes only current reward against one that enforces the five trajectory constraints, and check whether the constrained policy produces lower cumulative inter-layer conflict, lower retrospective regret, and less preference drift relative to a clear baseline without simply refusing all influence.

Watch

Extended reading notes

Core claim

Alignment cannot be defined solely as satisfying expressed preferences when those preferences are layered, dynamic, and constructed by interaction with the AI itself. The correct object is therefore a control problem over evolving human preference and belief states, constrained so that trajectories remain coherent across layers, reflectively endorsed from later preference states, bounded against system-driven shifts, epistemically sound, and empowering under uncertainty.

Load-bearing premise

That the five meta-preferences, together with still-unspecified distances, baselines, and tolerance numbers, can be made operational without collapsing into pure inaction or smuggling in unstated values.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper argues that standard AI alignment, which treats human preferences as fixed targets to be inferred and optimized, is inadequate because preferences are layered (A1), dynamic (A2), and constructed through interaction (A3). Drawing on psychology, behavioral economics, and constructivist theory, it introduces Constructive Alignment: alignment as a control problem over evolving preference trajectories rather than static preference satisfaction. Section 5 sketches a discrete-time control model with joint state (world x, preference θ, belief b), actions a, and interaction structure m, evolving under a transition kernel P. Alignment is cast as constrained optimization of an estimated well-being reward subject to five meta-preferences formalized as trajectory constraints: inner coherence, reflective endorsement, bounded influence, epistemic integrity, and empowerment under uncertainty. The claim is that governing influence so trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded, and empowering is the proper alignment target for persistent, personalized AI.

Significance. If the reframing holds, it productively shifts alignment research from reward recovery and static multi-objective aggregation toward trajectory-level control of preference formation—an increasingly urgent issue as systems become persistent and socially embedded. Strengths include a carefully documented empirical foundation for A1–A3, a clear diagnosis of why preference-satisfaction and related methods (IRL, RLHF, pluralistic aggregation, full-stack alignment) underspecify influence, and an explicit extension of Carroll et al.’s DR-MDP insight that preference change is a first-class control object. The paper is honest that §5 is a sketch, not a complete solution, and correctly identifies free modeling choices as open research questions. The contribution is primarily conceptual and agenda-setting rather than a finished formal or empirical result; its value lies in making normative commitments about acceptable preference change explicit and subject to analysis.

major comments (3)
  1. [§5.2–5.7] §5.2–5.7 (Constraints 5.1–5.4 and the empowerment discussion): The central claim that alignment becomes a well-posed control problem over preference trajectories rests on the five meta-preferences jointly rendering the underdetermined DR-MDP-style problem non-vacuous. Each constraint is defined only up to free choices—distance/disagreement measures D_coh, D_refl, d, ℰ; baseline π0; tolerances ε_coh, ε_refl, B, δ_max; horizon T and γ—left open in the “On measuring…” paragraphs. As the authors note when critiquing ParetoUD in Carroll et al., insufficient structure collapses to inaction or admits arbitrary influence. The manuscript does not exhibit any concrete instantiation that simultaneously (a) admits non-trivial improving policies and (b) blocks lock-in and value imposition. Without at least one worked example or existence argument, the claim that these constraints “make alignment a pr
  2. [§5.7] §5.7 (Empowerment Under Uncertainty): Unlike Constraints 5.1–5.4, empowerment is not given an explicit admissibility inequality or objective term. The text motivates preserving option sets over θ and b when preference estimates are uncertain, but supplies no formal object comparable to J_coh, J_refl, or J_inf. Because empowerment is listed among the five load-bearing meta-preferences that define Constructive Alignment, the asymmetry weakens the claim that the constraint set is jointly specified. Either supply a parallel constraint (e.g., a lower bound on mutual-information-style empowerment or on reachable evaluative diversity) or demote empowerment to a design principle rather than a formal meta-preference.
  3. [§5.5; Discussion] §5.5 and Discussion (baseline π0 and structural harm): Bounded influence is defined relative to a reference policy π0, with the text noting that “minimal intervention” is inappropriate in educational or harm-encoding settings. The Discussion further acknowledges that a no-intervention baseline is insufficient when existing trajectories encode discrimination, addiction, or violence. This is load-bearing: without a principled way to choose or justify π0 (or a family of baselines), the bounded-influence constraint can either freeze harmful status-quo preferences or smuggle in contested value judgments under the guise of “admissible influence.” The paper should either propose a concrete selection criterion for π0 or explicitly scope the current formalism to settings where a minimal-intervention baseline is normatively defensible.
minor comments (5)
  1. [Figure 1] Figure 1 caption and body: the layered preference depiction is helpful but the intermittent “interaction (yellow)” interventions are not linked to the formal objects a_t and m_t in §5.1; a brief mapping would improve continuity.
  2. [§5.1] §5.1: the transition is written (x_{t+1}, θ_{t+1}, b_{t+1}) ∼ P(· | x_t, θ_t, b_t, a_t, m_t) without stating whether m_t is chosen by the policy, by designers, or both. Clarifying the decision variables of the control problem would reduce ambiguity.
  3. [§4] Related Work (§4): Carroll et al. [160] is correctly credited for DR-MDPs; a short table or paragraph contrasting which of the five meta-preferences are already implicit in their recasting of IRL/RLHF/recommenders would sharpen the novelty claim.
  4. [§5.2] Notation: R̂⋆ appears with a combining character that may render inconsistently; prefer \hat{R}^\star or similar for camera-ready.
  5. [Keywords / Abstract] Keywords list “DR-MDP” but the acronym is only expanded later via citation; expand on first use in the abstract or introduction for readers outside that sub-literature.

Circularity Check

0 steps flagged · score 1.0 of 10

No load-bearing circularity: axioms rest on external literature, meta-preferences are stipulated normative constraints rather than derived predictions or uniqueness claims.

full rationale

The paper is a conceptual paradigm proposal, not a derivation of empirical predictions or forced uniqueness results. Axioms A1–A3 are supported by extensive external citations from psychology, behavioral economics, and constructivist theory (Piaget, Vygotsky, Slovic, Lichtenstein, Schwartz, Laibson, etc.); they are not defined in terms of the later meta-preferences. The five meta-preferences (Constraints 5.1–5.5) are explicitly introduced as motivated but non-unique higher-level constraints on an underdetermined control problem (DR-MDP-style), with free parameters (D_coh, d, ε, B, δ_max, π0) left open as modeling choices. The Discussion acknowledges they are “one possible set” and that further normative commitments are required. There is no fitted parameter re-labeled as a prediction, no self-definitional loop equating input and output, and no uniqueness theorem imported from the authors’ prior work. The single self-citation ([161]) is merely illustrative for epistemic integrity and is not load-bearing for the central reframing. Mild redefinition of “alignment success” as satisfaction of the authors’ meta-preferences is the contribution itself, not circularity by construction. Score 1 only for the ordinary presence of one non-central self-citation.

Assumptions & free parameters 6 free parameters · 6 assumptions · 2 invented entities

The load-bearing content is almost entirely conceptual: three empirical axioms about preferences, a control-theoretic state model, and five stipulated meta-preference constraints with free tolerance/baseline parameters. No fitted physical constants; the free parameters are the unspecified operational knobs of the framework itself. Invented entities are the paradigm name and the bundled meta-preference set as alignment criteria.

free parameters (6)
  • ε_coh (inner coherence tolerance)
    Admissible cumulative inter-layer conflict; left as free design choice with no calibration method (§5.3).
  • ε_refl (reflective endorsement tolerance)
    Admissible retrospective dissatisfaction from θ_T; free, horizon-relative (§5.4).
  • B and δ_max (influence budget and per-step bound)
    Cumulative and instantaneous preference-divergence limits vs baseline π0; free (§5.5).
  • baseline policy π0
    Counterfactual for measuring induced preference change; choice (minimal intervention vs human tutor etc.) is normative and free (§5.5).
  • horizon T and discount γ
    Trajectory evaluation window and conflict discounting; free modeling choices affecting all trajectory costs.
  • distance/disagreement measures D_coh, D_refl, d, ℰ
    Representation-dependent metrics left open (JS, Kendall τ, L2, KL, Brier, etc.); choice changes what the constraints mean (§5.3–5.6).
assumptions (6)
  • domain assumption A1: Preferences are layered (short-term wants, instrumental goals, identity, values).
    Stated as axiom in §2.1; supported by cited psychology/economics but treated as foundational for the control state θ.
  • domain assumption A2: Preferences are dynamic across generations, life stages, time, context, and action.
    §2.2; empirical literature used as premise that static targets are inadequate.
  • domain assumption A3: Preferences are constructed through interaction (experience, tools, media, measurement, social/algorithmic influence).
    §2.3; core premise that AI influence is unavoidable and must be governed.
  • ad hoc to paper Joint dynamics (x,θ,b) evolve under P(·|x,θ,b,a,m); preferences are state variables not fixed rewards.
    §5.1 modeling commitment extending DR-MDP-style ideas into the paper's layered constructive framing.
  • domain assumption True experienced well-being R*(τ) exists as normative target and can be imperfectly estimated as R̂* from interaction.
    §5.2; required for constrained optimization formulation despite unobservability.
  • ad hoc to paper The five meta-preferences are the right admissibility constraints for alignment.
    §5.3–5.7 and Discussion; authors present them as one possible empirically motivated set, not derived uniquely.
invented entities (2)
  • Constructive Alignment (paradigm)
    purpose: Name and organize alignment as control over preference trajectories rather than static satisfaction.
    New label for a synthesis of constructivism + preference dynamics + control; independent evidence is conceptual only.
  • Five meta-preferences as formal constraints (coherence, reflective endorsement, bounded influence, epistemic integrity, empowerment under uncertainty)
    purpose: Turn underdetermined preference-change control into constrained optimization with named admissibility conditions.
    Bundle is paper-specific; components draw on prior notions (meta-preferences, empowerment, DR-MDP influence) but the joint formal package is introduced here without external validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction." pith.science (2026). https://pith.science/paper/TPVLVF3C

@misc{pith2026260700001,
  author       = {Pith},
  title        = {Pith review of: Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TPVLVF3C}},
  note         = {Machine review of arXiv:2607.00001}
}
read the original abstract

Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with extensive empirical evidence showing that preferences are layered, dynamic, and constructed through interaction--particularly with adaptive technologies. As AI systems become more persistent, personalized, and socially embedded, they increasingly participate in shaping what people attend to, value, and endorse over time. We introduce Constructive Alignment, a paradigm that reframes alignment as a control problem over evolving human preference trajectories rather than static preference satisfaction. Drawing on behavioral economics, psychology, and constructivist social theory, we model preferences as layered state variables that evolve under interaction with AI systems. We formalize this view using a control-theoretic framework in which system actions and interaction design jointly influence both world states and human evaluative states. We argue that alignment is not primarily about controlling AI behavior, but about regulating how AI systems influence the evolution of human preferences--ensuring that value trajectories remain coherent, reflectively endorsed, epistemically grounded, bounded against manipulation, and empowering under uncertainty. Alignment thus becomes a problem of governing long-term value formation rather than simply satisfying static preferences.

Figures

Figures reproduced from arXiv: 2607.00001 by the authors.

Figure 1
Figure 1. Constructive Alignment argues for consideration of the dynamic and constructed nature of preferences. Preferences are depicted as layered, continuous processes that evolve over time (blue, purple, pink), rather than a single static objective to satisfy. Interaction (yellow), including with AI, intermittently intervenes, reshaping these layers, illustrating how preferences are constructed and updated through experien… view at source ↗

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Boundaries of Automation: A Theory of Persistent Human Participation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Human–AI collaboration may remain necessary even with highly capable AI because in some tasks the evaluative target is constituted through interaction, not fixed in advance.

Reference graph

Works this paper leans on

164 extracted references · 19 canonical work pages · cited by 1 Pith paper

  1. [1]

    Heritage, ‘I felt pure, unconditional love’: the people who marry their AI chat- bots, The Guardian (2025)

    S. Heritage, ‘I felt pure, unconditional love’: the people who marry their AI chat- bots, The Guardian (2025). URL: https://www.theguardian.com/tv-and-radio/2025/jul/12/ i-felt-pure-unconditional-love-the-people-who-marry-their-ai-chatbots

  2. [2]

    M. R. Meadi, T. Sillekens, S. Metselaar, A. v. Balkom, J. Bernstein, N. Batelaan, Exploring the Ethical Challenges of Conversational AI in Mental Health Care: Scoping Review, JMIR Mental Health 12 (2025) e60432. URL: https://mental.jmir.org/2025/1/e60432. doi:10.2196/60432

  3. [3]

    Greenfield, The Cambridge Analytica files: the story so far, The Guardian (2018)

    P. Greenfield, The Cambridge Analytica files: the story so far, The Guardian (2018). URL: https: //www.theguardian.com/news/2018/mar/26/the-cambridge-analytica-files-the-story-so-far

  4. [4]

    online brain

    J. Firth, J. Torous, B. Stubbs, J. A. Firth, G. Z. Steiner, L. Smith, M. Alvarez-Jimenez, J. Gleeson, D. Vancampfort, C. J. Armitage, J. Sarris, The “online brain”: how the Internet may be changing our cognition, World Psychiatry 18 (2019) 119–129. URL: https://onlinelibrary.wiley.com/doi/10. 1002/wps.20617. doi:10.1002/wps.20617

  5. [5]

    Franklin, H

    M. Franklin, H. Ashton, R. Gorman, S. Armstrong, Recognising the importance of preference change: A call for a coordinated multidisciplinary research effort in the age of AI, arXiv preprint arXiv:2203.10525 (2022)

  6. [6]

    Zhi-Xuan, M

    T. Zhi-Xuan, M. Carroll, M. Franklin, H. Ashton, Beyond Preferences in AI Alignment, Philosoph- ical Studies (2024). URL: http://arxiv.org/abs/2408.16984. doi:10.1007/s11098-024-02249-w , arXiv:2408.16984 [cs]

  7. [7]

    H. Shen, T. Knearem, R. Ghosh, K. Alkiek, K. Krishna, Y. Liu, Z. Ma, S. Petridis, Y.-H. Peng, L. Qiwei, S. Rakshit, C. Si, Y. Xie, J. P. Bigham, F. Bentley, J. Chai, Z. Lipton, Q. Mei, R. Mihalcea, M. Terry, D. Yang, M. R. Morris, P. Resnick, D. Jurgens, Position: Towards Bidirectional Human- AI Alignment, 2025. URL: http://arxiv.org/abs/2406.09264. doi:1...

  8. [8]

    Piaget, M

    J. Piaget, M. Cook, The origins of intelligence in children, volume 8, International universities press New York, 1952. Issue: 5

Show all 164 references
  1. [9]

    L. S. Vygotsky, M. Cole, Mind in society: the development of higher psychological processes, nachdr. ed., Harvard Univ. Press, Cambridge, Mass., 1981

  2. [10]

    L. A. Suchman, Plans and situated actions: The problem of human-machine communication, Cambridge university press, 1987

  3. [11]

    Hutchins, Cognition in the wild, A Bradford book, 8

    E. Hutchins, Cognition in the wild, A Bradford book, 8. pr ed., MIT Press, Cambridge, Mass., 2006

  4. [12]

    R. K. Sawyer (Ed.), The Cambridge handbook of the learning sciences, 1. publ., repr ed., Cambridge Univ. Press, Cambridge, 2009

  5. [13]

    Slovic, The construction of preference, American Psychologist 50 (1995) 364–371

    P. Slovic, The construction of preference, American Psychologist 50 (1995) 364–371. doi: 10. 1037/0003-066X.50.5.364, place: US

  6. [14]

    Lichtenstein, P

    S. Lichtenstein, P. Slovic (Eds.), The Construction of Preference, Cambridge University Press, Cambridge, 2006. URL: https://www.cambridge.org/core/books/construction-of-preference/ 994FE8DFB8D431338B2A009F25271FBC. doi:10.1017/CBO9780511618031

  7. [15]

    J. R. Bettman, M. F. Luce, J. W. Payne, Constructive Consumer Choice Processes, Journal of Consumer Research 25 (1998) 187–217. URL: https://academic.oup.com/jcr/article/25/3/187/ 1795625. doi:10.1086/209535

  8. [16]

    Ainslie, Specious reward: A behavioral theory of impulsiveness and impulse control, Psycho- logical Bulletin 82 (1975) 463–496

    G. Ainslie, Specious reward: A behavioral theory of impulsiveness and impulse control, Psycho- logical Bulletin 82 (1975) 463–496. doi:10.1037/h0076860

  9. [17]

    Laibson, Golden eggs and hyperbolic discounting, The Quarterly Journal of Economics 112 (1997) 443–478

    D. Laibson, Golden eggs and hyperbolic discounting, The Quarterly Journal of Economics 112 (1997) 443–478

  10. [18]

    Heidhues, P

    P. Heidhues, P. Strack, Identifying present bias from the timing of choices, American Economic Review 111 (2021) 2594–2622

  11. [19]

    J. J. Xiao, N. Porto, Present bias and financial behavior, FINANCIAL PLANNING REVIEW 2 (2019) e1048. URL: https://onlinelibrary.wiley.com/doi/10.1002/cfp2.1048. doi:10.1002/cfp2.1048

  12. [20]

    Y. Wang, F. A. Sloan, Present bias and health, Journal of risk and uncertainty 57 (2018) 177–198

  13. [21]

    A. W. Kruglanski, J. Y. Shah, A. Fishbach, R. Friedman, W. Y. Chun, D. Sleeth-Keppler, A theory of goal systems, in: The motivated mind, Routledge, 2018, pp. 207–250

  14. [22]

    A. W. Kruglanski, M. Chernikova, M. Babush, M. Dugas, B. M. Schumpe, The architecture of goal systems: Multifinality, equifinality, and counterfinality in means—end relations, in: Advances in motivation science, volume 2, Elsevier, 2015, pp. 69–98

  15. [23]

    A. W. Kruglanski, J. J. Bélanger, X. Chen, C. Köpetz, A. Pierro, L. Mannetti, The energetics of motivated cognition: A force-field analysis., Psychological Review 119 (2012) 1–20. URL: https://doi.apa.org/doi/10.1037/a0025488. doi:10.1037/a0025488

  16. [24]

    P. M. Gollwitzer, Implementation intentions: strong effects of simple plans., American psycholo- gist 54 (1999) 493

  17. [25]

    Conner (Ed.), Predicting health behaviour: research and practice with social cognition models,

    M. Conner (Ed.), Predicting health behaviour: research and practice with social cognition models,

  18. [26]

    Press, Maidenhead, 2009

    ed., repr ed., Open Univ. Press, Maidenhead, 2009

  19. [27]

    T. L. Webb, P. Sheeran, Does changing behavioral intentions engender behavior change? A meta-analysis of the experimental evidence, Psychological Bulletin 132 (2006) 249–268. doi: 10. 1037/0033-2909.132.2.249

  20. [28]

    R. R. Vallacher, D. M. Wegner, What do people think they’re doing? Action identification and human behavior, Psychological Review 94 (1987) 3–15. doi:10.1037/0033-295X.94.1.3

  21. [29]

    Fishbach, M

    A. Fishbach, M. J. Ferguson, The goal construct in social psychology, in: Social psychology: Handbook of basic principles, 2nd ed, The Guilford Press, New York, NY, US, 2007, pp. 490–515

  22. [30]

    H. G. Frankfurt, Freedom of the Will and the Concept of a Person, The Journal of Philosophy 68 (1971) 5–20. URL: https://www.jstor.org/stable/2024717. doi:10.2307/2024717

  23. [31]

    Stryker, P

    S. Stryker, P. J. Burke, The past, present, and future of an identity theory, Social psychology quarterly (2000) 284–297

  24. [32]

    G. A. Akerlof, R. E. Kranton, Economics and identity, The quarterly journal of economics 115 (2000) 715–753

  25. [33]

    Aquino, A

    K. Aquino, A. Reed II, The self-importance of moral identity, Journal of Personality and Social Psychology 83 (2002) 1423–1440. doi:10.1037/0022-3514.83.6.1423

  26. [34]

    C. S. Carver, M. F. Scheier, On the self-regulation of behavior, cambridge university press, 2001

  27. [35]

    D. P. McAdams, The psychology of life stories, Review of general psychology 5 (2001) 100–122

  28. [36]

    S. H. Schwartz, Universals in the content and structure of values: Theoretical advances and empirical tests in 20 countries, in: Advances in experimental social psychology, volume 25, Elsevier, 1992, pp. 1–65

  29. [37]

    S. H. Schwartz, An overview of the Schwartz theory of basic values, Online readings in Psychology and Culture 2 (2012) 11

  30. [38]

    S. H. Schwartz, J. Cieciuch, M. Vecchione, E. Davidov, R. Fischer, C. Beierlein, A. Ramos, M. Verkasalo, J.-E. Lönnqvist, K. Demirutku, Refining the theory of basic individual values., Journal of personality and social psychology 103 (2012) 663

  31. [39]

    Rokeach, The nature of human values., Free press, 1973

    M. Rokeach, The nature of human values., Free press, 1973

  32. [40]

    Weber, The Protestant Ethic and the Spirit of Capitalism [1904–5], na, 1930

    M. Weber, The Protestant Ethic and the Spirit of Capitalism [1904–5], na, 1930

  33. [41]

    Hitlin, J

    S. Hitlin, J. A. Piliavin, Values: Reviving a dormant concept, Annu. Rev. Sociol. 30 (2004) 359–393

  34. [42]

    Guiso, P

    L. Guiso, P. Sapienza, L. Zingales, Does culture affect economic outcomes?, Journal of Economic perspectives 20 (2006) 23–48

  35. [43]

    Tabellini, The scope of cooperation: Values and incentives, The Quarterly Journal of Economics 123 (2008) 905–950

    G. Tabellini, The scope of cooperation: Values and incentives, The Quarterly Journal of Economics 123 (2008) 905–950

  36. [44]

    R. H. Thaler, H. M. Shefrin, An economic theory of self-control, Journal of political Economy 89 (1981) 392–406

  37. [45]

    Fudenberg, D

    D. Fudenberg, D. K. Levine, A Dual-Self Model of Impulse Control, American Economic Review 96 (2006) 1449–1476. URL: https://www.aeaweb.org/articles?id=10.1257/aer.96.5.1449. doi:10.1257/ aer.96.5.1449

  38. [46]

    F. Gul, W. Pesendorfer, Temptation and Self-Control, Econometrica 69 (2001) 1403–1435. URL: https://www.jstor.org/stable/2692262

  39. [47]

    Amador, I

    M. Amador, I. Werning, G. Angeletos, Commitment vs. flexibility, Econometrica 74 (2006) 365–396

  40. [48]

    Bénabou, J

    R. Bénabou, J. Tirole, Identity, morals, and taboos: beliefs as assets, The Quarterly Journal of Economics 126 (2011) 805–855. doi:10.1093/qje/qjr002

  41. [49]

    Bénabou, L

    R. Bénabou, L. Henkel, Identity as self-image, Technical Report, National Bureau of Economic Research, 2025

  42. [50]

    Haerpfer, R

    C. Haerpfer, R. Inglehart, A. Moreno, C. Welzel, K. Kizilova, J. Diez-Medrano, M. Lagos, P. Norris, E. Ponarin, B. Puranen, World Values Survey Wave 7 (2017-2020) Cross-National Data-Set, 2020. URL: http://www.worldvaluessurvey.org/WVSDocumentationWV7.jsp. doi:10.14281/18241. 1

  43. [51]

    Inglehart, Modernization and postmodernization: Cultural, economic, and political change in 43 societies, Princeton university press, 2020

    R. Inglehart, Modernization and postmodernization: Cultural, economic, and political change in 43 societies, Princeton university press, 2020

  44. [52]

    P. B. Baltes, Theoretical propositions of life-span developmental psychology: On the dynamics between growth and decline., Developmental psychology 23 (1987) 611

  45. [53]

    Steinberg, A social neuroscience perspective on adolescent risk-taking, in: Biosocial theories of crime, Routledge, 2017, pp

    L. Steinberg, A social neuroscience perspective on adolescent risk-taking, in: Biosocial theories of crime, Routledge, 2017, pp. 435–463

  46. [54]

    L. H. Somerville, The teenage brain: Sensitivity to social evaluation, Current directions in psychological science 22 (2013) 121–127

  47. [55]

    Chein, D

    J. Chein, D. Albert, L. O’Brien, K. Uckert, L. Steinberg, Peers increase adolescent risk taking by enhancing activity in the brain’s reward circuitry (2011)

  48. [56]

    L. L. Carstensen, D. M. Isaacowitz, S. T. Charles, Taking time seriously: a theory of socioemotional selectivity., American psychologist 54 (1999) 165

  49. [57]

    B. W. Roberts, K. E. Walton, W. Viechtbauer, Patterns of mean-level change in personality traits across the life course: a meta-analysis of longitudinal studies., Psychological bulletin 132 (2006) 1

  50. [58]

    O’Donoghue, M

    T. O’Donoghue, M. Rabin, Doing It Now or Later, American Economic Review 89 (1999) 103–124. URL: https://pubs.aeaweb.org/doi/10.1257/aer.89.1.103. doi:10.1257/aer.89.1.103

  51. [59]

    Harris, D

    C. Harris, D. Laibson, Dynamic Choices of Hyperbolic Consumers, Econometrica 69 (2001) 935–957. URL: http://doi.wiley.com/10.1111/1468-0262.00225. doi:10.1111/1468-0262.00225

  52. [60]

    Loewenstein, Out of control: Visceral influences on behavior, Organizational behavior and human decision processes 65 (1996) 272–292

    G. Loewenstein, Out of control: Visceral influences on behavior, Organizational behavior and human decision processes 65 (1996) 272–292

  53. [61]

    G. F. Loewenstein, E. U. Weber, C. K. Hsee, N. Welch, Risk as feelings., Psychological bulletin 127 (2001) 267

  54. [62]

    Ariely, G

    D. Ariely, G. Loewenstein, The heat of the moment: The effect of sexual arousal on sexual decision making, Journal of behavioral decision making 19 (2006) 87–98

  55. [63]

    A. Mani, S. Mullainathan, E. Shafir, J. Zhao, Poverty impedes cognitive function, science 341 (2013) 976–980

  56. [64]

    B. Shiv, A. Fedorikhin, Heart and mind in conflict: The interplay of affect and cognition in consumer decision making, Journal of consumer Research 26 (1999) 278–292

  57. [65]

    Sweller, Cognitive load theory, in: Psychology of learning and motivation, volume 55, Elsevier, 2011, pp

    J. Sweller, Cognitive load theory, in: Psychology of learning and motivation, volume 55, Elsevier, 2011, pp. 37–76

  58. [66]

    Lieder, T

    F. Lieder, T. L. Griffiths, Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources, Behavioral and brain sciences 43 (2020) e1

  59. [67]

    J. W. Brehm, Postdecision changes in the desirability of alternatives., The Journal of Abnormal and Social Psychology 52 (1956) 384

  60. [68]

    Festinger, A theory of cognitive dissonance, A theory of cognitive dissonance, Stanford University Press, 1957

    L. Festinger, A theory of cognitive dissonance, A theory of cognitive dissonance, Stanford University Press, 1957. Pages: xi, 291

  61. [69]

    D. J. Bem, Self-perception theory, in: Advances in experimental social psychology, volume 6, Elsevier, 1972, pp. 1–62

  62. [70]

    Johansson, L

    P. Johansson, L. Hall, N. Chater, Preference Change through Choice, in: Neuroscience of Preference and Choice, Elsevier, 2012, pp. 121–141. URL: https://linkinghub.elsevier.com/retrieve/ pii/B9780123814319000061. doi:10.1016/B978-0-12-381431-9.00006-1

  63. [71]

    Van Veen, M

    V. Van Veen, M. K. Krug, J. W. Schooler, C. S. Carter, Neural activity predicts attitude change in cognitive dissonance, Nature Neuroscience 12 (2009) 1469–1474. URL: https://www.nature.com/ articles/nn.2413. doi:10.1038/nn.2413

  64. [72]

    M. K. Chen, J. L. Risen, How choice affects and reflects preferences: revisiting the free-choice paradigm., Journal of personality and social psychology 99 (2010) 573

  65. [73]

    Izuma, K

    K. Izuma, K. Murayama, Choice-induced preference change in the free-choice paradigm: a critical methodological review, Frontiers in psychology 4 (2013) 41

  66. [74]

    D. Lee, J. Daunizeau, Choosing what we like vs liking what we choose: How choice-induced preference change might actually be instrumental to decision-making, PloS one 15 (2020) e0231081

  67. [75]

    Sharot, B

    T. Sharot, B. De Martino, R. J. Dolan, How choice reveals and shapes expected hedonic outcome, Journal of Neuroscience 29 (2009) 3760–3765

  68. [76]

    Voigt, C

    K. Voigt, C. Murawski, S. Speer, S. Bode, Hard decisions shape the neural coding of preferences, Journal of Neuroscience 39 (2019) 718–726

  69. [77]

    B. M. Staw, Knee-deep in the big muddy: A study of escalating commitment to a chosen course of action, Organizational behavior and human performance 16 (1976) 27–44

  70. [78]

    H. R. Arkes, C. Blumer, The psychology of sunk cost, Organizational behavior and human decision processes 35 (1985) 124–140

  71. [79]

    S. Roth, T. Robbert, L. Straus, On the sunk-cost effect in economic decision-making: a meta- analytic review, Business research 8 (2015) 99–138

  72. [80]

    A. D. Martin, K. M. Quinn, Dynamic ideal point estimation via Markov chain Monte Carlo for the US Supreme Court, 1953–1999, Political analysis 10 (2002) 134–153

  73. [81]

    Loewenstein, Emotions in economic theory and economic behavior, American economic review 90 (2000) 426–432

    G. Loewenstein, Emotions in economic theory and economic behavior, American economic review 90 (2000) 426–432

  74. [82]

    J. D. Hamilton, A new approach to the economic analysis of nonstationary time series and the business cycle, Econometrica: Journal of the econometric society (1989) 357–384

  75. [83]

    Kőszegi, M

    B. Kőszegi, M. Rabin, A model of reference-dependent preferences, The Quarterly Journal of Economics 121 (2006) 1133–1165

  76. [84]

    Dewey, Human nature and conduct, volume 14, Southern Illinois University Press Carbondale, 1988

    J. Dewey, Human nature and conduct, volume 14, Southern Illinois University Press Carbondale, 1988

  77. [85]

    G. B. Saxe, Culture and cognitive development: Studies in mathematical understanding, Psychol- ogy Press, 2015

  78. [86]

    Heidegger, Basic writings: from Being and time (1927) to The task of thinking (1964) (1977)

    M. Heidegger, Basic writings: from Being and time (1927) to The task of thinking (1964) (1977)

  79. [87]

    J. J. Gibson, The theory of affordances:(1979), in: The people, place, and space reader, Routledge, 2014, pp. 56–60

  80. [88]

    Norman, The design of everyday things: Revised and expanded edition, Basic books, 2013

    D. Norman, The design of everyday things: Revised and expanded edition, Basic books, 2013

  81. [89]

    McLuhan, Understanding media: The extensions of man, MIT press, 1994

    M. McLuhan, Understanding media: The extensions of man, MIT press, 1994

  82. [90]

    H. H. Wilmer, L. E. Sherman, J. M. Chein, Smartphones and cognition: A review of research exploring the links between mobile technology habits and cognitive functioning, Frontiers in psychology 8 (2017) 605

  83. [91]

    Carr, The shallows: What the Internet is doing to our brains, WW Norton & Company, 2020

    N. Carr, The shallows: What the Internet is doing to our brains, WW Norton & Company, 2020

  84. [92]

    Kahneman, A

    D. Kahneman, A. Tversky, Choices, values, and frames., American psychologist 39 (1984) 341

  85. [93]

    Schuman, S

    H. Schuman, S. Presser, J. Ludwig, Context effects on survey responses to questions about abortion, Public Opinion Quarterly 45 (1981) 216–223

  86. [94]

    Tourangeau, L

    R. Tourangeau, L. J. Rips, K. Rasinski, The psychology of survey response (2000)

  87. [95]

    J. R. Busemeyer, P. D. Bruza, Quantum models of cognition and decision: principles and applica- tions, second edition ed., Cambridge University Press, Cambridge, United Kingdom ; New York, NY, USA, 2025

  88. [96]

    J. R. Busemeyer, E. M. Pothos, R. Franco, J. S. Trueblood, A quantum theoretical explanation for probability judgment errors., Psychological review 118 (2011) 193

  89. [97]

    D. Ragland, Quantum cognition: Bridging quantum mechan- ics and cognitive science, https://medium.com/@david.a.ragland/ quantum-cognition-bridging-quantum-mechanics-and-cognitive-science-5f5a07ea2724,

  90. [98]

    Accessed: 2025-05-20

  91. [99]

    I. S. Maksymov, G. Pogrebna, The Physics of Preference: Unravelling Imprecision of Human Preferences through Magnetisation Dynamics, Information 15 (2024) 413. URL: https://www. mdpi.com/2078-2489/15/7/413. doi:10.3390/info15070413

  92. [100]

    De Tarde, The laws of imitation, H

    G. De Tarde, The laws of imitation, H. Holt, 1903

  93. [101]

    M. S. Granovetter, The strength of weak ties, American journal of sociology 78 (1973) 1360–1380

  94. [102]

    Shirado, N

    H. Shirado, N. A. Christakis, Network engineering using autonomous agents increases cooperation in human groups, Iscience 23 (2020)

  95. [103]

    Bakshy, S

    E. Bakshy, S. Messing, L. A. Adamic, Exposure to ideologically diverse news and opinion on Facebook, Science 348 (2015) 1130–1132. URL: https://www.science.org/doi/10.1126/science. aaa1160. doi:10.1126/science.aaa1160

  96. [104]

    C. R. Sunstein, #Republic: divided democracy in the age of social media, paperback edition ed., Princeton University Press, Princeton, 2018

  97. [105]

    Pentland, Social physics: How good ideas spread-the lessons from a new science, Penguin, 2014

    A. Pentland, Social physics: How good ideas spread-the lessons from a new science, Penguin, 2014

  98. [106]

    Huszár, S

    F. Huszár, S. I. Ktena, C. O’Brien, L. Belli, A. Schlaikjer, M. Hardt, Algorithmic amplification of politics on Twitter, Proceedings of the national academy of sciences 119 (2022) e2025334119

  99. [107]

    A. D. I. Kramer, J. E. Guillory, J. T. Hancock, Experimental evidence of massive-scale emotional contagion through social networks, Proceedings of the National Academy of Sciences 111 (2014) 8788–8790. URL: https://www.pnas.org/doi/10.1073/pnas.1320040111. doi:10.1073/pnas. 1320040111

  100. [108]

    Allcott, L

    H. Allcott, L. Braghieri, S. Eichmeyer, M. Gentzkow, The Welfare Effects of Social Media, American Economic Review 110 (2020) 629–676. URL: https://www.aeaweb.org/articles?id=10. 1257/aer.20190658. doi:10.1257/aer.20190658

  101. [109]

    A. J. Chaney, B. M. Stewart, B. E. Engelhardt, How algorithmic confounding in recommendation systems increases homogeneity and decreases utility, 2018, pp. 224–232

  102. [110]

    Engeström, Expansive Learning at Work: Toward an activity theoretical reconceptualization, Journal of Education and Work 14 (2001) 133–156

    Y. Engeström, Expansive Learning at Work: Toward an activity theoretical reconceptualization, Journal of Education and Work 14 (2001) 133–156. URL: http://www.tandfonline.com/doi/abs/10. 1080/13639080020028747. doi:10.1080/13639080020028747

  103. [111]

    Engeström, A

    Y. Engeström, A. Sannino, From mediated actions to heterogenous coalitions: four generations of activity-theoretical studies of work and learning, Mind, culture, and activity 28 (2021) 4–23

  104. [112]

    Aristotle, The Nicomachean Ethics.(D

    M. Aristotle, The Nicomachean Ethics.(D. Ross, Trans.) (1998)

  105. [113]

    Kohlberg, The philosophy of moral development: Moral stages and the idea of justice (1981)

    L. Kohlberg, The philosophy of moral development: Moral stages and the idea of justice (1981)

  106. [114]

    Turiel, The development of social knowledge: Morality and convention, Cambridge University Press, 1983

    E. Turiel, The development of social knowledge: Morality and convention, Cambridge University Press, 1983

  107. [115]

    J. R. Rest, Moral development: Advances in research and theory (1986)

  108. [116]

    Gutmann, Democratic education (1987)

    A. Gutmann, Democratic education (1987)

  109. [117]

    S. J. Khader, Adaptive preferences and women’s empowerment, Oxford University Press, 2011

  110. [118]

    Friedman, C

    D. Friedman, C. Diem, Rational-Choice Theory, Feminist Critiques, and Gender Inequality, Theory on gender: Feminism on theory (1993) 91

  111. [119]

    I. M. Young, Justice and the Politics of Difference, Princeton university press, 1990

  112. [120]

    Volkow, M

    N. Volkow, M. Morales, The Brain on Drugs: From Reward to Addiction, Cell 162 (2015) 712–725. URL: https://linkinghub.elsevier.com/retrieve/pii/S0092867415009629. doi:10.1016/j. cell.2015.07.046

  113. [121]

    L. O. Gostin, Public health law: power, duty, restraint, volume 3, Univ of California Press, 2000

  114. [122]

    T. L. Beauchamp, J. F. Childress, Principles of biomedical ethics, Edicoes Loyola, 1994

  115. [123]

    R. H. Thaler, C. R. Sunstein, Nudge: improving decisions about health, wealth and happiness, revised edition, new international edition ed., Penguin Books, London New York Toronto Dublin Camberwell New Delhi Rosedale Johannesburg, 2009

  116. [124]

    Gabriel, Artificial intelligence, values, and alignment, Minds and machines 30 (2020) 411–437

    I. Gabriel, Artificial intelligence, values, and alignment, Minds and machines 30 (2020) 411–437

  117. [125]

    A. Y. Ng, S. Russell, Algorithms for inverse reinforcement learning., volume 1, 2000, p. 2

  118. [126]

    Hadfield-Menell, S

    D. Hadfield-Menell, S. J. Russell, P. Abbeel, A. Dragan, Cooperative inverse reinforcement learning, Advances in neural information processing systems 29 (2016)

  119. [127]

    P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, D. Amodei, Deep reinforcement learning from human preferences, Advances in neural information processing systems 30 (2017)

  120. [128]

    Russell, Human compatible: AI and the problem of control, Penguin Uk, 2019

    S. Russell, Human compatible: AI and the problem of control, Penguin Uk, 2019

  121. [129]

    A. R. Doshi, O. P. Hauser, Generative AI enhances individual creativity but reduces the collective diversity of novel content, Science Advances 10 (2024) eadn5290. URL: https://www.science.org/ doi/10.1126/sciadv.adn5290. doi:10.1126/sciadv.adn5290

  122. [130]

    Abbeel, A

    P. Abbeel, A. Y. Ng, Apprenticeship learning via inverse reinforcement learning, 2004, p. 1

  123. [131]

    D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, G. Irving, Fine- Tuning Language Models from Human Preferences, 2020. URL: http://arxiv.org/abs/1909.08593. doi:10.48550/arXiv.1909.08593, arXiv:1909.08593 [cs]

  124. [132]

    Krasheninnikov, R

    D. Krasheninnikov, R. Shah, H. v. Hoof, Combining Reward Information from Multi- ple Sources, 2021. URL: http://arxiv.org/abs/2103.12142. doi: 10.48550/arXiv.2103.12142, arXiv:2103.12142 [cs]

  125. [133]

    Ouyang, J

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welin- der, P. Christiano, J. Leike, R. Lowe, Training language models to follow instructions wi...

  126. [134]

    Amodei, C

    D. Amodei, C. Olah, J. Steinhardt, P. Christiano, J. Schulman, D. Mané, Concrete problems in AI safety, arXiv preprint arXiv:1606.06565 (2016)

  127. [135]

    Soares, B

    N. Soares, B. Fallenstein, S. Armstrong, E. Yudkowsky, Corrigibility., 2015

  128. [136]

    Christiano, B

    P. Christiano, B. Shlegeris, D. Amodei, Supervising strong learners by amplifying weak experts, arXiv preprint arXiv:1810.08575 (2018)

  129. [137]

    Irving, P

    G. Irving, P. Christiano, D. Amodei, AI safety via debate, arXiv preprint arXiv:1805.00899 (2018)

  130. [138]

    Leike, D

    J. Leike, D. Krueger, T. Everitt, M. Martic, V. Maini, S. Legg, Scalable agent alignment via reward modeling: a research direction, arXiv preprint arXiv:1811.07871 (2018)

  131. [139]

    Yudkowsky, Coherent extrapolated volition, Singularity Institute for Artificial Intelligence (2004)

    E. Yudkowsky, Coherent extrapolated volition, Singularity Institute for Artificial Intelligence (2004)

  132. [140]

    Evans, N

    O. Evans, N. D. Goodman, Learning the preferences of bounded agents, volume 6, 2015, pp. 2–1

  133. [141]

    R. Shah, N. Gundotra, P. Abbeel, A. D. Dragan, On the Feasibility of Learning, Rather than Assuming, Human Biases for Reward Inference, 2019. URL: http://arxiv.org/abs/1906.09624. doi:10. 48550/arXiv.1906.09624, arXiv:1906.09624 [cs]

  134. [142]

    Fickinger, S

    A. Fickinger, S. Zhuang, D. Hadfield-Menell, S. Russell, Multi-principal assistance games, arXiv preprint arXiv:2007.09540 (2020)

  135. [143]

    Kierans, A

    A. Kierans, A. Ghosh, H. Hazan, S. Dori-Hacohen, Quantifying misalignment between agents: Towards a sociotechnical understanding of alignment, volume 39, 2025, pp. 27365–27373

  136. [144]

    Conitzer, R

    V. Conitzer, R. Freedman, J. Heitzig, W. H. Holliday, B. M. Jacobs, N. Lambert, M. Mossé, E. Pacuit, S. Russell, H. Schoelkopf, E. Tewolde, W. S. Zwicker, Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback, 2024. URL: http://arxiv.org/abs/2404.10271...

  137. [145]

    Sorensen, J

    T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, A roadmap to pluralistic alignment, arXiv preprint arXiv:2402.05070 (2024)

  138. [146]

    S. Zhao, J. Dang, A. Grover, Group Preference Optimization: Few-Shot Alignment of Large Lan- guage Models, 2024. URL: http://arxiv.org/abs/2310.11523. doi:10.48550/arXiv.2310.11523, arXiv:2310.11523 [cs]

  139. [147]

    Srewa, T

    M. Srewa, T. Zhao, S. Elmalaki, PluralLLM: Pluralistic Alignment in LLMs via Federated Learning, in: Proceedings of the 3rd International Workshop on Human-Centered Sensing, Modeling, and Intelligent Systems, HumanSys ’25, Association for Computing Machinery, New York, NY, USA...

  140. [148]

    Srewa, T

    M. Srewa, T. Zhao, S. Elmalaki, A Systematic Evaluation of Preference Aggregation in Feder- ated RLHF for Pluralistic Alignment of LLMs, 2025. URL: https://openreview.net/forum?id= vfP16cLfH0

  141. [149]

    C. J. Bates, R. Bose, R. G. Keeney, V. A. Kazakova, Contractual AI: Toward More Aligned, Transparent, and Robust Dialogue Agents, volume 2, 2023, pp. 225–227

  142. [150]

    Levine, M

    S. Levine, M. Franklin, T. Zhi-Xuan, S. Y. Guyot, L. Wong, D. Kilov, Y. Choi, J. B. Tenenbaum, N. Goodman, S. Lazar, Resource Rational Contractualism Should Guide AI Alignment, arXiv preprint arXiv:2506.17434 (2025)

  143. [151]

    Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran- Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosui...

  144. [152]

    Huang, D

    S. Huang, D. Siddarth, L. Lovitt, T. I. Liao, E. Durmus, A. Tamkin, D. Ganguli, Collective constitutional ai: Aligning a language model with public input, 2024, pp. 1395–1417

  145. [153]

    Lazar, A

    S. Lazar, A. Nelson, AI safety on whose terms?, Science 381 (2023) 138–138

  146. [154]

    Kroll, J

    J. Kroll, J. Huey, S. Barocas, E. Felten, J. Reidenberg, D. Robinson, H. Yu, Accountable Algorithms, University of Pennsylvania Law Review 165 (2017) 633. URL: https://scholarship.law.upenn.edu/ penn_law_review/vol165/iss3/3

  147. [155]

    Edelman, T

    J. Edelman, T. Zhi-Xuan, R. Lowe, O. Klingefjord, V. Wang-Mascianica, M. Franklin, R. O. Kearns, E. Hain, A. Sarkar, M. Bakker, F. Barez, D. Duvenaud, J. Foerster, I. Gabriel, J. Gubbels, B. Good- man, A. Haupt, J. Heitzig, J. Jara-Ettinger, A. Kasirzadeh, J. R. Kirkpatrick, A...

  148. [156]

    J. R. Anthis, D. Asmar, K. R. Driggs-Campbell, A. Hardy, K. J. Meimandi, G. Keeling, M. Kochen- derfer, H. Liu, S. Liu, R. Martín-Martín, A. A. Rushdi, M. R. Schlichting, P. Stone, H. Sub- ramonyam, D. Yang, ICLR 2025 Workshop on Human-AI Coevolution, 2024. URL: https: //openr...

  149. [157]

    J. Li, T. Song, B. Xue, Y.-C. Lee, We Shape AI, and Thereafter AI Shape Us: Humans Align with AI through Social Influences, 2025. URL: https://openreview.net/forum?id=64rCWVC78p

  150. [158]

    Tanguy, R

    C. Tanguy, R. Janssens, T. Belpaeme, J. Dambre, Human Alignment: How Much Do We Adapt to LLMs?, in: Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), Association for Computational Linguistics, Vienna, Austria, 202...

  151. [159]

    J. Li, Y. Yang, Q. V. Liao, J. Zhang, Y.-C. Lee, As Confidence Aligns: Understanding the Effect of AI Confidence on Human Self-confidence in Human-AI Decision Making, in: Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, CHI ’25, Association for Com...

  152. [160]

    Mushkani, H

    R. Mushkani, H. Berard, S. Koseki, Negotiative Alignment: Embracing Disagreement to Achieve Fairer Outcomes – Insights from Urban Studies, 2025. URL: http://arxiv.org/abs/2503.12613. doi:10.48550/arXiv.2503.12613, arXiv:2503.12613 [cs]

  153. [161]

    Carroll, A

    M. Carroll, A. Foote, K. Feng, M. Williams, A. Dragan, W. B. Knox, S. Milli, CTRL-Rec: Control- ling Recommender Systems With Natural Language, 2025. URL: http://arxiv.org/abs/2510.12742. doi:10.48550/arXiv.2510.12742, arXiv:2510.12742 [cs]

  154. [162]

    Carroll, D

    M. Carroll, D. Foote, A. Siththaranjan, S. Russell, A. Dragan, Ai alignment with changing and influenceable reward functions, arXiv preprint arXiv:2405.17713 (2024)

  155. [163]

    C. Tran, K. Fasiang, M. Kanwal, E. O’Rourke, Starting from scratch again and again: Tracing the origins of high schoolers’ negative perceptions of block-based programming, in: Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26), Association f...

  156. [164]

    Salge, C

    C. Salge, C. Glackin, D. Polani, Empowerment – an introduction, 2013. URL: https://arxiv.org/ abs/1310.1863.arXiv:1310.1863

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.