Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read RLHF preference datasets are value-skewed, an audit claims: information-utility values dominate while prosocial and democratic values trail far behind.

desk verdict A credible first map of value distributions in RLHF data, but the headline cross-dataset percentages rest on an unvalidated classifier transfer. read the letter →

arxiv 2411.11937 v1 pith:NDMPVL2L submitted 2024-11-18 cs.LG cs.AI

classification cs.LGcs.AI
keywords humanvaluesRLHFdatasetsvalueaudittaxonomypreferencelearningAIalignmentdatasettransparencyRoBERTaclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Value Imprint is a two-phase technique for auditing which human values are actually embedded in the preference data used to fine-tune large language models with reinforcement learning from human feedback. The authors build a seven-category taxonomy of human values, hand-label 6,501 preference pairs from the Anthropic hh-rlhf dataset, and train a RoBERTa classifier to label the full Anthropic, OpenAI WebGPT, and Alpaca GPT-4 preference collections. Their central finding is that information-utility values, Wisdom/Knowledge and Information Seeking, dominate all three datasets, while prosocial and democratic values, especially Justice and Human/Animal Rights, are the least represented. The paper argues the framework is viable as an auditing tool, reporting roughly 80 percent classification accuracy and 84 percent agreement with human judgment on 500 classified examples. If correct, the work gives researchers a concrete way to see the value orientation of RLHF data before models are trained on it.

What carries the argument

The load-bearing mechanism is the two-phase Value Imprint pipeline. Phase one is a hierarchical human-values taxonomy: seven high-level categories (Information Seeking, Wisdom/Knowledge, Duty & Accountability, Civility & Tolerance, Empathy & Helpfulness, Well-being & Peace, Justice/Human & Animal Rights) derived from an integrated review of philosophy, axiology, and STS literature, with sub-values linked by hypernym-hyponym relations. Phase two is a transformer-based classifier: five researchers annotated 6,501 preference pairs using the taxonomy as a codebook, reaching an inter-annotator agreement of 0.85 (Krippendorff's alpha), and those labels trained a RoBERTa sequence classifier with weighted cross-entropy and class weights. That classifier is what produces the value distribution counts that form the paper's evidence, so the taxonomy's categories and the annotation quality are what carry the argument.

What would settle it

Have independent annotators label a random sample of a few hundred preferences from OpenAI WebGPT and Alpaca GPT-4-LLM using the paper's taxonomy, and compare the resulting value shares with the model's predictions; if the human-labeled Justice share is not near 0.04 percent for WebGPT, the central imbalance claim would fail as a measurement of those datasets.

Watch

Extended reading notes

Core claim

The paper's central claim is that RLHF preference datasets are not value-neutral: they operationalize a skewed distribution of human values, and the skew is systematic across three widely used datasets. Using a taxonomy of seven value families, the authors found that in the 6,501 ground-truth preferences from Anthropic hh-rlhf, Information Seeking (36.96%) and Wisdom/Knowledge (30.75%) were the dominant values, whereas Civility & Tolerance, Empathy & Helpfulness, Well-being & Peace, and Justice/Human & Animal Rights were each below 8 percent, with Justice at 3.12 percent. The classifier extended this pattern: Wisdom/Knowledge was the most common predicted value in all three datasets (78.17% of OpenAI WebGPT, 66.56% of Alpaca GPT-4-LLM, 33.84% of Anthropic chosen and 33.71% of Anthropic rejected preferences), while Justice & Human/Animal Rights was the least represented (0.04%, 0.17%, 1.76%, and 1.76%, respectively). The paper also reports that some 'chosen' responses in the Anthropic data contain unethical content, which it reads as evidence of the need for auditing before reward-model training.

Load-bearing premise

The model trained on 6,501 Anthropic preference pairs assigns accurate value labels to the OpenAI WebGPT and Alpaca GPT-4 datasets even though those datasets were collected in different formats with different domains and likely different label distributions.

Editorial extensions

If this is right

  • Any researcher can apply the released ground-truth labels and classified datasets to audit an RLHF corpus before training a reward model.
  • Models trained on these datasets are likely to be better calibrated for information-retrieval requests than for scenarios that require justice reasoning, empathy, or well-being support.
  • The presence of unethical chosen responses in the Anthropic data implies that preference datasets can encode harmful affordances even when annotators selected them as preferable, and audits can surface those cases.
  • Domain-specific value thresholds, such as requiring a medical LLM to reason about medical ethics, become measurable targets rather than vague aspirations.
  • The reported 80 percent accuracy and 84 percent human-agreement figures, if reproducible, make value auditing a practical complement to existing dataset documentation practices.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The headline cross-dataset percentages assume the classifier transfers from Anthropic-style pairs to WebGPT and Alpaca formats; a more direct test would be to relabel samples from those two datasets and compare the resulting distributions.
  • If the value imbalances are real, they may partially explain reward hacking and sycophancy failures: a reward model trained on mostly information-seeking preferences has little incentive to develop justice or well-being reasoning.
  • A natural extension is to use the same pipeline with a culturally adapted taxonomy on non-Western preference data, which the authors explicitly note their Western-oriented taxonomy is not built for.
  • The framework could be combined with post-training interventions: use audits to construct balanced or value-targeted preference sets, then measure whether classifier-visible value distributions shift downstream model behavior.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces Value Imprint, a two-phase framework for auditing the human values embedded in RLHF datasets. The authors first construct a seven-category human-values taxonomy from a literature review in philosophy, axiology, and STS, then use it to annotate 6,501 preferences from Anthropic's hh-rlhf dataset, reporting a Krippendorff alpha of 0.85. They then train a RoBERTa classifier on these labels and apply it to three datasets: Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM. The headline finding is that information-utility values (Wisdom/Knowledge and Information Seeking) dominate all three datasets, while prosocial and democratic values (Well-being, Justice, Human/Animal Rights) are the least represented. The authors contribute their taxonomy, ground-truth annotations, and classification outputs via a GitHub repository.

Significance. The paper addresses a genuine and under-studied problem: how to make the value content of RLHF datasets auditable. The annotation phase is a strength: the inter-annotator agreement of 0.85 is encouraging, and releasing the taxonomy and datasets is a useful contribution to the community. The proposed framework, if validated, would give researchers a practical tool for interrogating the value orientations of preference datasets. The empirical finding that RLHF corpora skew toward information-utility values and away from civic and prosocial values is potentially important for alignment research. However, the current evidence does not yet support the full strength of the three-dataset claim, because the classifier trained on Anthropic data is applied to WebGPT and Alpaca without a demonstrated validation of that transfer. The central framework is defensible, but the central comparative claim needs additional verification.

major comments (3)
  1. [Sections 3.4.2 and 4.2] The headline cross-dataset percentages are not independently verified. The RoBERTa classifier is trained exclusively on 6,501 annotated preferences from Anthropic hh-rlhf (Section 3.3) and then applied to OpenAI WebGPT Comparisons and Alpaca GPT-4-LLM without any held-out validation on those datasets or any reported domain adaptation. The only external check is a 500-item human evaluation reported as a single 84% agreement figure in Section 3.4.2, but the paper does not report how those 500 items were selected, whether they cover all three datasets, the per-dataset agreement, or the label distribution of the sample. Given the format differences (Human:/Assistant: dialogs versus question/answer_0 versus instruction/output) and the large shift for Wisdom/Knowledge (78.17% in WebGPT versus about 30% in the Anthropic training distribution), the classifier may be reproducing training-distribution priors and surface dialog structure rather than measuring the target datasets. The authors should provide per-dataset human evaluation, confusion matrices, or a classifier trained and evaluated on held-out data from each target dataset, or they should explicitly restrict the headline claim to the Anthropic dataset.
  2. [Section 3.1 and Section 4.2] Alpaca GPT-4-LLM is not an RLHF preference dataset in the same sense as hh-rlhf or WebGPT Comparisons. It contains GPT-4-generated instruction-output pairs, not human pairwise preference judgments. Treating it as one of 'all three RLHF datasets' conflates distinct data-generating processes and weakens the comparative claim. The authors should either re-scope the claim to 'instruction-tuning and RLHF datasets' or provide a clear justification for why the Alpaca corpus is included in an audit of RLHF preferences, and adjust the abstract and conclusions accordingly.
  3. [Sections 3.4.2 and 4.1.2] The evaluation of the classifier is underspecified. Section 3.4.2 reports an 'accuracy score range of 80%' and Section 4.1.2 reports F1 scores for selected classes (e.g., Empathy & Helpfulness 0.629, Well-being & Peace 0.649), but the paper does not report the test-set accuracy with confidence intervals, the macro- or weighted-average F1, a confusion matrix, or per-class support. Because the aggregate percentages in Section 4.2 are computed from classifier outputs, the uncertainty in those aggregates should be quantified. The 500-item human evaluation also needs a detailed protocol, including selection procedure, number of annotators, agreement metric, and per-dataset results, before '84%' can be interpreted as evidence of transferability.
minor comments (6)
  1. [Section 3.3] The description of the ground-truth annotation should state explicitly how the 6,501 preferences were sampled from Anthropic hh-rlhf (e.g., random, stratified, or other), and whether each annotated unit is a single response, a prompt-response pair, or a chosen/rejected pair. This information affects the interpretation of the label distribution and the classifier training.
  2. [Section 3.1] The WebGPT analysis uses only the answer_0 column and drops the comparison structure that defines the dataset. The authors should justify this choice and clarify whether answer_0 is the preferred answer, the first answer, or something else, since this affects what 'human values embedded in WebGPT' means.
  3. [Figure 3] The heatmap in Figure 3 lacks axis labels, a color scale, and a legend, so the reader cannot read the quantitative comparisons that the text reports. The figure should be self-contained or be replaced by a table.
  4. [Appendix E] The sentence 'It contains 169,352 per row. resulting in a combined 338,704 if treated independently' is incomplete and inconsistent with the train/test counts reported in Section 3.1. Please correct the wording and the arithmetic.
  5. [References] A few reference entries have inconsistent journal names (e.g., 'Nous' versus 'Noûs' in entries [69], [134], [138], and [145]). Please unify and verify the bibliographic details.
  6. [Appendix D] The limitation in Appendix D that preferences often embody multiple values and that the model assigns a single dominant-value label is important, but it is not carried into the abstract or conclusions. The statements there present the value distributions as definitive facts rather than as dominant-value interpretations, which overstates the precision of the audit.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the findings are empirical measurements from human annotation and a held-out-evaluated classifier; cross-dataset transfer is a validity risk, not a circularity.

full rationale

The paper's derivation chain is empirical rather than definitional. The human-values taxonomy is constructed from a literature review (Section 3.2, Appendix B), the 6,501 ground-truth labels are produced by qualitative human annotation with a reported inter-annotator agreement of 0.85 Krippendorff's Alpha (Section 3.3), and the RoBERTa classifier is evaluated on a held-out 20% test split with reported accuracy and per-class F1 scores (Section 3.4.2, Section 4.1.2). The dominant-value percentages across the three datasets are model outputs on data, not quantities fitted into existence by the framework. The cross-dataset application to WebGPT Comparisons and Alpaca GPT-4-LLM is a genuine external-validity concern because the classifier was trained only on Anthropic hh-rlhf and no per-dataset validation is reported; however, this is a robustness and generalization threat, not circularity, because the WebGPT and Alpaca predictions are not equal to the training labels by construction. The human evaluation of 500 classification outputs provides an independent check, though its sampling details are underreported. Appendix D candidly concedes that preferences can embody multiple values and that determining the dominant value is subjective, which further supports treating the cross-dataset percentages as approximate rather than as definitionally forced. The self-citations in the paper ([14], [15], [25], [26], [40]) appear only in related-work or background context and are not load-bearing for the central claim; no uniqueness theorem is imported, no ansatz is smuggled in via self-citation, and no fitted parameter is renamed as a prediction. Therefore the paper does not exhibit significant circularity.

Assumptions & free parameters 1 free parameters · 4 assumptions · 1 invented entities

The central measurement depends on the taxonomy's coverage and the single-label annotation assumption, on the representativeness of the 6,501 sampled preferences, and on the transfer of a classifier trained on Anthropic data to other datasets. The taxonomy categories are a constructed codebook without external validation, which is the largest conceptual burden.

free parameters (1)
  • Classifier hyperparameters = batch size 64, max sequence length 128, 8 epochs, early stopping patience 2
    Reported in Section 3.4.2. They are chosen training settings that affect model quality but are not fitted to derive the headline value percentages; they are listed for completeness.
assumptions (4)
  • domain assumption The seven-category human values taxonomy adequately captures the values expressed in RLHF preferences.
    The taxonomy is built from a Western philosophical literature review and author judgment; Section 3.2 and Appendix B. The headline distribution is only as valid as the taxonomy's coverage and boundaries.
  • domain assumption Each RLHF preference can be assigned one dominant human value.
    Limitations in Appendix D acknowledge that preferences often embody multiple values and that determining the dominant value is subjective; annotation protocol in Section 3.3 forces a single label.
  • domain assumption RoBERTa's pretrained representations transfer to the value classification task across RLHF datasets.
    Invoked in Section 3.4.2 when the model trained on Anthropic annotations is applied to WebGPT and Alpaca without domain adaptation.
  • domain assumption The sampled 6,501 preferences are representative of the Anthropic hh-rlhf dataset.
    Appendix E.2 states the ground truth was curated through random sampling, but the sampling strategy and seed are not documented; class imbalance may reflect sampling or annotation choices.
invented entities (1)
  • Seven-category human values taxonomy
    purpose: Serves as the annotation codebook and defines the label space for the classifier; all measured value distributions are expressed in these categories.
    The taxonomy is derived from the authors' literature review and thematic analysis (Appendix B) and is not validated against an external benchmark or an independent value survey. The boundary between Information Seeking and Wisdom/Knowledge is particularly judgment-dependent and directly drives the main finding.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets." pith.science (2026). https://pith.science/paper/NDMPVL2L

@misc{pith2026241111937,
  author       = {Pith},
  title        = {Pith review of: Value Imprint: A Technique for Auditing the Human Values Embedded in RLHF Datasets},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NDMPVL2L}},
  note         = {Machine review of arXiv:2411.11937}
}
read the original abstract

LLMs are increasingly fine-tuned using RLHF datasets to align them with human preferences and values. However, very limited research has investigated which specific human values are operationalized through these datasets. In this paper, we introduce Value Imprint, a framework for auditing and classifying the human values embedded within RLHF datasets. To investigate the viability of this framework, we conducted three case study experiments by auditing the Anthropic/hh-rlhf, OpenAI WebGPT Comparisons, and Alpaca GPT-4-LLM datasets to examine the human values embedded within them. Our analysis involved a two-phase process. During the first phase, we developed a taxonomy of human values through an integrated review of prior works from philosophy, axiology, and ethics. Then, we applied this taxonomy to annotate 6,501 RLHF preferences. During the second phase, we employed the labels generated from the annotation as ground truth data for training a transformer-based machine learning model to audit and classify the three RLHF datasets. Through this approach, we discovered that information-utility values, including Wisdom/Knowledge and Information Seeking, were the most dominant human values within all three RLHF datasets. In contrast, prosocial and democratic values, including Well-being, Justice, and Human/Animal Rights, were the least represented human values. These findings have significant implications for developing language models that align with societal values and norms. We contribute our datasets to support further research in this area.

Figures

Figures reproduced from arXiv: 2411.11937 by the authors.

Figure 1
Figure 1. Value Imprint is a technique for auditing the human values embedded within RLHF datasets using an AI-focused human values taxonomy. to foster the measurement of the ethical judgment of language models. Birhane et al. [39] also introduced a technique for annotating the values embedded within machine learning research papers; however, did not focus on examining RLHF datasets. Obi & Gray [40] examined values engineers … view at source ↗
Figure 2
Figure 2. This image presents a visual version of the taxonomy that supported our audit. [See Table [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. This heatmap compares how the human values embedded within the three RLHF datasets [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AI Alignment at Your Discretion

    cs.AI 2025-02 conditional novelty 7.0 of 10

    The paper formalizes alignment discretion and shows empirically that annotators and models exercise substantial, often arbitrary, and mutually divergent discretion when applying alignment principles.

Reference graph

Works this paper leans on

164 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [1]

    Deep reinforcement learning from human preferences,

    P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei, “Deep reinforcement learning from human preferences,”Advances in neural information processing systems, vol. 30, 2017

  2. [2]

    Training a helpful and harmless assistant with reinforcement learning from human feedback,

    Y . Bai, A. Jones, K. Ndousse, A. Askell, A. Chen, N. DasSarma, D. Drain, S. Fort, D. Ganguli, T. Henighan et al., “Training a helpful and harmless assistant with reinforcement learning from human feedback,” arXiv preprint arXiv:2204.05862, 2022

  3. [3]

    Policy shaping: Integrating human feedback with reinforcement learning,

    S. Griffith, K. Subramanian, J. Scholz, C. L. Isbell, and A. L. Thomaz, “Policy shaping: Integrating human feedback with reinforcement learning,” Advances in neural information processing systems, vol. 26, 2013

  4. [4]

    Training language models to follow instructions with human feedback,

    L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al. , “Training language models to follow instructions with human feedback,” Advances in neural information processing systems, vol. 35, pp. 27 730–27 744, 2022

  5. [5]

    Constitutional ai: Harmlessness from ai feedback,

    Y . Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirho- seini, C. McKinnon et al., “Constitutional ai: Harmlessness from ai feedback,” arXiv preprint arXiv:2212.08073, 2022

  6. [6]

    Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback,

    T. Yu, Y . Yao, H. Zhang, T. He, Y . Han, G. Cui, J. Hu, Z. Liu, H.-T. Zheng, M. Sunet al., “Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 13 807–13 816

  7. [7]

    Aligning large multimodal models with factually augmented rlhf,

    Z. Sun, S. Shen, S. Cao, H. Liu, C. Li, Y . Shen, C. Gan, L.-Y . Gui, Y .-X. Wang, Y . Yang et al., “Aligning large multimodal models with factually augmented rlhf,” arXiv preprint arXiv:2309.14525, 2023

  8. [8]

    The silence of the llms: Cross-lingual analysis of political bias and false information prevalence in chatgpt, google bard, and bing chat,

    A. Urman and M. Makhortykh, “The silence of the llms: Cross-lingual analysis of political bias and false information prevalence in chatgpt, google bard, and bing chat,” 2023

Show all 164 references
  1. [9]

    Chatgpt: Jack of all trades, master of none,

    J. Koco ´n, I. Cichecki, O. Kaszyca, M. Kochanek, D. Szydło, J. Baran, J. Bielaniewicz, M. Gruza, A. Janz, K. Kanclerzet al., “Chatgpt: Jack of all trades, master of none,”Information Fusion, p. 101861, 2023

  2. [10]

    Process for adapting language models to society (palms) with values-targeted datasets,

    I. Solaiman and C. Dennison, “Process for adapting language models to society (palms) with values-targeted datasets,” Advances in Neural Information Processing Systems, vol. 34, pp. 5861–5873, 2021

  3. [11]

    Embedding democratic values into social media ais via societal objective functions,

    C. Jia, M. S. Lam, M. C. Mai, J. Hancock, and M. S. Bernstein, “Embedding democratic values into social media ais via societal objective functions,” arXiv preprint arXiv:2307.13912, 2023

  4. [12]

    Building human values into recommender systems: An interdisciplinary synthesis,

    J. Stray, A. Halevy, P. Assar, D. Hadfield-Menell, C. Boutilier, A. Ashar, C. Bakalar, L. Beattie, M. Ekstrand, C. Leibowicz et al., “Building human values into recommender systems: An interdisciplinary synthesis,” ACM Transactions on Recommender Systems, 2022

  5. [13]

    Embedding societal values into social media algorithms,

    M. Bernstein, A. Christin, J. Hancock, T. Hashimoto, C. Jia, M. Lam, N. Meister, N. Persily, T. Piccardi, M. Saveski et al., “Embedding societal values into social media algorithms,” Journal of Online Trust and Safety, vol. 2, no. 1, 2023. 11

  6. [14]

    Let’s talk about socio-technical angst: Tracing the history and evolution of dark patterns on twitter from 2010-2021,

    I. Obi, C. M. Gray, S. S. Chivukula, J.-N. Duane, J. Johns, M. Will, Z. Li, and T. Carlock, “Let’s talk about socio-technical angst: Tracing the history and evolution of dark patterns on twitter from 2010-2021,” arXiv preprint arXiv:2207.10563, 2022

  7. [15]

    Speculative vulnerabil- ity: Uncovering the temporalities of vulnerability in people’s experiences of the pandemic,

    J. S. Seberger, I. Obi, M. Loukil, W. Liao, D. J. Wild, and S. Patil, “Speculative vulnerabil- ity: Uncovering the temporalities of vulnerability in people’s experiences of the pandemic,” Proceedings of the ACM on Human-Computer Interaction, vol. 6, no. CSCW2, pp. 1–27, 2022

  8. [16]

    Aligning ai with shared human values,

    D. Hendrycks, C. Burns, S. Basart, A. Critch, J. Li, D. Song, and J. Steinhardt, “Aligning ai with shared human values,” arXiv preprint arXiv:2008.02275, 2020

  9. [17]

    Learning norms from stories: A prior for value aligned agents,

    M. S. A. Nahian, S. Frazier, M. Riedl, and B. Harrison, “Learning norms from stories: A prior for value aligned agents,” in Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 2020, pp. 124–130

  10. [18]

    Aligning to social norms and values in interactive narratives,

    P. Ammanabrolu, L. Jiang, M. Sap, H. Hajishirzi, and Y . Choi, “Aligning to social norms and values in interactive narratives,”arXiv preprint arXiv:2205.01975, 2022

  11. [19]

    Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties,

    T. Sorensen, L. Jiang, J. Hwang, S. Levine, V . Pyatkin, P. West, N. Dziri, X. Lu, K. Rao, C. Bhagavatula et al., “Value kaleidoscope: Engaging ai with pluralistic human values, rights, and duties,” arXiv preprint arXiv:2309.00779, 2023

  12. [20]

    Reflexive design for fairness and other human values in formal models,

    B. Fish and L. Stark, “Reflexive design for fairness and other human values in formal models,” in Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 2021, pp. 89–99

  13. [21]

    A human rights-based approach to responsible ai,

    V . Prabhakaran, M. Mitchell, T. Gebru, and I. Gabriel, “A human rights-based approach to responsible ai,” arXiv preprint arXiv:2210.02667, 2022

  14. [22]

    Making intelligence: Ethical values in iq and ml bench- marks,

    B. Blili-Hamelin and L. Hancox-Li, “Making intelligence: Ethical values in iq and ml bench- marks,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Trans- parency, 2023, pp. 271–284

  15. [23]

    Co-designing checklists to understand organizational challenges and opportunities around fairness in ai,

    M. A. Madaio, L. Stark, J. Wortman Vaughan, and H. Wallach, “Co-designing checklists to understand organizational challenges and opportunities around fairness in ai,” in Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–14

  16. [24]

    Value sensitive design and infor- mation systems,

    B. Friedman, P. H. Kahn, A. Borning, and A. Huldtgren, “Value sensitive design and infor- mation systems,” Early engagement and new technologies: Opening up the laboratory , pp. 55–95, 2013

  17. [25]

    Building an ethics-focused action plan: Roles, process moves, and trajectories,

    C. M. Gray, I. Obi, S. S. Chivukula, Z. Li, T. Carlock, M. Will, A. C. Pivonka, J. Johns, B. Rigsbee, A. R. Menon et al., “Building an ethics-focused action plan: Roles, process moves, and trajectories,” 2024

  18. [26]

    Co-designing ethical supports for technology practitioners,

    Z. Li, I. Obi, S. S. Chivukula, M. Will, J. Johns, A. C. Pivonka, T. Carlock, A. R. Menon, A. Bharadwaj, and C. M. Gray, “Co-designing ethical supports for technology practitioners,” in 2023 IEEE International Symposium on Ethics in Engineering, Science, and Technology (ETHICS...

  19. [27]

    Integrating quantitative and qualitative reasoning for value alignment,

    J. Szabo, J. M. Such, N. Criado, and S. Modgil, “Integrating quantitative and qualitative reasoning for value alignment,” in European Conference on Multi-Agent Systems. Springer, 2022, pp. 383–402

  20. [28]

    A survey on bias and fairness in machine learning,

    N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM computing surveys (CSUR), vol. 54, no. 6, pp. 1–35, 2021

  21. [29]

    Barocas, M

    S. Barocas, M. Hardt, and A. Narayanan, Fairness and machine learning: Limitations and opportunities. MIT press, 2023

  22. [30]

    Unsolved problems in ml safety,

    D. Hendrycks, N. Carlini, J. Schulman, and J. Steinhardt, “Unsolved problems in ml safety,” arXiv preprint arXiv:2109.13916, 2021

  23. [31]

    Gender and racial bias in visual question answering datasets,

    Y . Hirota, Y . Nakashima, and N. Garcia, “Gender and racial bias in visual question answering datasets,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022, pp. 1280–1292. 12

  24. [32]

    Data and its (dis) contents: A survey of dataset development and use in machine learning research,

    A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis) contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11, 2021

  25. [33]

    A hunt for the snark: Annotator diversity in data practices,

    S. Kapania, A. S. Taylor, and D. Wang, “A hunt for the snark: Annotator diversity in data practices,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 2023, pp. 1–15

  26. [34]

    Uncurated image-text datasets: Shedding light on demographic bias,

    N. Garcia, Y . Hirota, Y . Wu, and Y . Nakashima, “Uncurated image-text datasets: Shedding light on demographic bias,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 6957–6966

  27. [35]

    Bold: Dataset and metrics for measuring biases in open-ended language generation,

    J. Dhamala, T. Sun, V . Kumar, S. Krishna, Y . Pruksachatkun, K.-W. Chang, and R. Gupta, “Bold: Dataset and metrics for measuring biases in open-ended language generation,” in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 2021, pp. 862–872

  28. [36]

    Augmented datasheets for speech datasets and ethical decision-making,

    O. Papakyriakopoulos, A. S. G. Choi, W. Thong, D. Zhao, J. Andrews, R. Bourke, A. Xiang, and A. Koenecke, “Augmented datasheets for speech datasets and ethical decision-making,” in Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency, 2023, pp. 881–904

  29. [37]

    Data cards: Purposeful and transparent dataset documentation for responsible ai,

    M. Pushkarna, A. Zaldivar, and O. Kjartansson, “Data cards: Purposeful and transparent dataset documentation for responsible ai,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022, pp. 1776–1826

  30. [38]

    A framework for deprecating datasets: Standardizing documentation, identification, and communication,

    A. S. Luccioni, F. Corry, H. Sridharan, M. Ananny, J. Schultz, and K. Crawford, “A framework for deprecating datasets: Standardizing documentation, identification, and communication,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022...

  31. [39]

    The values encoded in machine learning research,

    A. Birhane, P. Kalluri, D. Card, W. Agnew, R. Dotan, and M. Bao, “The values encoded in machine learning research,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, 2022, pp. 173–184

  32. [40]

    Auditing practitioner judgment for algorithmic fairness implications,

    I. Obi and C. M. Gray, “Auditing practitioner judgment for algorithmic fairness implications,” in 2023 IEEE International Symposium on Ethics in Engineering, Science, and Technology (ETHICS). IEEE, 2023, pp. 01–05

  33. [41]

    Webgpt: Browser-assisted question-answering with human feedback,

    R. Nakano, J. Hilton, S. Balaji, J. Wu, L. Ouyang, C. Kim, C. Hesse, S. Jain, V . Kosaraju, W. Saunders et al., “Webgpt: Browser-assisted question-answering with human feedback,” arXiv preprint arXiv:2112.09332, 2021

  34. [42]

    Instruction tuning with gpt-4,

    B. Peng, C. Li, P. He, M. Galley, and J. Gao, “Instruction tuning with gpt-4,”arXiv preprint arXiv:2304.03277, 2023

  35. [43]

    The significance of a life’s shape,

    D. Dorsey, “The significance of a life’s shape,” Ethics, vol. 125, no. 2, pp. 303–330, 2015

  36. [44]

    Against time bias,

    P. Greene and M. Sullivan, “Against time bias,” Ethics, vol. 125, no. 4, pp. 947–970, 2015

  37. [45]

    Welfare as success,

    S. Keller, “Welfare as success,” Noûs, vol. 43, no. 4, pp. 656–683, 2009

  38. [46]

    Climate ethics in a dark and dangerous time,

    S. M. Gardiner, “Climate ethics in a dark and dangerous time,” Ethics, vol. 127, no. 2, pp. 430–465, 2017

  39. [47]

    Well-being and value,

    J. Goldsworthy, “Well-being and value,” Utilitas, vol. 4, no. 1, pp. 1–26, 1992

  40. [48]

    Discounting, climate change, and the ecological fallacy,

    M. Rendall, “Discounting, climate change, and the ecological fallacy,”Ethics, vol. 129, no. 3, pp. 441–463, 2019

  41. [49]

    Consequentialism and collective action,

    B. Hedden, “Consequentialism and collective action,” Ethics, vol. 130, no. 4, pp. 530–554, 2020

  42. [50]

    Law and social order,

    R. Hardin, “Law and social order,” Philosophical Issues, vol. 11, pp. 61–85, 2001. 13

  43. [51]

    Which desires are relevant to well-being?

    C. Heathwood, “Which desires are relevant to well-being?” Noûs, vol. 53, no. 3, pp. 664–688, 2019

  44. [52]

    Happiness, the self and human flourishing,

    D. M. Haybron, “Happiness, the self and human flourishing,”Utilitas, vol. 20, no. 1, pp. 21–49, 2008

  45. [53]

    Addiction and the self,

    H. Pickard, “Addiction and the self,” Noûs, vol. 55, no. 4, pp. 737–761, 2021

  46. [54]

    A set of solutions to parfit’s problems,

    S. Rachels, “A set of solutions to parfit’s problems,” Noûs, vol. 35, no. 2, pp. 214–238, 2001

  47. [55]

    The “we” in the “me

    B. Prainsack, “The “we” in the “me” solidarity and health care in the era of personalized medicine,” Science, Technology, & Human Values, vol. 43, no. 1, pp. 21–44, 2018

  48. [56]

    Ethics, public policy, and global warming,

    D. Jamieson, “Ethics, public policy, and global warming,”Global Bioethics, vol. 5, no. 1, pp. 31–42, 1992

  49. [57]

    “now is a time for optimism

    J. Rüppel, ““now is a time for optimism”: the politics of personalized medicine in mental health research,” Science, Technology, & Human Values, vol. 44, no. 4, pp. 581–611, 2019

  50. [58]

    The sum of averages: An egyptology-proof average view,

    K. Grill, “The sum of averages: An egyptology-proof average view,”Utilitas, vol. 35, no. 2, pp. 103–118, 2023

  51. [59]

    A fresh start for the objective-list theory of well-being,

    G. Fletcher, “A fresh start for the objective-list theory of well-being,”Utilitas, vol. 25, no. 2, pp. 206–220, 2013

  52. [60]

    Anti-natalism, pollyannaism, and asymmetry: A defence of cheery optimism,

    M. Hauskeller, “Anti-natalism, pollyannaism, and asymmetry: A defence of cheery optimism,” The Journal of Value Inquiry, vol. 56, no. 1, pp. 21–35, 2022

  53. [61]

    Utilitarianism and heuristics,

    B. Gesang, “Utilitarianism and heuristics,” The Journal of Value Inquiry, vol. 55, pp. 705–723, 2021

  54. [62]

    The expressive function of healthcare,

    J. Go, “The expressive function of healthcare,” The Journal of Ethics , vol. 27, no. 3, pp. 329–353, 2023

  55. [63]

    How to protect children? a pragmatic approach: On state interven- tion and children’s welfare,

    R. Gutwald and M. Reder, “How to protect children? a pragmatic approach: On state interven- tion and children’s welfare,”The Journal of Ethics, vol. 27, no. 1, pp. 77–95, 2023

  56. [64]

    Liberty, security, and fairness,

    G. Cullity, “Liberty, security, and fairness,”The Journal of Ethics, vol. 25, no. 2, pp. 141–159, 2021

  57. [65]

    The epistemic and informational requirements of utilitarianism,

    H. Breakey, “The epistemic and informational requirements of utilitarianism,”Utilitas, vol. 21, no. 1, pp. 72–99, 2009

  58. [66]

    Consequentialism and respect: Two strategies for justifying act utilitarianism,

    B. Eggleston, “Consequentialism and respect: Two strategies for justifying act utilitarianism,” Utilitas, vol. 32, no. 1, pp. 1–18, 2020

  59. [67]

    True and useful: On the structure of a two level normative theory,

    F. Feldman, “True and useful: On the structure of a two level normative theory,” Utilitas, vol. 24, no. 2, pp. 151–171, 2012

  60. [68]

    Consequential evaluation and practical reason,

    A. Sen, “Consequential evaluation and practical reason,”The Journal of Philosophy, vol. 97, no. 9, pp. 477–502, 2000

  61. [69]

    Multiple-act consequentialism,

    J. Mendola, “Multiple-act consequentialism,” Nous, vol. 40, no. 3, pp. 395–427, 2006

  62. [70]

    Knowing how to know,

    A. W. Branscomb, “Knowing how to know,”Science, Technology, & Human Values, vol. 6, no. 3, pp. 5–9, 1981

  63. [71]

    Reducing the burden of decision in digital democracy applications: A compara- tive analysis of six decision-making software,

    M. Deseriis, “Reducing the burden of decision in digital democracy applications: A compara- tive analysis of six decision-making software,” Science, Technology, & Human Values, vol. 48, no. 2, pp. 401–427, 2023

  64. [72]

    Why agents must claim rights: A reply,

    A. Gewirth, “Why agents must claim rights: A reply,” The Journal of Philosophy, vol. 79, no. 7, pp. 403–410, 1982

  65. [73]

    The egalitarianism of human rights,

    A. Buchanan, “The egalitarianism of human rights,” Ethics, vol. 120, no. 4, pp. 679–710, 2010. 14

  66. [74]

    Abilities and the sources of unfreedom,

    A. T. Schmidt, “Abilities and the sources of unfreedom,”Ethics, vol. 127, no. 1, pp. 179–207, 2016

  67. [75]

    Respect and the basis of equality,

    I. Carter, “Respect and the basis of equality,” Ethics, vol. 121, no. 3, pp. 538–571, 2011

  68. [76]

    The value of autonomy and autonomy of the will,

    S. Darwall, “The value of autonomy and autonomy of the will,” Ethics, vol. 116, no. 2, pp. 263–284, 2006

  69. [77]

    The justification of human rights and the basic right to justification: A reflexive approach,

    R. Forst, “The justification of human rights and the basic right to justification: A reflexive approach,” Ethics, vol. 120, no. 4, pp. 711–740, 2010

  70. [78]

    Valuing autonomy and respecting persons: Manipulation, seduction, and the basis of moral constraints,

    S. Buss, “Valuing autonomy and respecting persons: Manipulation, seduction, and the basis of moral constraints,” Ethics, vol. 115, no. 2, pp. 195–235, 2005

  71. [79]

    Autonomous action: Self-determination in the passive mode,

    ——, “Autonomous action: Self-determination in the passive mode,”Ethics, vol. 122, no. 4, pp. 647–691, 2012

  72. [80]

    Self-determination, revolution, and intervention,

    A. Buchanan, “Self-determination, revolution, and intervention,”Ethics, vol. 126, no. 2, pp. 447–473, 2016

  73. [81]

    The historical injustice problem for political liberalism,

    E. I. Kelly, “The historical injustice problem for political liberalism,”Ethics, vol. 128, no. 1, pp. 75–94, 2017

  74. [82]

    Autonomy and manipulated freedom,

    T. Kapitan, “Autonomy and manipulated freedom,”Philosophical Perspectives, vol. 14, pp. 81–103, 2000

  75. [83]

    Inequality: a complex, individualistic, and comparative notion,

    L. S. Temkin, “Inequality: a complex, individualistic, and comparative notion,” Philosophical Issues, vol. 11, pp. 327–353, 2001

  76. [84]

    The universal scope of positive duties correlative to human rights,

    M. Capriati, “The universal scope of positive duties correlative to human rights,” Utilitas, vol. 30, no. 3, pp. 355–378, 2018

  77. [85]

    Calibrating qalys to respect equality of persons,

    D. Franklin, “Calibrating qalys to respect equality of persons,” Utilitas, vol. 29, no. 1, pp. 65–87, 2017

  78. [86]

    Equality and priority,

    D. McKerlie, “Equality and priority,” Utilitas, vol. 6, no. 1, pp. 25–42, 1994

  79. [87]

    Ownership and justice for animals,

    A. Cochrane, “Ownership and justice for animals,” Utilitas, vol. 21, no. 4, pp. 424–442, 2009

  80. [88]

    Free agency,

    G. Watson, “Free agency,” in Agency And Responsiblity. Routledge, 2018, pp. 92–106

  81. [89]

    What do we want from a theory of justice?

    S. Amartya, “What do we want from a theory of justice?” in Theories of Justice. Routledge, 2017, pp. 27–50

  82. [90]

    The badness of death for sociable cattle,

    D. Story, “The badness of death for sociable cattle,”The Journal of Value Inquiry, pp. 1–20, 2023

  83. [91]

    Amnesties and forgiveness,

    P. Lenta, “Amnesties and forgiveness,” The Journal of Value Inquiry , vol. 57, no. 2, pp. 277–294, 2023

  84. [92]

    Amnesty and false beliefs,

    J. Espindola, “Amnesty and false beliefs,” The Journal of Value Inquiry, vol. 56, no. 3, pp. 431–449, 2022

  85. [93]

    Animal suffering and moral salience: A defense of kant’s indirect view,

    M. C. Altman, “Animal suffering and moral salience: A defense of kant’s indirect view,”The Journal of Value Inquiry, vol. 53, pp. 275–288, 2019

  86. [94]

    The human right to subsistence and the collective duty to aid,

    V . Igneski, “The human right to subsistence and the collective duty to aid,”The Journal of Value Inquiry, vol. 51, pp. 33–50, 2017

  87. [95]

    Quiet resistance: The value of personal defiance,

    T. Fakhoury, “Quiet resistance: The value of personal defiance,”The Journal of Ethics, vol. 25, no. 3, pp. 403–422, 2021

  88. [96]

    A lockean theory of climate justice for food security,

    A. Inoue, “A lockean theory of climate justice for food security,”The Journal of Ethics, vol. 27, no. 2, pp. 151–172, 2023. 15

  89. [97]

    John stuart mill and the conflicts of equality,

    S. O. Hansson, “John stuart mill and the conflicts of equality,”The Journal of Ethics, vol. 26, no. 3, pp. 433–453, 2022

  90. [98]

    Structural injustice and ethical consumption,

    M. Peacock, “Structural injustice and ethical consumption,” The Journal of Ethics, vol. 27, no. 2, pp. 191–210, 2023

  91. [99]

    Betraying animals,

    S. Cooke, “Betraying animals,” The Journal of Ethics, vol. 23, no. 2, pp. 183–200, 2019

  92. [100]

    Egalitarian provision of necessary medical treatment,

    R. C. Hughes, “Egalitarian provision of necessary medical treatment,” The Journal of Ethics, vol. 24, no. 1, pp. 55–78, 2020

  93. [101]

    Animal pain: What it is and why it matters,

    B. E. Rollin, “Animal pain: What it is and why it matters,” The Journal of ethics, vol. 15, pp. 425–437, 2011

  94. [102]

    Geoengineering justice: The role of recognition,

    M. Hourdequin, “Geoengineering justice: The role of recognition,” Science, Technology, & Human Values, vol. 44, no. 3, pp. 448–477, 2019

  95. [103]

    Times thirty: Access, maintenance, and justice,

    R. N. Crooks, “Times thirty: Access, maintenance, and justice,” Science, Technology, & Human Values, vol. 44, no. 1, pp. 118–142, 2019

  96. [104]

    On the emergence of science and justice,

    J. Reardon, “On the emergence of science and justice,” Science, Technology, & Human Values, vol. 38, no. 2, pp. 176–200, 2013

  97. [105]

    The source of responsibility,

    R. Clarke, “The source of responsibility,” Ethics, vol. 133, no. 2, p. 163–188, Jan 2023

  98. [106]

    Law,‘ought’, and ‘can’,

    F. Wilmot-Smith, “Law,‘ought’, and ‘can’,” Ethics, vol. 133, no. 4, pp. 529–557, 2023

  99. [107]

    Trading social visibility for economic amenability: data- based value translation on a “health and fitness platform

    C. Ochs, B. Büttner, and J. Lamla, “Trading social visibility for economic amenability: data- based value translation on a “health and fitness platform”,” Science, Technology, & Human Values, vol. 46, no. 3, pp. 480–506, 2021

  100. [108]

    How privacy rights engender direct doxastic duties,

    L. A. Munch, “How privacy rights engender direct doxastic duties,” The Journal of Value Inquiry, vol. 56, no. 4, pp. 547–562, 2022

  101. [109]

    Is technology value-neutral?

    B. Miller, “Is technology value-neutral?” Science, Technology, & Human Values, vol. 46, no. 1, pp. 53–80, 2021

  102. [110]

    Legibility as a design principle: Surfacing values in sensing technologies,

    H. Robbins, T. Stone, J. Bolte, and J. van den Hoven, “Legibility as a design principle: Surfacing values in sensing technologies,” Science, Technology, & Human Values, vol. 46, no. 5, pp. 1104–1135, 2021

  103. [111]

    Balancing uncertain risks and benefits in human subjects research,

    R. Barke, “Balancing uncertain risks and benefits in human subjects research,” Science, Technology, & Human Values, vol. 34, no. 3, pp. 337–364, 2009

  104. [112]

    Values levers: Building ethics into design,

    K. Shilton, “Values levers: Building ethics into design,”Science, Technology, & Human Values, vol. 38, no. 3, pp. 374–397, 2013

  105. [113]

    Privacy worlds: Exploring values and design in the development of the tor anonymity network,

    B. Collier and J. Stewart, “Privacy worlds: Exploring values and design in the development of the tor anonymity network,” Science, Technology, & Human Values, vol. 47, no. 5, pp. 910–936, 2022

  106. [114]

    Green design tools: building values and politics into material choices,

    A. Kokai, A. Iles, and C. M. Rosen, “Green design tools: building values and politics into material choices,” Science, Technology, & Human Values, vol. 46, no. 6, pp. 1139–1171, 2021

  107. [115]

    Materializing morality: Design ethics and technological mediation,

    P.-P. Verbeek, “Materializing morality: Design ethics and technological mediation,”Science, Technology, & Human Values, vol. 31, no. 3, pp. 361–380, 2006

  108. [116]

    Of trolleys and self-driving cars: What machine ethicists can and cannot learn from trolleyology,

    P. Königs, “Of trolleys and self-driving cars: What machine ethicists can and cannot learn from trolleyology,”Utilitas, vol. 35, no. 1, pp. 70–87, 2023

  109. [117]

    Kant and the trolley,

    S. Kahn, “Kant and the trolley,” The Journal of Value Inquiry, pp. 1–11, 2021

  110. [118]

    Employers have a duty of beneficence to design for meaningful work: a general argument and logistics warehouses as a case study,

    J. Smids, H. Berkers, P. Le Blanc, S. Rispens, and S. Nyholm, “Employers have a duty of beneficence to design for meaningful work: a general argument and logistics warehouses as a case study,”The Journal of Ethics, pp. 1–28, 2023. 16

  111. [119]

    To believe, or not to believe–that is not the (only) question: The hybrid view of privacy,

    L. Munch and J. Mainz, “To believe, or not to believe–that is not the (only) question: The hybrid view of privacy,”The Journal of Ethics, vol. 27, no. 3, pp. 245–261, 2023

  112. [120]

    Redefining ability, saving educational meritocracy,

    T. Harel Ben Shahar, “Redefining ability, saving educational meritocracy,” The Journal of Ethics, vol. 27, no. 3, pp. 263–283, 2023

  113. [121]

    Wisdom, expertise, and the application of ethics,

    D. Nelkin, “Wisdom, expertise, and the application of ethics,” Science, Technology, & Human Values, vol. 6, no. 1, pp. 16–17, 1981

  114. [122]

    Emotions and practical reason: Rethinking evaluation and motivation,

    B. W. Helm, “Emotions and practical reason: Rethinking evaluation and motivation,”Noûs, vol. 35, no. 2, pp. 190–213, 2001

  115. [123]

    Virtue and embodied skill: Refining the virtue-skill analogy,

    D. Vigani, “Virtue and embodied skill: Refining the virtue-skill analogy,”The Journal of Value Inquiry, vol. 55, pp. 251–268, 2021

  116. [124]

    Adversity, conflict, wisdom,

    M. Brady, M. Ardelt, M. Plews-Ogan, and S. Pope, “Adversity, conflict, wisdom,”The Journal of Value Inquiry, vol. 53, pp. 463–465, 2019

  117. [125]

    Why suffering is essential to wisdom,

    M. S. Brady, “Why suffering is essential to wisdom,”The Journal of Value Inquiry, vol. 53, pp. 467–469, 2019

  118. [126]

    Wisdom, suffering, and humility,

    J. Baehr, “Wisdom, suffering, and humility,”The Journal of Value Inquiry, vol. 53, no. 3, pp. 397–413, 2019

  119. [127]

    More on the more life experience model: What we have learned (so far),

    J. Glück, S. Bluck, and N. M. Weststrate, “More on the more life experience model: What we have learned (so far),” The Journal of Value Inquiry, vol. 53, pp. 349–370, 2019

  120. [128]

    Prudence and authenticity: Intrapersonal conflicts of value,

    D. O. Brink, “Prudence and authenticity: Intrapersonal conflicts of value,” The Philosophical Review, vol. 112, no. 2, pp. 215–245, 2003

  121. [129]

    Prudence for changing selves,

    K. Bykvist, “Prudence for changing selves,” Utilitas, vol. 18, no. 3, pp. 264–283, 2006

  122. [130]

    A theory of the comic as insight,

    K. Lash, “A theory of the comic as insight,” The Journal of Philosophy, vol. 45, no. 5, pp. 113–121, 1948

  123. [131]

    Knowledge and purpose,

    H. W. Johnstone, “Knowledge and purpose,”The Journal of Philosophy, vol. 47, no. 17, pp. 493–500, 1950

  124. [132]

    The intelligence of virtue and skill,

    W. Small, “The intelligence of virtue and skill,” The Journal of Value Inquiry, vol. 55, pp. 229–249, 2021

  125. [133]

    Virtues as skills, and the virtues of self-regulation,

    M. Stichter, “Virtues as skills, and the virtues of self-regulation,”The Journal of Value Inquiry, vol. 55, no. 2, pp. 355–369, 2021

  126. [134]

    Trust and the rationality of toleration,

    R. H. Dees, “Trust and the rationality of toleration,” Nous, vol. 32, no. 1, pp. 82–98, 1998

  127. [135]

    Reasonableness, intellectual modesty, and reciprocity in political justification,

    R. Leland and H. Van Wietmarschen, “Reasonableness, intellectual modesty, and reciprocity in political justification,” Ethics, vol. 122, no. 4, pp. 721–747, 2012

  128. [136]

    Respect for persons and the moral force of socially constructed norms,

    L. Valentini, “Respect for persons and the moral force of socially constructed norms,”Noûs, vol. 55, no. 2, pp. 385–408, 2021

  129. [137]

    Toleration as the balance between liberty and security,

    A. E. Galeotti and F. Liveriero, “Toleration as the balance between liberty and security,”The Journal of Ethics, vol. 25, no. 2, pp. 161–179, 2021

  130. [138]

    Shared agency and rational cooperation,

    C. McMahon, “Shared agency and rational cooperation,” Nous, vol. 39, no. 2, pp. 284–308, 2005

  131. [139]

    Racism as civic vice,

    J. Fischer, “Racism as civic vice,” Ethics, vol. 131, no. 3, pp. 539–570, 2021

  132. [140]

    Modesty as a virtue of attention,

    N. Bommarito, “Modesty as a virtue of attention,” Philosophical Review, vol. 122, no. 1, pp. 93–117, 2013

  133. [141]

    Manners as desire management,

    A. Berninger, “Manners as desire management,” The Journal of Value Inquiry, vol. 55, no. 1, pp. 155–173, 2021. 17

  134. [142]

    John stuart mill’s harm principle and free speech: expanding the notion of harm,

    M. C. Bell, “John stuart mill’s harm principle and free speech: expanding the notion of harm,” Utilitas, vol. 33, no. 2, pp. 162–179, 2021

  135. [143]

    A puzzle concerning gratitude and accountability,

    R. H. Wallace, “A puzzle concerning gratitude and accountability,” The Journal of Ethics , vol. 26, no. 3, pp. 455–480, 2022

  136. [144]

    Is there a right to respect?

    M. O. Fiocco, “Is there a right to respect?” Utilitas, vol. 24, no. 4, pp. 502–524, 2012

  137. [145]

    Schopenhauer on the content of compassion,

    C. Marshall, “Schopenhauer on the content of compassion,” Noûs, vol. 55, no. 4, pp. 782–799, 2021

  138. [146]

    Defending limits on the sacrifices we ought to make for others,

    V . Igneski, “Defending limits on the sacrifices we ought to make for others,”Utilitas, vol. 20, no. 4, pp. 424–446, 2008

  139. [147]

    The value of humanity,

    S. Buss, “The value of humanity,” Journal of Philosophy, vol. 109, no. 5, p. 341–377, 2012

  140. [148]

    The evolution of human altruism,

    P. Kitcher, “The evolution of human altruism,”The Journal of Philosophy, vol. 90, no. 10, pp. 497–516, 1993

  141. [149]

    Benevolence toward efforts,

    S. G. Smith, “Benevolence toward efforts,” The Journal of Value Inquiry, pp. 1–15, 2023

  142. [150]

    Does dyadic gratitude make sense? the lived experience and conceptual delineation of gratitude in absence of a benefactor,

    N. Hebbink, A. Schinkel, and D. de Ruyter, “Does dyadic gratitude make sense? the lived experience and conceptual delineation of gratitude in absence of a benefactor,” The Journal of Value Inquiry, pp. 1–20, 2023

  143. [151]

    Group gratitude: a taxonomy,

    J. Cockayne and G. Salter, “Group gratitude: a taxonomy,”The Journal of Value Inquiry, pp. 1–22, 2023

  144. [152]

    An existential foundation for an ethics of care in heidegger’s being and time,

    R. Stevens, “An existential foundation for an ethics of care in heidegger’s being and time,”The Journal of Ethics, vol. 26, no. 3, pp. 415–431, 2022

  145. [153]

    An overview of the schwartz theory of basic values,

    S. H. Schwartz, “An overview of the schwartz theory of basic values,” Online readings in Psychology and Culture, vol. 2, no. 1, p. 11, 2012. 18 Appendix A Data Access The datasets used for this research are hosted on GitHub. https://github.com/hv-rsrch/valueimprint. We named t...

  146. [154]

    In contrast, Schwartz’s values (e.g., Self-Direction, Stimulation, Hedonism) are broad and more focused on general human motivations and behavior

    Contextual Specificity: The values identified in our paper (e.g., Information Seeking, Wis- dom/Knowledge, Duty & Accountability) are more directly applicable to human-AI interac- tions and decision-making processes. In contrast, Schwartz’s values (e.g., Self-Direction, Stimul...

  147. [155]

    Schwartz’s theory, developed before the current AI era, does not explicitly address these technological factors

    Technological Relevance: Our framework includes values like Information Seeking and Wisdom/Knowledge, which are particularly relevant in the context of AI as information processing and knowledge generation systems. Schwartz’s theory, developed before the current AI era, does n...

  148. [156]

    Also, the Duty & Accountability value in our framework is particularly relevant to ongoing discussions about AI transparency and responsibility and preventing AI harms

    Accountability and Transparency: Our framework includes the Civility/Tolerance human value, which is helpful for content moderation and monitoring how AI systems might reshape societal norms and values. Also, the Duty & Accountability value in our framework is particularly rel...

  149. [157]

    helpful and harmless

    Operational Focus: The human values in this paper, such as Information Seeking, Em- pathy/Helpfulness, and Duty & Accountability, have a more operational focus, directly applicable to AI functionalities and behaviors. Schwartz’s value theory, while helpful in studying human so...

  150. [158]

    The end goal is to foster a being that thrives in the world

    Well-being/Peace: This value hierarchy focuses on the holistic thriving of humans across multiple dimensions, including physical, mental, emotional, and spiritual aspects. The end goal is to foster a being that thrives in the world. The sub-values within this category include ...

  151. [159]

    The emphasis here is on using information to achieve immediate outcomes

    Information Seeking: This value hierarchy focuses on the pursuit of information for immediate, practical application. The emphasis here is on using information to achieve immediate outcomes. For example, asking for directions on how to get to the airport from their current loc...

  152. [160]

    Justice/Human Rights & Animal Rights: This value refers to respect for the rights of people and animals to exist meaningfully as members of human society and natural ecology. The values within this group include human rights, animal rights, equality, impartiality, fairness, eq...

  153. [161]

    Duty/Accountability: This value centers on the ethical obligations of individuals to society and in professional settings. Some of the values within this category include non-maleficence, law-abiding, privacy, confidentiality, integrity, accountability, trustworthiness, reliab...

  154. [162]

    It involves the pursuit of knowledge for its own sake

    Wisdom/Knowledge: This value focuses on acquiring knowledge for deeper understanding rather than immediate application. It involves the pursuit of knowledge for its own sake. An example of this involves seeking to understand the processes that lead to rain formation or learnin...

  155. [163]

    Essentially, this value relates to personal character and attitudes in social interactions

    Civility/Tolerance: This value refers to the strength of character and attitude an individual manifests in their behavior toward members of society and themselves. Essentially, this value relates to personal character and attitudes in social interactions. Some of the values wi...

  156. [164]

    It involves understanding the context and plight of the human or animal to provide assistance to help them navigate that situation

    Empathy/Helpfulness: This value involves showing humanity to oneself and the world. It involves understanding the context and plight of the human or animal to provide assistance to help them navigate that situation. Some of the values within this category include benevolence, ...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.