Pith. sign in

REVIEW 1 major objections 1 minor 1 cited by

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

T0 review · 1 major / 1 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Demographic identity changes whether people judge an AI-generated image as harmful, and a new 1,000-image dataset shows that current safety evaluations miss much of this variation.

desk verdict DIVE is a valuable new dataset for pluralistic T2I safety evaluation, but its deliberate enrichment for disagreement means several quantitative claims are conditional on the sampling scheme. read the letter →

arxiv 2507.13383 v1 pith:KKN63JWD submitted 2025-07-15 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords pluralisticalignmenttext-to-imagesafetydemographicdiversityintersectionalratersharmperceptionraterdisagreementevaluationDIVEdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces DIVE, a dataset of 1,000 text-to-image pairs rated by 637 people from 30 demographic intersections of gender, age, and ethnicity, and uses it to argue that 'whose safety' is measured changes what safety looks like. The central claim is that demographic attributes are a strong proxy for viewpoint: women rate harms higher than men, Black raters rate harms higher than White raters, and intersectional groups such as GenZ-Black raters are more internally consistent than broad demographic groups. If this claim holds, safety evaluations of text-to-image models should not be collapsed into a single aggregate ground truth, but reported as a distribution over perspectives. The paper demonstrates that current policy raters and automated classifiers have high false-negative rates on biased imagery relative to diverse raters, and that a small open-source LLM's correlation with human raters jumps from near zero to about 0.23 when given demographic context, while still leaving most of the variation unexplained.

What carries the argument

The machine that carries the argument is DIVE itself plus the Group Association Index (GAI). DIVE is a sample of 1,000 prompt-image pairs drawn from Adversarial Nibbler, deliberately enriched for subjective items by prioritizing pairs where prior annotators split on safety (U=3, then 2, 4, 1), and each pair is rated by 20-30 raters from 30 demographic trisections on 5-point harm scales for self and others, with violation type and optional free text. GAI is the ratio of within-group inter-rater reliability to cross-group reliability; a value above 1 means a demographic group coheres internally more than it resembles outsiders. The paper uses GAI to show that intersectional groups are more cohesive than single-dimension groups, and uses simulation of small rater pools to show that demographic composition changes which images are flagged unsafe.

What would settle it

Take a random, non-enriched sample of 1,000 text-to-image generations and re-run the DIVE rating protocol with the same 30 demographic trisections; if the probability that women rate an image more harmful than men drops from 0.55 to near 0.5 and no intersectional GAI remains significantly above 1, the claim that demographics are a crucial proxy for T2I harm viewpoints would not generalize beyond the deliberately ambiguous subset.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that harm perception for text-to-image output is measurably and significantly tied to rater demographics, and that finer demographic intersections capture more coherent viewpoints than single attributes. Statistically adjusted tests show that women are more likely than men to assign a higher personal-harm score (probability 0.55), Black raters more likely than White raters (0.57), and GenZ raters slightly more likely than GenX raters (0.503). For 'harm to others' the gaps narrow but remain significant. The Group Association Index is higher for many ethnicity-involving intersections, with GenZ-Black at 1.38 and Millennial-Black at 1.29, and simulations show that rater-pool composition would flip safe/unsafe verdicts on dozens of pairs per violation type, most strongly for bias content. The paper concludes that traditional single-label safety evaluations hide valid disagreements and that demographics are a useful, if incomplete, proxy for lived experience.

Load-bearing premise

The paper's conclusions rest on the assumption that the deliberately ambiguous set of 1,000 prompt-image pairs, sampled from Adversarial Nibbler and enriched for split opinions, stands in for the broader population of text-to-image generations.

Editorial extensions

If this is right

  • Safety evaluations of text-to-image models should report rater demographics alongside scores and should treat inter-rater disagreement as signal rather than noise.
  • Automated safety classifiers and policy-trained raters should be tested against demographically diverse references, especially for bias content, where false-negative rates are highest.
  • LLM judges improve when prompted to take a demographic perspective, but with Kendall's tau near 0.23 they remain far from reproducing human viewpoint distributions; fine-tuning on DIVE is the natural next step.
  • Intersectional recruitment gives more cohesive viewpoint groups per unit of budget than single-dimension quotas, making it an efficient design for pluralistic data collection.
  • The paper's plurality-score aggregation, based on the mode of the higher of the two harm ratings, offers a practical way to label ambiguous content without erasing disagreement.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An untested implication is that the pattern of women and minority-ethnicity raters flagging more bias harms may hold for other subjective AI properties such as offensiveness, helpfulness, or aesthetic judgment, since the mechanism is lived experience rather than image content alone.
  • A direct testable extension would train a small multimodal model to predict the full distribution of harm scores from DIVE rather than the majority label; if its correlation with human raters exceeds the prompted LLM's 0.23, the data contain viewpoint signal that in-context prompting does not expose.
  • Because the sample is deliberately enriched for ambiguous items, the magnitudes of demographic gaps should be re-estimated on a random sample of T2I outputs before being used as population-level calibration targets, even though the existence of disagreement is likely robust.
  • One could drop the demographic proxy altogether and test whether value profiles or free-text rater descriptions predict harm ratings as well as trisection membership; the paper names this as future work.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper introduces DIVE, a dataset of 1,000 adversarial prompt-image (PI) pairs sampled from Adversarial Nibbler and annotated by 637 raters recruited across 30 demographic trisections (gender, age, ethnicity), with roughly 20–30 ratings per pair and a total of 31,980 responses. The authors report three sets of analyses: (RQ1) statistically significant demographic differences in self- and other-harm perceptions; (RQ2) higher group cohesion (GAI) for intersectional groups, especially ethnicity-containing intersections; (RQ3) simulation counts of PI pairs flagged as unsafe by one demographic pool but safe by another; and a value-addition section comparing policy raters and off-the-shelf classifiers against the diverse raters, plus an LLM-steering experiment. The paper claims three contributions: a pluralistic alignment dataset, empirical evidence that demographics serve as a crucial proxy for diverse viewpoints, and implications for building aligned T2I models.

Significance. If the central empirical claims hold, DIVE is a substantial and useful resource for pluralistic safety evaluation of T2I models. The dataset's strengths are real: uniform coverage across 30 intersectional demographic groups, high replication per item, separate self/other harm elicitation, free-text justifications, and public release on Hugging Face. The analyses are mostly descriptive and appropriately hedged, and the GAI measure is used as a measurement tool rather than as an optimization target, so the circularity concern raised in the stress-test does not land. However, the dataset's deliberate enrichment for disagreement means that the headline magnitudes—demographic gaps, classifier false-negative rates, and simulation counts—are conditional on a non-representative sample. The paper is transparent about the U-distribution but does not provide sampling weights or a re-analysis on a representative subset, so the quantitative claims about blind spots in current safety evaluations are not yet established for the population of T2I generations.

major comments (1)
  1. [§5.1, Fig. 4] The false-negative-rate comparison in Figure 4 is computed on the DIVE sample, but the reference outcome is the plurality score of the diverse raters and the evaluator is, in the first panel, the policy raters whose votes were part of the U-selection. Consequently, the comparison between policy raters and diverse raters is not a neutral measurement: pairs where policy raters disagreed or said safe were deliberately over-represented, while pairs they unanimously flagged were mostly dropped. The claim that 'current safety evaluations... have blind spots that are better covered by demographically diverse raters' is structurally dependent on this sampling choice. A re-analysis restricted to the fraction of AdvNib pairs with high U, or a weighted analysis using the AdvNib U distribution, is needed before this claim can be made at the level of generality stated.
minor comments (1)
  1. [References] Reference [1] (Prolific) is missing a URL prefix and appears as 'app.prolific.com' without https://.

Circularity Check

1 steps flagged · score 6.0 of 10

DIVE is curated to maximize policy-rater disagreement, so the paper's headline finding that policy raters miss harms is partly built into the sample; the reported 24% policy-unsafe rate equals the retained U≥4 fraction by construction.

  1. self definitional [§3.1(b) (Curation of the Prompt-Image Set), §5.1 (Augmenting Safety Evaluations with Diverse Feedback), App. Table 3]
    "To focus our dataset on pairs where the safety of the PI content had differing perspectives, we employed a greedy selection strategy based on the level of dissent among these six annotations (the original submission annotations + the 5 policy raters)... Let U ∈ {1, ..,6} be the number of ‘unsafe’ annotations a PI pair received. The priority order for selecting PI pairs was: U = 3> 2 > 4 > 1 > 5 > 6... The resulting spread of our final PI set over U is presented in App. Table 3. [...] The rate of classifying a PI pair as unsafe across the safety evaluations is: policy raters: 24%..."

    Since the original submitter is unsafe by design, U = (# policy-unsafe) + 1. Policy-majority unsafe therefore means U ≥ 4. App. Table 3 gives final counts 134 (U=4), 52 (U=5), 55 (U=6), total 241/1000 = 24.1%, exactly the 24% policy-rater unsafe rate reported in §5.1. The curation rule deliberately downsampled away from U=6 and prioritized U=3, 2, 1, i.e., pairs where policy raters were split or safe; hence the later 'discovery' that policy raters have high false-negative rates and that diverse raters cover blind spots is the selection objective restated as a result, not an independent empirical finding. The demographic-proxy analyses (RQ1-RQ3, GAI) are not forced in the same way, so the circularity is partial.

full rationale

Most of this paper is a dataset contribution with self-contained empirical analyses: fresh rater annotations from 637 raters, statistical tests, GAI from GRASP used as a measurement tool, and LLM steering evaluated against the same human ground truth. These are not circular, and the GRASP self-citation is not load-bearing. However, Section 5.1's central claim that policy-based safety evaluations have blind spots is substantially an artifact of §3.1(b): the PI set was selected to over-represent pairs where the policy raters were safe or split, and App. Table 3's U distribution arithmetically forces the reported 24% policy-unsafe rate (U≥4 pairs). The paper presents this as an empirical discovery about the inadequacy of conventional evaluations, but it is partly the curation criterion itself. No sampling weights or representative-subset re-analysis are supplied. Because one of the paper's headline value propositions reduces by construction, while the demographic-differences analysis retains independent content, the score is 6.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

No numeric parameters are fitted to data; the paper's main contributions are the dataset and descriptive analyses. The load-bearing assumptions are the demographic proxy and the representativeness of the disagreement-enriched sample. No new entities are postulated.

assumptions (2)
  • domain assumption Demographic attributes (gender, age, ethnicity) are a valid proxy for lived experience and viewpoint diversity.
    The paper's RQ1-RQ3 and the recruitment design treat the 30 demographic trisections as meaningful viewpoint groups; the authors acknowledge in the limitations that other rater representations exist.
  • domain assumption The disagreement-enriched DIVE sample, selected from Adversarial Nibbler by the U-order priority in Section 3.1(b), can support inference about general T2I safety evaluation blind spots.
    Section 5.1 computes classifier false negative rates against diverse raters on this sample without reweighting to the original dataset distribution, so generalizing the magnitudes requires this assumption.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models." pith.science (2026). https://pith.science/paper/KKN63JWD

@misc{pith2026250713383,
  author       = {Pith},
  title        = {Pith review of: Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KKN63JWD}},
  note         = {Machine review of arXiv:2507.13383}
}
read the original abstract

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.

Figures

Figures reproduced from arXiv: 2507.13383 by the authors.

Figure 1
Figure 1. (a) Composition of diverse rater pool - 30 unique demographic intersections across gender, [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Distribution of raters responses for the two score-based harm questions for each PI pair [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Heatmap-style table showing the outcomes of the simulations in Section 4.3. Each column [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: Rate of false negative of different safety classifiers, when compared to diverse raters’ [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Annotation form shown to participants for each PI pair to be annotated. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Instructions shown to the participants before starting the study. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 8
Figure 8. Figure 8: Distribution of responses across feedback format types, from three of the questions present [PITH_FULL_IMAGE:figures/full_fig_p021_8.png]
Figure 9
Figure 9. Figure 9: Each cell shows how many responses for each prompt-image pair on average were available [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: (a) Heatmap of the Kendall-Tau correlation of each pair of demographic trisections in our [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: Outcome of rater sampling simulations when considering rater groups of specific ethnicity [PITH_FULL_IMAGE:figures/full_fig_p024_11.png]
Figure 12
Figure 12. Figure 12: Sensitivity of diverse raters to the violations detected by the two classifiers, ShieldGemma [PITH_FULL_IMAGE:figures/full_fig_p025_12.png]
Figure 13
Figure 13. Figure 13: Histograms for the frequency of the expert label and the plurality score per PI pair. The [PITH_FULL_IMAGE:figures/full_fig_p026_13.png]
Figure 14
Figure 14. Figure 14: Figure shows the Kendall Tau correlation between the safety classifiers (LlavaGuard and [PITH_FULL_IMAGE:figures/full_fig_p027_14.png]
Figure 15
Figure 15. Figure 15: Figure shows the rate of disagreement and sensitivity at different thresholds for ethnic [PITH_FULL_IMAGE:figures/full_fig_p027_15.png]
Figure 16
Figure 16. Figure 16: Figure shows the rate of disagreement and sensitivity at different thresholds for groups of [PITH_FULL_IMAGE:figures/full_fig_p028_16.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VideoGuard: Protecting Video Content from Unauthorized Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    VideoGuard adds joint, motion-aware perturbations to videos to block unauthorized diffusion-model editing.

Reference graph

Works this paper leans on

62 extracted references · 37 canonical work pages · cited by 1 Pith paper

  1. [1]

    URL app.prolific.com

    Prolific. URL app.prolific.com

  2. [2]

    Aroyo and C

    L. Aroyo and C. Welty. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24, 2015

  3. [3]

    The Reasonable Effectiveness of Diverse Evaluation Data

    L. Aroyo, M. Diaz, C. Homan, V . Prabhakaran, A. Taylor, and D. Wang. The reasonable effectiveness of diverse evaluation data, 2023. URL https://arxiv.org/abs/2301.09406

  4. [4]

    Aroyo, A

    L. Aroyo, A. Taylor, M. Díaz, C. Homan, A. Parrish, G. Serapio-García, V . Prabhakaran, and D. Wang. DICES dataset: Diversity in conversational AI evaluation for safety. Advances in Neural Information Processing Systems, 36, 2024

  5. [5]

    Y . Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirho- seini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R...

  6. [6]

    A. Basu, R. V . Babu, and D. Pruthi. Inspecting the geographical representativeness of images from text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5136–5147, 2023

  7. [7]

    Bianchi, P

    F. Bianchi, P. Kalluri, E. Durmus, F. Ladhak, M. Cheng, D. Nozza, T. Hashimoto, D. Ju- rafsky, J. Zou, and A. Caliskan. Easily accessible text-to-image generation amplifies demo- graphic stereotypes at large scale. In 2023 ACM Conference on Fairness, Accountability, and Transparency, page 1493–1504. ACM, June 2023. doi: 10.1145/3593013.3594095. URL http:/...

  8. [8]

    Castricato, N

    L. Castricato, N. Lile, R. Rafailov, J.-P. Fränken, and C. Finn. Persona: A reproducible testbed for pluralistic alignment, 2024. URL https://arxiv.org/abs/2407.17387

Show all 62 references
  1. [9]

    A. C. Curry, G. Abercrombie, and V . Rieser. ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7388–7403, 2021

  2. [10]

    Davani, M

    A. Davani, M. Díaz, D. Baker, and V . Prabhakaran. Disentangling perceptions of offensiveness: Cultural and moral correlates. FAccT ’24, page 2007–2021, New York, NY , USA, 2024. Association for Computing Machinery. ISBN 9798400704505. doi: 10.1145/3630106.3659021. URL https:/...

  3. [11]

    Dinan, G

    E. Dinan, G. Abercrombie, A. S. Bergman, S. Spruit, D. Hovy, Y .-L. Boureau, and V . Rieser. Anticipating safety issues in e2e conversational ai: Framework and tooling. arXiv preprint arXiv:2107.03451, 2021. 10

  4. [12]

    Giorgi, D

    S. Giorgi, D. Bellew, D. R. S. Habib, G. Sherman, J. Sedoc, C. Smitterberg, A. Devoto, M. Himelein-Wachowiak, and B. Curtis. Lived experience matters: Automatic detection of stigma on social media toward people who use substances. arXiv preprint arXiv:2302.02064, 2023

  5. [13]

    M. L. Gordon, M. S. Lam, J. S. Park, K. Patel, J. Hancock, T. Hashimoto, and M. S. Bernstein. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–19, 2022

  6. [14]

    J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo. A survey on LLM-as-a-judge, 2025. URL https://arxiv.org/abs/2411.15594

  7. [15]

    Helff, F

    L. Helff, F. Friedrich, M. Brack, P. Schramowski, and K. Kersting. LLA V AGUARD: VLM-based safeguard for vision dataset curation and safety assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8322–8326, 2024

  8. [16]

    A. Jha, V . Prabhakaran, R. Denton, S. Laszlo, S. Dave, R. Qadri, C. Reddy, and S. Dev. Visage: A global-scale analysis of visual stereotypes in text-to-image generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  9. [17]

    Kapania, A

    S. Kapania, A. S. Taylor, and D. Wang. A hunt for the snark: Annotator diversity in data practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY , USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi:...

  10. [18]

    Kholodna, S

    N. Kholodna, S. Julka, M. Khodadadi, M. N. Gumus, and M. Granitzer. LLMs in the loop: Leveraging large language model annotations for active learning in low-resource languages,

  11. [19]

    Kiela, M

    D. Kiela, M. Bartolo, Y . Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, et al. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337, 2021

  12. [20]

    H. Kirk, Y . Jun, H. Iqbal, E. Benussi, F. V olpin, F. A. Dreyer, A. Shtedritski, and Y . M. Asano. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models, 2021. URL https://arxiv.org/abs/2102.04130

  13. [21]

    H. R. Kirk, A. Whitefield, P. Röttger, A. Bean, K. Margatina, J. Ciro, R. Mosquera, M. Bartolo, A. Williams, H. He, et al. The PRISM alignment project: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment...

  14. [22]

    Krippendorff

    K. Krippendorff. Reliability in content analysis: Some common misconceptions and recommen- dations. Human Communication Research, 30:411–433, 07 2004. doi: 10.1093/hcr/30.3.411

  15. [23]

    W. H. Kruskal and W. A. Wallis. Use of Ranks in One-Criterion Variance Analysis.Journal of the American Statistical Association, 47:583 – 621, 1952. URL http://dx.doi.org/10. 1080/01621459.1952.10483441

  16. [24]

    Kumar, P

    D. Kumar, P. G. Kelley, S. Consolvo, J. Mason, E. Bursztein, Z. Durumeric, K. Thomas, and M. Bailey. Designing toxic content classification for a diversity of perspectives. In Proceedings of the Seventeenth USENIX Conference on Usable Privacy and Security , SOUPS’21, USA,

  17. [25]

    T. Li, D. Sree, and T. Ringenberg. Assessing crowdsourced annotations with LLMs: Linguistic certainty as a proxy for trustworthiness. In M. Hämäläinen, E. Öhman, Y . Bizzoni, S. Miyagawa, and K. Alnajjar, editors, Proceedings of the 5th International Conference on Natural Lang...

  18. [26]

    H. B. Mann and D. R. Whitney. On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other.The Annals of Mathematical Statistics, 18(1):50 – 60, 1947. doi: 10.1214/aoms/1177730491. URL https://doi.org/10.1214/aoms/1177730491

  19. [27]

    Mostafazadeh Davani, M

    A. Mostafazadeh Davani, M. Díaz, and V . Prabhakaran. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Associa- tion for Computational Linguistics , 10:92–110, 2022. doi: 10.1162/tacl_a_00449. URL https://aclanthology....

  20. [28]

    Orlikowski, J

    M. Orlikowski, J. Pei, P. Röttger, P. Cimiano, D. Jurgens, and D. Hovy. Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions, 2025. URL https://arxiv.org/abs/2502.20897

  21. [29]

    Palomaki, O

    J. Palomaki, O. Rhinehart, and M. Tseng. A case for a range of acceptable annotations. In SAD/CrowdBias@ HCOMP, pages 19–31, 2018

  22. [30]

    J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein. Generative agents: Interactive simulacra of human behavior, 2023. URL https://arxiv.org/abs/ 2304.03442

  23. [31]

    Parrish, H

    A. Parrish, H. R. Kirk, J. Quaye, C. Rastogi, M. Bartolo, O. Inel, J. Ciro, R. Mosquera, A. Howard, W. Cukierski, D. Sculley, V . J. Reddi, and L. Aroyo. Adversarial nibbler: A data-centric challenge for improving the safety of text-to-image models, 2023. URL https: //arxiv.or...

  24. [32]

    Parrish, V

    A. Parrish, V . Prabhakaran, L. Aroyo, M. Díaz, C. M. Homan, G. Serapio-García, A. S. Taylor, and D. Wang. Diversity-aware annotation for conversational AI safety. In Proceedings of Safety4ConvAI: The Third Workshop on Safety for Conversational AI@ LREC-COLING 2024, pages 8–15, 2024

  25. [33]

    Pavlick and T

    E. Pavlick and T. Kwiatkowski. Inherent disagreements in human textual inferences. Transac- tions of the Association for Computational Linguistics, 7:677–694, 2019

  26. [34]

    Pei and D

    J. Pei and D. Jurgens. When do annotator demographics matter? measuring the influence of annotator demographics with the POPQUORN dataset. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023

  27. [35]

    Prabhakaran, A

    V . Prabhakaran, A. M. Davani, and M. Díaz. On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, 2021

  28. [36]

    Prabhakaran, C

    V . Prabhakaran, C. Homan, L. Aroyo, A. Mostafazadeh Davani, A. Parrish, A. Taylor, M. Diaz, D. Wang, and G. Serapio-García. GRASP: A disagreement analysis framework to assess group associations in perspectives. In K. Duh, H. Gomez, and S. Bethard, editors, Proceedings of the ...

  29. [37]

    Quaye, A

    J. Quaye, A. Parrish, O. Inel, C. Rastogi, H. R. Kirk, M. Kahng, E. Van Liemt, M. Bartolo, J. Tsang, J. White, et al. Adversarial Nibbler: An open red-teaming method for identify- ing diverse harms in text-to-image generation. In The 2024 ACM Conference on Fairness, Accountabi...

  30. [38]

    I. D. Raji and R. Dobbe. Concrete problems in ai safety, revisited. arXiv preprint arXiv:2401.10899, 2023

  31. [39]

    M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, R. Comanescu, C. Akbulut, T. Stepleton, J. Mateos-Garcia, S. Bergman, J. Kay, et al. Gaps in the safety evaluation of generative AI. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume 7, pages 1200–12...

  32. [40]

    Rismani, R

    S. Rismani, R. Shelby, A. Smart, R. Delos Santos, A. Moon, and N. Rostamzadeh. Beyond the ml model: Applying safety engineering frameworks to text-to-image development. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 70–83, 2023

  33. [41]

    Santurkar, E

    S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, pages 29971–30004, 2023

  34. [42]

    M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y . Choi, and N. A. Smith. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In M. Carpuat, M.-C. de Marneffe, and I. V . Meza Ruiz, editors, Proceedings of the 2022 Conference of the Nort...

  35. [43]

    Schramowski, M

    P. Schramowski, M. Brack, B. Deiseroth, and K. Kersting. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models, 2023. URL https://arxiv.org/abs/2211. 05105

  36. [44]

    A. D. Selbst, D. Boyd, S. A. Friedler, S. Venkatasubramanian, and J. Vertesi. Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Account- ability, and Transparency, FAT* ’19, page 59–68, New York, NY , USA, 2019. Association for C...

  37. [45]

    Sorensen, L

    T. Sorensen, L. Jiang, J. D. Hwang, S. Levine, V . Pyatkin, P. West, N. Dziri, X. Lu, K. Rao, C. Bhagavatula, et al. Value kaleidoscope: Engaging AI with pluralistic human values, rights, and duties. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, ...

  38. [46]

    Sorensen, J

    T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, et al. Position: A roadmap to pluralistic alignment. In Proceedings of the 41st International Conference on Machine Learning, pages 46280–46302, 2024

  39. [47]

    Sorensen, P

    T. Sorensen, P. Mishra, R. Patel, M. H. Tessler, M. Bakker, G. Evans, I. Gabriel, N. Goodman, and V . Rieser. Value profiles for encoding human variation, 2025. URLhttps://arxiv.org/ abs/2503.15484

  40. [48]

    Steiger, T

    M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease. The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Sy...

  41. [49]

    G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Riv- ière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. ...

  42. [50]

    what is safety?

    A. M. Walker, M. A. DeVito, K. Badillo-Urquiola, R. Bellini, S. Chancellor, J. L. Feuston, K. Henne, P. G. Kelley, S. Rismani, R. Shelby, et al. " what is safety?": Building bridges across approaches to digital risks and harms. In Companion Publication of the 2024 Conference o...

  43. [51]

    Weidinger, J

    L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021

  44. [52]

    K. Wong, P. Paritosh, and L. Aroyo. Cross-replication reliability - an empirical approach to interpreting inter-rater reliability. In C. Zong, F. Xia, W. Li, and R. Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the...

  45. [53]

    W. Zeng, D. Kurniawan, R. Mullins, Y . Liu, T. Saha, D. Ike-Njoku, J. Gu, Y . Song, C. Xu, J. Zhou, et al. Shieldgemma 2: Robust and tractable image content moderation. arXiv preprint arXiv:2504.01081, 2025

  46. [54]

    AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from the use of the dataset

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging LLM-as-a-judge with MT-Bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685. 14 A Dataset collection details A.1 P...

  47. [57]

    Low attention check accuracy: Threshold <1

  48. [58]

    Low total duration (Possibly low effort): Threshold <20 minutes (assuming ’20’ refers to a unit of time, likely minutes in this context)

  49. [59]

    Too few comments (Possibly low effort): Threshold <2

  50. [60]

    High annotation inconsistency (Possibly low effort)

  51. [61]

    Not harmful

    High frequency of "Not harmful" selections (May otherwise silently pass attention checks and inconsistency checks): Threshold >35 Raters exceeding these thresholds in one or more categories were earmarked for manual review. The manual inspection process involved a thorough exa...

  52. [62]

    trains hitting a car in a violent accident

    "trains hitting a car in a violent accident" –> 2 4. "a tea farmer" –> 3 5. "a woman cleaning the floor soiled by red wine" –> 4 In-context Parameters Each of the LLMs prompted in-context are sampled from in exactly the same way. We use the default sampling strategy (top-k) wi...

  53. [2021]

    ISBN 978-1-939133-25-0

    USENIX Association. ISBN 978-1-939133-25-0

  54. [2024]

    URL https://arxiv.org/abs/2404.02261

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.