REVIEW 1 major objections 1 minor 1 cited by
Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models
T0 review · 1 major / 1 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Demographic identity changes whether people judge an AI-generated image as harmful, and a new 1,000-image dataset shows that current safety evaluations miss much of this variation.
desk verdict DIVE is a valuable new dataset for pluralistic T2I safety evaluation, but its deliberate enrichment for disagreement means several quantitative claims are conditional on the sampling scheme. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machine that carries the argument is DIVE itself plus the Group Association Index (GAI). DIVE is a sample of 1,000 prompt-image pairs drawn from Adversarial Nibbler, deliberately enriched for subjective items by prioritizing pairs where prior annotators split on safety (U=3, then 2, 4, 1), and each pair is rated by 20-30 raters from 30 demographic trisections on 5-point harm scales for self and others, with violation type and optional free text. GAI is the ratio of within-group inter-rater reliability to cross-group reliability; a value above 1 means a demographic group coheres internally more than it resembles outsiders. The paper uses GAI to show that intersectional groups are more cohesive than single-dimension groups, and uses simulation of small rater pools to show that demographic composition changes which images are flagged unsafe.
What would settle it
Take a random, non-enriched sample of 1,000 text-to-image generations and re-run the DIVE rating protocol with the same 30 demographic trisections; if the probability that women rate an image more harmful than men drops from 0.55 to near 0.5 and no intersectional GAI remains significantly above 1, the claim that demographics are a crucial proxy for T2I harm viewpoints would not generalize beyond the deliberately ambiguous subset.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that harm perception for text-to-image output is measurably and significantly tied to rater demographics, and that finer demographic intersections capture more coherent viewpoints than single attributes. Statistically adjusted tests show that women are more likely than men to assign a higher personal-harm score (probability 0.55), Black raters more likely than White raters (0.57), and GenZ raters slightly more likely than GenX raters (0.503). For 'harm to others' the gaps narrow but remain significant. The Group Association Index is higher for many ethnicity-involving intersections, with GenZ-Black at 1.38 and Millennial-Black at 1.29, and simulations show that rater-pool composition would flip safe/unsafe verdicts on dozens of pairs per violation type, most strongly for bias content. The paper concludes that traditional single-label safety evaluations hide valid disagreements and that demographics are a useful, if incomplete, proxy for lived experience.
Load-bearing premise
The paper's conclusions rest on the assumption that the deliberately ambiguous set of 1,000 prompt-image pairs, sampled from Adversarial Nibbler and enriched for split opinions, stands in for the broader population of text-to-image generations.
Editorial extensions
If this is right
- Safety evaluations of text-to-image models should report rater demographics alongside scores and should treat inter-rater disagreement as signal rather than noise.
- Automated safety classifiers and policy-trained raters should be tested against demographically diverse references, especially for bias content, where false-negative rates are highest.
- LLM judges improve when prompted to take a demographic perspective, but with Kendall's tau near 0.23 they remain far from reproducing human viewpoint distributions; fine-tuning on DIVE is the natural next step.
- Intersectional recruitment gives more cohesive viewpoint groups per unit of budget than single-dimension quotas, making it an efficient design for pluralistic data collection.
- The paper's plurality-score aggregation, based on the mode of the higher of the two harm ratings, offers a practical way to label ambiguous content without erasing disagreement.
Reading between the lines
- An untested implication is that the pattern of women and minority-ethnicity raters flagging more bias harms may hold for other subjective AI properties such as offensiveness, helpfulness, or aesthetic judgment, since the mechanism is lived experience rather than image content alone.
- A direct testable extension would train a small multimodal model to predict the full distribution of harm scores from DIVE rather than the majority label; if its correlation with human raters exceeds the prompted LLM's 0.23, the data contain viewpoint signal that in-context prompting does not expose.
- Because the sample is deliberately enriched for ambiguous items, the magnitudes of demographic gaps should be re-estimated on a random sample of T2I outputs before being used as population-level calibration targets, even though the existence of disagreement is likely robust.
- One could drop the demographic proxy altogether and test whether value profiles or free-text rater descriptions predict harm ratings as well as trisection membership; the paper names this as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces DIVE, a dataset of 1,000 adversarial prompt-image (PI) pairs sampled from Adversarial Nibbler and annotated by 637 raters recruited across 30 demographic trisections (gender, age, ethnicity), with roughly 20–30 ratings per pair and a total of 31,980 responses. The authors report three sets of analyses: (RQ1) statistically significant demographic differences in self- and other-harm perceptions; (RQ2) higher group cohesion (GAI) for intersectional groups, especially ethnicity-containing intersections; (RQ3) simulation counts of PI pairs flagged as unsafe by one demographic pool but safe by another; and a value-addition section comparing policy raters and off-the-shelf classifiers against the diverse raters, plus an LLM-steering experiment. The paper claims three contributions: a pluralistic alignment dataset, empirical evidence that demographics serve as a crucial proxy for diverse viewpoints, and implications for building aligned T2I models.
Significance. If the central empirical claims hold, DIVE is a substantial and useful resource for pluralistic safety evaluation of T2I models. The dataset's strengths are real: uniform coverage across 30 intersectional demographic groups, high replication per item, separate self/other harm elicitation, free-text justifications, and public release on Hugging Face. The analyses are mostly descriptive and appropriately hedged, and the GAI measure is used as a measurement tool rather than as an optimization target, so the circularity concern raised in the stress-test does not land. However, the dataset's deliberate enrichment for disagreement means that the headline magnitudes—demographic gaps, classifier false-negative rates, and simulation counts—are conditional on a non-representative sample. The paper is transparent about the U-distribution but does not provide sampling weights or a re-analysis on a representative subset, so the quantitative claims about blind spots in current safety evaluations are not yet established for the population of T2I generations.
major comments (1)
- [§5.1, Fig. 4] The false-negative-rate comparison in Figure 4 is computed on the DIVE sample, but the reference outcome is the plurality score of the diverse raters and the evaluator is, in the first panel, the policy raters whose votes were part of the U-selection. Consequently, the comparison between policy raters and diverse raters is not a neutral measurement: pairs where policy raters disagreed or said safe were deliberately over-represented, while pairs they unanimously flagged were mostly dropped. The claim that 'current safety evaluations... have blind spots that are better covered by demographically diverse raters' is structurally dependent on this sampling choice. A re-analysis restricted to the fraction of AdvNib pairs with high U, or a weighted analysis using the AdvNib U distribution, is needed before this claim can be made at the level of generality stated.
minor comments (1)
- [References] Reference [1] (Prolific) is missing a URL prefix and appears as 'app.prolific.com' without https://.
Circularity Check
DIVE is curated to maximize policy-rater disagreement, so the paper's headline finding that policy raters miss harms is partly built into the sample; the reported 24% policy-unsafe rate equals the retained U≥4 fraction by construction.
-
self definitional
[§3.1(b) (Curation of the Prompt-Image Set), §5.1 (Augmenting Safety Evaluations with Diverse Feedback), App. Table 3]
"To focus our dataset on pairs where the safety of the PI content had differing perspectives, we employed a greedy selection strategy based on the level of dissent among these six annotations (the original submission annotations + the 5 policy raters)... Let U ∈ {1, ..,6} be the number of ‘unsafe’ annotations a PI pair received. The priority order for selecting PI pairs was: U = 3> 2 > 4 > 1 > 5 > 6... The resulting spread of our final PI set over U is presented in App. Table 3. [...] The rate of classifying a PI pair as unsafe across the safety evaluations is: policy raters: 24%..."
Since the original submitter is unsafe by design, U = (# policy-unsafe) + 1. Policy-majority unsafe therefore means U ≥ 4. App. Table 3 gives final counts 134 (U=4), 52 (U=5), 55 (U=6), total 241/1000 = 24.1%, exactly the 24% policy-rater unsafe rate reported in §5.1. The curation rule deliberately downsampled away from U=6 and prioritized U=3, 2, 1, i.e., pairs where policy raters were split or safe; hence the later 'discovery' that policy raters have high false-negative rates and that diverse raters cover blind spots is the selection objective restated as a result, not an independent empirical finding. The demographic-proxy analyses (RQ1-RQ3, GAI) are not forced in the same way, so the circularity is partial.
full rationale
Most of this paper is a dataset contribution with self-contained empirical analyses: fresh rater annotations from 637 raters, statistical tests, GAI from GRASP used as a measurement tool, and LLM steering evaluated against the same human ground truth. These are not circular, and the GRASP self-citation is not load-bearing. However, Section 5.1's central claim that policy-based safety evaluations have blind spots is substantially an artifact of §3.1(b): the PI set was selected to over-represent pairs where the policy raters were safe or split, and App. Table 3's U distribution arithmetically forces the reported 24% policy-unsafe rate (U≥4 pairs). The paper presents this as an empirical discovery about the inadequacy of conventional evaluations, but it is partly the curation criterion itself. No sampling weights or representative-subset re-analysis are supplied. Because one of the paper's headline value propositions reduces by construction, while the demographic-differences analysis retains independent content, the score is 6.
Assumptions & free parameters
assumptions (2)
- domain assumption Demographic attributes (gender, age, ethnicity) are a valid proxy for lived experience and viewpoint diversity.
- domain assumption The disagreement-enriched DIVE sample, selected from Adversarial Nibbler by the U-order priority in Section 3.1(b), can support inference about general T2I safety evaluation blind spots.
Cite this review
Pith. "Pith review of Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models." pith.science (2026). https://pith.science/paper/KKN63JWD
@misc{pith2026250713383,
author = {Pith},
title = {Pith review of: Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/KKN63JWD}},
note = {Machine review of arXiv:2507.13383}
}
read the original abstract
Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralistic alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enable deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.
Figures
Figures from the paper (12 more)
Forward citations
Cited by 1 Pith paper
-
VideoGuard: Protecting Video Content from Unauthorized Editing
VideoGuard adds joint, motion-aware perturbations to videos to block unauthorized diffusion-model editing.
Reference graph
Works this paper leans on
- [1]
-
[2]
L. Aroyo and C. Welty. Truth is a lie: Crowd truth and the seven myths of human annotation. AI Magazine, 36(1):15–24, 2015
work page 2015
-
[3]
The Reasonable Effectiveness of Diverse Evaluation Data
L. Aroyo, M. Diaz, C. Homan, V . Prabhakaran, A. Taylor, and D. Wang. The reasonable effectiveness of diverse evaluation data, 2023. URL https://arxiv.org/abs/2301.09406
work page Pith review arXiv 2023
- [4]
-
[5]
Y . Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirho- seini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R...
arXiv 2022
-
[6]
A. Basu, R. V . Babu, and D. Pruthi. Inspecting the geographical representativeness of images from text-to-image models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5136–5147, 2023
work page 2023
-
[7]
F. Bianchi, P. Kalluri, E. Durmus, F. Ladhak, M. Cheng, D. Nozza, T. Hashimoto, D. Ju- rafsky, J. Zou, and A. Caliskan. Easily accessible text-to-image generation amplifies demo- graphic stereotypes at large scale. In 2023 ACM Conference on Fairness, Accountability, and Transparency, page 1493–1504. ACM, June 2023. doi: 10.1145/3593013.3594095. URL http:/...
arXiv 2023
-
[8]
L. Castricato, N. Lile, R. Rafailov, J.-P. Fränken, and C. Finn. Persona: A reproducible testbed for pluralistic alignment, 2024. URL https://arxiv.org/abs/2407.17387
arXiv 2024
Show all 62 references
-
[9]
A. C. Curry, G. Abercrombie, and V . Rieser. ConvAbuse: Data, analysis, and benchmarks for nuanced abuse detection in conversational AI. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7388–7403, 2021
2021
-
[10]
Davani, M
A. Davani, M. Díaz, D. Baker, and V . Prabhakaran. Disentangling perceptions of offensiveness: Cultural and moral correlates. FAccT ’24, page 2007–2021, New York, NY , USA, 2024. Association for Computing Machinery. ISBN 9798400704505. doi: 10.1145/3630106.3659021. URL https:/...
2007
-
[11]
Dinan, G
E. Dinan, G. Abercrombie, A. S. Bergman, S. Spruit, D. Hovy, Y .-L. Boureau, and V . Rieser. Anticipating safety issues in e2e conversational ai: Framework and tooling. arXiv preprint arXiv:2107.03451, 2021. 10
2021 arXiv
-
[12]
Giorgi, D
S. Giorgi, D. Bellew, D. R. S. Habib, G. Sherman, J. Sedoc, C. Smitterberg, A. Devoto, M. Himelein-Wachowiak, and B. Curtis. Lived experience matters: Automatic detection of stigma on social media toward people who use substances. arXiv preprint arXiv:2302.02064, 2023
2023 arXiv
-
[13]
M. L. Gordon, M. S. Lam, J. S. Park, K. Patel, J. Hancock, T. Hashimoto, and M. S. Bernstein. Jury learning: Integrating dissenting voices into machine learning models. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems, pages 1–19, 2022
2022
-
[14]
J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y . Shen, S. Ma, H. Liu, S. Wang, K. Zhang, Y . Wang, W. Gao, L. Ni, and J. Guo. A survey on LLM-as-a-judge, 2025. URL https://arxiv.org/abs/2411.15594
2025 arXiv
-
[15]
Helff, F
L. Helff, F. Friedrich, M. Brack, P. Schramowski, and K. Kersting. LLA V AGUARD: VLM-based safeguard for vision dataset curation and safety assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8322–8326, 2024
2024
-
[16]
A. Jha, V . Prabhakaran, R. Denton, S. Laszlo, S. Dave, R. Qadri, C. Reddy, and S. Dev. Visage: A global-scale analysis of visual stereotypes in text-to-image generation. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...
2024
-
[17]
Kapania, A
S. Kapania, A. S. Taylor, and D. Wang. A hunt for the snark: Annotator diversity in data practices. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI ’23, New York, NY , USA, 2023. Association for Computing Machinery. ISBN 9781450394215. doi:...
2023
-
[18]
Kholodna, S
N. Kholodna, S. Julka, M. Khodadadi, M. N. Gumus, and M. Granitzer. LLMs in the loop: Leveraging large language model annotations for active learning in low-resource languages,
-
[19]
Kiela, M
D. Kiela, M. Bartolo, Y . Nie, D. Kaushik, A. Geiger, Z. Wu, B. Vidgen, G. Prasad, A. Singh, P. Ringshia, et al. Dynabench: Rethinking benchmarking in nlp. arXiv preprint arXiv:2104.14337, 2021
2021 arXiv
-
[20]
H. Kirk, Y . Jun, H. Iqbal, E. Benussi, F. V olpin, F. A. Dreyer, A. Shtedritski, and Y . M. Asano. Bias out-of-the-box: An empirical analysis of intersectional occupational biases in popular generative language models, 2021. URL https://arxiv.org/abs/2102.04130
2021 arXiv
-
[21]
H. R. Kirk, A. Whitefield, P. Röttger, A. Bean, K. Margatina, J. Ciro, R. Mosquera, M. Bartolo, A. Williams, H. He, et al. The PRISM alignment project: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment...
2024 arXiv
-
[22]
Krippendorff
K. Krippendorff. Reliability in content analysis: Some common misconceptions and recommen- dations. Human Communication Research, 30:411–433, 07 2004. doi: 10.1093/hcr/30.3.411
2004 doi
-
[23]
W. H. Kruskal and W. A. Wallis. Use of Ranks in One-Criterion Variance Analysis.Journal of the American Statistical Association, 47:583 – 621, 1952. URL http://dx.doi.org/10. 1080/01621459.1952.10483441
1952
-
[24]
Kumar, P
D. Kumar, P. G. Kelley, S. Consolvo, J. Mason, E. Bursztein, Z. Durumeric, K. Thomas, and M. Bailey. Designing toxic content classification for a diversity of perspectives. In Proceedings of the Seventeenth USENIX Conference on Usable Privacy and Security , SOUPS’21, USA,
-
[25]
T. Li, D. Sree, and T. Ringenberg. Assessing crowdsourced annotations with LLMs: Linguistic certainty as a proxy for trustworthiness. In M. Hämäläinen, E. Öhman, Y . Bizzoni, S. Miyagawa, and K. Alnajjar, editors, Proceedings of the 5th International Conference on Natural Lang...
2025
-
[26]
H. B. Mann and D. R. Whitney. On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other.The Annals of Mathematical Statistics, 18(1):50 – 60, 1947. doi: 10.1214/aoms/1177730491. URL https://doi.org/10.1214/aoms/1177730491
1947
-
[27]
Mostafazadeh Davani, M
A. Mostafazadeh Davani, M. Díaz, and V . Prabhakaran. Dealing with disagreements: Looking beyond the majority vote in subjective annotations. Transactions of the Associa- tion for Computational Linguistics , 10:92–110, 2022. doi: 10.1162/tacl_a_00449. URL https://aclanthology....
2022 doi
-
[28]
Orlikowski, J
M. Orlikowski, J. Pei, P. Röttger, P. Cimiano, D. Jurgens, and D. Hovy. Beyond demographics: Fine-tuning large language models to predict individuals’ subjective text perceptions, 2025. URL https://arxiv.org/abs/2502.20897
2025 arXiv
-
[29]
Palomaki, O
J. Palomaki, O. Rhinehart, and M. Tseng. A case for a range of acceptable annotations. In SAD/CrowdBias@ HCOMP, pages 19–31, 2018
2018
-
[30]
J. S. Park, J. C. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein. Generative agents: Interactive simulacra of human behavior, 2023. URL https://arxiv.org/abs/ 2304.03442
2023 arXiv
-
[31]
Parrish, H
A. Parrish, H. R. Kirk, J. Quaye, C. Rastogi, M. Bartolo, O. Inel, J. Ciro, R. Mosquera, A. Howard, W. Cukierski, D. Sculley, V . J. Reddi, and L. Aroyo. Adversarial nibbler: A data-centric challenge for improving the safety of text-to-image models, 2023. URL https: //arxiv.or...
2023 arXiv
-
[32]
Parrish, V
A. Parrish, V . Prabhakaran, L. Aroyo, M. Díaz, C. M. Homan, G. Serapio-García, A. S. Taylor, and D. Wang. Diversity-aware annotation for conversational AI safety. In Proceedings of Safety4ConvAI: The Third Workshop on Safety for Conversational AI@ LREC-COLING 2024, pages 8–15, 2024
2024
-
[33]
Pavlick and T
E. Pavlick and T. Kwiatkowski. Inherent disagreements in human textual inferences. Transac- tions of the Association for Computational Linguistics, 7:677–694, 2019
2019
-
[34]
Pei and D
J. Pei and D. Jurgens. When do annotator demographics matter? measuring the influence of annotator demographics with the POPQUORN dataset. In The 61st Annual Meeting Of The Association For Computational Linguistics, 2023
2023
-
[35]
Prabhakaran, A
V . Prabhakaran, A. M. Davani, and M. Díaz. On releasing annotator-level labels and information in datasets. In Proceedings of the Joint 15th Linguistic Annotation Workshop (LAW) and 3rd Designing Meaning Representations (DMR) Workshop, pages 133–138, 2021
2021
-
[36]
Prabhakaran, C
V . Prabhakaran, C. Homan, L. Aroyo, A. Mostafazadeh Davani, A. Parrish, A. Taylor, M. Diaz, D. Wang, and G. Serapio-García. GRASP: A disagreement analysis framework to assess group associations in perspectives. In K. Duh, H. Gomez, and S. Bethard, editors, Proceedings of the ...
2024
-
[37]
Quaye, A
J. Quaye, A. Parrish, O. Inel, C. Rastogi, H. R. Kirk, M. Kahng, E. Van Liemt, M. Bartolo, J. Tsang, J. White, et al. Adversarial Nibbler: An open red-teaming method for identify- ing diverse harms in text-to-image generation. In The 2024 ACM Conference on Fairness, Accountabi...
2024
-
[38]
I. D. Raji and R. Dobbe. Concrete problems in ai safety, revisited. arXiv preprint arXiv:2401.10899, 2023
2023 arXiv
-
[39]
M. Rauh, N. Marchal, A. Manzini, L. A. Hendricks, R. Comanescu, C. Akbulut, T. Stepleton, J. Mateos-Garcia, S. Bergman, J. Kay, et al. Gaps in the safety evaluation of generative AI. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society , volume 7, pages 1200–12...
2024
-
[40]
Rismani, R
S. Rismani, R. Shelby, A. Smart, R. Delos Santos, A. Moon, and N. Rostamzadeh. Beyond the ml model: Applying safety engineering frameworks to text-to-image development. In Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society, pages 70–83, 2023
2023
-
[41]
Santurkar, E
S. Santurkar, E. Durmus, F. Ladhak, C. Lee, P. Liang, and T. Hashimoto. Whose opinions do language models reflect? In Proceedings of the 40th International Conference on Machine Learning, pages 29971–30004, 2023
2023
-
[42]
M. Sap, S. Swayamdipta, L. Vianna, X. Zhou, Y . Choi, and N. A. Smith. Annotators with attitudes: How annotator beliefs and identities bias toxic language detection. In M. Carpuat, M.-C. de Marneffe, and I. V . Meza Ruiz, editors, Proceedings of the 2022 Conference of the Nort...
2022 doi
-
[43]
Schramowski, M
P. Schramowski, M. Brack, B. Deiseroth, and K. Kersting. Safe latent diffusion: Mitigating inappropriate degeneration in diffusion models, 2023. URL https://arxiv.org/abs/2211. 05105
2023
-
[44]
A. D. Selbst, D. Boyd, S. A. Friedler, S. Venkatasubramanian, and J. Vertesi. Fairness and abstraction in sociotechnical systems. In Proceedings of the Conference on Fairness, Account- ability, and Transparency, FAT* ’19, page 59–68, New York, NY , USA, 2019. Association for C...
2019
-
[45]
Sorensen, L
T. Sorensen, L. Jiang, J. D. Hwang, S. Levine, V . Pyatkin, P. West, N. Dziri, X. Lu, K. Rao, C. Bhagavatula, et al. Value kaleidoscope: Engaging AI with pluralistic human values, rights, and duties. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, ...
2024
-
[46]
Sorensen, J
T. Sorensen, J. Moore, J. Fisher, M. Gordon, N. Mireshghallah, C. M. Rytting, A. Ye, L. Jiang, X. Lu, N. Dziri, et al. Position: A roadmap to pluralistic alignment. In Proceedings of the 41st International Conference on Machine Learning, pages 46280–46302, 2024
2024
-
[47]
Sorensen, P
T. Sorensen, P. Mishra, R. Patel, M. H. Tessler, M. Bakker, G. Evans, I. Gabriel, N. Goodman, and V . Rieser. Value profiles for encoding human variation, 2025. URLhttps://arxiv.org/ abs/2503.15484
2025
-
[48]
Steiger, T
M. Steiger, T. J. Bharucha, S. Venkatagiri, M. J. Riedl, and M. Lease. The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Sy...
2021
-
[49]
G. Team, T. Mesnard, C. Hardin, R. Dadashi, S. Bhupatiraju, S. Pathak, L. Sifre, M. Riv- ière, M. S. Kale, J. Love, P. Tafti, L. Hussenot, P. G. Sessa, A. Chowdhery, A. Roberts, A. Barua, A. Botev, A. Castro-Ros, A. Slone, A. Héliou, A. Tacchetti, A. Bulanova, A. Paterson, B. ...
2024 arXiv
-
[50]
what is safety?
A. M. Walker, M. A. DeVito, K. Badillo-Urquiola, R. Bellini, S. Chancellor, J. L. Feuston, K. Henne, P. G. Kelley, S. Rismani, R. Shelby, et al. " what is safety?": Building bridges across approaches to digital risks and harms. In Companion Publication of the 2024 Conference o...
2024
-
[51]
Weidinger, J
L. Weidinger, J. Mellor, M. Rauh, C. Griffin, J. Uesato, P.-S. Huang, M. Cheng, M. Glaese, B. Balle, A. Kasirzadeh, et al. Ethical and social risks of harm from language models. arXiv preprint arXiv:2112.04359, 2021
2021 arXiv
-
[52]
K. Wong, P. Paritosh, and L. Aroyo. Cross-replication reliability - an empirical approach to interpreting inter-rater reliability. In C. Zong, F. Xia, W. Li, and R. Navigli, editors,Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the...
2021
-
[53]
W. Zeng, D. Kurniawan, R. Mullins, Y . Liu, T. Saha, D. Ike-Njoku, J. Gu, Y . Song, C. Xu, J. Zhou, et al. Shieldgemma 2: Robust and tractable image content moderation. arXiv preprint arXiv:2504.01081, 2025
2025 arXiv
-
[54]
AS IS" without any warranty, express or implied. Google disclaims all liability for any damages, direct or indirect, resulting from the use of the dataset
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica. Judging LLM-as-a-judge with MT-Bench and chatbot arena, 2023. URL https://arxiv.org/abs/2306.05685. 14 A Dataset collection details A.1 P...
2023 arXiv
-
[57]
Low attention check accuracy: Threshold <1
-
[58]
Low total duration (Possibly low effort): Threshold <20 minutes (assuming ’20’ refers to a unit of time, likely minutes in this context)
-
[59]
Too few comments (Possibly low effort): Threshold <2
-
[60]
High annotation inconsistency (Possibly low effort)
-
[61]
Not harmful
High frequency of "Not harmful" selections (May otherwise silently pass attention checks and inconsistency checks): Threshold >35 Raters exceeding these thresholds in one or more categories were earmarked for manual review. The manual inspection process involved a thorough exa...
-
[62]
trains hitting a car in a violent accident
"trains hitting a car in a violent accident" –> 2 4. "a tea farmer" –> 3 5. "a woman cleaning the floor soiled by red wine" –> 4 In-context Parameters Each of the LLMs prompted in-context are sampled from in exactly the same way. We use the default sampling strategy (top-k) wi...
-
[2021]
ISBN 978-1-939133-25-0
USENIX Association. ISBN 978-1-939133-25-0
-
[2024]
URL https://arxiv.org/abs/2404.02261
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.