Pith. sign in

REVIEW 2 cited by

Which Demographics do LLMs Default to During Annotation?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.08820 v3 pith:7BMRGDWQ submitted 2024-10-11 cs.CL

classification cs.CL
keywords demographicsannotationllmspromptsannotationsannotatorsdatademographic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Demographics and cultural background of annotators influence the labels they assign in text annotation -- for instance, an elderly woman might find it offensive to read a message addressed to a "bro", but a male teenager might find it appropriate. It is therefore important to acknowledge label variations to not under-represent members of a society. Two research directions developed out of this observation in the context of using large language models (LLM) for data annotations, namely (1) studying biases and inherent knowledge of LLMs and (2) injecting diversity in the output by manipulating the prompt with demographic information. We combine these two strands of research and ask the question to which demographics an LLM resorts to when no demographics is given. To answer this question, we evaluate which attributes of human annotators LLMs inherently mimic. Furthermore, we compare non-demographic conditioned prompts and placebo-conditioned prompts (e.g., "you are an annotator who lives in house number 5") to demographics-conditioned prompts ("You are a 45 year old man and an expert on politeness annotation. How do you rate {instance}"). We study these questions for politeness and offensiveness annotations on the POPQUORN data set, a corpus created in a controlled manner to investigate human label variations based on demographics which has not been used for LLM-based analyses so far. We observe notable influences related to gender, race, and age in demographic prompting, which contrasts with previous studies that found no such effects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

    cs.CL 2026-07 conditional novelty 6.5 of 10

    Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.

  2. Assessing the Reliability of LLMs Annotations in the Context of Demographic Bias and Model Explanation

    cs.CL 2025-07 reject novelty 5.0 of 10

    Demographics explain only about 8% of variance in sexism annotations, persona-prompted LLMs do not reliably beat baselines, and SHAP-guided highlighting helps smaller models.

Pith tools