Pith. sign in

REVIEW 2 cited by

Stereotype or Personalization? User Identity Biases Chatbot Recommendations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05613 v2 pith:IME7FEYQ submitted 2024-10-08 cs.CL

classification cs.CL
keywords userrecommendationsidentityllmsrevealedbiascasesgenerate
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While personalized recommendations are often desired by users, it can be difficult in practice to distinguish cases of bias from cases of personalization: we find that models generate racially stereotypical recommendations regardless of whether the user revealed their identity intentionally through explicit indications or unintentionally through implicit cues. We demonstrate that when people use large language models (LLMs) to generate recommendations, the LLMs produce responses that reflect both what the user wants and who the user is. We argue that chatbots ought to transparently indicate when recommendations are influenced by a user's revealed identity characteristics, but observe that they currently fail to do so. Our experiments show that even though a user's revealed identity significantly influences model recommendations (p < 0.001), model responses obfuscate this fact in response to user queries. This bias and lack of transparency occurs consistently across multiple popular consumer LLMs and for four American racial groups.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs

    cs.CY 2025-02 conditional novelty 7.0 of 10

    The paper introduces DiffAware and CtxtAware metrics, an 8-benchmark suite with 16,000 questions, and shows ten LLMs are less able to recognize legitimate group differences than current fairness benchmarks suggest.

  2. Whose Boat Does it Float? Improving Personalization in Preference Tuning via Inferred User Personas

    cs.CL 2025-01 conditional novelty 6.0 of 10

    A two-stage pipeline, Persona Inference and Persona Tailoring, augments preference datasets with LLM-inferred user personas and trains models to tailor responses to them, improving personalization over standard DPO.

Pith tools