Pith. sign in

REVIEW 10 cited by

"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09219 v5 pith:AGVVVHYG submitted 2023-10-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords biaseslettersllm-generatedapplicationbiasgenderharmsprofessional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) have recently emerged as an effective tool to assist individuals in writing various types of content, including professional documents such as recommendation letters. Though bringing convenience, this application also introduces unprecedented fairness concerns. Model-generated reference letters might be directly used by users in professional scenarios. If underlying biases exist in these model-constructed letters, using them without scrutinization could lead to direct societal harms, such as sabotaging application success rates for female applicants. In light of this pressing issue, it is imminent and necessary to comprehensively study fairness issues and associated harms in this real-world use case. In this paper, we critically examine gender biases in LLM-generated reference letters. Drawing inspiration from social science findings, we design evaluation methods to manifest biases through 2 dimensions: (1) biases in language style and (2) biases in lexical content. We further investigate the extent of bias propagation by analyzing the hallucination bias of models, a term that we define to be bias exacerbation in model-hallucinated contents. Through benchmarking evaluation on 2 popular LLMs- ChatGPT and Alpaca, we reveal significant gender biases in LLM-generated recommendation letters. Our findings not only warn against using LLMs for this application without scrutinization, but also illuminate the importance of thoroughly studying hidden biases and harms in LLM-generated professional documents.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Feeding a model's own explanation of its biased story output back into a rewritten prompt improves demographic parity by 2% to 20%.

  2. McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A new Chinese bias benchmark with 4,077 instances and five tasks indicates larger language models are less biased than smaller ones when bias is measured through understanding tasks.

  3. From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.

  4. DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Caste-based stereotypes are measurably present in widely used LLMs, with the largest bias appearing when Dalits and Shudras are compared with dominant castes.

  5. Meflex: A Multi-agent Scaffolding System for Entrepreneurial Ideation Iteration via Nonlinear Business Plan Writing

    cs.HC 2026-02 conditional novelty 5.0 of 10

    A nonlinear, LLM-scaffolded business-plan writing tool with reflection and meta-reflection improves perceived usability and helps students iterate ideas in a 30-participant study.

  6. Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks

    cs.CL 2025-07 conditional novelty 5.0 of 10

    A name-substituted variant of the BBQ benchmark shows that LLMs retain nationality stereotypes even when explicit labels are removed, with smaller models showing more bias and lower accuracy.

  7. A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across U.S. and Chinese survey questions, DeepSeek, GPT-4o, Qwen2.5, and Llama-3.3 all show demographic overgeneralization, with no consistent home-field advantage for the Chinese model.

  8. The Biased Samaritan: LLM biases in Perceived Kindness

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Across ten commercial LLMs, demographic groups other than white male middle-aged were rated as more likely to help, while the control 'person' condition aligned with that majority baseline.

  9. Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A Chinese gender-bias corpus and shared-task benchmark show detection and classification are feasible, while automatic mitigation remains weak.

  10. On the Surprising Efficacy of LLMs for Penetration-Testing

    cs.CR 2025-07 conditional novelty 3.0 of 10

    A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.

Pith tools