REVIEW 10 cited by
"Kelly is a Warm Person, Joseph is a Role Model": Gender Biases in LLM-Generated Reference Letters
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have recently emerged as an effective tool to assist individuals in writing various types of content, including professional documents such as recommendation letters. Though bringing convenience, this application also introduces unprecedented fairness concerns. Model-generated reference letters might be directly used by users in professional scenarios. If underlying biases exist in these model-constructed letters, using them without scrutinization could lead to direct societal harms, such as sabotaging application success rates for female applicants. In light of this pressing issue, it is imminent and necessary to comprehensively study fairness issues and associated harms in this real-world use case. In this paper, we critically examine gender biases in LLM-generated reference letters. Drawing inspiration from social science findings, we design evaluation methods to manifest biases through 2 dimensions: (1) biases in language style and (2) biases in lexical content. We further investigate the extent of bias propagation by analyzing the hallucination bias of models, a term that we define to be bias exacerbation in model-hallucinated contents. Through benchmarking evaluation on 2 popular LLMs- ChatGPT and Alpaca, we reveal significant gender biases in LLM-generated recommendation letters. Our findings not only warn against using LLMs for this application without scrutinization, but also illuminate the importance of thoroughly studying hidden biases and harms in LLM-generated professional documents.
Forward citations
Cited by 10 Pith papers
-
Mitigation of Gender and Ethnicity Bias in AI-Generated Stories through Model Explanations
Feeding a model's own explanation of its biased story output back into a rewritten prompt improves demographic parity by 2% to 20%.
-
McBE: A Multi-task Chinese Bias Evaluation Benchmark for Large Language Models
A new Chinese bias benchmark with 4,077 instances and five tasks indicates larger language models are less biased than smaller ones when bias is measured through understanding tasks.
-
From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
Dual-character narrative prompts reveal gender biases in six multimodal LLMs that are largely invisible in single-character evaluations, and GENRES provides a structured benchmark to measure them.
-
DECASTE: Unveiling Caste Stereotypes in Large Language Models through Multi-Dimensional Bias Analysis
Caste-based stereotypes are measurably present in widely used LLMs, with the largest bias appearing when Dalits and Shudras are compared with dominant castes.
-
Meflex: A Multi-agent Scaffolding System for Entrepreneurial Ideation Iteration via Nonlinear Business Plan Writing
A nonlinear, LLM-scaffolded business-plan writing tool with reflection and meta-reflection improves perceived usability and helps students iterate ideas in a 30-participant study.
-
Obscured but Not Erased: Evaluating Nationality Bias in LLMs via Name-Based Bias Benchmarks
A name-substituted variant of the BBQ benchmark shows that LLMs retain nationality stereotypes even when explicit labels are removed, with smaller models showing more bias and lower accuracy.
-
A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies
Across U.S. and Chinese survey questions, DeepSeek, GPT-4o, Qwen2.5, and Llama-3.3 all show demographic overgeneralization, with no consistent home-field advantage for the Chinese model.
-
The Biased Samaritan: LLM biases in Perceived Kindness
Across ten commercial LLMs, demographic groups other than white male middle-aged were rated as more likely to help, while the control 'person' condition aligned with that majority baseline.
-
Overview of the NLPCC 2025 Shared Task: Gender Bias Mitigation Challenge
A Chinese gender-bias corpus and shared-task benchmark show detection and classification are feasible, while automatic mitigation remains weak.
-
On the Surprising Efficacy of LLMs for Penetration-Testing
A critical review arguing that LLMs are surprisingly effective for penetration testing because the task is largely pattern-matching, while noting serious reliability, safety, and cost barriers to autonomous use.
Discussion (0). Continue with ORCID to comment.