Pith. sign in

REVIEW 3 cited by

GenderAlign: An Alignment Dataset for Mitigating Gender Bias in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.13925 v3 pith:UXG7SVDJ submitted 2024-06-20 cs.CL cs.AI

classification cs.CLcs.AI
keywords genderbiasalignmentllmsbiasesdatasetgenderalignavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) are prone to generating content that exhibits gender biases, raising significant ethical concerns. Alignment, the process of fine-tuning LLMs to better align with desired behaviors, is recognized as an effective approach to mitigate gender biases. Although proprietary LLMs have made significant strides in mitigating gender bias, their alignment datasets are not publicly available. The commonly used and publicly available alignment dataset, HH-RLHF, still exhibits gender bias to some extent. There is a lack of publicly available alignment datasets specifically designed to address gender bias. Hence, we developed a new dataset named GenderAlign, aiming at mitigating a comprehensive set of gender biases in LLMs. This dataset comprises 8k single-turn dialogues, each paired with a "chosen" and a "rejected" response. Compared to the "rejected" responses, the "chosen" responses demonstrate lower levels of gender bias and higher quality. Furthermore, we categorized the gender biases in the "rejected" responses of GenderAlign into 4 principal categories. The experimental results show the effectiveness of GenderAlign in reducing gender bias in LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

    cs.AI 2025-11 conditional novelty 6.0 of 10

    LLMs solve logic-grid puzzles more accurately when the solution matches gender stereotypes than when it contradicts them, revealing implicit bias in deductive reasoning.

  2. PANDA -- Paired Anti-hate Narratives Dataset from Asia: Using an LLM-as-a-Judge to Create the First Chinese Counterspeech Dataset

    cs.CL 2025-01 conditional novelty 5.0 of 10

    The first Chinese-language counterspeech dataset of paired hate speech and counterspeech instances, created via an LLM-as-a-Judge pipeline with only partial human verification.

  3. The Fair Game: Auditing & Debiasing AI Algorithms Over Time

    cs.AI 2025-08 unverdicted novelty 4.0 of 10

    Proposes 'Fair Game', a reinforcement-learning loop in which an auditor's bias criteria, updatable over time, steer a debiasing agent that adapts an ML model's predictions.

Pith tools