Pith. sign in

REVIEW 5 cited by

Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.15585 v1 pith:EVK6RRMH submitted 2024-01-28 cs.CL

classification cs.CL
keywords llmspredictionsreasoningtasksunscalablewordsmodelbias
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

There exist both scalable tasks, like reading comprehension and fact-checking, where model performance improves with model size, and unscalable tasks, like arithmetic reasoning and symbolic reasoning, where model performance does not necessarily improve with model size. Large language models (LLMs) equipped with Chain-of-Thought (CoT) prompting are able to make accurate incremental predictions even on unscalable tasks. Unfortunately, despite their exceptional reasoning abilities, LLMs tend to internalize and reproduce discriminatory societal biases. Whether CoT can provide discriminatory or egalitarian rationalizations for the implicit information in unscalable tasks remains an open question. In this study, we examine the impact of LLMs' step-by-step predictions on gender bias in unscalable tasks. For this purpose, we construct a benchmark for an unscalable task where the LLM is given a list of words comprising feminine, masculine, and gendered occupational words, and is required to count the number of feminine and masculine words. In our CoT prompts, we require the LLM to explicitly indicate whether each word in the word list is a feminine or masculine before making the final predictions. With counting and handling the meaning of words, this benchmark has characteristics of both arithmetic reasoning and symbolic reasoning. Experimental results in English show that without step-by-step prediction, most LLMs make socially biased predictions, despite the task being as simple as counting words. Interestingly, CoT prompting reduces this unconscious social bias in LLMs and encourages fair predictions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Guiding LLM Decision-Making with Fairness Reward Models

    cs.LG 2025-07 conditional novelty 7.0 of 10

    A single process-level reward model, trained on weakly labeled biased versus unbiased reasoning, transfers across tasks and models to reduce equalized odds gaps in LLM decision-making.

  2. Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases

    cs.CL 2025-09 conditional novelty 6.0 of 10

    LLMs shift their gendered pronoun choices, toward 'they' and away from 'he', when prompts signal a gender-bias evaluation, so measured bias is highly dependent on prompt framing.

  3. More or Less Wrong: A Benchmark for Directional Bias in LLM Comparative Reasoning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Comparative words in prompts can shift LLM answers toward the framed direction in simple arithmetic comparisons, with demographic terms amplifying the effect.

  4. BiasFilter: An Inference-Time Debiasing Framework for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    BiasFilter filters low-fairness segments during LLM generation using a reward model trained on a GPT-4-scored preference dataset, cutting bias on CEB and FairMT.

  5. Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

    cs.CL 2025-08 unverdicted novelty 3.0 of 10

    The submission's abstract promises an LLM safety survey, but the provided body is the opening page of an unrelated arithmetic-dynamics paper, so the artifact is internally inconsistent.

Pith tools