Pith. sign in

REVIEW 5 cited by

Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14457 v1 pith:XBOFKHVE submitted 2025-01-24 cs.CL

classification cs.CL
keywords biasgenderllmseditingneuronneuronscapabilitiesmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) often exhibit gender bias, posing challenges for their safe deployment. Existing methods to mitigate bias lack a comprehensive understanding of its mechanisms or compromise the model's core capabilities. To address these issues, we propose the CommonWords dataset, to systematically evaluate gender bias in LLMs. Our analysis reveals pervasive bias across models and identifies specific neuron circuits, including gender neurons and general neurons, responsible for this behavior. Notably, editing even a small number of general neurons can disrupt the model's overall capabilities due to hierarchical neuron interactions. Based on these insights, we propose an interpretable neuron editing method that combines logit-based and causal-based strategies to selectively target biased neurons. Experiments on five LLMs demonstrate that our method effectively reduces gender bias while preserving the model's original capabilities, outperforming existing fine-tuning and editing approaches. Our findings contribute a novel dataset, a detailed analysis of bias mechanisms, and a practical solution for mitigating gender bias in LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Alignment on Llama 3 reduces explicit bias but amplifies implicit bias, because aligned models no longer represent 'black' and 'white' as racial concepts in ambiguous contexts.

  2. Demographic Prompting at Scale: When More Attributes Hurt LLM--Human Agreement

    cs.CL 2026-07 conditional novelty 6.5 of 10

    Across five subjective tasks and five open-source LLMs, demographic prompting improves human agreement only for 1–3 high-signal, directionally coherent attributes and degrades under the full attribute set.

  3. A Mechanistic Analysis of Gender Sensitivity in Dense Retrieval Models

    cs.IR 2026-08 conditional novelty 6.0 of 10

    Gender sensitivity in dense retrieval models originates in input embeddings and is carried by a small set of late-layer attention heads that jointly encode gender and term-matching signals.

  4. Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement

    cs.AI 2026-01 conditional novelty 6.0 of 10

    Amplifying neurons that encode bias awareness, found with 45 yes/no questions, reduces gender/race/religion bias in three LLMs while roughly preserving general reasoning.

  5. Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding

    cs.CL 2026-08 conditional novelty 5.0 of 10

    A new concept-centric GeoQA benchmark probes LLMs on abstraction, compositionality, and grounding of direction, distance, and topology, finding good abstraction, weak grounding, and a Mistral-family deficit.

Pith tools