Pith. sign in

REVIEW 2 cited by

Large Language Model Bias Mitigation from the Perspective of Knowledge Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.09341 v2 pith:ACKGARTV submitted 2024-05-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords debiasingfairnessknowledgeexistingbiaseditablefastfine-grained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Existing debiasing methods inevitably make unreasonable or undesired predictions as they are designated and evaluated to achieve parity across different social groups but leave aside individual facts, resulting in modified existing knowledge. In this paper, we first establish a new bias mitigation benchmark BiasKE leveraging existing and additional constructed datasets, which systematically assesses debiasing performance by complementary metrics on fairness, specificity, and generalization. Meanwhile, we propose a novel debiasing method, Fairness Stamp (FAST), which enables editable fairness through fine-grained calibration on individual biased knowledge. Comprehensive experiments demonstrate that FAST surpasses state-of-the-art baselines with remarkable debiasing performance while not hampering overall model capability for knowledge preservation, highlighting the prospect of fine-grained debiasing strategies for editable fairness in LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering

    cs.CL 2026-07 accept novelty 7.0 of 10

    Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.

  2. Know-MRI: A Knowledge Mechanisms Revealer&Interpreter for Large Language Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Know-MRI combines eleven existing LLM interpretation methods into one extensible toolkit with automatic input-to-method matching and dual UI and code interfaces.

Pith tools