REVIEW 9 cited by
Decoding Biases: Automated Methods and LLM Judges for Gender Bias Detection in Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large Language Models (LLMs) have excelled at language understanding and generating human-level text. However, even with supervised training and human alignment, these LLMs are susceptible to adversarial attacks where malicious users can prompt the model to generate undesirable text. LLMs also inherently encode potential biases that can cause various harmful effects during interactions. Bias evaluation metrics lack standards as well as consensus and existing methods often rely on human-generated templates and annotations which are expensive and labor intensive. In this work, we train models to automatically create adversarial prompts to elicit biased responses from target LLMs. We present LLM- based bias evaluation metrics and also analyze several existing automatic evaluation methods and metrics. We analyze the various nuances of model responses, identify the strengths and weaknesses of model families, and assess where evaluation methods fall short. We compare these metrics to human evaluation and validate that the LLM-as-a-Judge metric aligns with human judgement on bias in response generation.
Forward citations
Cited by 9 Pith papers
-
LLMs on Trial: Evaluating Judicial Fairness for Large Language Models
A new 177,100-case benchmark shows that 16 LLMs systematically vary criminal sentences based on extra-legal demographic and procedural details, revealing pervasive judicial unfairness.
-
Bias, Accuracy, and Trust: Gender-Diverse Perspectives on Large Language Models
Gender-diverse users perceive ChatGPT's gender bias differently, with non-binary/transgender participants reporting condescending and stereotypical responses, and men reporting higher trust.
-
Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities
Across 25,000 stories from five LLMs, an LLM judge rated stories mentioning intellectual disabilities as more infantile, paternalistic, dependent, and inspirational than stories without the label.
-
A Close Reading Approach to Gender Narrative Biases in AI-Generated Stories
A close reading of 15 AI-generated stories finds that even when character counts are balanced, narrative roles, descriptions, and plot dynamics remain gender-stereotyped (e.g., every villain is male).
-
Mental Health Equity in LLMs: Leveraging Multi-Hop Question Answering to Detect Amplified and Silenced Perspectives
A multi-hop QA probe of four LLMs claims intersectional mental-health bias and 66-94% debiasing, but the bias metric is undefined and no conventional baseline is tested.
-
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
A literature review organizes LLM development into a six-phase software engineering lifecycle and identifies challenges and research directions for each phase.
-
Towards Fair Rankings: Leveraging LLMs for Gender Bias Detection and Measurement
LLM-based three-class gender labeling agrees with human annotations better than the lexical NFaiRR score, and the proposed CWEx metric combines neutral exposure with male-female exposure disparity for ranking fairness...
-
LFTF: Locating First and Then Fine-Tuning for Mitigating Gender Bias in Large Language Models
A block-localizing fine-tuning method for gender debiasing is presented, but its stated loss is inconsistent with its reported behavior and the evaluation tables contain duplicate rows.
-
Do Biased Models Have Biased Thoughts?
The manuscript is internally inconsistent: the abstract describes an LLM fairness experiment while the body is a different paper on pilot-wave quantum mechanics, so no coherent result can be assessed.
Discussion (0). Continue with ORCID to comment.