REVIEW 4 cited by
BiasAlert: A Plug-and-play Tool for Social Bias Detection in LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Evaluating the bias in Large Language Models (LLMs) becomes increasingly crucial with their rapid development. However, existing evaluation methods rely on fixed-form outputs and cannot adapt to the flexible open-text generation scenarios of LLMs (e.g., sentence completion and question answering). To address this, we introduce BiasAlert, a plug-and-play tool designed to detect social bias in open-text generations of LLMs. BiasAlert integrates external human knowledge with inherent reasoning capabilities to detect bias reliably. Extensive experiments demonstrate that BiasAlert significantly outperforms existing state-of-the-art methods like GPT4-as-A-Judge in detecting bias. Furthermore, through application studies, we demonstrate the utility of BiasAlert in reliable LLM bias evaluation and bias mitigation across various scenarios. Model and code will be publicly released.
Forward citations
Cited by 4 Pith papers
-
BiasFilter: An Inference-Time Debiasing Framework for Large Language Models
BiasFilter filters low-fairness segments during LLM generation using a reward model trained on a GPT-4-scored preference dataset, cutting bias on CEB and FairMT.
-
Large Language Models in Architecture Studio: A Framework for Learning Outcomes
A conceptual framework maps LLM-based interventions onto architecture studio challenges and Bloom's taxonomy across self-, peer-, and teacher-led learning.
-
Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
A supervised fine-tuning plus difficulty-filtered reinforcement learning recipe improves video temporal grounding on three benchmarks, with datasets and models released.
-
Detection, Classification, and Mitigation of Gender Bias in Large Language Models
A Chinese gender-bias system using SFT, chain-of-thought, and DPO with GPT-4-generated preference pairs reports top validation scores and first place on all three NLPCC 2025 subtasks.
Discussion (0). Continue with ORCID to comment.