Pith. sign in

REVIEW 2 cited by

Towards detecting unanticipated bias in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.02650 v1 pith:WOXQKWO5 submitted 2024-04-03 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords biasesllmsmodelsdetectinglanguageresearchdecisionsimpact
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Over the last year, Large Language Models (LLMs) like ChatGPT have become widely available and have exhibited fairness issues similar to those in previous machine learning systems. Current research is primarily focused on analyzing and quantifying these biases in training data and their impact on the decisions of these models, alongside developing mitigation strategies. This research largely targets well-known biases related to gender, race, ethnicity, and language. However, it is clear that LLMs are also affected by other, less obvious implicit biases. The complex and often opaque nature of these models makes detecting such biases challenging, yet this is crucial due to their potential negative impact in various applications. In this paper, we explore new avenues for detecting these unanticipated biases in LLMs, focusing specifically on Uncertainty Quantification and Explainable AI methods. These approaches aim to assess the certainty of model decisions and to make the internal decision-making processes of LLMs more transparent, thereby identifying and understanding biases that are not immediately apparent. Through this research, we aim to contribute to the development of fairer and more transparent AI systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. STEREODISCO: Discovering Stereotypicality in LLMs

    cs.AI 2026-07 conditional novelty 7.0 of 10

    LLMs encode stereotypes along recoverable geometric axes in attention heads, and two tested LLMs share more stereotype content with each other than with documented human stereotypes.

  2. Musical ethnocentrism in Large Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    LLMs like ChatGPT and Mixtral show a strong Western bias when asked to name top musical contributors and to rate the musical cultures of countries.

Pith tools