Pith. sign in

REVIEW 3 major objections 4 minor 97 references

Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Placing demographic information in system prompts rather than user prompts consistently increases bias in outputs of six commercial LLMs.

desk verdict Useful measurement study, but the position claim is confounded because system and user prompts differ in wording, not just placement. read the letter →

arxiv 2505.21091 v3 pith:SNBMPCBW submitted 2025-05-27 cs.CY cs.AIcs.CL

classification cs.CYcs.AIcs.CL
keywords systempromptsbiasinlargelanguagemodelsrepresentationalallocativepromptplacementAIauditingtransparencyalgorithmicsupplychains
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that where demographic information appears in an LLM prompt—in the hidden system prompt rather than the user's own message—changes what the model outputs, in ways that can disadvantage protected groups. Across six commercial models and 50 demographic descriptors, putting audience identity in the system prompt consistently produced larger differences in negative sentiment when describing demographic groups (representational bias), and the gap widened with model size. In resource-allocation ranking tasks, system-prompt placement also shifted priority rankings away from baseline more than user-prompt placement, and direct system-versus-user comparisons revealed distinct ranking patterns that baseline comparisons miss. The authors claim these effects matter because system-prompt layers are opaque, hidden from end-users and often from deployers, so biased behavior caused by prompt position is hard to detect or correct. The paper concludes that system-prompt analysis belongs in AI auditing.

What carries the argument

The operating mechanism is the prompt-position contrast built into five carefully constructed conditions that vary only how demographic identity is conveyed—stated explicitly or implied through preferences, behaviors, and values, placed either in the system prompt or the user prompt, with a no-information default. Bias is quantified for representation by a min-max negative-sentiment range across audience personas for each described group (average over 50 groups), and for allocation by Kendall's τ rank correlation against baseline rankings and between placement conditions. These two instruments convert an otherwise qualitative concern about 'hidden instructions' into measurable, comparable quantities.

What would settle it

Run the two experiments with word-for-word matched prompt text that differs only in placement—for example, system prompt 'You are talking to a Christian. Answer in a way true to this identity.' versus user prompt 'I am a Christian. Answer in a way true to this identity.'—and compare ΔBias and Kendall's τ; if the system-versus-user differences disappear, the claimed position effect is an artifact of unmatched instructions.

Watch

Extended reading notes

Core claim

The central claim is that position is a causal channel for bias: the same demographic identity statement, placed in the system prompt instead of the user prompt, produces measurably more biased descriptions and different resource-allocation rankings. The paper supports this with a five-condition design (default, explicit system, implicit system, explicit user, implicit user) applied to six commercially deployed LLMs. It reports that system prompts generate consistently higher audience bias in negative sentiment than user prompts across all models, with the difference increasing with model size (peaking at ΔBias = 0.335 for Claude-3.5-Sonnet), and that system prompts tend to produce larger deviations from baseline priority rankings than user prompts. The authors frame this as a supply-chain transparency problem: because system prompts are layered and inaccessible, stakeholders cannot attribute or audit the bias.

Load-bearing premise

The central comparison assumes the system and user prompt conditions differ only in where the demographic identity sits, but the system condition carries an extra directive ('Answer their questions in a way that stays true to the nature of this identity') that the user condition lacks, so the reported differences could come from that wording rather than from position.

Editorial extensions

If this is right

  • System-prompt layers become a required audit target: model evaluations that only examine user-visible behavior will miss a measurable source of demographic bias.
  • Larger models amplify the placement effect, so scaling model capability without auditing the system layer may increase rather than reduce representational harm.
  • Direct system-versus-user comparisons are needed because comparisons to a baseline alone can hide reordering differences that still change allocation outcomes.
  • Both explicit demographic statements and implicit signals (assumed preferences or values) show the effect, so models infer and weight identity even when it is not declared outright.
  • In deployed services, changing where audience information is configured—developer console versus user profile text—could alter who gets priority in ranking-based decisions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: if position itself is the active ingredient, the same bias channel likely applies to any sensitive system-prompt content, not just the 50 demographic descriptors tested; legal disclaimers, institutional policies, or safety rules embedded at system level could similarly shift outputs in unmonitored ways.
  • Inference: a direct test of the causal claim would match wording exactly across conditions; if the effect persists with identical text in both positions, it would confirm that instruction-hierarchy weighting, rather than content, drives the bias.
  • Inference: the results connect to model sycophancy: system-prompt personas may push models to mirror the assumed worldview of the audience, and measuring this separately from negative sentiment would be a natural follow-up.
  • Inference: for auditors, the practical implication is that LLM supply-chain contracts should require disclosure and versioning of layered system prompts, since the paper shows that such prompts are behaviorally consequential but invisible.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper investigates whether placing demographic information about the user in a system prompt versus a user prompt changes LLM behavior. Across six commercial LLMs and 50 demographic descriptors, the authors measure two outcomes: (RQ1) the range of negative sentiment in model-generated descriptions of demographic groups, and (RQ2) deviations in resource-allocation rankings from a baseline. They report that system-prompt placement leads to higher sentiment-range bias and greater ranking deviations, and they argue this constitutes a hidden source of representational and allocative bias that should be incorporated into AI auditing. The paper includes a large empirical study with public code, two new datasets, and a discussion of transparency in AI supply chains.

Significance. If the central claim were established—that prompt position, independent of content, systematically increases bias—this would be a significant contribution to fairness and transparency auditing, since system prompts are often invisible to end-users and deployers. The paper's strengths include the breadth of the evaluation (six commercial models, 50 groups, 40 allocation scenarios), the use of deterministic decoding (temperature=0), and the release of code and datasets. The resource-allocation scenario dataset and the analysis of direct system-versus-user ranking divergence address a relevant and understudied question. However, the experimental design as reported does not isolate position from instruction content, and several empirical claims are overstated relative to the data. The contribution's value depends on whether these issues can be resolved with matched-content conditions and more careful claims.

major comments (3)
  1. [Table 2, Section 3.3] The 'position' effect is confounded with instruction content. The System Prompt Explicit Condition reads 'You are talking to {persona}. Answer their questions in a way that stays true to the nature of this identity,' while the User Prompt Explicit Condition reads only 'I am {persona}.' The System Prompt Implicit Condition likewise appends 'Answer their questions in a way that stays true to the nature of this identity,' which the User Prompt Implicit Condition omits. In addition, the user conditions use 'You are a helpful assistant' as the system prompt, while the system conditions do not include this baseline. Consequently, the observed differences in Figs. 5–7 could be driven by the extra imperative to 'stay true' to the persona, or by the absence of a 'helpful assistant' role, rather than by the position of the demographic information itself. The RQ1/RQ2 causal conclusions and the title 'Position is Power' therefore do not follow from this design.
  2. [Abstract; Section 4.1.2; Figure 5; Table 3] The abstract and Section 4.1.2 claim that 'system prompts consistently generate higher bias in demographic descriptions... across all models,' but the reported data contradict this. Figure 5's caption explicitly notes the exception of GPT-4o-mini, and Table 3 shows a negative ΔBias (−0.003) for GPT-4o-mini in the explicit small-models condition, and −0.041 for Gemini-1.5-Flash-8B in the implicit small-models condition, as well as a zero difference for GPT-4o in the implicit large-models condition. Since the paper's headline contribution is the consistency of the effect, these counterexamples must be acknowledged and the claims revised to describe the direction, magnitude, and exceptions appropriately.
  3. [Section 4.2, Figures 6–7] The allocative-bias results are also less uniform than the abstract suggests. For explicit prompts, Figure 6a shows that Claude-3.5-Haiku deviates less from baseline under system prompts than under user prompts. For implicit prompts, Figure 6b shows that the two largest models (Gemini-1.5-Pro and Claude-3.5-Sonnet) deviate more under user prompts, reversing the direction seen in smaller models. The abstract's hedged 'can produce greater deviations' is acceptable, but the conclusion section (Section 6) states as a general finding that 'system prompts tended to cause greater deviations from baseline rankings compared to user prompts,' which is not supported by the data across conditions and models. The claims need to be matched to the observed patterns.
minor comments (4)
  1. [Table 2] In the System Prompt Implicit Condition, the text reads 'a person that likes likes {like},' which appears to be a typo; the duplicate 'likes' should be removed.
  2. [Section 3.4] The notation Biassystem and Biasuser is used to define ΔBias, but the definitions of these terms are not explicitly stated; please define them as the averages of B_audience,j over all described groups j for each condition.
  3. [Section 4.1.2, Figure 12] The text says 'Fig. 12 shows the same trends for implicit prompting conditions,' but Figure 12's caption states that user prompts consistently produce lower bias ranges 'except in Gemini-1.5-Flash-8B, and GPT-4o.' This is not the same trend, and the inconsistency should be reconciled.
  4. [Section 5.4] The limitations section acknowledges that the study does not make normative judgments about the observed differences, but the abstract and conclusion frame the findings as 'biases' and 'harms.' Please clarify the normative status of the measured sentiment-range differences, particularly whether higher min–max range necessarily implies harm.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the bias measurements are direct empirical observations with no fitted parameters or self-citation chain supporting the central claim.

full rationale

The paper's central quantities are measured, not derived: Bias_condition is the mean over described groups of max-min negative sentiment ranges computed from model outputs (Section 3.4), and RQ2 uses Kendall's tau between directly observed rankings (Section 3.5). No parameter is fitted to the outcome and then renamed a prediction; the same prompt templates and the same GPT-4o-generated implicit descriptors are held fixed across the system- and user-placement conditions, so placement comparisons are not forced by construction. The use of GPT-4o to construct the implicit descriptor dataset and allocation scenarios is a data-generation step whose outputs are inputs to both compared conditions; it does not make the measured system-vs-user difference equal to its own input. Self-citations (e.g., [23] for 'accountability horizon', [42] for toxicity control, [85]/[97] for fairness criteria) are background or metric-motivation citations and are not load-bearing for the empirical RQ1/RQ2 results. The content mismatch between system and user prompt conditions (Table 2) is a potential validity confound, but a confound is an experimental-design concern, not a circularity in which an output is equivalent to an input by construction. Under the stated hard rules, no circular step can be quoted, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 0 invented entities

The central claim rests on the validity of sentiment-based bias measurement, the representativeness of the manually and GPT-4o-selected stimuli, and the assumption that prompt conditions are matched except for placement. The latter is violated by an extra instruction in the system prompt condition.

assumptions (5)
  • domain assumption Sentiment scores from roberta-base-sentiment provide a valid measure of representational bias in LLM-generated text.
    The paper uses min/max negative sentiment range as its bias metric; if the sentiment analyzer is biased or insensitive to relevant text features, the measured differences may not reflect model bias. Section 3.4.
  • domain assumption The min-max spread of sentiment across audience conditions is an appropriate measure of bias.
    This metric captures variability of sentiment across audiences, but variability is not necessarily harmful bias; it could be appropriate personalization or sycophancy. Section 3.4.
  • domain assumption The system and user prompt conditions differ only in the placement of demographic information.
    In fact, Table 2 shows an extra instruction in the system prompt condition, so this assumption is violated.
  • domain assumption The 50 GDPR-based demographic descriptors and 40 allocation scenarios are representative of protected groups and decision contexts.
    Selection involved researcher choice and GPT-4o assistance, which may introduce selection bias. Appendix A.
  • domain assumption Commercial API outputs with temperature=0 are stable and the models tested approximate current deployment conditions.
    API access is opaque; model versions and system prompts may change. Section 3.2.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)." pith.science (2026). https://pith.science/paper/SNBMPCBW

@misc{pith2026250521091,
  author       = {Pith},
  title        = {Pith review of: Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SNBMPCBW}},
  note         = {Machine review of arXiv:2505.21091}
}
read the original abstract

System prompts in Large Language Models (LLMs) are predefined directives that guide model behaviour, taking precedence over user inputs in text processing and generation. LLM deployers increasingly use them to ensure consistent responses across contexts. While model providers set a foundation of system prompts, deployers and third-party developers can append additional prompts without visibility into others' additions, while this layered implementation remains entirely hidden from end-users. As system prompts become more complex, they can directly or indirectly introduce unaccounted for side effects. This lack of transparency raises fundamental questions about how the position of information in different directives shapes model outputs. As such, this work examines how the placement of information affects model behaviour. To this end, we compare how models process demographic information in system versus user prompts across six commercially available LLMs and 50 demographic groups. Our analysis reveals significant biases, manifesting in differences in user representation and decision-making scenarios. Since these variations stem from inaccessible and opaque system-level configurations, they risk representational, allocative and potential other biases and downstream harms beyond the user's ability to detect or correct. Our findings draw attention to these critical issues, which have the potential to perpetuate harms if left unexamined. Further, we argue that system prompt analysis must be incorporated into AI auditing processes, particularly as customisable system prompts become increasingly prevalent in commercial AI deployments.

Figures

Figures reproduced from arXiv: 2505.21091 by the authors.

Figure 1
Figure 1. [Influence of Prompt Placement on AI Model Bias] Comparison of two model outputs by Claude-3.5-Haiku. The [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. [AI Supply Chain Prompt Hierarchy and Visi [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Example Prompt for an Allocation Decision: Organ Transplant Scenario. Prompting is in the Explicit User Condition [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (12 more)
Figure 4
Figure 4. Figure 4: [Negative Sentiment Compared Between System [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: [Audience Bias by model size and prompt condi [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: [Model ranking correlation against baseline, lower [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: [Model ranking correlation of system prompts [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: [Description Bias Between Explicit System and User Prompts for Claude models] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p020_8.png]
Figure 9
Figure 9. Figure 9: [Description Bias Between Explicit System and User Prompts for Gemini models] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p021_9.png]
Figure 10
Figure 10. Figure 10: [Description Bias Between Explicit System and User Prompts for GPT models] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p022_10.png]
Figure 11
Figure 11. Figure 11: [Description Bias Between Implicit System and User Prompts for Claude-3.5-Sonnet] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p023_11.png]
Figure 12
Figure 12. Figure 12: [Audience bias by model size and prompt condition, higher values indicate larger ranges in negative sentiment] [PITH_FULL_IMAGE:figures/full_fig_p023_12.png]
Figure 13
Figure 13. Figure 13: [Description Bias Between Implicit System and User Prompts for Claude-3.5-Haiku] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p024_13.png]
Figure 14
Figure 14. Figure 14: [Description Bias Between Implicit System and User Prompts for Gemini models] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p025_14.png]
Figure 15
Figure 15. Figure 15: [Description Bias Between Implicit System and User Prompts for GPT models] The heatmap compares negative [PITH_FULL_IMAGE:figures/full_fig_p026_15.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

97 extracted references · 16 canonical work pages

  1. [1]

    GPT-4o System Card

    2024. GPT-4o System Card. https://openai.com/index/gpt-4o-system-card/

  2. [2]

    Introducing the next generation of Claude

    2024. Introducing the next generation of Claude. https://www.anthropic.com/ news/claude-3-family

  3. [3]

    Memory and new controls for ChatGPT

    2024. Memory and new controls for ChatGPT. https://openai.com/index/ memory-and-new-controls-for-chatgpt/

  4. [4]

    Model Spec (2024/05/08)

    2024. Model Spec (2024/05/08). https://cdn.openai.com/spec/model-spec-2024- 05-08.html/#follow-the-chain-of-command

  5. [5]

    Gemini API

    2025. Gemini API. https://ai.google.dev/gemini-api/docs

  6. [6]

    Friedler, C

    Mohsen Abbasi, Sorelle A. Friedler, C. Scheidegger, and Suresh Venkatasubra- manian. 2019. Fairness in representation: quantifying stereotyping as a repre- sentational harm. (2019), 801–809. https://doi.org/10.1137/1.9781611975673.90

  7. [7]

    Daron Acemoglu. 2024. Harms of AI. In The Oxford Handbook of AI Governance . Oxford University Press. https://doi.org/10.1093/oxfordhb/9780197579329.013. 65

  8. [8]

    Amith Ananthram, Elias Stengel-Eskin, Carl Vondrick, Mohit Bansal, and Kathleen McKeown. 2024. See It from My Perspective: Diagnosing the West- ern Cultural Bias of Large Vision-Language Models in Image Understanding. https://doi.org/10.48550/arXiv.2406.11665

Show all 97 references
  1. [9]

    Sodiq Odetunde Babatunde, Opeyemi Abayomi Odejide, Tolulope Esther Edun- jobi, and Damilola Oluwaseun Ogundipe. 2024. THE ROLE OF AI IN MARKET- ING PERSONALIZATION: A THEORETICAL EXPLORATION OF CONSUMER ENGAGEMENT STRATEGIES. International Journal of Management & En- trepreneu...

  2. [10]

    Agathe Balayn, Mireia Yurrita, Fanny Rancourt, Fabio Casati, and Ujwal Gadi- raju. 2025. Unpacking Trust Dynamics in the LLM Supply Chain: An Empirical Exploration to Foster Trustworthy LLM Production & Use. In Proceedings of the 2025 CHI Conference on Human Factors in Computi...

  3. [11]

    Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification. In Findings of the Association for Computational Linguistics: EMNLP 2020. Association for Computational Li...

  4. [12]

    Solon Barocas, Kate Crawford, Aaron Shapiro, and Hanna Wallach. 2017. The problem with bias: Allocative versus representational harms in machine learning. In 9th Annual conference of the special interest group for computing, information and society. New York, NY

  5. [13]

    2023.Fairness and Machine Learning: Limitations and Opportunities

    Solon Barocas, Moritz Hardt, and Arvind Narayanan. 2023.Fairness and Machine Learning: Limitations and Opportunities . MIT Press

  6. [14]

    Solon Barocas and Andrew D. Selbst. 2016. Big Data’s Disparate Impact. https: //doi.org/10.2139/ssrn.2477899

  7. [15]

    Rick Battle and Teja Gollapudi. 2024. The Unreasonable Effectiveness of Eccen- tric Automatic Prompts. https://doi.org/10.48550/arXiv.2402.10949

  8. [17]

    Agata Blasiak, Jeffrey Khong, and Theodore Kee. 2020. CURATE.AI: Opti- mizing Personalized Medicine with Artificial Intelligence. SLAS TECHNOL- OGY: Translating Life Sciences Innovation 25, 2 (April 2020), 95–105. https: //doi.org/10.1177/2472630319890316

  9. [18]

    Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach. 2020. Lan- guage (Technology) is Power: A Critical Survey of "Bias" in NLP. https: //doi.org/10.48550/arXiv.2005.14050

  10. [19]

    Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S

    Rishi Bommasani, Drew A. Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S. Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, Erik Brynjolfsson, Shyamal Buch, Dallas Card, Rodrigo Castellon, Niladri Chatterji, Annie Chen, Kathleen Creel, Jare...

  11. [20]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...

  12. [22]

    Jennifer Cobbe and Jatinder Singh. 2021. Artificial intelligence as a service: Legal responsibilities, liabilities, and policy challenges.Computer Law & Security Review 42 (Sept. 2021), 105573. https://doi.org/10.1016/j.clsr.2021.105573

  13. [24]

    Sasha Costanza-Chock, Inioluwa Deborah Raji, and Joy Buolamwini. 2022. Who Audits the Auditors? Recommendations from a field scan of the algorithmic au- diting ecosystem. InProceedings of the 2022 ACM Conference on Fairness, Account- ability, and Transparency (FAccT ’22). Asso...

  14. [25]

    Kate Crawford. 2016. Opinion | Artificial Intelligence’s White Guy Problem. The New York Times (June 2016). https://www.nytimes.com/2016/06/26/opinion/ sunday/artificial-intelligences-white-guy-problem.html

  15. [26]

    Hannah Devinney, Jenny Björklund, and Henrik Björklund. 2022. Theories of “Gender” in NLP Bias Research. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’22) . Association for Com- puting Machinery, New York, NY, USA, 2083–2102. h...

  16. [27]

    Jwala Dhamala, Tony Sun, Varun Kumar, Satyapriya Krishna, Yada Pruk- sachatkun, Kai-Wei Chang, and Rahul Gupta. 2021. BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Tr...

  17. [29]

    Jingchao Fang, Nikos Arechiga, Keiichi Namikoshi, Nayeli Bravo, Candice Hogan, and David A. Shamma. 2024. On LLM Wizards: Identifying Large Language Models’ Behaviors for Wizard of Oz Experiments. In Proceedings of the ACM International Conference on Intelligent Virtual Agents...

  18. [30]

    Sourojit Ghosh, Pranav Narayanan Venkit, Sanjana Gautam, Shomir Wilson, and Aylin Caliskan. 2024. Do Generative AI Models Output Harm while Represent- ing Non-Western Cultures: Evidence from A Community-Centered Approach. Proceedings of the AAAI/ACM Conference on AI, Ethics, a...

  19. [31]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, A...

  20. [32]

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. 2024. Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs. https://doi.org/10.48550/arXiv. 2311.04892

  21. [33]

    Rishav Hada, Safiya Husain, Varun Gumma, Harshita Diddee, Aditya Yadavalli, Agrima Seth, Nidhi Kulkarni, Ujwal Gadiraju, Aditya Vashistha, Vivek Seshadri, and Kalika Bali. 2024. Akal Badi ya Bias: An Exploratory Study of Gender Bias in Hindi Language Technology. In The 2024 AC...

  22. [34]

    Moritz Hardt, Eric Price, and Nathan Srebro. 2016. Equality of Opportunity in Supervised Learning. https://arxiv.org/abs/1610.02413

  23. [35]

    Ruidan He, Linlin Liu, Hai Ye, Qingyu Tan, Bosheng Ding, Liying Cheng, Jia-Wei Low, Lidong Bing, and Luo Si. 2021. On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation. https://doi.org/10.48550/ arXiv.2106.03164

  24. [36]

    Lily Hu and Issa Kohler-Hausmann. 2020. What’s Sex Got To Do With Fair Ma- chine Learning?. InProceedings of the 2020 Conference on Fairness, Accountability, and Transparency. 513–513. https://doi.org/10.1145/3351095.3375674

  25. [37]

    Hancock, and Mor Naaman

    Maurice Jakesch, Jeffrey T. Hancock, and Mor Naaman. 2023. Human heuristics for AI-generated language are flawed. Proceedings of the National Academy of Sciences 120, 11 (2023). https://doi.org/10.1073/pnas.2208839120

  26. [38]

    Seyyed Ahmad Javadi, Chris Norval, Richard Cloete, and Jatinder Singh. 2021. Monitoring AI Services for Misuse. In Proceedings of the 2021 AAAI/ACM Con- ference on AI, Ethics, and Society . ACM, Virtual Event USA, 597–607. https: //doi.org/10.1145/3461702.3462566

  27. [39]

    Guangyuan Jiang, Manjie Xu, Song-Chun Zhu, Wenjuan Han, Chi Zhang, and Yixin Zhu. 2023. Evaluating and inducing personality in pre-trained language models. In Proceedings of the 37th International Conference on Neural Information Processing Systems (New Orleans, LA, USA) (NIPS...

  28. [40]

    Zhifeng Jiang, Zhihua Jin, and Guoliang He. 2025. Safeguarding System Prompts for LLMs. https://doi.org/10.48550/arXiv.2412.13426

  29. [41]

    Jared Katzman, Angelina Wang, Morgan Scheuerman, Su Lin Blodgett, Kristen Laird, Hanna Wallach, and Solon Barocas. 2023. Taxonomizing and Measuring Representational Harms: A Look at Image Tagging. Proceedings of the AAAI Conference on Artificial Intelligence 37, 12 (June 2023)...

  30. [42]

    Elisabeth Kirsten, Ivan Habernal, Vedant Nanda, and Muhammad Bilal Za- far. 2025. The Impact of Inference Acceleration on Bias of LLMs. In Proceed- ings of the 2025 Conference of the Nations of the Americas Chapter of the As- sociation for Computational Linguistics: Human Lang...

  31. [43]

    Ryan Koo, Minhwa Lee, Vipul Raheja, Jong Inn Park, Zae Myung Kim, and Dongyeop Kang. 2024. Benchmarking Cognitive Biases in Large Language Models as Evaluators. InFindings of the Association for Computational Linguistics: ACL 2024, Lun-Wei Ku, Andre Martins, and Vivek Srikumar...

  32. [44]

    Adriano Koshiyama, Emre Kazim, Philip Treleaven, Pete Rai, Lukasz Szpruch, Giles Pavey, Ghazi Ahamat, Franziska Leutner, Randy Goebel, Andrew Knight, Janet Adams, Christina Hitrova, Jeremy Barnett, Parashkev Nachev, David Barber, Tomas Chamorro-Premuzic, Konstantin Klemmer, Mi...

  33. [45]

    Bushra Kundi, Christo El Morr, Rachel Gorman, and Ena Dua. 2023. Artificial intelligence and bias: a scoping review. AI and Society (2023), 199–215

  34. [46]

    Nils Köbis and Luca D. Mossink. 2021. Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry. Computers in Human Behavior 114 (2021), 106553. https://doi.org/10.1016/j.chb.2020.106553

  35. [47]

    Ehsan Latif and Xiaoming Zhai. 2024. Fine-tuning ChatGPT for automatic scoring. Computers and Education: Artificial Intelligence 6 (June 2024), 100210. https://doi.org/10.1016/j.caeai.2024.100210

  36. [49]

    Michelle Seng Ah Lee and Jat Singh. 2021. The Landscape and Gaps in Open Source Fairness Toolkits. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (Yokohama, Japan) (CHI ’21) . Association for Computing Machinery, New York, NY, USA, Article 699,...

  37. [50]

    Seongyun Lee, Sue Hyun Park, Seungone Kim, and Minjoon Seo. 2024. Aligning to Thousands of Preferences via System Message Generalization. https://doi. org/10.48550/arXiv.2405.17977

  38. [52]

    Alina Leidinger and Richard Rogers. 2024. How Are LLMs Mitigating Stereo- typing Harms? Learning from Search Engine Studies. https://arxiv.org/abs/ 2407.11733

  39. [53]

    Artificial Intelligence as a Service

    Kornel Lewicki, Michelle Seng Ah Lee, Jennifer Cobbe, and Jatinder Singh. 2023. Out of Context: Investigating the Bias and Fairness Concerns of “Artificial Intelligence as a Service”. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23) . A...

  40. [54]

    Kenneth Li, Tianle Liu, Naomi Bashkansky, David Bau, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg. 2024. Measuring and Controlling Instruction (In)Stability in Language Model Dialogs. https://doi.org/10.48550/ arXiv.2402.10962

  41. [55]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. RoBERTa: A Robustly Optimized BERT Pretraining Approach. https://arxiv.org/abs/1907. 11692

  42. [56]

    Bowman, and Rachel Rudinger

    Chandler May, Alex Wang, Shikha Bordia, Samuel R. Bowman, and Rachel Rudinger. 2019. On Measuring Social Biases in Sentence Encoders. https: //arxiv.org/abs/1903.10561

  43. [57]

    Katelyn Mei, Sonia Fereidooni, and Aylin Caliskan. 2023. Bias Against 93 Stigma- tized Groups in Masked Language Models and Downstream Sentiment Classifica- tion Tasks. In2023 ACM Conference on Fairness, Accountability, and Transparency. ACM, Chicago IL USA, 1699–1710. https:/...

  44. [58]

    Mazda Moayeri, Elham Tabassi, and Soheil Feizi. 2024. WorldBench: Quantifying Geographic Disparities in LLM Factual Recall. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. ACM, Rio de Janeiro Brazil, 1211–

  45. [59]

    Norman Mu, Sarah Chen, Zifan Wang, Sizhe Chen, David Karamardian, Lulwa Aljeraisy, Basel Alomair, Dan Hendrycks, and David Wagner. 2024. Can LLMs Follow Simple Rules? https://doi.org/10.48550/arXiv.2311.04235

  46. [60]

    Mir Murtaza, Yamna Ahmed, Jawwad Ahmed Shamsi, Fahad Sherwani, and Mariam Usman. 2022. AI-Based Personalized E-Learning Systems: Issues, Chal- lenges, and Solutions. IEEE Access 10 (2022), 81323–81342. https://doi.org/10. 1109/ACCESS.2022.3193938

  47. [61]

    Ayesha Nadeem, Babak Abedin, and Olivera Marjanovic. 2020. Gender bias in AI: a review of contributing factors and mitigating strategies. In ACIS 2020 Proceedings. AIS Electronic Library (AISeL), 1–12. https://www.acis2020.org/

  48. [62]

    Maayan Nahmias, Yifat Perel. 2021. The Oversight of Content Moderation by AI: Impact Assessments and Their Limitations. Harvard Journal on Legislation 58 (2021), 145

  49. [63]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs) FAccT ’25, June 23–26, 2025, Athens, Gre...

  50. [64]

    I’m fully who I am

    Anaelia Ovalle, Palash Goyal, Jwala Dhamala, Zachary Jaggers, Kai-Wei Chang, Aram Galstyan, Richard Zemel, and Rahul Gupta. 2023. “I’m fully who I am”: To- wards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation. In Proceedings of the 20...

  51. [65]

    Sinead O’Connor and Helen Liu. 2024. Gender bias perpetuation and mitigation in AI technologies: challenges and opportunities. AI & SOCIETY 39, 4 (Aug. 2024), 2045–2057. https://doi.org/10.1007/s00146-023-01675-4

  52. [66]

    Ye Sul Park. 2024. White Default: Examining Racialized Biases Behind AI- Generated Images. Art Education 77, 4 (July 2024), 36–45. https://doi.org/10. 1080/00043125.2024.2330340

  53. [67]

    Parliamant and Council of the European Union. 2016. Regulation (EU) 2016/679 of the European Parliament and of the Council. https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX: 32016R0679&from=EN#d1e2051-1-1

  54. [68]

    Christopher Potts, Zhengxuan Wu, Atticus Geiger, and Douwe Kiela. 2021. DynaSent: A Dynamic Benchmark for Sentiment Analysis. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Langu...

  55. [70]

    Yanzhao Qin, Tao Zhang, Tao Zhang, Yanjun Shen, Wenjing Luo, Haoze Sun, Yan Zhang, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang, and Bin Cui. 2024. SysBench: Can Large Language Models Follow System Messages? https://doi.org/10.48550/arXiv.2408.10943

  56. [71]

    Inioluwa Deborah Raji and Joy Buolamwini. 2019. Actionable Auditing: Investi- gating the Impact of Publicly Naming Biased Performance Results of Commercial AI Products. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. ACM, Honolulu HI USA, 429–435. ht...

  57. [73]

    Brianna Richardson and Juan E. Gilbert. 2021. A Framework for Fairness: A Systematic Review of Existing Fair AI Solutions. https://doi.org/10.48550/arXiv. 2112.05700

  58. [74]

    Shahnewaz Karim Sakib and Anindya Bijoy Das. 2024. Challenging Fairness: A Comprehensive Exploration of Bias in LLM-Based Recommendations. https: //doi.org/10.48550/arXiv.2409.10825

  59. [75]

    Selbst, Danah Boyd, Sorelle A

    Andrew D. Selbst, Danah Boyd, Sorelle A. Friedler, Suresh Venkatasubramanian, and Janet Vertesi. 2019. Fairness and Abstraction in Sociotechnical Systems. In Proceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19). Association for Computing Machi...

  60. [76]

    Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, a...

  61. [77]

    Renee Shelby, Shalaleh Rismani, Kathryn Henne, AJung Moon, Negar Ros- tamzadeh, Paul Nicholas, N’Mah Yilla-Akbari, Jess Gallegos, Andrew Smart, Emilio Garcia, and Gurleen Virk. 2023. Sociotechnical Harms of Algorith- mic Systems: Scoping a Taxonomy for Harm Reduction. In Proce...

  62. [78]

    Tianhao Shen, Renren Jin, Yufei Huang, Chuang Liu, Weilong Dong, Zishan Guo, Xinwei Wu, Yan Liu, and Deyi Xiong. 2023. Large Language Model Alignment: A Survey. https://arxiv.org/abs/2309.15025

  63. [79]

    Emily Sheng, Kai-Wei Chang, Premkumar Natarajan, and Nanyun Peng. 2019. The Woman Worked as a Babysitter: On Biases in Language Generation. https: //arxiv.org/abs/1909.01326

  64. [80]

    Hari Shrawgi, Prasanjit Rath, Tushar Singhal, and Sandipan Dandapat. 2024. Uncovering Stereotypes in Large Language Models: A Task Complexity-based Approach. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume ...

  65. [81]

    Bangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Lo- geswaran, Moontae Lee, Dallas Card, and David Jurgens. 2024. You don’t need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instrum...

  66. [82]

    Jatinder Singh, Jennifer Cobbe, and Chris Norval. 2019. Decision Provenance: Harnessing Data Flow for Accountable Systems.IEEE Access 7 (2019), 6562–6574. https://doi.org/10.1109/ACCESS.2018.2887201

  67. [83]

    I’m sorry to hear that

    Eric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani, and Adina Williams. 2022. "I’m sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset. https://doi.org/10.48550/arXiv.2205. 09209

  68. [84]

    Nathalie A. Smuha. 2021. Beyond the Individual: Governing AI’s Societal Harm. https://papers.ssrn.com/abstract=3941956

  69. [85]

    Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar

    Till Speicher, Hoda Heidari, Nina Grgic-Hlaca, Krishna P. Gummadi, Adish Singla, Adrian Weller, and Muhammad Bilal Zafar. 2018. A Unified Approach to Quantifying Algorithmic Unfairness: Measuring Individual & Group Unfairness via Inequality Indices. In Proceedings of the 24th ...

  70. [86]

    Harini Suresh and John V. Guttag. 2021. A Framework for Understanding Sources of Harm throughout the Machine Learning Life Cycle. In Equity and Access in Algorithms, Mechanisms, and Optimization . 1–9. https://doi.org/10. 1145/3465416.3483305

  71. [87]

    Gray, Emma Pierson, and Karen Levy

    Harini Suresh, Emily Tseng, Meg Young, Mary L. Gray, Emma Pierson, and Karen Levy. 2024. Participation in the age of foundation models. In The 2024 ACM Conference on Fairness, Accountability, and Transparency. 1609–1621. https: //doi.org/10.1145/3630106.3658992

  72. [88]

    Chris Sweeney and Maryam Najafian. 2020. Reducing sentiment polarity for demographic attributes in word embeddings using adversarial learning. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (Barcelona, Spain) (FAT* ’20). Association for Com...

  73. [89]

    Alex Tamkin, Amanda Askell, Liane Lovitt, Esin Durmus, Nicholas Joseph, Shauna Kravec, Karina Nguyen, Jared Kaplan, and Deep Ganguli. 2023. Eval- uating and Mitigating Discrimination in Language Model Decisions. http: //arxiv.org/abs/2312.03689

  74. [90]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, Dan Bikel, Lukas Blecher, Cristian Canton Ferrer, Moya Chen, Guillem Cucurull, David Esiobu, Jude Fernandes, Jeremy Fu, W...

  75. [91]

    Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, and Alex Beutel. 2024. The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions. http://arxiv.org/abs/2404.13208

  76. [92]

    Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li

    Boxin Wang, Weixin Chen, Hengzhi Pei, Chulin Xie, Mintong Kang, Chenhui Zhang, Chejian Xu, Zidi Xiong, Ritik Dutta, Rylan Schaeffer, Sang T. Truong, Simran Arora, Mantas Mazeika, Dan Hendrycks, Zinan Lin, Yu Cheng, Sanmi Koyejo, Dawn Song, and Bo Li. 2024. DecodingTrust: A Com...

  77. [93]

    Yuan Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao, and Yi Fang. 2024. Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers. https://arxiv.org/abs/2404.03192

  78. [94]

    Laura Weidinger, Jonathan Uesato, Maribeth Rauh, Conor Griffin, Po-Sen Huang, John Mellor, Amelia Glaese, Myra Cheng, Borja Balle, Atoosa Kasirzadeh, Court- ney Biles, Sasha Brown, Zac Kenton, Will Hawkins, Tom Stepleton, Abeba Birhane, Lisa Anne Hendricks, Laura Rimell, Willi...

  79. [95]

    AI supply chain

    David Gray Widder and Dawn Nafus. 2023. Dislocated accountabilities in the “AI supply chain”: Modularity and developers’ notions of responsibility. Big Data Soc. 10, 1 (Jan. 2023). FAccT ’25, June 23–26, 2025, Athens, Greece Neumann et al

  80. [96]

    Bowen Xu, Shaoyu Wu, Kai Liu, and Lulu Hu. 2024. Mixture-of-Instructions: Comprehensive Alignment of a Large Language Model through the Mixture of Diverse System Prompting Instructions. https://doi.org/10.48550/arXiv.2404. 18410

  81. [97]

    Gummadi, and Adrian Weller

    Muhammad Bilal Zafar, Isabel Valera, Manuel Gomez Rodriguez, Krishna P. Gummadi, and Adrian Weller. 2017. From Parity to Preference-based Notions of Fairness in Classification. https://arxiv.org/abs/1707.00010

  82. [98]

    Zamfirescu-Pereira, Richmond Y

    J.D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, and Qian Yang

  83. [99]

    Zhehao Zhang, Ryan A. Rossi, Branislav Kveton, Yijia Shao, Diyi Yang, Hamed Zamani, Franck Dernoncourt, Joe Barrow, Tong Yu, Sungchul Kim, Ruiyi Zhang, Jiuxiang Gu, Tyler Derr, Hongjie Chen, Junda Wu, Xiang Chen, Zichao Wang, Subrata Mitra, Nedim Lipka, Nesreen Ahmed, and Yu W...

  84. [100]

    Dora Zhao, Angelina Wang, and Olga Russakovsky. 2021. Understanding and Evaluating Racial Biases in Image Captioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) . 14830–14840

  85. [101]

    In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems

    Why Johnny Can’t Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. ACM, Hamburg Germany, 1–21. https://doi.org/10.1145/ 3544548.3581388

  86. [102]

    Yu, and Lichao Sun

    Ce Zhou, Qian Li, Chen Li, Jun Yu, Yixin Liu, Guangjing Wang, Kai Zhang, Cheng Ji, Qiben Yan, Lifang He, Hao Peng, Jianxin Li, Jia Wu, Ziwei Liu, Pengtao Xie, Caiming Xiong, Jian Pei, Philip S. Yu, and Lichao Sun. 2024. A comprehensive survey on pretrained foundation models: a...

  87. [103]

    Lei Zhu, Xinjiang Wang, Wayne Zhang, and Rynson W. H. Lau. 2024. RelayAt- tention for Efficient Large Language Model Serving with Long System Prompts. https://doi.org/10.48550/arXiv.2402.14808 Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (L...

  88. [104]

    A Helpful Assistant

    Mingqian Zheng, Jiaxin Pei, and David Jurgens. 2023. Is "A Helpful Assistant" the Best Role for Large Language Models? A Systematic Evaluation of Social Roles in System Prompts. https://doi.org/10.48550/arXiv.2311.10054

  89. [2023]

    https://arxiv

    Towards Understanding Sycophancy in Language Models. https://arxiv. org/abs/2310.13548

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.