Pith. sign in

REVIEW 3 major objections 6 minor 26 references

Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)

T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This paper reports that GPT-4o produced biased answers to about a third of adversarial palliative care questions and about a quarter of identity-swapped scenario pairs, and it introduces the PCAD datasets for auditing such bias.

desk verdict A useful first domain-specific bias audit with two new public datasets, but the headline bias rates rest on non-independent author ratings with near-zero interrater agreement, so treat the rates as provisional. read the letter →

arxiv 2502.08073 v1 pith:K3HWAD2I submitted 2025-02-12 cs.CY

classification cs.CY
keywords palliativecarelargelanguagemodelsGPT-4oalgorithmicbiasadversarialdatasetcounterfactualfairnesshealthequityend-of-life
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Using two new adversarial datasets, this paper asks whether GPT-4o reproduces known inequities in palliative care: less access, poorer pain management, and assumptions tied to ethnicity, age, and diagnosis. Three palliative care physicians rated the model's responses with validated bias rubrics. The result is that a pooled 33% of 100 deliberately provocative questions and 26% of 84 counterfactual scenario pairs were judged biased. The most common failure was accepting a biased premise without challenging it, and in paired scenarios the model sometimes suggested withholding care based on identity. If the measurement holds, clinicians and regulators cannot assume LLM advice about end-of-life care is neutral.

What carries the argument

The load-bearing instrument is the Palliative Care Adversarial Dataset (PCAD), built in two parts. PCAD-Direct contains 100 short, intentionally provocative questions that smuggle in a biased premise (for example, asking whether Black patients need fewer opioids because of a higher pain threshold), testing whether the model pushes back. PCAD-Counterfactual contains 84 scenario pairs that differ only by an identity attribute, operationalizing counterfactual fairness: if the ideal answer should be the same but the model treats the pair differently, that asymmetry is bias. Responses were generated from GPT-4o with temperature set to 0 and scored independently by three palliative care physicians using validated rubrics that classify bias into six dimensions, including allowing a biased premise, omitting structural explanations, and potential for withholding care.

What would settle it

Take the same 184-item dataset, have a larger panel of palliative care clinicians score it, and have a subset re-score a random sample after a few weeks; if consensus bias rates fall near zero, or the same rater cannot reproduce their own scores, the claim that GPT-4o's typical palliative care output is biased would not hold.

Watch

Extended reading notes

Core claim

The paper's central claim is that GPT-4o perpetuates bias in palliative care responses rather than merely reflecting neutral medical knowledge. For PCAD-Direct, 100 adversarial questions built on biased premises, the pooled bias rate across three raters was 0.33 (95% CI: 0.28, 0.38); the most common bias dimension was "allows biased premise" (0.47; 95% CI: 0.39, 0.55). For PCAD-Counterfactual, 84 pairs of scenarios identical except for age, ethnicity, or diagnosis, the pooled bias rate was 0.26 (95% CI: 0.20, 0.31), with "potential for withholding" as the most common source (0.25; 95% CI: 0.18, 0.34). Bias rates were not statistically different across the four care dimensions or three identity axes. The authors conclude that these biased outputs could contribute to inequitable clinical decision-making.

Load-bearing premise

The whole measurement rests on the assumption that three physicians with poor-to-fair agreement are rating the same underlying property; if their disagreements reflect rubric ambiguity rather than real bias in the model's answers, the pooled rates are not a stable estimate.

Editorial extensions

If this is right

  • A clinician who asks GPT-4o for palliative care guidance can expect a biased answer in roughly one in three direct questions and one in four paired comparisons, so outputs need human review before influencing decisions.
  • Because failure to challenge biased premises was the most common direct-question flaw, the model can normalize stereotypes when a user already holds them.
  • The "potential for withholding" result means some responses could steer clinicians away from offering care or resources to patients based on age, ethnicity, or diagnosis.
  • Bias did not concentrate in one identity axis or care dimension, so fixes targeted at a single topic or single group would be incomplete.
  • The two PCAD datasets are reusable probes for auditing other large language models and tracking debiasing progress.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial: the low interrater agreement (Krippendorff's alpha 0.26 for direct and 0.09 for counterfactual bias presence) suggests the pooled rates are softer than they look; the "any-vote" rates of 0.55 and 0.57 may better reflect the range of defensible judgments.
  • Editorial: because the study pinned one model version (gpt-4o-2024-05-13) at one temperature setting, the exact rates are a snapshot; the PCAD instruments could be rerun on current models to see whether bias persists or shifts.
  • Editorial: an adversarial design that intentionally plants biased premises will naturally inflate the "allows biased premise" label; the clinically important question is whether the same failure appears in ordinary, non-adversarial queries.
  • Editorial: extending the axes to gender, disability, sexual identity, or socioeconomic status may reveal patterns that ethnicity, age, and diagnosis do not capture.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces two adversarial datasets, PCAD-Direct (100 questions) and PCAD-Counterfactual (84 paired scenarios), targeting four palliative-care dimensions and three identity axes, and uses them to probe GPT-4o. Three palliative-care physicians rated the model responses with rubrics adapted from Pfohl et al. The central reported result is that pooled bias rates are 0.33 for adversarial questions and 0.26 for counterfactual pairs, with 'allows biased premise' and 'potential for withholding' as the most frequent bias dimensions; the authors conclude that GPT-4o perpetuates bias in palliative care and that PCAD is a novel evaluation tool.

Significance. If the measurement were trustworthy, this would be a clinically important and timely result: it would be the first systematic demonstration of LLM bias in palliative care, and the public PCAD datasets would be reusable resources for auditing future models. The study has real strengths: the datasets are openly deposited, the prompts and decoding settings are reported, the rubrics are externally developed and validated, and the paper is unusually transparent about interrater reliability and about the dependence of some bias categories on the adversarial design. The significance, however, is currently conditional because the headline rates rest on pooled ratings with very low interrater agreement and on statistical tests that do not respect the nesting of ratings within items. These are fixable, but they are not presentation issues.

major comments (3)
  1. [Results, 'Interrater reliability'; Supplemental Tables 13 and 14] The primary measurement is a pooled bias rate, but Krippendorff's alpha is 0.26 for PCAD-Direct bias presence and 0.09 for PCAD-Counterfactual binary bias presence, and Fleiss' kappa ranges only from slight to fair. With this level of disagreement, the pooled per-rating rate is not an interpretable estimate of GPT-4o's bias: the raters are evidently applying the rubric with different thresholds, so the pooled number is an average over incompatible criteria. The paper should present majority-vote rates as the primary outcome and should demonstrate, through item-level agreement or consensus adjudication, that the reported rates are not an artifact of pooling.
  2. [Methods, 'Statistical analyses'] The pooled rates treat each rating as an independent sample, which inflates the effective sample size threefold (300 ratings for PCAD-Direct and 252 for PCAD-Counterfactual, but only 100 and 84 items, respectively). The bootstrap confidence intervals (e.g., 0.28 to 0.38 and 0.20 to 0.31) and the Kruskal-Wallis and Mann-Whitney tests on pooled ratings therefore ignore within-item correlation across raters and are likely too narrow. The authors should recompute all pooled estimates and tests with item-level clustering or cluster-bootstrap methods; without that reanalysis, the reported 'consistency' across dimensions and axes and even the headline rates are not statistically supported.
  3. [Methods, 'LLM-response evaluation'; Discussion, limitations] The three raters, listed as ON, FM, and SS, are all co-authors of the study, and the grading standardisation session was facilitated by the primary author; the limitations section does not disclose this non-independence. Because the rubric is subjective, the low interrater reliability is a serious concern, and the large gap between any-vote and majority-vote rates in the counterfactual data (0.57 vs. 0.15, Results section) shows that the pooled estimate is highly sensitive to the rater threshold. An independent rater panel, or at minimum a sensitivity analysis excluding co-author ratings, is needed before a claim that 'GPT-4o perpetuates bias' can be supported by these data.
minor comments (6)
  1. [Supplemental Table 4] The confidence interval for 'Inaccurate for axes of identity' is reported as (0.03, 0.01), which is impossible; the lower and upper bounds should be checked.
  2. [Supplemental Table 14] The Fleiss' kappa row for 'Should differ' reports a lower bound of -0.7, which is likely a typo for -0.07; please correct and re-check all interval bounds in this table.
  3. [Results, adversarial questions and Figure 4C] The ethnicity-specific rate is reported as 0.34 (95% CI: 0, 0.26, 0.44); the CI is malformed and appears to include zero, which is inconsistent with the claim that this is the highest rate among axes.
  4. [Methods, 'Experimental setting'] The text states that top_p was set to 1 and later says top_p was set to default; please clarify which value was actually used.
  5. [Introduction and Methods] The manuscript says the LLM responses were assessed 'across six dimensions' in one place and across four care dimensions elsewhere; the terminology should be made consistent by distinguishing care dimensions from bias dimensions.
  6. [Figure 3A] The caption says 'Ideal answers should differ,' but the reported 91% refers to cases where the ideal answers were judged not to differ; the caption and surrounding text should be reworded to avoid this apparent inversion.

Circularity Check

1 steps flagged · score 2.0 of 10

No construction-level circularity: the reported bias rates are measured with Pfohl et al.'s external rubric; the sole self-citation (ref 17) is minor and not load-bearing.

  1. other [Methods, 'Datasets' subsection (identity-axis selection; citation 17)]
    "our analysis focused on ethnicity, age, and diagnosis. These three were selected based on evidence from our literature review.17"

    Reference 17 is the corresponding author's own MSc thesis, the paper's only self-citation. It supports the choice of identity axes built into PCAD, which is an input to dataset construction rather than a derived result. The bias rates themselves are measured outcomes (a bias-free model would have scored near zero), and the same axes are independently supported by external references 3-11 in the introduction. The self-citation is therefore minor and not load-bearing for the central claim that GPT-4o perpetuates bias.

full rationale

This is an external measurement study rather than a derivation: PCAD questions are constructed, GPT-4o responses are generated, three palliative care physicians grade them with rubrics taken from Pfohl et al. (an external group), and the pooled rates are descriptive statistics of those grades. No fitted parameter is later renamed as a prediction, and no reported rate is equivalent to an input by construction: a model that rejected every biased premise would have scored near zero on 'allows biased premise' (measured 0.47), so the 0.33 and 0.26 pooled rates were not forced. The manuscript's own limitation passages all weigh as correctness risk rather than circularity: (1) low inter-rater reliability is disclosed in Results ('Reliability was rated as "poor" based on Krippendorff's alpha (α < 0.67)') and in the Discussion; (2) the dependence of 'allows biased premise' on the adversarial design is disclosed ('This high prevalence likely stems from the dataset's adversarial design, which introduced intentionally biased premises to test the model's ability to detect and reject them'). Two weaknesses are not disclosed and further reduce confidence: the three graders (ON, FM, SS) are also co-authors and dataset designers, so the gold-standard labels are not independent of the hypothesis being tested, and the statistical analysis 'treats each rating as an independent sample' even though ratings are nested within questions and raters. These are measurement-validity and statistical concerns, not construction-level circularity. The only self-citation (ref 17, the corresponding author's MSc thesis) supports selection of identity axes and is corroborated by external references 3-11, keeping the circularity score in the 0-2 'no significant circularity' band.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

The central measurement rests on four main assumptions: the choice of identity axes, the validity of the adapted rubrics, the representativeness of a single model for the broader claim, and the statistical handling of clustered ratings. No free parameters are fitted; the results are empirical rates rather than model predictions.

assumptions (4)
  • domain assumption Ethnicity, age, and diagnosis are key axes of identity that shape palliative care inequity.
    Selected based on a literature review (reference 17), but the set is not exhaustive; the study does not test other axes such as disability, gender, or homelessness.
  • domain assumption The Pfohl et al. rubrics are valid and appropriate for palliative care contexts.
    The rubrics were previously validated in a general medical LLM bias toolbox; the authors adapt examples to palliative care but do not revalidate them in this domain.
  • domain assumption GPT-4o's July 2024 behavior is representative of large language models generally.
    The title generalizes to 'LLMs', but only one model was tested at one temperature and one prompt in July 2024.
  • domain assumption Pooled bias rates treating each grader-question rating as an independent sample are valid.
    Ratings are clustered within questions and raters; the independence assumption is violated, likely making the bootstrap confidence intervals too narrow.
invented entities (1)
  • Palliative Care Adversarial Dataset (PCAD) independent evidence
    purpose: Benchmark datasets for probing LLM bias in palliative care, divided into PCAD-Direct and PCAD-Counterfactual.
    The datasets are publicly released on Figshare (dx.doi.org/10.6084/m9.figshare.28396016) and can be reused by other researchers for future audits.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)." pith.science (2026). https://pith.science/paper/K3HWAD2I

@misc{pith2026250208073,
  author       = {Pith},
  title        = {Pith review of: Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/K3HWAD2I}},
  note         = {Machine review of arXiv:2502.08073}
}
read the original abstract

Bias and inequity in palliative care disproportionately affect marginalised groups. Large language models (LLMs), such as GPT-4o, hold potential to enhance care but risk perpetuating biases present in their training data. This study aimed to systematically evaluate whether GPT-4o propagates biases in palliative care responses using adversarially designed datasets. In July 2024, GPT-4o was probed using the Palliative Care Adversarial Dataset (PCAD), and responses were evaluated by three palliative care experts in Canada and the United Kingdom using validated bias rubrics. The PCAD comprised PCAD-Direct (100 adversarial questions) and PCAD-Counterfactual (84 paired scenarios). These datasets targeted four care dimensions (access to care, pain management, advance care planning, and place of death preferences) and three identity axes (ethnicity, age, and diagnosis). Bias was detected in a substantial proportion of responses. For adversarial questions, the pooled bias rate was 0.33 (95% confidence interval [CI]: 0.28, 0.38); "allows biased premise" was the most frequently identified source of bias (0.47; 95% CI: 0.39, 0.55), such as failing to challenge stereotypes. For counterfactual scenarios, the pooled bias rate was 0.26 (95% CI: 0.20, 0.31), with "potential for withholding" as the most frequently identified source of bias (0.25; 95% CI: 0.18, 0.34), such as withholding interventions based on identity. Bias rates were consistent across care dimensions and identity axes. GPT-4o perpetuates biases in palliative care, with implications for clinical decision-making and equity. The PCAD datasets provide novel tools to assess and address LLM bias in palliative care.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 21 canonical work pages

  1. [1]

    Palliative care, https://www.who.int/news-room/fact-sheets/detail/palliative-care (accessed 9 July 2024)

  2. [2]

    Definition of INEQUITY, https://www.merriam-webster.com/dictionary/inequity (accessed 9 July 2024)

  3. [3]

    Social Inequalities in Palliative Care for Cancer Patients in the United States: A Structured Review

    Elk R, Felder TM, Cayir E, et al. Social Inequalities in Palliative Care for Cancer Patients in the United States: A Structured Review. Semin Oncol Nurs 2018; 34: 303–315

  4. [4]

    Clarke G, Chapman E, Crooks J, et al. Does ethnicity affect pain management for people with advanced disease? A mixed methods cross-national systematic review of ‘very high’ Human Development Index English-speaking countries. BMC Palliat Care 2022; 21: 46

  5. [5]

    How does ethnicity affect presence of advance care planning in care records for individuals with advanced disease? A mixed-methods systematic review

    Crooks J, Trotter S, Patient Public Involvement Consortium, et al. How does ethnicity affect presence of advance care planning in care records for individuals with advanced disease? A mixed-methods systematic review. BMC Palliat Care 2023; 22: 43

  6. [6]

    Association between Chinese or South Asian ethnicity and end-of-life care in Ontario, Canada

    Yarnell CJ, Fu L, Bonares MJ, et al. Association between Chinese or South Asian ethnicity and end-of-life care in Ontario, Canada. CMAJ 2020; 192: E266–E274

  7. [7]

    Hospice care access inequalities: a systematic review and narrative synthesis

    Tobin J, Rogers A, Winterburn I, et al. Hospice care access inequalities: a systematic review and narrative synthesis. BMJ Support Palliat Care 2022; 12: 142–151

  8. [8]

    Examining Age Inequalities in Operationalized Components of Advance Care Planning: Truncation of the ACP Process With Age

    Prater LC, Wickizer T, Bose-Brill S. Examining Age Inequalities in Operationalized Components of Advance Care Planning: Truncation of the ACP Process With Age. J Pain Symptom Manage 2019; 57: 731–737

Show all 26 references
  1. [9]

    Equity in the provision of palliative care in the UK: review of evidence, https://eprints.lse.ac.uk/61550/1/equity_in_the_provision_of_paliative_care.pdf (2015)

    Dixon J, King D, Matosevic T, et al. Equity in the provision of palliative care in the UK: review of evidence, https://eprints.lse.ac.uk/61550/1/equity_in_the_provision_of_paliative_care.pdf (2015)

  2. [10]

    Changing patterns in place of cancer death in England: a population-based study

    Gao W, Ho YK, Verne J, et al. Changing patterns in place of cancer death in England: a population-based study. PLoS Med 2013; 10: e1001410

  3. [11]

    Comparison of terminally ill cancer- vs

    Stiel S, Heckel M, Seifert A, et al. Comparison of terminally ill cancer- vs. non-cancer patients in specialized palliative home care in Germany – a single service analysis. BMC Palliat Care 2015; 14: 34

  4. [12]

    Incorporating artificial intelligence in palliative care: opportunities and challenges

    Munive Jesus U. Incorporating artificial intelligence in palliative care: opportunities and challenges. Hosp Palliat Med Int J 2024; 7: 81–83

  5. [13]

    The future landscape of large language models in medicine

    Clusmann J, Kolbinger FR, Muti HS, et al. The future landscape of large language models in medicine. Commun Med 2023; 3: 141

  6. [14]

    Large language models show human-like content biases in transmission chain experiments

    Acerbi A, Stubbersfield JM. Large language models show human-like content biases in transmission chain experiments. Proc Natl Acad Sci U S A 2023; 120: e2313790120

  7. [15]

    Large language models propagate race-based medicine

    Omiye JA, Lester JC, Spichak S, et al. Large language models propagate race-based medicine. NPJ Digit Med 2023; 6: 195

  8. [16]

    A toolbox for surfacing health equity harms and biases in large language models

    Pfohl SR, Cole-Lewis H, Sayres R, et al. A toolbox for surfacing health equity harms and biases in large language models. Nat Med 2024; 1–11. 15

  9. [17]

    Evaluation of Biases in Large Language Models in Palliative Care

    Akhras N. Evaluation of Biases in Large Language Models in Palliative Care . MSc (Thesis), King’s College London, 2024

  10. [18]

    A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models

    Pfohl SR, Cole-Lewis H, Sayres R, et al. A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models. arXiv [cs.CY] , http://arxiv.org/abs/2403.12025 (2024)

  11. [19]

    The TRIPOD-LLM reporting guideline for studies using large language models

    Gallifant J, Afshar M, Ameen S, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med 2025; 31: 60–69

  12. [20]

    Counterfactual Fairness

    Kusner MJ, Loftus JR, Russell C, et al. Counterfactual Fairness. arXiv [stat.ML] , http://arxiv.org/abs/1703.06856 (2017)

  13. [21]

    Towards Expert-Level Medical Question Answering with Large Language Models

    Singhal K, Tu T, Gottweis J, et al. Towards Expert-Level Medical Question Answering with Large Language Models. arXiv [cs.CL] , http://arxiv.org/abs/2305.09617 (2023)

  14. [22]

    K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient

    Marzi G, Balzano M, Marchiori D. K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient. MethodsX 2024; 12: 102545

  15. [23]

    The measurement of observer agreement for categorical data

    Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics 1977; 33: 159–174

  16. [24]

    GPTBIAS: A comprehensive framework for evaluating bias in large language models

    Zhao J, Fang M, Pan S, et al. GPTBIAS: A comprehensive framework for evaluating bias in large language models. arXiv [cs.CL] , http://arxiv.org/abs/2312.06315 (2023)

  17. [25]

    Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models

    Ferrara E. Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models. arXiv [cs.CY] , http://arxiv.org/abs/2304.03738 (2023)

  18. [26]

    The Equitable AI Research Roundtable (EARR): Towards Community-Based Decision Making in Responsible AI Development

    Smith-Loud J, Smart A, Neal D, et al. The Equitable AI Research Roundtable (EARR): Towards Community-Based Decision Making in Responsible AI Development. arXiv [cs.AI] , http://arxiv.org/abs/2303.08177 (2023). 16 Supplemental Materials Supplemental Table 1. Independent Assessm...

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.