REVIEW 3 major objections 6 minor 26 references
Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)
T0 review · 3 major / 6 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read This paper reports that GPT-4o produced biased answers to about a third of adversarial palliative care questions and about a quarter of identity-swapped scenario pairs, and it introduces the PCAD datasets for auditing such bias.
desk verdict A useful first domain-specific bias audit with two new public datasets, but the headline bias rates rest on non-independent author ratings with near-zero interrater agreement, so treat the rates as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing instrument is the Palliative Care Adversarial Dataset (PCAD), built in two parts. PCAD-Direct contains 100 short, intentionally provocative questions that smuggle in a biased premise (for example, asking whether Black patients need fewer opioids because of a higher pain threshold), testing whether the model pushes back. PCAD-Counterfactual contains 84 scenario pairs that differ only by an identity attribute, operationalizing counterfactual fairness: if the ideal answer should be the same but the model treats the pair differently, that asymmetry is bias. Responses were generated from GPT-4o with temperature set to 0 and scored independently by three palliative care physicians using validated rubrics that classify bias into six dimensions, including allowing a biased premise, omitting structural explanations, and potential for withholding care.
What would settle it
Take the same 184-item dataset, have a larger panel of palliative care clinicians score it, and have a subset re-score a random sample after a few weeks; if consensus bias rates fall near zero, or the same rater cannot reproduce their own scores, the claim that GPT-4o's typical palliative care output is biased would not hold.
Extended reading notes
Core claim
The paper's central claim is that GPT-4o perpetuates bias in palliative care responses rather than merely reflecting neutral medical knowledge. For PCAD-Direct, 100 adversarial questions built on biased premises, the pooled bias rate across three raters was 0.33 (95% CI: 0.28, 0.38); the most common bias dimension was "allows biased premise" (0.47; 95% CI: 0.39, 0.55). For PCAD-Counterfactual, 84 pairs of scenarios identical except for age, ethnicity, or diagnosis, the pooled bias rate was 0.26 (95% CI: 0.20, 0.31), with "potential for withholding" as the most common source (0.25; 95% CI: 0.18, 0.34). Bias rates were not statistically different across the four care dimensions or three identity axes. The authors conclude that these biased outputs could contribute to inequitable clinical decision-making.
Load-bearing premise
The whole measurement rests on the assumption that three physicians with poor-to-fair agreement are rating the same underlying property; if their disagreements reflect rubric ambiguity rather than real bias in the model's answers, the pooled rates are not a stable estimate.
Editorial extensions
If this is right
- A clinician who asks GPT-4o for palliative care guidance can expect a biased answer in roughly one in three direct questions and one in four paired comparisons, so outputs need human review before influencing decisions.
- Because failure to challenge biased premises was the most common direct-question flaw, the model can normalize stereotypes when a user already holds them.
- The "potential for withholding" result means some responses could steer clinicians away from offering care or resources to patients based on age, ethnicity, or diagnosis.
- Bias did not concentrate in one identity axis or care dimension, so fixes targeted at a single topic or single group would be incomplete.
- The two PCAD datasets are reusable probes for auditing other large language models and tracking debiasing progress.
Reading between the lines
- Editorial: the low interrater agreement (Krippendorff's alpha 0.26 for direct and 0.09 for counterfactual bias presence) suggests the pooled rates are softer than they look; the "any-vote" rates of 0.55 and 0.57 may better reflect the range of defensible judgments.
- Editorial: because the study pinned one model version (gpt-4o-2024-05-13) at one temperature setting, the exact rates are a snapshot; the PCAD instruments could be rerun on current models to see whether bias persists or shifts.
- Editorial: an adversarial design that intentionally plants biased premises will naturally inflate the "allows biased premise" label; the clinically important question is whether the same failure appears in ordinary, non-adversarial queries.
- Editorial: extending the axes to gender, disability, sexual identity, or socioeconomic status may reveal patterns that ethnicity, age, and diagnosis do not capture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces two adversarial datasets, PCAD-Direct (100 questions) and PCAD-Counterfactual (84 paired scenarios), targeting four palliative-care dimensions and three identity axes, and uses them to probe GPT-4o. Three palliative-care physicians rated the model responses with rubrics adapted from Pfohl et al. The central reported result is that pooled bias rates are 0.33 for adversarial questions and 0.26 for counterfactual pairs, with 'allows biased premise' and 'potential for withholding' as the most frequent bias dimensions; the authors conclude that GPT-4o perpetuates bias in palliative care and that PCAD is a novel evaluation tool.
Significance. If the measurement were trustworthy, this would be a clinically important and timely result: it would be the first systematic demonstration of LLM bias in palliative care, and the public PCAD datasets would be reusable resources for auditing future models. The study has real strengths: the datasets are openly deposited, the prompts and decoding settings are reported, the rubrics are externally developed and validated, and the paper is unusually transparent about interrater reliability and about the dependence of some bias categories on the adversarial design. The significance, however, is currently conditional because the headline rates rest on pooled ratings with very low interrater agreement and on statistical tests that do not respect the nesting of ratings within items. These are fixable, but they are not presentation issues.
major comments (3)
- [Results, 'Interrater reliability'; Supplemental Tables 13 and 14] The primary measurement is a pooled bias rate, but Krippendorff's alpha is 0.26 for PCAD-Direct bias presence and 0.09 for PCAD-Counterfactual binary bias presence, and Fleiss' kappa ranges only from slight to fair. With this level of disagreement, the pooled per-rating rate is not an interpretable estimate of GPT-4o's bias: the raters are evidently applying the rubric with different thresholds, so the pooled number is an average over incompatible criteria. The paper should present majority-vote rates as the primary outcome and should demonstrate, through item-level agreement or consensus adjudication, that the reported rates are not an artifact of pooling.
- [Methods, 'Statistical analyses'] The pooled rates treat each rating as an independent sample, which inflates the effective sample size threefold (300 ratings for PCAD-Direct and 252 for PCAD-Counterfactual, but only 100 and 84 items, respectively). The bootstrap confidence intervals (e.g., 0.28 to 0.38 and 0.20 to 0.31) and the Kruskal-Wallis and Mann-Whitney tests on pooled ratings therefore ignore within-item correlation across raters and are likely too narrow. The authors should recompute all pooled estimates and tests with item-level clustering or cluster-bootstrap methods; without that reanalysis, the reported 'consistency' across dimensions and axes and even the headline rates are not statistically supported.
- [Methods, 'LLM-response evaluation'; Discussion, limitations] The three raters, listed as ON, FM, and SS, are all co-authors of the study, and the grading standardisation session was facilitated by the primary author; the limitations section does not disclose this non-independence. Because the rubric is subjective, the low interrater reliability is a serious concern, and the large gap between any-vote and majority-vote rates in the counterfactual data (0.57 vs. 0.15, Results section) shows that the pooled estimate is highly sensitive to the rater threshold. An independent rater panel, or at minimum a sensitivity analysis excluding co-author ratings, is needed before a claim that 'GPT-4o perpetuates bias' can be supported by these data.
minor comments (6)
- [Supplemental Table 4] The confidence interval for 'Inaccurate for axes of identity' is reported as (0.03, 0.01), which is impossible; the lower and upper bounds should be checked.
- [Supplemental Table 14] The Fleiss' kappa row for 'Should differ' reports a lower bound of -0.7, which is likely a typo for -0.07; please correct and re-check all interval bounds in this table.
- [Results, adversarial questions and Figure 4C] The ethnicity-specific rate is reported as 0.34 (95% CI: 0, 0.26, 0.44); the CI is malformed and appears to include zero, which is inconsistent with the claim that this is the highest rate among axes.
- [Methods, 'Experimental setting'] The text states that top_p was set to 1 and later says top_p was set to default; please clarify which value was actually used.
- [Introduction and Methods] The manuscript says the LLM responses were assessed 'across six dimensions' in one place and across four care dimensions elsewhere; the terminology should be made consistent by distinguishing care dimensions from bias dimensions.
- [Figure 3A] The caption says 'Ideal answers should differ,' but the reported 91% refers to cases where the ideal answers were judged not to differ; the caption and surrounding text should be reworded to avoid this apparent inversion.
Circularity Check
No construction-level circularity: the reported bias rates are measured with Pfohl et al.'s external rubric; the sole self-citation (ref 17) is minor and not load-bearing.
-
other
[Methods, 'Datasets' subsection (identity-axis selection; citation 17)]
"our analysis focused on ethnicity, age, and diagnosis. These three were selected based on evidence from our literature review.17"
Reference 17 is the corresponding author's own MSc thesis, the paper's only self-citation. It supports the choice of identity axes built into PCAD, which is an input to dataset construction rather than a derived result. The bias rates themselves are measured outcomes (a bias-free model would have scored near zero), and the same axes are independently supported by external references 3-11 in the introduction. The self-citation is therefore minor and not load-bearing for the central claim that GPT-4o perpetuates bias.
full rationale
This is an external measurement study rather than a derivation: PCAD questions are constructed, GPT-4o responses are generated, three palliative care physicians grade them with rubrics taken from Pfohl et al. (an external group), and the pooled rates are descriptive statistics of those grades. No fitted parameter is later renamed as a prediction, and no reported rate is equivalent to an input by construction: a model that rejected every biased premise would have scored near zero on 'allows biased premise' (measured 0.47), so the 0.33 and 0.26 pooled rates were not forced. The manuscript's own limitation passages all weigh as correctness risk rather than circularity: (1) low inter-rater reliability is disclosed in Results ('Reliability was rated as "poor" based on Krippendorff's alpha (α < 0.67)') and in the Discussion; (2) the dependence of 'allows biased premise' on the adversarial design is disclosed ('This high prevalence likely stems from the dataset's adversarial design, which introduced intentionally biased premises to test the model's ability to detect and reject them'). Two weaknesses are not disclosed and further reduce confidence: the three graders (ON, FM, SS) are also co-authors and dataset designers, so the gold-standard labels are not independent of the hypothesis being tested, and the statistical analysis 'treats each rating as an independent sample' even though ratings are nested within questions and raters. These are measurement-validity and statistical concerns, not construction-level circularity. The only self-citation (ref 17, the corresponding author's MSc thesis) supports selection of identity axes and is corroborated by external references 3-11, keeping the circularity score in the 0-2 'no significant circularity' band.
Assumptions & free parameters
assumptions (4)
- domain assumption Ethnicity, age, and diagnosis are key axes of identity that shape palliative care inequity.
- domain assumption The Pfohl et al. rubrics are valid and appropriate for palliative care contexts.
- domain assumption GPT-4o's July 2024 behavior is representative of large language models generally.
- domain assumption Pooled bias rates treating each grader-question rating as an independent sample are valid.
invented entities (1)
-
Palliative Care Adversarial Dataset (PCAD)
independent evidence
Cite this review
Pith. "Pith review of Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)." pith.science (2026). https://pith.science/paper/K3HWAD2I
@misc{pith2026250208073,
author = {Pith},
title = {Pith review of: Large language models perpetuate bias in palliative care: development and analysis of the Palliative Care Adversarial Dataset (PCAD)},
year = {2026},
howpublished = {\url{https://pith.science/paper/K3HWAD2I}},
note = {Machine review of arXiv:2502.08073}
}
read the original abstract
Bias and inequity in palliative care disproportionately affect marginalised groups. Large language models (LLMs), such as GPT-4o, hold potential to enhance care but risk perpetuating biases present in their training data. This study aimed to systematically evaluate whether GPT-4o propagates biases in palliative care responses using adversarially designed datasets. In July 2024, GPT-4o was probed using the Palliative Care Adversarial Dataset (PCAD), and responses were evaluated by three palliative care experts in Canada and the United Kingdom using validated bias rubrics. The PCAD comprised PCAD-Direct (100 adversarial questions) and PCAD-Counterfactual (84 paired scenarios). These datasets targeted four care dimensions (access to care, pain management, advance care planning, and place of death preferences) and three identity axes (ethnicity, age, and diagnosis). Bias was detected in a substantial proportion of responses. For adversarial questions, the pooled bias rate was 0.33 (95% confidence interval [CI]: 0.28, 0.38); "allows biased premise" was the most frequently identified source of bias (0.47; 95% CI: 0.39, 0.55), such as failing to challenge stereotypes. For counterfactual scenarios, the pooled bias rate was 0.26 (95% CI: 0.20, 0.31), with "potential for withholding" as the most frequently identified source of bias (0.25; 95% CI: 0.18, 0.34), such as withholding interventions based on identity. Bias rates were consistent across care dimensions and identity axes. GPT-4o perpetuates biases in palliative care, with implications for clinical decision-making and equity. The PCAD datasets provide novel tools to assess and address LLM bias in palliative care.
Reference graph
Works this paper leans on
-
[1]
Palliative care, https://www.who.int/news-room/fact-sheets/detail/palliative-care (accessed 9 July 2024)
work page 2024
-
[2]
Definition of INEQUITY, https://www.merriam-webster.com/dictionary/inequity (accessed 9 July 2024)
work page 2024
-
[3]
Social Inequalities in Palliative Care for Cancer Patients in the United States: A Structured Review
Elk R, Felder TM, Cayir E, et al. Social Inequalities in Palliative Care for Cancer Patients in the United States: A Structured Review. Semin Oncol Nurs 2018; 34: 303–315
work page 2018
-
[4]
Clarke G, Chapman E, Crooks J, et al. Does ethnicity affect pain management for people with advanced disease? A mixed methods cross-national systematic review of ‘very high’ Human Development Index English-speaking countries. BMC Palliat Care 2022; 21: 46
work page 2022
-
[5]
Crooks J, Trotter S, Patient Public Involvement Consortium, et al. How does ethnicity affect presence of advance care planning in care records for individuals with advanced disease? A mixed-methods systematic review. BMC Palliat Care 2023; 22: 43
work page 2023
-
[6]
Association between Chinese or South Asian ethnicity and end-of-life care in Ontario, Canada
Yarnell CJ, Fu L, Bonares MJ, et al. Association between Chinese or South Asian ethnicity and end-of-life care in Ontario, Canada. CMAJ 2020; 192: E266–E274
work page 2020
-
[7]
Hospice care access inequalities: a systematic review and narrative synthesis
Tobin J, Rogers A, Winterburn I, et al. Hospice care access inequalities: a systematic review and narrative synthesis. BMJ Support Palliat Care 2022; 12: 142–151
work page 2022
-
[8]
Prater LC, Wickizer T, Bose-Brill S. Examining Age Inequalities in Operationalized Components of Advance Care Planning: Truncation of the ACP Process With Age. J Pain Symptom Manage 2019; 57: 731–737
work page 2019
Show all 26 references
-
[9]
Equity in the provision of palliative care in the UK: review of evidence, https://eprints.lse.ac.uk/61550/1/equity_in_the_provision_of_paliative_care.pdf (2015)
Dixon J, King D, Matosevic T, et al. Equity in the provision of palliative care in the UK: review of evidence, https://eprints.lse.ac.uk/61550/1/equity_in_the_provision_of_paliative_care.pdf (2015)
2015
-
[10]
Changing patterns in place of cancer death in England: a population-based study
Gao W, Ho YK, Verne J, et al. Changing patterns in place of cancer death in England: a population-based study. PLoS Med 2013; 10: e1001410
2013
-
[11]
Comparison of terminally ill cancer- vs
Stiel S, Heckel M, Seifert A, et al. Comparison of terminally ill cancer- vs. non-cancer patients in specialized palliative home care in Germany – a single service analysis. BMC Palliat Care 2015; 14: 34
2015
-
[12]
Incorporating artificial intelligence in palliative care: opportunities and challenges
Munive Jesus U. Incorporating artificial intelligence in palliative care: opportunities and challenges. Hosp Palliat Med Int J 2024; 7: 81–83
2024
-
[13]
The future landscape of large language models in medicine
Clusmann J, Kolbinger FR, Muti HS, et al. The future landscape of large language models in medicine. Commun Med 2023; 3: 141
2023
-
[14]
Large language models show human-like content biases in transmission chain experiments
Acerbi A, Stubbersfield JM. Large language models show human-like content biases in transmission chain experiments. Proc Natl Acad Sci U S A 2023; 120: e2313790120
2023
-
[15]
Large language models propagate race-based medicine
Omiye JA, Lester JC, Spichak S, et al. Large language models propagate race-based medicine. NPJ Digit Med 2023; 6: 195
2023
-
[16]
A toolbox for surfacing health equity harms and biases in large language models
Pfohl SR, Cole-Lewis H, Sayres R, et al. A toolbox for surfacing health equity harms and biases in large language models. Nat Med 2024; 1–11. 15
2024
-
[17]
Evaluation of Biases in Large Language Models in Palliative Care
Akhras N. Evaluation of Biases in Large Language Models in Palliative Care . MSc (Thesis), King’s College London, 2024
2024
-
[18]
A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models
Pfohl SR, Cole-Lewis H, Sayres R, et al. A Toolbox for Surfacing Health Equity Harms and Biases in Large Language Models. arXiv [cs.CY] , http://arxiv.org/abs/2403.12025 (2024)
2024 arXiv
-
[19]
The TRIPOD-LLM reporting guideline for studies using large language models
Gallifant J, Afshar M, Ameen S, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med 2025; 31: 60–69
2025
-
[20]
Counterfactual Fairness
Kusner MJ, Loftus JR, Russell C, et al. Counterfactual Fairness. arXiv [stat.ML] , http://arxiv.org/abs/1703.06856 (2017)
2017 arXiv
-
[21]
Towards Expert-Level Medical Question Answering with Large Language Models
Singhal K, Tu T, Gottweis J, et al. Towards Expert-Level Medical Question Answering with Large Language Models. arXiv [cs.CL] , http://arxiv.org/abs/2305.09617 (2023)
2023 arXiv
-
[22]
K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient
Marzi G, Balzano M, Marchiori D. K-Alpha Calculator–Krippendorff’s Alpha Calculator: A user-friendly tool for computing Krippendorff's Alpha inter-rater reliability coefficient. MethodsX 2024; 12: 102545
2024
-
[23]
The measurement of observer agreement for categorical data
Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics 1977; 33: 159–174
1977
-
[24]
GPTBIAS: A comprehensive framework for evaluating bias in large language models
Zhao J, Fang M, Pan S, et al. GPTBIAS: A comprehensive framework for evaluating bias in large language models. arXiv [cs.CL] , http://arxiv.org/abs/2312.06315 (2023)
2023 arXiv
-
[25]
Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models
Ferrara E. Should ChatGPT be Biased? Challenges and Risks of Bias in Large Language Models. arXiv [cs.CY] , http://arxiv.org/abs/2304.03738 (2023)
2023 arXiv
-
[26]
The Equitable AI Research Roundtable (EARR): Towards Community-Based Decision Making in Responsible AI Development
Smith-Loud J, Smart A, Neal D, et al. The Equitable AI Research Roundtable (EARR): Towards Community-Based Decision Making in Responsible AI Development. arXiv [cs.AI] , http://arxiv.org/abs/2303.08177 (2023). 16 Supplemental Materials Supplemental Table 1. Independent Assessm...
2023 arXiv
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.