Pith. sign in

REVIEW 3 cited by

DiversityMedQA: Assessing Demographic Biases in Medical Diagnosis using Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.01497 v2 pith:7Q4Z7YPD submitted 2024-09-02 cs.CL

classification cs.CL
keywords medicaldemographicdiversitymedqaacrossbenchmarkbiasesdiagnosislanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As large language models (LLMs) gain traction in healthcare, concerns about their susceptibility to demographic biases are growing. We introduce {DiversityMedQA}, a novel benchmark designed to assess LLM responses to medical queries across diverse patient demographics, such as gender and ethnicity. By perturbing questions from the MedQA dataset, which comprises medical board exam questions, we created a benchmark that captures the nuanced differences in medical diagnosis across varying patient profiles. Our findings reveal notable discrepancies in model performance when tested against these demographic variations. Furthermore, to ensure the perturbations were accurate, we also propose a filtering strategy that validates each perturbation. By releasing DiversityMedQA, we provide a resource for evaluating and mitigating demographic bias in LLM medical diagnoses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CXXCrafter: An LLM-Based Agent for Automated C/C++ Open Source Software Building

    cs.SE 2025-05 conditional novelty 7.0 of 10

    An LLM-driven agent, CXXCrafter, automatically builds 587 of 752 C/C++ open-source projects (78%), beating default build commands (39%) and bare LLMs (32 to 38%).

  2. Ensuring Safety and Trust: Analyzing the Risks of Large Language Models in Medicine

    cs.CL 2024-11 conditional novelty 6.0 of 10

    MedGuard, a 1,000-question expert-verified benchmark, shows current medical LLMs lag human physicians on safety across fairness, privacy, robustness, resilience, and truthfulness.

  3. One Size Fits None: Rethinking Fairness in Medical AI

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Across mortality, graft failure, and triage prediction, subgroup analysis shows that aggregate accuracy hides clinically relevant disparities, including lower PRC for Black and female patients.

Pith tools