REVIEW 3 major objections 5 minor 1 cited by
Simulating Ethics: Using LLM Debate Panels to Model Deliberation on Medical Dilemmas
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read The paper claims that ADEPT, a system of LLM personas debating a ventilator-triage dilemma, shows that changing which ethical perspectives sit on the panel materially changes how the debate unfolds and how votes land, even when the facts…
desk verdict A transparent, honest proof-of-concept for LLM persona deliberation panels, but the central causal claim about panel composition rests on a single stochastic run per condition. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is ADEPT's three-phase deliberation loop: opening statements, rebuttals, and a secret ballot, each logged for audit. Personas are defined by a structured YAML schema that fixes their ethical principle, approach, core questions, decision criteria, deliberation style, forbidden moves, and citations, so each panel member acts as a 'moral lens' rather than a generic debater. The controlled input is a fixed scenario with four pre-specified allocation options, and the experimental instrument is the two-panel comparison: four personas are shared, two are swapped, and every prompt, response, and vote is recorded as an inspectable artefact.
What would settle it
Run both panels repeatedly, say twenty times each with different random seeds, and compare how often each continuing persona changes its vote. If the four shared personas shift their positions just as often when the panel membership is unchanged as when it is swapped, the claim that panel composition drives the outcome would be falsified.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that panel composition is a decisive variable in simulated ethical deliberation. With every factual input held fixed, swapping two of the six personas redirected the debate's attention toward moral injury, legal risk, and public trust, and it changed the final positions of the four personas who appeared in both debates: the Front-Line ICU Nurse and the Consequentialist moved from Option 1 to Option 2, while the Virtue Ethicist moved from Option 2 to Option 3. The stated upshot is that ADEPT offers 'a transparent, replicable workflow for running and analysing multi-agent AI debates in bioethics' and evidence that 'the moral perspectives included in such panels can materially change the outcome even when the factual inputs remain constant.'
Load-bearing premise
The load-bearing premise is that the differences between the two debates were caused by swapping the two personas, rather than by random variation in the language model, since each panel was run only once at temperature 0.7.
Editorial extensions
If this is right
- If ADEPT works as claimed, deliberative bodies could probe how the same clinical facts yield different ethical recommendations depending on which moral perspectives are represented.
- The auditable transcript and vote log would let ethicists, regulators, and clinicians trace exactly which arguments carried which votes, rather than only seeing a final recommendation.
- Replacing personas can surface concerns that would otherwise stay implicit, such as legal exposure under human-rights law or the risk that prognosis scores disadvantage disabled patients.
- A single majority policy can hide divergent justifications, so consensus in AI deliberation should be reported together with the coalition that produced it.
Reading between the lines
- Because each panel ran once at temperature 0.7, the observed differences could in principle be random model variation; running the same panels many times with different seeds would test whether the vote shifts are reliably caused by composition.
- A natural extension is to use ADEPT-style panels as an 'ethical red team' for draft policies, comparing the objections AI personas raise with those of real ethics committees before a guideline is adopted.
- Comparing the same personas across different underlying language models would separate persona-level effects from base-model leanings, which the paper flags as an open question.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ADEPT, an LLM-orchestrated multi-agent framework that stages structured ethical debates among personas defined by explicit ethical-theory and stakeholder specifications. It demonstrates the system on a ventilator-triage scenario using two six-persona panels that differ only in two members: Debate 1 includes a Catholic Bioethicist and a Care Ethicist, while Debate 2 substitutes a Deontologist and a Legal Arbiter. Both debates return the same majority policy outcome (Option 2, Clinical + Equity Weighted Lottery, by 4–2), but the supporting coalitions, argumentative themes, and some individual votes differ. The paper claims three contributions: a transparent, replicable workflow; evidence that panel composition can materially change outcomes even when factual inputs are fixed; and an analysis of implications and future directions for AI-mediated ethical deliberation.
Significance. If the causal claim about panel composition were adequately supported, ADEPT would be a useful and genuinely transparent tool for exploring normative pluralism, with clear value for ethics pedagogy, policy prototyping, and the study of deliberative dynamics. The paper's strengths include publicly available code and full debate transcripts, detailed persona specifications in Appendix B, direct quotation from the generated transcripts, and an unusually candid limitations section that identifies epistemic-reliability and black-box concerns. However, the central empirical contribution is not established by the current evidence: the comparison rests on a single stochastic run per panel condition, and several headline statements in the abstract and conclusion overstate what the data show. The paper is best read as a qualitative proof of concept, and the revision should either add the repeated-run evidence needed for the causal claim or explicitly scale the claim back to that scope.
major comments (3)
- [§3.1, §3.2.1, Tables 3–4] The paper's central causal claim—that replacing two personas redirected the debate and changed continuing personas' positions—rests on one debate run per panel at temperature 0.7, with no seeds, repeated runs, or statistical controls. Because every generation step is stochastic and the debate is a multi-turn chain, the observed differences in coalitions and thematic emphasis are indistinguishable from run-to-run noise. This concern is load-bearing for contribution (ii), which asserts that moral perspectives 'can materially change the outcome.' As written, the manuscript does not provide sufficient evidence for that claim; it would need multiple runs per condition with reported variability, or a clear reframing of the result as a single illustrative case study rather than evidence of a systematic effect.
- [Abstract, §4.1, Tables 3–4] The abstract states that the altered membership 'changed four continuing personas' final positions' and that the work provides 'evidence that the moral perspectives included in such panels can materially change the outcome.' This is not what the reported data show: the final policy outcome is identical in both debates (Option 2, 4–2), and the vote tables indicate that the Disability-Rights Advocate voted for Option 2 in both debates, so only three continuing personas changed their option. The abstract and Section 4.1 should be corrected to state the actual outcome and the actual number of vote changers, and the language of 'changed the outcome' should be replaced with a more precise description of changed coalitions and justifications.
- [§5.3, §7] The limitations section acknowledges that the study includes 'only two debate iterations' and that the consistent majority for Option 2 'could be influenced by subtle, uninstructed inclinations within the foundational model itself.' This directly undercuts the conclusion's claim that the comparative analysis revealed 'tangible shifts in ... the final policy preferences of the simulated committee.' The conclusion should be aligned with the limitations: at most, the paper shows that a single pair of runs produced different argumentative trajectories, not that panel composition reliably shifts deliberative outcomes. This is a logical inconsistency between the paper's explicit caveats and its summary claims.
minor comments (5)
- [§3.1] The phrase 'opening statements → rebuftals → secret ballot' contains a typo; 'rebuftals' should be 'rebuttals.'
- [§3.2] The subsection numbering jumps from 3.2.2 to 3.2.4, with no 3.2.3; the numbering should be corrected or a placeholder section added.
- [Tables 3 and 4] The table titles contain the typo 'Vote Talley'; this should read 'Vote Tally.'
- [§3.2.4] The phrase 'how varying ethical perspectives influences debate dynamics' should use the plural verb 'influence' to agree with 'perspectives.'
- [§4.3] The illustrative examples are drawn only from Debate 1, which limits their usefulness for the comparative claim; including at least one parallel exchange from Debate 2 would let readers assess the claimed shift in argumentative style directly.
Circularity Check
No circular derivation: the debate outputs are observed LLM transcripts, not quantities derived from fitted parameters or self-citation chains; the main threat to the paper's causal claim is stochastic single-run sampling, which is an internal-validity issue, not circularity.
full rationale
ADEPT's two debates produce observable transcripts and vote tallies that are not derived from fitted parameters, nor does the paper invoke a uniqueness theorem or load-bearing self-citation. The personas are hand-authored inputs (Section 3.2.2, Appendix B) and the LLM's behavior is stochastic (temperature 0.7, Section 3.1), so the reported shifts in arguments and coalitions are empirical outputs rather than analytic consequences of the design. One could say the new personas' arguments are unsurprising given their explicitly specified decision criteria (e.g., the Deontologist's criterion that perfect duties override consequences), but the paper reports these as observed deliberation, not as a prediction derived from an equation; it is a demonstration of prompt-following rather than a circular reduction. The single-run design and the abstract's 'four continuing personas' phrasing are accuracy and validity concerns (absence of replication, and a vote-table discrepancy), not circularity. Self-citations are to the author's own GitHub code and transcript files, which are standard data-availability links and do not carry the argument. Score 1 reflects a minor interpretive concern, not a circular derivation.
Assumptions & free parameters
free parameters (2)
- temperature =
0.7
- number of debate runs per panel =
1
assumptions (4)
- domain assumption LLM personas can faithfully embody distinct ethical frameworks
- domain assumption The two debates are comparable experimental conditions differing only in the two replaced personas
- domain assumption The LLM-assisted qualitative analysis plus human verification yields accurate interpretations
- domain assumption The scenario and options are representative of real allocation dilemmas
Cite this review
Pith. "Pith review of Simulating Ethics: Using LLM Debate Panels to Model Deliberation on Medical Dilemmas." pith.science (2026). https://pith.science/paper/4JZNJZ36
@misc{pith2026250521112,
author = {Pith},
title = {Pith review of: Simulating Ethics: Using LLM Debate Panels to Model Deliberation on Medical Dilemmas},
year = {2026},
howpublished = {\url{https://pith.science/paper/4JZNJZ36}},
note = {Machine review of arXiv:2505.21112}
}
read the original abstract
This paper introduces ADEPT, a system using Large Language Model (LLM) personas to simulate multi-perspective ethical debates. ADEPT assembles panels of 'AI personas', each embodying a distinct ethical framework or stakeholder perspective (like a deontologist, consequentialist, or disability rights advocate), to deliberate on complex moral issues. Its application is demonstrated through a scenario about prioritizing patients for a limited number of ventilators inspired by real-world challenges in allocating scarce medical resources. Two debates, each with six LLM personas, were conducted; they only differed in the moral viewpoints represented: one included a Catholic bioethicist and a care theorist, the other substituted a rule-based Kantian philosopher and a legal adviser. Both panels ultimately favoured the same policy -- a lottery system weighted for clinical need and fairness, crucially avoiding the withdrawal of ventilators for reallocation. However, each panel reached that conclusion through different lines of argument, and their voting coalitions shifted once duty- and rights-based voices were present. Examination of the debate transcripts shows that the altered membership redirected attention toward moral injury, legal risk and public trust, which in turn changed four continuing personas' final positions. The work offers three contributions: (i) a transparent, replicable workflow for running and analysing multi-agent AI debates in bioethics; (ii) evidence that the moral perspectives included in such panels can materially change the outcome even when the factual inputs remain constant; and (iii) an analysis of the implications and future directions for such AI-mediated approaches to ethical deliberation and policy.
Figures
Forward citations
Cited by 1 Pith paper
-
Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
Fine-tuning on speaker-attributed, action-tagged transcripts from public meetings lets LLM agents mimic government meeting participants well enough that human judges often cannot tell them from real people.
Reference graph
Works this paper leans on
-
[1]
Improving Factuality and Reasoning in Language Models through Multiagent Debate [Internet]
Du Y, Li S, Torralba A, Tenenbaum JB, Mordatch I. Improving Factuality and Reasoning in Language Models through Multiagent Debate [Internet]. arXiv; 2023 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2305.14325
arXiv 2023
-
[2]
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate [Internet]
Zhang Y, Yang X, Feng S, Wang D, Zhang Y, Song K. Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate [Internet]. arXiv; 2024 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2408.04472
arXiv 2024
-
[3]
AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation [Internet]
Wu Q, Bansal G, Zhang J, Wu Y, Li B, Zhu E, et al. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation [Internet]. arXiv; 2023 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2308.08155
arXiv 2023
-
[4]
Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs [Internet]
Smit A, Duckworth P, Grinsztajn N, Barrett TD, Pretorius A. Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs [Internet]. arXiv; 2024 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2311.17371
arXiv 2024
-
[5]
AI can help humans find common ground in democratic deliberation
Tessler MH, Bakker MA, Jarrett D, Sheahan H, Chadwick MJ, Koster R, et al. AI can help humans find common ground in democratic deliberation. Science. 2024 Oct 18;386(6719):eadq2852
work page 2024
-
[6]
LLM-Consensus: Multi- Agent Debate for Visual Misinformation Detection [Internet]
Lakara K, Channing G, Sock J, Rupprecht C, Torr P, Collomosse J, et al. LLM-Consensus: Multi- Agent Debate for Visual Misinformation Detection [Internet]. arXiv; 2025 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2410.20140
arXiv 2025
-
[7]
Principles of Biomedical Ethics
Childress JF, Beauchamp TL. Principles of Biomedical Ethics. New York Oxford: Oxford University Press Inc; 2019. 512 p
work page 2019
-
[8]
Health care ethics consultation: nature, goals, and competencies
Aulisio MP, Arnold RM, Youngner SJ. Health care ethics consultation: nature, goals, and competencies. A position paper from the Society for Health and Human Values-Society for Bioethics Consultation Task Force on Standards for Bioethics Consultation. Ann Intern Med. 2000 Jul 4;133(1):59–69
work page 2000
Show all 30 references
-
[9]
Health Care Ethics Consultation: An Update on Core Competencies and Emerging Standards from the American Society for Bioethics and Humanities’ Core Competencies Update Task Force
Tarzian AJ, ASBH Core Competencies Update Task Force1. Health Care Ethics Consultation: An Update on Core Competencies and Emerging Standards from the American Society for Bioethics and Humanities’ Core Competencies Update Task Force. The American Journal of Bioethics. 2013 Fe...
2013
-
[10]
Ethics Committees in Health Care Institutions [Internet]
American Medical Association. Ethics Committees in Health Care Institutions [Internet]. AMA
-
[11]
Roles and responsibilities of clinical ethics committees in priority setting
Magelssen M, Miljeteig I, Pedersen R, Førde R. Roles and responsibilities of clinical ethics committees in priority setting. BMC Medical Ethics. 2017 Dec 1;18(1):68
2017
-
[12]
Timmers M, van Dijck JTJM, van Wijk RPJ, Legrand V, van Veen E, Maas AIR, et al. How do 66 European institutional review boards approve one protocol for an international prospective observational study on traumatic brain injury? Experiences from the CENTER-TBI study. BMC Med E...
2020
-
[13]
A review of clinical ethics consultations in a regional healthcare system over a two-year timeframe
Anderson G, Hodge J, Fox D, Jutila S, McCarty C. A review of clinical ethics consultations in a regional healthcare system over a two-year timeframe. BMC Medical Ethics. 2024 Nov 9;25(1):127
2024
-
[14]
Clinical ethics consultations in psychiatric compared to non-psychiatric medical settings: characteristics and outcomes
Löbbing T, Carvalho Fernando S, Driessen M, Schulz M, Behrens J, Kobert KKB. Clinical ethics consultations in psychiatric compared to non-psychiatric medical settings: characteristics and outcomes. Heliyon. 2019 Jan 1;5(1):e01192
2019
-
[15]
Can digital tools foster ethical deliberation? Humanit Soc Sci Commun
Sleigh J, Hubbs S, Blasimme A, Vayena E. Can digital tools foster ethical deliberation? Humanit Soc Sci Commun. 2024 Jan 17;11(1):1–10
2024
-
[16]
If Multi-Agent Debate is the Answer, What is the Question? [Internet]
Zhang H, Cui Z, Wang X, Zhang Q, Wang Z, Wu D, et al. If Multi-Agent Debate is the Answer, What is the Question? [Internet]. arXiv; 2025 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2502.08788
2025 arXiv
-
[17]
Integrating Artificial Intelligence into Citizens’ Assemblies: Benefits, Concerns and Future Pathways
McKinney S. Integrating Artificial Intelligence into Citizens’ Assemblies: Benefits, Concerns and Future Pathways. Journal of Deliberative Democracy [Internet]. 2024 Jul 18 [cited 2025 May 22];20(1). Available from: https://delibdemjournal.org/article/id/1556/
2024
-
[18]
ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate [Internet]
Chan CM, Chen W, Su Y, Yu J, Xue W, Zhang S, et al. ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate [Internet]. arXiv; 2023 [cited 2025 May 22]. Available from: http://arxiv.org/abs/2308.07201
2023 arXiv
-
[19]
Tipping the Balance
Triem H, Ding Y. “Tipping the Balance”: Human Intervention in Large Language Model Multi- Agent Debate. Proceedings of the Association for Information Science and Technology. 2024;61(1):361–73
2024
-
[20]
LLM-Deliberation: Evaluating LLMs with Interactive Multi-Agent Negotiation Game [Internet]
Abdelnabi S, Gomaa A, Sivaprasad S, Schönherr L, Fritz M. LLM-Deliberation: Evaluating LLMs with Interactive Multi-Agent Negotiation Game [Internet]. OpenReview.net; 2023 [cited 2025 May 22]. Available from: https://openreview.net/forum?id=cfL8zApofK
2023
-
[21]
ADEPT [Internet]
Zohny H. ADEPT [Internet]. 2025a [cited 2025 May 27]. Available from: https://github.com/hazemzohny/ADEPT
2025
-
[22]
ADEPT Debate 1 [Internet]
Zohny H. ADEPT Debate 1 [Internet]. 2025b. Available from: https://github.com/hazemzohny/ADEPT/blob/main/debate_outputs/debate_1.txt
-
[23]
ADEPT Debate 2 [Internet]
Zohny H. ADEPT Debate 2 [Internet]. 2025c. Available from: https://github.com/hazemzohny/ADEPT/blob/main/debate_outputs/debate_2.txt
-
[24]
Fair Allocation of Scarce Medical Resources in the Time of Covid-19
Emanuel EJ, Persad G, Upshur R, Thome B, Parker M, Glickman A, et al. Fair Allocation of Scarce Medical Resources in the Time of Covid-19. New England Journal of Medicine. 2020 May 21;382(21):2049–55
2020
-
[25]
Too Many Patients…A Framework to Guide Statewide Allocation of Scarce Mechanical Ventilation During Disasters
Daugherty Biddison EL, Faden R, Gwon HS, Mareiniss DP, Regenberg AC, Schoch-Spana M, et al. Too Many Patients…A Framework to Guide Statewide Allocation of Scarce Mechanical Ventilation During Disasters. Chest. 2019;155(4):848–54
2019
-
[26]
Community health services two-hour urgent community response standard [Internet]
NHS. Community health services two-hour urgent community response standard [Internet]. 2021; Available from: https://www.england.nhs.uk/publication/community-health-services-two- hour-urgent-community-response-standard-guidance/
2021
-
[27]
LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding [Internet]
Chew R, Bollenbacher J, Wenger M, Speer J, Kim A. LLM-Assisted Content Analysis: Using Large Language Models to Support Deductive Coding [Internet]. arXiv; 2023 [cited 2025 May 24]. Available from: http://arxiv.org/abs/2306.14924
2023 arXiv
-
[28]
LLM-in-the-loop: Leveraging Large Language Model for Thematic Analysis [Internet]
Dai SC, Xiong A, Ku LW. LLM-in-the-loop: Leveraging Large Language Model for Thematic Analysis [Internet]. arXiv; 2023 [cited 2025 May 24]. Available from: http://arxiv.org/abs/2310.15100
2023 arXiv
-
[29]
LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue- Specific Characteristics to Enhance Contextual Understanding [Internet]
Na Y, Feng S. LLM-Assisted Automated Deductive Coding of Dialogue Data: Leveraging Dialogue- Specific Characteristics to Enhance Contextual Understanding [Internet]. arXiv; 2025 [cited 2025 May 24]. Available from: http://arxiv.org/abs/2504.19734
2025 arXiv
-
[2025]
Available from: https://code-medical-ethics.ama-assn.org/ethics-opinions/ethics- committees-health-care-institutions
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.