REVIEW 3 major objections 1 cited by
OpenHospital is a live hospital arena where physician agents evolve collective intelligence by interacting with dynamic patient agents, improving clinical metrics while lowering token cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 20:58 UTC pith:PNN5QQVL
load-bearing objection Useful hospital multi-agent arena with real metric trajectories, but the CI / data-in-agent-self story is mostly oracle-guided individual improvement. the 3 major comments →
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
As physician agents process successive batches of cases inside OpenHospital under a closed-loop reflection mechanism, their Examination Precision, Diagnostic Accuracy, and Treatment Plan Alignment all rise while total input tokens fall, and cooperative behaviors such as peer consultation emerge spontaneously; the arena therefore both evolves and quantifies LLM-based collective intelligence.
What carries the argument
The data-in-agent-self paradigm: physician agents receive no static case files; they must interact with patient agents (treated as dynamic entities) to obtain clinical information, forcing knowledge integration, multi-agent debate, and measurable evolution tracked by examination, diagnosis, treatment, and token metrics.
Load-bearing premise
The evolution claim rests on the premise that post-case self-critique against ground-truth diagnoses plus synthetic patients validated only by other language models produces genuine collective intelligence rather than guided individual improvement.
What would settle it
Train the same physician agents for the same number of batches using only interaction logs and no ground-truth reflection signal; if examination precision, diagnostic accuracy, treatment alignment, and spontaneous cross-department consultations fail to improve, the central claim is false.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OpenHospital, an interactive multi-agent arena for evolving and benchmarking LLM-based collective intelligence (CI) in a clinical setting. Physician agents interact with synthetic patient agents under a “data-in-agent-self” paradigm, using a multi-stage pipeline (DeepSeek-v3.1) to generate 12,000 patient records with comorbidities and long-tail diseases across 19 departments. A baseline of 38 physician agents (Agent-Kernel, Qwen3-Next-80B) is trained over 22 batches of ~500 cases each; after every case agents perform multi-dimensional self-critique against ground-truth diagnoses, examinations and treatment guidelines. Reported gains are Examination Precision 45.05%→61.31%, Diagnostic Accuracy 48.11%→57.34%, Treatment Plan Alignment 58.49%→61.52%, with declining total input tokens; qualitative case studies illustrate refined individual reasoning and spontaneous cross-departmental consultation. The authors claim the arena both fosters genuine CI and supplies rigorous dual metrics of medical proficiency and system efficiency.
Significance. If the central claim holds—that physician–patient interaction plus closed-loop reflection produces genuine, quantifiable collective intelligence rather than ordinary supervised improvement—the work would supply a reusable, privacy-safe evolutionary arena and a multi-dimensional clinical benchmark that current static MAS evaluations lack. Strengths include a large synthetic comorbidity-rich dataset, explicit multi-metric tracking (examination, diagnosis, treatment, tokens), open-source Agent-Kernel linkage, and qualitative evidence of peer consultation. These contributions would be useful to the multi-agent and medical-AI communities even if the strongest CI interpretation requires further controls.
major comments (3)
- §4.2 Baseline / closed-loop reflection: after every case agents receive multi-dimensional self-critique that explicitly synthesizes diagnostic accuracy against ground truth, examination efficiency and therapeutic safety. The same ground truth defines the three Medical Capability metrics in §4.1. Without an ablation that removes or degrades this oracle signal (or a pure interaction-only control), the reported trajectories (Fig. 4) and token decline (Fig. 5) are equally consistent with ordinary supervised individual improvement; the “data-in-agent-self CI” interpretation therefore remains untested and load-bearing for the central claim.
- §4.3 and Fig. 6 (cooperative behaviors): the action space already includes multi-agent consultation (§4.2). The single qualitative example of Infectious-Diseases → Cardiology consultation does not establish that cooperation is spontaneous or necessary rather than an enabled primitive. A non-consulting or single-agent control is required to support the claim that OpenHospital’s collaborative necessity drives emergent CI.
- §3 and §4.1 evaluation: both patient-agent validation (Medical Consistency 4.4113, Accuracy/Relevance/Persona scores) and Treatment Plan Alignment rely on LLM-as-judge (GPT-5.2 / Baichuan-M2). No human clinician inter-rater reliability, no error bars or statistical tests on the 22-batch trajectories, and no comparison against a non-LLM clinical gold standard are reported. This weakens the claim that the metrics constitute a “rigorous” benchmark of medical proficiency.
Circularity Check
Metric gains and 'CI evolution' reduce by construction to ground-truth-guided reflection that optimizes the same labels used for evaluation; cooperation is enabled by the pre-defined action space.
specific steps
-
fitted input called prediction
[Section 4.2 Baseline / Experimental Setup]
"Central to this baseline is a closed-loop reflection mechanism designed to drive autonomous evolution; after each case, agents engage in a multi-dimensional self-critique that synthesizes diagnostic accuracy against ground truth, examination efficiency, and therapeutic safety to bridge efficacy gaps. By integrating these diagnostic, investigative, and treatment reflections into a unified feedback loop, the agents systematically accumulate clinical experience and optimize their decision-making logic over time."
The reflection step injects the identical ground-truth labels that later define Diagnostic Accuracy, Examination Precision and Treatment Plan Alignment. Metric gains across the 22 batches are therefore the direct, expected product of supervised optimization against those labels, not an independent prediction of emergent CI from interaction data alone. Calling the resulting trajectories 'evolution of collective intelligence via data-in-agent-self' renames a fitted supervised loop as a first-principles result.
-
self definitional
[Section 4.1 Evaluation Metrics + Section 4.3 Evaluation Results]
"Examination Precision assesses the relevance and necessity of ordered tests. Defined as |E_pred ∩ E_std| / |E_pred| … Diagnostic Accuracy measures the correctness of the final consensus diagnosis. Formally, for a case i with ground truth D_true, the score is 1 if the agent's diagnosis D_pred = D_true … Treatment Plan Alignment evaluates therapeutic quality against gold-standard guidelines … As Figure 4 illustrates, the agents exhibit consistent upward trends … Examination Precision … 45.05% to 61.31% … Diagnostic Accuracy … 48.11% to 57.34% … Treatment Plan Alignment … 58.49% to 61.52%."
The three medical-capability metrics are defined directly in terms of the same ground-truth examinations, diagnoses and guidelines that the reflection mechanism optimizes after every case. Reporting improvement on those metrics therefore restates the success of the supervised critique loop; the quantities being 'predicted' (or claimed to emerge) are definitionally the quantities being fitted.
-
other
[Section 4.2 Baseline action space + Section 4.4 Case Studies / Figure 6]
"These agents operate within a sophisticated action space that encompasses patient perception, targeted inquiry, diagnostic examination, multi-agent consultation, and knowledge retrieval … analysis of the diagnostic process reveals the spontaneous emergence of sophisticated cooperative behaviors … the agent proactively initiates a consultation with the Cardiology Department … This interaction highlights the collaborative necessity intrinsic to OpenHospital"
Multi-agent consultation is an explicitly enumerated primitive of the action space. Observing agents invoke that primitive and then labeling the behavior 'spontaneous emergence of collective intelligence' is circular: the capability is present by construction of the environment rather than discovered from interaction data.
full rationale
The paper's central claim—that OpenHospital's data-in-agent-self interactions foster genuine collective intelligence, evidenced by rising Examination Precision / Diagnostic Accuracy / Treatment Plan Alignment and spontaneous peer consultation—is not an independent prediction. The baseline's closed-loop reflection explicitly critiques every case against ground-truth diagnoses, examination standards and therapeutic guidelines (the identical quantities that define the three medical metrics). Consequently the observed trajectories (45%→61%, 48%→57%, 58%→61%) and token reduction are the expected outcome of supervised individual improvement rather than emergent CI arising solely from physician–patient interaction. Cooperative behaviors are likewise licensed a priori by an action space that already contains multi-agent consultation. No ablation removes the oracle signal, so the 'evolution of CI' claim collapses to the training loop by construction. Patient-agent synthesis and LLM-as-judge validation introduce milder self-consistency circularity but are secondary. Self-citation of Agent-Kernel is present yet not load-bearing for the uniqueness of the result. Overall partial circularity of the fitted-input / self-definitional kind, score 6.
Axiom & Free-Parameter Ledger
free parameters (3)
- training batches / cases per batch =
22 batches × ~500 cases
- physician agent count and departmental distribution =
38 agents / 19 departments
- base LLM choice for agents =
Qwen3-Next-80B-A3B-Instruct
axioms (3)
- domain assumption LLM-generated synthetic patients with multi-stage refinement under epidemiological constraints are clinically coherent enough to serve as noumena for CI evolution.
- ad hoc to paper Closed-loop multi-dimensional self-critique against ground truth after each case produces data-in-agent-self collective intelligence rather than ordinary supervised improvement.
- domain assumption Examination Precision, Diagnostic Accuracy and LLM-judged Treatment Plan Alignment together constitute a robust quantitative measure of collective intelligence.
invented entities (3)
-
data-in-agent-self paradigm
no independent evidence
-
OpenHospital arena (thing-in-itself)
no independent evidence
-
Four pillars of realistic patient simulation (Clinical Correctness, Persona Diversity, Linguistic Fluency, Behavioral Realism)
no independent evidence
read the original abstract
Large Language Model (LLM)-based Collective Intelligence (CI) presents a promising approach to overcoming the data wall and continuously boosting the capabilities of LLM agents. However, there is currently no dedicated arena for evolving and benchmarking LLM-based CI. To address this gap, we introduce OpenHospital, an interactive arena where physician agents can evolve CI through interactions with patient agents. This arena employs a data-in-agent-self paradigm that rapidly enhances agent capabilities and provides robust evaluation metrics for benchmarking both medical proficiency and system efficiency. Experiments demonstrate the effectiveness of OpenHospital in both fostering and quantifying CI.
Forward citations
Cited by 1 Pith paper
-
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
EHR-derived standardized patients and dual-track evaluation reveal LLMs trail clinicians by 37.28 points on full psychiatric encounters, with mental-status assessment the main bottleneck.
Reference graph
Works this paper leans on
-
[1]
Citysim: Modeling urban behaviors and city dynamics with large-scale llm-driven agent simulation.Preprint, arXiv:2506.21805. Ziru Chen, Shijie Chen, Yuting Ning, Qianheng Zhang, Boshi Wang, Botao Yu, Yifei Li, Zeyi Liao, Chen Wei, Zitong Lu, Vishal Dey, Mingyi Xue, Frazier N. Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu...
-
[2]
Scienceagentbench: Toward rigorous as- sessment of language agents for data-driven scientific discovery.Preprint, arXiv:2410.05080. DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingx- uan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, and 1 oth- ers
-
[3]
Deepseek-v3 technical report.Preprint, arXiv:2412.19437. Chengfeng Dou, Chong Liu, Fan Yang, Fei Li, Jiyuan Jia, Mingyang Chen, Qiang Ju, Shuai Wang, Shunya Dang, Tianpeng Li, Xiangrong Zeng, Yijie Zhou, Chenzheng Zhu, Da Pan, Fei Deng, Guangwei Ai, Guosheng Dong, Hongda Zhang, Jinyang Tai, and 14 others
-
[4]
Baichuan-m2: Scaling medi- cal capability with large verifier system.Preprint, arXiv:2509.02208. Bofei Gao, Feifan Song, Zhe Yang, Zefan Cai, Yibo Miao, Qingxiu Dong, Lei Li, Chenghao Ma, Liang Chen, Runxin Xu, Zhengyang Tang, Benyou Wang, Daoguang Zan, Shanghaoran Quan, Ge Zhang, Lei Sha, Yichang Zhang, Xuancheng Ren, Tianyu Liu, and Baobao Chang
-
[5]
Alireza Ghafarollahi and Markus J
Omni-math: A univer- sal olympiad level mathematic benchmark for large language models.Preprint, arXiv:2410.07985. Alireza Ghafarollahi and Markus J. Buehler
-
[6]
Saeedeh Ghanadbashi and Fatemeh Golpayegani
Sci- agents: Automating scientific discovery through multi-agent intelligent graph reasoning.Preprint, arXiv:2409.05556. Saeedeh Ghanadbashi and Fatemeh Golpayegani
-
[7]
Ontology-enhanced decision-making for autonomous agents in dynamic and partially observable environ- ments.Preprint, arXiv:2405.17691. Pareesa Ameneh Golnari, Adarsh Kumarappan, Wen Wen, Xiaoyu Liu, Gabriel Ryan, Yuting Sun, Shengyu Fu, and Elsie Nallipogu
-
[8]
Devbench: A realistic, developer-informed benchmark for code generation models.Preprint, arXiv:2601.11895. Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryu- taro Tanno, Khaled Saab, Dan Popovici, Jacob Blum, 8 Fan Zhang, Katherine Chou, Avinatan Hassidim, Bu- rak ...
-
[9]
Towards an ai co-scientist.Preprint, arXiv:2502.18864. Sirui Hong, Mingchen Zhuge, Jonathan Chen, Xiawu Zheng, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, Chenyu Ran, Lingfeng Xiao, Chenglin Wu, and Jürgen Schmidhuber
-
[10]
Agents’ room: Nar- rative generation through multi-step collaboration. Preprint, arXiv:2410.02603. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan
-
[11]
Swe-bench: Can language mod- els resolve real-world github issues?Preprint, arXiv:2310.06770. Immanuel Kant. 1781.Critique of Pure Reason. Johann Friedrich Hartknoch, Riga, Russian Empire. Junkai Li, Yunghwei Lai, Weitao Li, Jingyi Ren, Meng Zhang, Xinhui Kang, Siyu Wang, Peng Li, Ya-Qin Zhang, Weizhi Ma, and Yang Liu
-
[12]
Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qing- min Liao
Agent hospi- tal: A simulacrum of hospital with evolvable medical agents.Preprint, arXiv:2405.02957. Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qing- min Liao
-
[13]
Jianhao Lin, Lexuan Sun, and Yixin Yan
Econagent: Large language model- empowered agents for simulating macroeconomic activities.Preprint, arXiv:2310.10436. Jianhao Lin, Lexuan Sun, and Yixin Yan
-
[14]
Simu- lating macroeconomic expectations using llm agents. Preprint, arXiv:2505.17648. Zhao Mandi, Shreeya Jain, and Shuran Song
-
[15]
Roco: Dialectic multi-robot collaboration with large language models.Preprint, arXiv:2307.04738. Yuren Mao, Peigen Liu, Xinjian Wang, Rui Ding, Jing Miao, Hui Zou, Mingjie Qi, Wanxiang Luo, Longbin Lai, Kai Wang, Zhengping Qian, Peilun Yang, Yun- jun Gao, and Ying Zhang
-
[16]
Agent-kernel: A microkernel multi-agent system framework for adap- tive social simulation powered by llms.Preprint, arXiv:2512.01610. OpenAI
-
[17]
Generative agents: Interac- tive simulacra of human behavior.Preprint, arXiv:2304.03442. Jinghua Piao, Yuwei Yan, Jun Zhang, Nian Li, Junbo Yan, Xiaochong Lan, Zhihong Lu, Zhiheng Zheng, Jing Yi Wang, Di Zhou, Chen Gao, Fengli Xu, Fang Zhang, Ke Rong, Jun Su, and Yong Li
-
[18]
Agentsociety: Large-scale simulation of llm-driven generative agents advances understanding of human behaviors and society.Preprint, arXiv:2502.08691. Chen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, and Maosong Sun
-
[19]
Jun Sashihara, Yukihisa Fujita, Kota Nakamura, Masahiro Kuwahara, and Teruaki Hayashi
Chatdev: Communica- tive agents for software development.Preprint, arXiv:2307.07924. Jun Sashihara, Yukihisa Fujita, Kota Nakamura, Masahiro Kuwahara, and Teruaki Hayashi
-
[20]
Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor
Llm-based multi-agent system for simulating strate- gic and goal-oriented data marketplaces.Preprint, arXiv:2511.13233. Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis, Jeffrey Jopling, and Michael Moor
-
[21]
Agentclinic: a multimodal agent benchmark to evalu- ate ai in simulated clinical environments.Preprint, arXiv:2405.07960. Haoyang Su, Renqi Chen, Shixiang Tang, Zhenfei Yin, Xinzhe Zheng, Jinzhe Li, Biqing Qi, Qi Wu, Hui Li, Wanli Ouyang, Philip Torr, Bowen Zhou, and Nan- qing Dong
-
[22]
Many heads are better than one: Improved scientific idea generation by a llm-based multi-agent system.Preprint, arXiv:2410.09403. Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W White, Doug Burger, and Chi Wang
-
[23]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Zhou, and 1 others
Autogen: Enabling next-gen llm ap- plications via multi-agent conversation.Preprint, arXiv:2308.08155. An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Zhou, and 1 others
-
[24]
Bangguo Yu, Qihao Yuan, Kailai Li, Hamidreza Kasaei, and Ming Cao
Qwen3 techni- cal report.Preprint, arXiv:2505.09388. Bangguo Yu, Qihao Yuan, Kailai Li, Hamidreza Kasaei, and Ming Cao. 2025a. Co-navgpt: Multi-robot coop- erative visual semantic navigation using vision lan- guage models.Preprint, arXiv:2310.07937. Tian Yu, Ken Shi, Zixin Zhao, and Gerald Penn. 2025b. Multi-agent based character simulation for story writ...
Pith/arXiv arXiv 2025
-
[25]
Simulating classroom education with LLM- empowered agents. InProceedings of the 2025 Con- ference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (V olume 1: Long Papers), 9 pages 10364–10379, Albuquerque, New Mexico. As- sociation for Computational Linguistics. Xuhui Zhou, Hao Zhu, Leen...
2025
-
[26]
Sotopia: Interactive evaluation for social intelligence in language agents.Preprint, arXiv:2310.11667. Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhe Wang, Zhenhailong Wang, Cheng Qian, Xiangru Tang, Heng Ji, and Jiaxuan You
-
[27]
Multiagentbench: Evaluating the col- laboration and competition of llm agents.Preprint, arXiv:2503.01935. 10
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.