REVIEW 21 cited by
Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The recent rapid development of large language models (LLMs) has sparked a new wave of technological revolution in medical artificial intelligence (AI). While LLMs are designed to understand and generate text like a human, autonomous agents that utilize LLMs as their "brain" have exhibited capabilities beyond text processing such as planning, reflection, and using tools by enabling their "bodies" to interact with the environment. We introduce a simulacrum of hospital called Agent Hospital that simulates the entire process of treating illness, in which all patients, nurses, and doctors are LLM-powered autonomous agents. Within the simulacrum, doctor agents are able to evolve by treating a large number of patient agents without the need to label training data manually. After treating tens of thousands of patient agents in the simulacrum (human doctors may take several years in the real world), the evolved doctor agents outperform state-of-the-art medical agent methods on the MedQA benchmark comprising US Medical Licensing Examination (USMLE) test questions. Our methods of simulacrum construction and agent evolution have the potential in benefiting a broad range of applications beyond medical AI.
Forward citations
Cited by 21 Pith papers
-
MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
MedGame converts static clinical case reports into structured interactive storytelling games and shows that fine-tuned open-source LLMs approach commercial performance on a new 5,000-case benchmark.
-
Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning
An agent with structured CT evidence packages and guideline-guided control iteratively escalates imaging phases only when current evidence is judged insufficient for diagnosis.
-
Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography
A multimodal framework using pulmonary findings, LLM-generated reasoning, and cardiac subvolume features achieves AUC 0.919 for CVD screening and 0.838 for CVD mortality from LDCT on NLST.
-
TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit
TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.
-
DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
DynamiCare is a multi-agent LLM framework that runs multi-round diagnostic dialogues with a dynamically adjusted specialist team, evaluated on a new 500-patient benchmark built from MIMIC-III.
-
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.
-
AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
AutoCT achieves test ROC-AUC 0.753, 0.639, and 0.702 on Phase I/II/III trial approval prediction using 100-sample subsets and LLM-generated features.
-
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
A 5,000-prompt medical safety benchmark reveals that decentralized LLM multi-agent teams resist a malicious insider agent better than shared-pool teams, and a personality-screening defense partially restores safety.
-
SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator
This paper introduces AutoSafe, an automated pipeline that generates agent risk scenarios, samples safe actions via self-reflection, and fine-tunes LLM agents to improve safety on synthetic and real-world benchmarks.
-
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
CoPlan is a contestable, human-in-the-loop care planning interface that combines role-based AI argumentation with human review before final plan generation.
-
OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence
OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.
-
InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training
Rubric-based incremental RL with LLM-generated case-specific checklists lifts Qwen3-4B's HealthBench-Hard score from 7.0 to 27.5 with 2k samples, and improves InfoBench instruction-following from 42.0 to 82.9.
-
Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics
An LLM-powered multi-agent school with dual experience/knowledge memory increasingly reproduces an expert-curated classroom script, with the full memory configuration scoring highest.
-
Baichuan-M2: Scaling Medical Capability with Large Verifier System
Baichuan-M2, a 32B medical LLM trained with a patient simulator and a clinical rubric generator as RL verifiers, reports state-of-the-art HealthBench scores (60.1 overall, 34.7 hard), ahead of all open-source models.
-
Automated Clinical Problem Detection from SOAP Notes using a Collaborative Multi-Agent LLM Architecture
A multi-agent LLM system where a manager assembles specialist agents to debate SOAP notes slightly outperforms a single LLM at detecting heart failure, kidney injury, and sepsis, with a macro F1 gain of about 0.01.
-
MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation
A memory-augmented LLM planner that stores and retrieves page-level summaries from past trajectories improves success rates on mobile GUI task benchmarks.
-
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.
-
Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
A scoping review of 557 studies finds medical agentic AI is technically promising but overwhelmingly validated on benchmarks, simulations, and retrospective data rather than in real clinical workflows.
-
Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies
A survey that categorizes deep reinforcement learning scaling strategies into data, network, and training budget dimensions and outlines challenges for scaling DRL systems.
-
InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation
InqEduAgent fits a Gaussian process to simulated collaboration gains and then uses a Pareto front to pick learning partners, reporting small average gains over random pairing on six CMMLU domains.
-
A Survey on Agent Workflow -- Status and Future
A review that classifies 24 agent workflow systems along functional and architectural axes and argues for standardization, optimization, and security work.
Discussion (0). Sign in to comment.