Pith. sign in

REVIEW 21 cited by

Agent Hospital: A Simulacrum of Hospital with Evolvable Medical Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.02957 v3 pith:EV3WXX23 submitted 2024-05-05 cs.AI

classification cs.AI
keywords agentsmedicalsimulacrumagenthospitalllmstreatingautonomous
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent rapid development of large language models (LLMs) has sparked a new wave of technological revolution in medical artificial intelligence (AI). While LLMs are designed to understand and generate text like a human, autonomous agents that utilize LLMs as their "brain" have exhibited capabilities beyond text processing such as planning, reflection, and using tools by enabling their "bodies" to interact with the environment. We introduce a simulacrum of hospital called Agent Hospital that simulates the entire process of treating illness, in which all patients, nurses, and doctors are LLM-powered autonomous agents. Within the simulacrum, doctor agents are able to evolve by treating a large number of patient agents without the need to label training data manually. After treating tens of thousands of patient agents in the simulacrum (human doctors may take several years in the real world), the evolved doctor agents outperform state-of-the-art medical agent methods on the MedQA benchmark comprising US Medical Licensing Examination (USMLE) test questions. Our methods of simulacrum construction and agent evolution have the potential in benefiting a broad range of applications beyond medical AI.

Discussion (0). Sign in to comment.

Forward citations

Cited by 21 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 36 citations worldwide. Full citation record

  1. MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education

    cs.CL 2026-07 conditional novelty 6.0 of 10

    MedGame converts static clinical case reports into structured interactive storytelling games and shows that fine-tuned open-source LLMs approach commercial performance on a new 5,000-case benchmark.

  2. Policy-Driven CT-Agent: Modeling Phase-Aware Diagnostic Control for Clinically Consistent CT Reasoning

    cs.LG 2026-07 conditional novelty 6.0 of 10

    An agent with structured CT evidence packages and guideline-guided control iteratively escalates imaging phases only when current evidence is judged insufficient for diagnosis.

  3. Explainable Cross-Disease Reasoning for Cardiovascular Risk Assessment from Low-Dose Computed Tomography

    cs.CV 2025-11 conditional novelty 6.0 of 10

    A multimodal framework using pulmonary findings, LLM-generated reasoning, and cardiac subvolume features achieves AUC 0.919 for CVD screening and 0.838 for CVD mortality from LDCT on NLST.

  4. TinyTroupe: An LLM-powered Multiagent Persona Simulation Toolkit

    cs.MA 2025-07 conditional novelty 6.0 of 10

    TinyTroupe provides a toolkit for fine-grained persona-based LLM multi-agent simulations with built-in support for population sampling, experimentation, and validation.

  5. DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making

    cs.AI 2025-07 conditional novelty 6.0 of 10

    DynamiCare is a multi-agent LLM framework that runs multi-round diagnostic dialogues with a dynamically adjusted specialist team, evaluated on a new 500-patient benchmark built from MIMIC-III.

  6. Modeling Earth-Scale Human-Like Societies with One Billion Agents

    cs.MA 2025-06 conditional novelty 6.0 of 10

    Light Society scales LLM-agent social simulations to one billion agents by substituting most LLM interactions with a distilled surrogate model.

  7. AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents

    cs.LG 2025-06 reject novelty 6.0 of 10

    AutoCT achieves test ROC-AUC 0.753, 0.639, and 0.702 on Phase I/II/III trial approval prediction using 100-sample subsets and LLM-generated features.

  8. MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems

    cs.MA 2025-05 conditional novelty 6.0 of 10

    A 5,000-prompt medical safety benchmark reveals that decentralized LLM multi-agent teams resist a malicious insider agent better than shared-pool teams, and a personality-screening defense partially restores safety.

  9. SafeAgent: Safeguarding LLM Agents via an Automated Risk Simulator

    cs.AI 2025-05 conditional novelty 6.0 of 10

    This paper introduces AutoSafe, an automated pipeline that generates agent risk scenarios, samples safe actions via self-reflection, and fine-tunes LLM agents to improve safety on synthetic and real-world benchmarks.

  10. CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs

    cs.AI 2026-08 conditional novelty 5.0 of 10

    CoPlan is a contestable, human-in-the-loop care planning interface that combines role-based AI argumentation with human review before final plan generation.

  11. OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence

    cs.AI 2026-03 conditional novelty 5.0 of 10

    OpenHospital is an interactive physician-patient multi-agent arena that improves clinical metrics via ground-truth reflection and reports cooperative behaviors as evidence of evolving LLM collective intelligence.

  12. InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

    cs.CL 2025-10 conditional novelty 5.0 of 10

    Rubric-based incremental RL with LLM-generated case-specific checklists lifts Qwen3-4B's HealthBench-Hard score from 7.0 to 27.5 with 2k samples, and improves InfoBench instruction-following from 42.0 to 82.9.

  13. Evolution in Simulation: AI-Agent School with Dual Memory for High-Fidelity Educational Dynamics

    cs.AI 2025-10 reject novelty 5.0 of 10

    An LLM-powered multi-agent school with dual experience/knowledge memory increasingly reproduces an expert-curated classroom script, with the full memory configuration scoring highest.

  14. Baichuan-M2: Scaling Medical Capability with Large Verifier System

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Baichuan-M2, a 32B medical LLM trained with a patient simulator and a clinical rubric generator as RL verifiers, reports state-of-the-art HealthBench scores (60.1 overall, 34.7 hard), ahead of all open-source models.

  15. Automated Clinical Problem Detection from SOAP Notes using a Collaborative Multi-Agent LLM Architecture

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A multi-agent LLM system where a manager assembles specialist agents to debate SOAP notes slightly outperforms a single LLM at detecting heart failure, kidney injury, and sepsis, with a macro F1 gain of about 0.01.

  16. MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A memory-augmented LLM planner that stores and retrieves page-level summaries from past trajectories improves success rates on mobile GUI task benchmarks.

  17. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5 of 10

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

  18. Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A scoping review of 557 studies finds medical agentic AI is technically promising but overwhelmingly validated on benchmarks, simulations, and retrospective data rather than in real clinical workflows.

  19. Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies

    cs.LG 2025-08 conditional novelty 4.0 of 10

    A survey that categorizes deep reinforcement learning scaling strategies into data, network, and training budget dimensions and outlines challenges for scaling DRL systems.

  20. InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation

    cs.AI 2025-08 reject novelty 4.0 of 10

    InqEduAgent fits a Gaussian process to simulated collaboration gains and then uses a Pareto front to pick learning partners, reporting small average gains over random pairing on six CMMLU domains.

  21. A Survey on Agent Workflow -- Status and Future

    cs.AI 2025-08 conditional novelty 3.0 of 10

    A review that classifies 24 agent workflow systems along functional and architectural axes and argues for standardization, optimization, and security work.

Pith tools