Pith. sign in

REVIEW 5 cited by

Polaris: A Safety-focused LLM Constellation Architecture for Healthcare

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13313 v1 pith:EJHXZNFE submitted 2024-03-20 cs.AI cs.CL

classification cs.AIcs.CL
keywords healthcareagentssystemmedicalnursesconstellationconversationspolaris
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We develop Polaris, the first safety-focused LLM constellation for real-time patient-AI healthcare conversations. Unlike prior LLM works in healthcare focusing on tasks like question answering, our work specifically focuses on long multi-turn voice conversations. Our one-trillion parameter constellation system is composed of several multibillion parameter LLMs as co-operative agents: a stateful primary agent that focuses on driving an engaging conversation and several specialist support agents focused on healthcare tasks performed by nurses to increase safety and reduce hallucinations. We develop a sophisticated training protocol for iterative co-training of the agents that optimize for diverse objectives. We train our models on proprietary data, clinical care plans, healthcare regulatory documents, medical manuals, and other medical reasoning documents. We align our models to speak like medical professionals, using organic healthcare conversations and simulated ones between patient actors and experienced nurses. This allows our system to express unique capabilities such as rapport building, trust building, empathy and bedside manner. Finally, we present the first comprehensive clinician evaluation of an LLM system for healthcare. We recruited over 1100 U.S. licensed nurses and over 130 U.S. licensed physicians to perform end-to-end conversational evaluations of our system by posing as patients and rating the system on several measures. We demonstrate Polaris performs on par with human nurses on aggregate across dimensions such as medical safety, clinical readiness, conversational quality, and bedside manner. Additionally, we conduct a challenging task-based evaluation of the individual specialist support agents, where we demonstrate our LLM agents significantly outperform a much larger general-purpose LLM (GPT-4) as well as from its own medium-size class (LLaMA-2 70B).

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ClinConsensus: A Physician-Calibrated Benchmark for Evaluating Clinical Rubric Coverage in Chinese Medical LLMs

    cs.CL 2026-03 conditional novelty 6.0 of 10

    ClinConsensus, a Chinese medical benchmark with physician-calibrated thresholded rubric coverage, shows LLMs achieve 39.6–52.1% rubric accuracy but only 17.8–32.9% CACS@10, exposing a large coverage gap.

  2. Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A multi-agent LLM assistant with DSPy-optimized prompts improved Romanian doctors' written communication quality and patient satisfaction in a live telemedicine deployment, but the evaluation is confounded by self-sel...

  3. MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems

    cs.MA 2025-05 conditional novelty 6.0 of 10

    A 5,000-prompt medical safety benchmark reveals that decentralized LLM multi-agent teams resist a malicious insider agent better than shared-pool teams, and a personality-screening defense partially restores safety.

  4. AI Agents for Conversational Patient Triage: Preliminary Simulation-Based Evaluation with Real-World EHR Data

    cs.CL 2025-06 reject novelty 5.0 of 10

    A patient simulator built from EHR vignettes was rated consistent with those vignettes in 97.7% of 519 conversations by two clinicians, while the AI triage system's top three diagnoses contained the most likely diagno...

  5. Robustness tests for biomedical foundation models should tailor to specifications

    cs.SE 2025-02 conditional novelty 5.0 of 10

    The paper proposes task-tailored 'robustness specifications' to guide robustness testing of biomedical foundation models across their lifecycle.

Pith tools