{"id":"b65d0420-1108-4617-b756-b72e797693ed","arxiv_id":"2608.09861","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A video-based medical AI system matched or outperformed primary care physicians in a randomized 100-scenario simulated telehealth study, while patients still preferred human doctors for rapport.","lead":"Researchers at Google tested a new voice-and-video version of their medical AI assistant, AMIE, against real primary care doctors in 100 simulated medical visits. When expert doctors scored the consultations, the AI matched or beat the human doctors on taking medical history, making diagnoses, and guiding patients through physical exam steps over video.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AMIE's diagnostic and clinical-reasoning advantage may stem from an offline multi-draft post-questionnaire, not from real-time video consultation; separating live from offline scoring would settle the central claim.","rationale":"The reader's weakest assumption was the fairness of the PCP baseline, focusing on patient actors, actable scenarios, and PCP cameras being off. That concern is real and is explicitly acknowledged in the manuscript's limitations. I focused instead on a distinct, less-acknowledged asymmetry: AMIE's post-encounter questionnaire uses multi-draft synthesis with additional inference-time compute, while PCPs receive no analogous boost. Because clinical evaluators saw the post-questionnaire alongside the live consultation, the reported superiority in diagnosis, clinical reasoning, and treatment planning may be substantially driven by this offline step. This matters because those domains are pillars of the paper's 'expert-level real-time consultation' claim. The two concerns are related in that both question whether the comparison supports the headline, but they are technically different: the reader's concern is about external validity of the human baseline, while mine is about construct validity of the real-time metric. Neither invalidates the study's descriptive value, and both are addressable by design or re-analysis, so the reader's CONDITIONAL verdict remains appropriate without a verdict change.","tokens_in":46124,"tokens_out":5739,"duration_ms":57596,"concrete_test":"Re-score all 100 consultations with clinical evaluators (or a fresh panel) using only the video and transcript, with the post-questionnaire redacted, for the top-1 DDx accuracy and the diagnosis/clinical-reasoning and treatment-planning rubric items; compare AMIE (Video) vs PCP (Video) on these live-only scores. If the AMIE advantage shrinks to non-significance, the real-time expert-level claim must be re-scoped to the post-encounter reasoning stage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that AMIE (Video) performs at expert level in real-time clinical video consultations. The headline comparisons, however, may not measure purely real-time performance. Section 2.3.1 states that after each encounter AMIE answers a structured post-questionnaire while 'leveraging inference-time compute scaling via multi-draft synthesis' to improve recommendation quality; PCPs complete the same questionnaire without that capability. Section 2.3.3 says clinical evaluators assessed consultation recordings, transcripts, and post-questionnaire responses. Results in Section 3.1 — top-1 diagnosis 91% vs 77% (p=0.039), clinical reasoning 90% vs 76% (p=2.0e-5), treatment planning 79% vs 67% (p=8.9e-5) — are therefore consistent with an advantage produced in the offline, latency-unconstrained post-encounter step rather than during the live video dialogue. The limitations section (5.1) does not flag this asymmetry; it notes only that the post-encounter step uses inference-time scaling. If the diagnostic and clinical-reasoning advantage disappears when only the live video/transcript is scored, the 'expert-level real-time video consultation' claim oversells what was measured: it conflates real-time consultation skill with extra offline compute allocated to one arm.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents AMIE (Video), a Gemini-based multi-agent system for real-time audio-visual clinical consultations, and reports a randomized OSCE-style study comparing AMIE (Video), AMIE (Text), and 10 primary care physicians (PCPs) across 100 actable clinical scenarios enacted by 15 professional patient actors. A separate panel of 20 clinical evaluators rated consultations using general and case-specific rubrics, and patient actors provided communication and modality preferences. The central claim is that AMIE (Video) achieves expert-level performance in real-time video consultations, with clinical evaluators rating it on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination, while PCPs retain an edge in rapport and partnership building. The paper also contributes a taxonomy of telehealth audio-visual cues and an automated evaluation suite used to guide system development.","tokens_in":46334,"tokens_out":4903,"duration_ms":47756,"significance":"If the headline claim holds, this would be a notable milestone for medical AI: a controlled, randomized comparison of a real-time video-capable system against board-certified PCPs with independent clinical evaluators, pre-specified rubrics, scenario-level blocking, bootstrap confidence intervals, and FDR-corrected tests. The design has real strengths: the scenario packs were developed externally, the evaluator panel was independent of the consulting PCPs, and the multi-arm structure with a text-only ablation is appropriate for isolating modality effects. The paper also ships a useful taxonomy of audio-visual telehealth competencies. However, two load-bearing asymmetries prevent the current evidence from fully supporting the abstract's 'expert-level real-time video consultation' claim: the post-encounter questionnaire gives AMIE an offline inference-compute advantage that is not available to the PCPs, and the PCP baseline was forced to keep cameras off, which differentially affects communication-related ratings. The central claim is defensible in principle, but it needs re-analysis or reframing before publication.","major_comments":[{"comment":"The headline diagnostic and clinical-reasoning comparisons do not isolate real-time consultation skill. Section 2.3.1 states that after each encounter AMIE answers a structured post-questionnaire while 'leveraging inference-time compute scaling via multi-draft synthesis' to improve recommendation quality, whereas PCPs complete the same questionnaire without that capability. Section 2.3.3 and Appendix A.6.7 state that clinical evaluators assessed consultation recordings, transcripts, and post-questionnaire responses, and the top-1 diagnostic accuracy reported in Section 3.1 is based on the generated differential diagnosis, which comes from the post-questionnaire. The large observed advantages for AMIE (Video) over PCPs in top-1 diagnosis (91% vs 77%), clinical reasoning (90% vs 76%), and treatment planning (79% vs 67%) are therefore consistent with an offline, latency-unconstrained advantage rather than a real-time video consultation advantage. The paper should report scores based only on the live video/transcript, or at minimum provide a sensitivity analysis that separates the post-questionnaire contribution and a discussion of how much of the gap survives when the post-encounter step is removed. Without this, the abstract's claim of 'expert-level AI in real-time clinical video consultations' exceeds what the data support.","section":"§2.3.1, §2.3.3, §3.1"},{"comment":"The human baseline is not a fair comparator for communication-related claims because PCPs were required to keep their cameras off, as described in Section 2.3.1 and Appendix A.6.1. This removes gaze, facial expression, and other non-verbal channels that are normal components of video consultations, and Section 5.1 acknowledges that it reduces the ecological validity of the human baseline. Several headline findings concern communication, empathy, rapport, and 'on par or better' ratings from both clinical evaluators and patient actors; the camera-off requirement may therefore inflate AMIE's relative performance on those dimensions. The authors should either include a visible-physician condition, perform a sensitivity analysis that excludes or down-weights communication items, or substantially temper the communication and rapport comparisons in the abstract and Section 3.1. The current discussion, while candid about the limitation, does not carry the caveat into the interpretation of the main comparative claims.","section":"§2.3.1, §A.6.1, §5.1"}],"minor_comments":[{"comment":"The text says 'While not significant, directionally AMIE (Video) scored higher than AMIE (Text) on all criteria, as reflected by the rating for \"Happy to see again\" (91% vs 74%, p=0.038)', but a p-value of 0.038 is below the conventional 0.05 threshold; clarify whether this p-value survived FDR correction and state the significance criterion explicitly.","section":"§3.2"},{"comment":"The text refers to 'six case-specific rubric domains' but the case-specific rubrics described in Appendix A.6.9 comprise five domains (history taking, perception and examination, clinical reasoning, treatment planning, communication) plus an overall score; please reconcile the count in the figure and prose.","section":"§3.1, Figure 5B"},{"comment":"The statement that mean conversation duration was comparable (8.94 minutes for AMIE Video vs 9.30 minutes for PCPs) is presented without confidence intervals or a statistical test; please add these if the similarity is intended as a claim, or mark it as descriptive only.","section":"§2.3.1"},{"comment":"The abstract reports '30 primary care physicians (PCPs)' while the main text distinguishes 10 consulting PCPs and 20 clinical evaluator PCPs; please clarify this in the abstract to avoid conflating the two roles.","section":"Abstract, §2.3"}],"recommendation":"major_revision","confidential_remarks":"The central confound identified in Major Comment 1 is the most important issue. If the authors can re-analyze the data to show that the diagnostic and clinical-reasoning advantages persist when only the live video/transcript is scored, the paper would be much stronger. The camera-off baseline also deserves more prominent treatment in the abstract and conclusions. The paper is from Google Research with declared competing interests; the disclosure is adequate, but the framing of 'expert-level' should be calibrated to the actual comparison conditions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is the first randomized OSCE to show a video-based AI consultation system matching or beating PCPs, and the evaluation is genuinely well built. But the headline “expert-level real-time video consultation” is not fully supported, because the clinical evaluators scored the post-questionnaire as well as the live encounter, and the AI got extra offline compute on that questionnaire.\n\nWhat is new and good: the asynchronous three-agent architecture (Talker, Planner, Perception) is a real engineering contribution, cutting mean turn latency from roughly 21 seconds to 2.6 seconds. The telehealth taxonomy and automated evaluation suite are thoughtful and will be useful to the field. The OSCE itself is strong: 100 clinical scenarios, professional patient actors, independent PCP evaluators, scenario-level blocking, FDR-corrected Wilcoxon tests, and bootstrap confidence intervals. The reported effects are large — overall rubric 83% vs 68%, top-1 diagnosis 91% vs 77%, perception and examination 74% vs 47%.\n\nThe main soft spot is the offline post-questionnaire. Section 2.3.1 states that after each encounter AMIE uses multi-draft synthesis under inference-time compute scaling, while PCPs complete the same questionnaire without that boost. Section 2.3.3 says evaluators assessed consultation recordings, transcripts, and post-questionnaire responses. So the diagnostic accuracy and clinical reasoning numbers, which drive the “expert-level” claim, could mostly reflect offline compute rather than real-time consultation skill. The limitations section acknowledges the post-encounter step but does not flag this asymmetry as a threat to the central comparison. That is a load-bearing caveat.\n\nOther soft spots: PCPs had cameras off, scenarios were restricted to actable presentations, and there are no inter-rater reliability metrics and no released data or code. These are secondary and mostly acknowledged in the paper.\n\nWho is this for: anyone working on conversational medical AI, telehealth evaluation, or multi-agent systems for clinical interaction. It deserves a serious referee, not a desk reject. A careful referee should ask the authors to separate live-only scoring from post-questionnaire scoring, and to state plainly how much of the advantage comes from each. If the live-only advantage holds, the claim is real. If not, the paper remains a solid architecture and evaluation contribution, but not an expert-level demonstration.","headline":"Strong, well-designed OSCE study showing a video-based AI beating PCPs on simulated telehealth encounters, but the central 'real-time expert-level' claim is partly confounded by an offline post-questionnaire step where the AI gets extra compute.","tokens_in":47050,"tokens_out":2232,"would_cite":true,"duration_ms":22177,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AMIE (Video), a real-time video consultation AI, was rated by clinicians on par with or better than primary care physicians in simulated telehealth visits.","keywords":["medical AI","video consultations","telehealth","multi-agent system","OSCE","audio-visual perception","clinical diagnosis","patient actors"],"falsifier":"Run the same 100 cases with PCP cameras on and with unscripted real patients or real telehealth recordings, then compare top-1 diagnostic match and overall rubric scores; if PCPs reach AMIE's 91% top-1 or 83% overall scores, the expert-level claim is falsified.","tokens_in":45912,"feed_emoji":"🩺","tokens_out":9831,"duration_ms":77050,"temperature":0.7,"pith_summary":"This paper argues that an AI system, AMIE in a video configuration, can conduct real-time clinical video consultations at a level that independent clinical evaluators rate on par with or better than primary care physicians. The evidence is a randomized OSCE study of 100 scripted telehealth scenarios, enacted by 15 professional patient actors, with 10 board-certified PCPs consulting via video and 20 independent PCPs evaluating the recordings. AMIE (Video) scored 83% overall on case-specific rubrics versus 68% for PCPs ($p=1.3\\times10^{-9}$), matched the reference diagnosis at top-1 in 91% of cases versus 77% ($p=0.039$), and was rated higher on physical observation and guided examination. Patient actors preferred AMIE (Video) over text chat for communication, convenience, and feeling understood, while PCPs retained a non-significant preference for rapport and partnership. If the result transfers beyond actors and scripted cases, it would be the first demonstration of expert-level AI in real-time video consultations, a step toward AI that can see and hear patients rather than only read text.","feed_headline":"AI video doctor rated on par with primary care physicians","feed_subtitle":"Clinical evaluators gave AMIE (Video) 83% versus 68% for physicians across 100 telehealth scenarios.","key_machinery":"The central mechanism is an asynchronous three-agent harness. A Talker agent responds to the patient at low latency; a Planner agent maintains clinical goals, a running differential, and management milestones; and a Perception agent watches and listens over a longer video window while keeping a persistent memory of audio-visual cues. The agents run in parallel, so the Talker can reply within about 2.6 seconds per turn while the Planner and Perception agents continue deeper reasoning and perception in the background. This decoupling is what the paper argues makes expert-level real-time video consultation possible, and automated ablations show the Perception and Planner agents each meaningfully raise rubric scores over the Talker-only baseline.","core_discovery":"The paper's central claim is that AMIE in its video configuration—AMIE (Video)—is the first AI system to demonstrate expert-level performance in real-time clinical video consultations. In a randomized OSCE with 100 telehealth scenarios, 15 professional patient actors, 10 consulting board-certified PCPs, and 20 independent PCP evaluators, clinical evaluators rated AMIE (Video) at 83% overall on case-specific rubrics versus 68% for PCPs ($p=1.3\\times10^{-9}$), and AMIE matched the reference diagnosis at top-1 in 91% of cases versus 77% for PCPs ($p=0.039$). The largest advantage was in perception and examination (74% vs 47%) and guided physical examination (72% vs 39%). Patient actors preferred AMIE (Video) over text chat for communication, convenience, and feeling understood, while PCPs were preferred, though without statistical significance, for rapport and partnership. The authors state the system is not ready for real-world deployment, with limitations in fine anatomical precision, subtle affective cues, high-frequency movements, and the use of actors rather than real patients.","pith_inferences":["Editorial inference: In a real clinic, PCPs would keep cameras on, which may narrow or erase the reported rapport and empathy gaps, making the expert-level framing depend on the cameras-off design.","Editorial inference: Because AMIE (Text) already scores near AMIE (Video) on diagnostic and management rubrics, the paper's own data suggest the video modality mainly buys perception, examination, and patient experience rather than raw diagnostic accuracy.","Editorial inference: A natural next experiment would use full-duplex interaction, letting the AI speak up when a patient performs a maneuver incorrectly; this directly targets the missed-correction failures the paper documents.","Editorial inference: Since dermatologic and other non-actable cases were excluded, testing on real patient video with visible lesions would be a fast way to see whether the perception advantage generalizes beyond actors."],"forward_implications":["If the central claim holds, video-based medical AI can be evaluated against physicians on the modality patients actually use for remote care, not just text chat.","The perception-and-examination gain (74% vs 47%) suggests the video configuration specifically recovers information that text-only AI discards.","The asynchronous architecture cuts mean turn latency from 21.4 seconds to 2.6 seconds, making real-time dialogue practical while retaining clinical reasoning gains.","The comparable top-3 diagnostic accuracy (98% vs 90%) implies the AI's edge is in ranking the right diagnosis first rather than in carrying a broader differential.","Patient actors preferred AMIE (Video) over text chat for communicative effectiveness, convenience, and feeling understood, while PCPs remained preferred for rapport and partnership, so human connection is a separate axis."],"supporting_citations":[{"why":"supplies the virtual OSCE framework, general rubrics, and the text-based conversational diagnostic baseline that AMIE (Video) extends to video.","marker":"[11]"},{"why":"provides the prior audio-visual medical AI system, its 20 proof-of-concept scenarios, and the sequential harness whose latency is improved upon.","marker":"[19]"},{"why":"contributes the management reasoning agent design and multi-draft synthesis used by the Planner and post-questionnaire inference scaling.","marker":"[10]"},{"why":"establishes the multimodal text-plus-image conversational diagnostic baseline that the video configuration aims to surpass.","marker":"[12]"},{"why":"supplies the telehealth-specific OSCE rubrics for audio-visual observation and guided physical examination.","marker":"[26]"},{"why":"defines the differential diagnosis accuracy framing used for the top-k comparison with PCPs.","marker":"[9]"},{"why":"provides the virtual physical examination framework used in the taxonomy of audio-visual cues.","marker":"[24]"}],"fun_headline_variants":["AI video consult matches or beats physicians in 100 scenarios","First AI to hit expert-level in live video medical consults","AMIE video AI rated on par with docs in telehealth trial","Video AI outshines physicians in simulated clinical exams","AI achieves doctor-level skill in real-time video consultations"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on the assumption that board-certified PCPs consulting through video with their cameras off, in scripted encounters with professional actors, are a fair proxy for real clinical performance; if real patients, real environments, or a visible physician change the comparison, the expert-level result may not transfer.","fun_headline_variants_meta":{"raw":{"variants":["AI video consult matches or beats physicians in 100 scenarios","First AI to hit expert-level in live video medical consults","AMIE video AI rated on par with docs in telehealth trial","Video AI outshines physicians in simulated clinical exams","AI achieves doctor-level skill in real-time video consultations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000284,"raw_usage":{"total_tokens":1752,"prompt_tokens":1098,"completion_tokens":654,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":714,"completion_tokens_details":{"reasoning_tokens":573}},"tokens_in":714,"tokens_out":654,"duration_ms":6449,"temperature":1.0,"reasoning_tokens":573,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:21:33.703552+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 100 cases with PCP cameras on and with unscripted real patients or real telehealth recordings, then compare top-1 diagnostic match and overall rubric scores; if PCPs reach AMIE's 91% top-1 or 83% overall scores, the expert-level claim is falsified.","supporting_citations":[{"cited_title":"Journal of telemedicine and telecare , volume=","cited_arxiv_id":null,"evidence_quote":"supplies the telehealth-specific OSCE rubrics for audio-visual observation and guided physical examination."},{"cited_title":"Fundamental Clinical Skills , publisher =","cited_arxiv_id":null,"evidence_quote":"provides the virtual physical examination framework used in the taxonomy of audio-visual cues."}],"review_version":2}