{"id":"5ef2c3b6-26fd-4277-9bbb-3f76489b00a4","arxiv_id":"2608.09024","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An EHR-grounded benchmark generates traceable triage dialogues and evaluates AI models on information elicitation, ESI prediction, and patient-facing communication.","lead":"This paper presents EHR2Dial-Triage, a benchmark that simulates emergency department triage conversations from electronic health records and traces every patient statement to its source record and the turn when it appears. It provides a controlled setting for testing whether AI clinicians ask the right questions, use the answers, and communicate triage decisions clearly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The event–turn lineage is only as strong as the unvalidated LLM verifier; no accuracy or inter-annotator agreement is reported, so the central tracing claim and all acquisition metrics rest on an untested component.","rationale":"The reader's weakest assumption correctly identifies the unvalidated patient simulator and verifier. I agree that external validation is missing, but I would sharpen the concern: the verifier is the load-bearing component for the paper's most distinctive claim. The patient simulator determines whether the dialogues are clinically realistic, but even perfect realism would not make the benchmark's lineage trustworthy if the verifier cannot reliably map open-ended utterances to normalized facts and source events. Conversely, a validated verifier would preserve the structural contribution even if patient realism is imperfect, because the benchmark's stated purpose is source-grounded tracing and controlled model comparison rather than deployment-ready simulation. The paper explicitly says its Section 4 analyses 'do not establish clinical validity,' but it does not flag the absence of any verifier accuracy measurement, which is a different and more specific gap. The internal aggregate checks—role-consistent realization, progressive disclosure, and predictive relevance—are useful, but they are indirect: a biased verifier could inflate coverage, distort entropy trajectories, and even affect the under-triage confidence finding if it systematically misattributes disclosures. The concrete test I propose would settle this directly by measuring verifier-human agreement on fact extraction, source linking, and first-turn assignment. If agreement is high, the central claim is supported; if it is low, the benchmark's headline metrics need recomputation and the paper should report a corrected lineage before the benchmark is relied upon. The reader's CONDITIONAL verdict already reflects the need for such validation, so my stress-test does not move the verdict; it specifies the condition that must be met.","tokens_in":31342,"tokens_out":4606,"duration_ms":47529,"concrete_test":"Sample 100 completed dialogues stratified by encounter and clinician model, and have two emergency nurses independently annotate each patient utterance for (i) the atomic facts expressed, (ii) the supporting EHR source event for each fact, and (iii) the first turn at which each fact becomes explicit. Compute verifier agreement with the consensus clinician annotation using fact-level F1, source-link accuracy, and first-turn agreement. If fact-level F1 is below 0.9 or source-link accuracy is below 0.95, re-run the headline acquisition metrics with the clinician-corrected lineage; the magnitude of any shift determines whether the event–turn tracing claim survives. Include 20 adversarial utterances containing one supported and one unsupported detail to test whether the verifier rejects or partitions partial support correctly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's core contribution is exact, machine-readable event–turn lineage, and that lineage is produced by the verifier described in Section 3.3. Every accepted disclosure must be mapped to a normalized atomic fact, linked to a supporting EHR source event, and assigned a first accepted turn. The paper reports no evaluation of this verifier: there is no precision/recall for fact extraction, no accuracy for source-event linking, no inter-annotator agreement against human annotation, and no audit of partial-support cases where an utterance contains both supported and unsupported details. Section 4's internal validation shows aggregate patterns consistent with the role-based design, but those patterns would also be observed if the verifier systematically over-accepted near-matches, merged distinct facts, or attributed a fact to the wrong source event. Because FACTCOV, HistoryMedCov, reader-entropy trajectories, question-yield, and the safety analysis in Section 5.4 are all computed from the verifier's ledger, an unmeasured verifier error rate directly threatens the central claim. The patient-simulator realism concern is secondary: even an idealized simulator would not rescue the lineage if the verifier mislabels disclosures.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EHR2Dial-Triage, a framework and benchmark for emergency-department triage conversations grounded in MIMIC-IV-ED. Each encounter is partitioned into clinician-visible information C_i, patient-reportable information P_i, and future/evaluation-only information; two role-specific LLMs generate alternating dialogues under these asymmetric information constraints. A verifier maps accepted disclosures to normalized atomic facts, links each fact to its supporting EHR event, and records the first accepted turn, producing what the paper calls event–turn lineage. The paper validates internal properties (role-consistent realization, bounded interaction scale, progressive disclosure, predictive relevance of disclosed facts) and then evaluates LLMs on two tasks: information elicitation as triage clinician, and dialogue-based ESI prediction with patient-facing closing generation, including an under-triage and confidence safety analysis.","tokens_in":1525,"tokens_out":1425,"duration_ms":68100,"significance":"If the event–turn lineage is reliable, EHR2Dial-Triage fills a genuine gap: it is a reusable, temporally partitioned conversational triage benchmark that allows exact, machine-readable tracing of when triage-relevant evidence enters a dialogue, and it separates information acquisition from downstream prediction and communication. Strengths include the released generator and frozen corpus, the explicit role/time information boundaries, the careful acknowledgment in Section 4 that the internal analyses are not clinical validation, and the clinically meaningful safety analysis of confidence under severe under-triage. The main significance is conditional: the benchmark's signature claim of auditable lineage depends on a verifier whose accuracy is not reported, so the contribution cannot yet be fully assessed.","major_comments":[{"comment":"The central claim of exact, machine-readable event–turn lineage depends entirely on the verifier described in Section 3.3, but the paper reports no precision/recall for fact extraction, no accuracy for source-event linking, no inter-annotator agreement against human annotation, and no audit of partial-support cases where an utterance contains both supported and unsupported details. All downstream quantities—FACTCOV, HistoryMedCov, reader-entropy trajectories, QYield, and the grounding metrics in Appendix C.3—are computed from the verifier's ledger. If the verifier systematically over-accepts near-matches or attributes a disclosure to the wrong source event, the benchmark's central contribution would not measure what it claims. Please add a verifier evaluation on a human-annotated sample covering fact extraction, event linking, first-accepted-turn assignment, and partial support, and report error rates; a sensitivity analysis showing how coverage and entropy change under simulated verifier errors would further establish robustness.","section":"§3.3, Eq. (6)"},{"comment":"The paper explicitly states in Section 4 that the internal analyses 'do not establish clinical validity or readiness for deployment.' That scoping is appropriate, but patient-simulator realism is load-bearing for the information-acquisition metrics in Experiment I: if the patient model does not approximate how real patients disclose triage-relevant information, then question-yield and coverage comparisons are difficult to interpret as measures of clinical information seeking. The manuscript currently offers no comparison against real triage conversations, recorded triage notes, or clinician judgment. I am not asking for deployment-level validation, but a small clinician or human-annotation study, or an explicit quantitative comparison to real triage documentation, would substantially strengthen the claim that the benchmark studies triage information acquisition rather than only simulated-patient behavior.","section":"§4 opening; §5.2"},{"comment":"The reader-entropy numbers in Section 4.3 and Table 3 appear inconsistent. Section 4.3 reports mean predictive entropy decreasing from 1.61 nats initially to approximately 1.04 nats at Round 10, while Table 3 reports Entropy@10 between 0.69 and 0.76 nats for what is described as the same fixed, patient-disjoint multinomial logistic-regression reader. If the readers, cohorts, or feature representations differ between the two analyses, the text must state this; if they are the same, there is a numerical error in one of the computations. This discrepancy matters because Figure 4 is the primary evidence that progressively disclosed information is predictive of the recorded ESI label.","section":"§4.3 vs Table 3 (§5.2)"}],"minor_comments":[{"comment":"Several headings in the case study are preceded by the literal artifact '/clipboard-lis' (e.g., before 'Information revealed in the conversation' and 'Turn-wise accumulated grounded state'); this appears to be a copy-paste artifact and should be removed.","section":"Appendix A"},{"comment":"Question yield (QYield) is defined in Equation (9) but never reported in the results section; either report it in Experiment I or remove the definition to avoid an unused metric.","section":"Appendix C.1"},{"comment":"The question-target assignment procedure is described as a 'fixed target annotation procedure' but no reliability measure (e.g., annotator agreement on a sample) is reported; since question-target distributions are a main descriptive result, a small agreement check would be useful.","section":"§5.2, Figure 5"},{"comment":"The caption states 'Best open-weight results are bold,' but the table also contains proprietary models; clarify whether bolding is restricted to the open-weight block or applied across all models.","section":"Table 3"},{"comment":"The phrase 'across models, and patient personas' contains an unnecessary comma; the writing is generally clear, but the paper would benefit from a final pass for such mechanical errors.","section":"Abstract and Section 1"}],"recommendation":"major_revision","confidential_remarks":"The verifier-validation issue is the make-or-break point: it is fixable within the manuscript's scope, but without it the paper's signature contribution cannot be assessed. I would ask the authors to provide verifier accuracy and human agreement, and to reconcile the Section 4.3 / Table 3 entropy discrepancy. If the verifier cannot be validated on a sample, the benchmark should be described as providing approximate rather than exact lineage. The patient-simulator realism concern is secondary but should also be addressed at least minimally."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: EHR2Dial-Triage is a genuinely useful benchmark artifact, and the event–turn lineage plus the three-way temporal partition are new relative to PatientSim, Note2Chat, and TriageSim. But the central claim of exact, machine-readable tracing rests on a verifier the paper never validates, so I would treat the acquisition metrics as provisional until that is addressed.\n\nWhat the paper does well: it separates information acquisition from downstream ESI use, creates a clear role-based realization of who introduces what, and the internal validation in Section 4 is honest. The authors explicitly say the internal checks do not establish clinical validity or deployment readiness. The role-consistency numbers (95.1% of vitals introduced by clinician, 98.8% of chief complaints by patient) and the progressive-disclosure analysis are consistent with the design. The under-triage confidence finding in Figure 6 is a genuinely useful safety observation, and the code is released, which makes the resource testable by others.\n\nThe soft spots are real but fixable. The stress-test note is right: the verifier is load-bearing, and there is no reported precision/recall, no inter-annotator agreement, no audit of partial-support cases, and no source-linking accuracy. The role-consistency patterns could in principle be produced by a systematically over-accepting verifier, so they do not independently validate the lineage. Every acquisition metric, reader-entropy trajectory, and the safety analysis inherits the verifier's error rate. That is a missing evaluation, not a design flaw, but the phrase \"exact tracing\" in the abstract and Section 3.3 is stronger than what is demonstrated.\n\nSecondary issues: the patient simulator is not externally validated against real triage conversations or clinician judgment, which the authors acknowledge. The model comparisons in Table 3 have no error bars, and the cohort is deliberately acuity-enriched, so absolute performance numbers should not be read as population estimates. The grounding evaluator for closing generation is also a fixed LLM verifier without reported accuracy.\n\nBottom line: this is a solid benchmark paper for researchers who build or evaluate clinical dialogue agents. It deserves a serious referee. The referee should require verifier evaluation (precision/recall and inter-annotator agreement), error bars or variance estimates, and a clear statement of what the acquisition metrics mean if the verifier is imperfect.","headline":"A useful new benchmark for conversational triage with genuinely new event–turn lineage, but the exact-tracing claim rests on an unvalidated verifier; worth reviewing and fixing.","tokens_in":32063,"tokens_out":2146,"would_cite":true,"duration_ms":20828,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EHR2Dial-Triage is a benchmark that makes the evidence-gathering process of emergency triage auditable by linking every accepted patient disclosure to the EHR event that supports it and the first turn at which it appears.","keywords":["emergency department triage","conversational agents","EHR grounding","event-turn lineage","Emergency Severity Index","information elicitation","patient simulation","MIMIC-IV-ED"],"falsifier":"Run the verifier on a set of real triage conversations recorded in an emergency department (with patient consent) and have clinicians annotate each disclosure's supporting EHR event and first relevant turn; if agreement between the verifier's lineage and clinicians' annotations is low, or if substituting real transcripts for the simulated patient responses changes fact-coverage and entropy trajectories materially, the acquisition metrics would not measure what they claim.","tokens_in":31140,"feed_emoji":"🏥","tokens_out":7629,"duration_ms":66420,"temperature":0.7,"pith_summary":"Emergency-department triage is usually benchmarked as acuity prediction from a fixed snapshot, even though real triage is a conversation in which clinicians gather missing details. This paper introduces EHR2Dial-Triage, an agentic conversation-generation framework and benchmark built from a public emergency-department EHR database. Its central proposal is to run triage dialogues under explicit role and time boundaries: the clinician sees only arrival information, the patient can disclose only record-supported facts, and events that happen after triage stay hidden. Every accepted disclosure is logged with the EHR event that supports it and the first dialogue turn at which it becomes available, producing a machine-readable ledger of when and how evidence entered the conversation. If the simulation is faithful, this lets researchers evaluate information elicitation, evidence use, five-level acuity prediction, and patient communication as separate capabilities rather than as one final prediction.","feed_headline":"Benchmark links every triage disclosure to its EHR source and turn","feed_subtitle":"Lets researchers measure how and when conversational agents gather the evidence behind emergency acuity decisions.","key_machinery":"The load-bearing object is the event-turn lineage ledger. Each accepted disclosure is reduced to a normalized atomic fact; the ledger records its fact identifier, supporting EHR source, round index, and speaker role, and the dialogue-visible state accumulates as $Z_{i,t}=Z_{i,t-1}\\cup\\Delta Z_{i,t}$. The information partition $C_i/P_i$ and the ten-round budget then turn this ledger into metrics: fact coverage, history and medication coverage, fixed-reader entropy, question yield, and grounding precision. This is what allows the benchmark to measure not only whether a model asked useful questions, but exactly which record-supported facts entered the conversation, when, and from which role.","core_discovery":"The paper claims that EHR2Dial-Triage turns conversational triage into an auditable evidence-acquisition process. Each encounter is partitioned into clinician-visible information $C_i$, patient-reportable information $P_i$, and future-hidden evaluation-only information. Two role-specific language models alternate under asymmetric access: the clinician conditions on $C_i$ and the dialogue history, the patient on $P_i$ and a non-clinical persona, and a verifier maps each accepted disclosure to a normalized atomic fact, its supporting EHR source, and the first accepted speaker turn. Over a frozen corpus of 4,041 conversations, the paper reports that patient facts are disclosed progressively (reaching 86% of triage-important facts by round ten), that cumulative disclosures reduce an ESI reader's entropy and raise recorded-ESI accuracy from a 42% majority baseline to 59% top-1, and that different models lead on information acquisition versus final ESI prediction versus grounded communication. The authors' central conclusion is that conversational triage should be evaluated as a multi-stage process of acquisition, reasoning, and communication, not by final acuity alone.","pith_inferences":["The event-turn ledger could be reused as training signal for question-asking policies, since it attributes each new fact to the clinician turn that preceded it, making credit assignment for information gain explicit.","With the same ledger, one could run counterfactual analyses the paper does not report, such as holding the dialogue fixed but delaying a fact's first turn to ask whether the ESI reader's distribution or confidence would change materially.","The temporal-boundary design could transfer to other time-critical clinical decisions, such as admission or sepsis screening, where the core question is how quickly and from which sources evidence must be acquired.","If the framework were run on real transcribed triage conversations with expert annotation of source events, the acquisition metrics could be calibrated against human behavior; the paper does not provide such validation."],"forward_implications":["Clinician models can be compared on what they ask and elicit, not only on the acuity label they output, because every acquired fact carries a source and a turn.","The same completed dialogue can be used to separate the ability to gather evidence from the ability to use it: a fixed ESI reader converts acquired facts into an acuity distribution, while a separate evaluation tests prediction and closing generation on identical transcripts.","Temporal partitioning lets researchers check whether a system respects what could have been known at triage time, using future-hidden events to detect post-triage leakage.","Safety analyses can quantify under-triage severity and confidence calibration, such as how often a model assigns ESI 4 or 5 to an encounter recorded as ESI 1 or 2 and whether reported confidence drops when it does.","Because only facts originating in $P_i$ get acquisition credit, the metrics reward eliciting patient-reportable evidence rather than restating clinician-visible arrival data."],"supporting_citations":[{"why":"Supplies the source EHR encounters, triage records, prior diagnoses, and medication reconciliation used to construct $C_i$, $P_i$, and the recorded ESI labels.","marker":"Johnson et al., 2023"},{"why":"Demonstrates persona-conditioned patient simulation from EHR profiles, the method the paper extends with event-level provenance.","marker":"Kyung et al., 2025"},{"why":"Provides the note-to-dialogue EHR grounding paradigm that EHR2Dial-Triage extends with temporal and event-turn linkage.","marker":"Zhou et al., 2026"},{"why":"Prior conversational-triage work using an EHR-derived patient simulator, the setting this benchmark differentiates from through explicit lineage.","marker":"Rashidian et al., 2025"},{"why":"TriageSim, a conversational triage framework from structured ED records, serves as the comparison point for temporal tracing and role boundaries.","marker":"Srirag et al., 2026"},{"why":"Standardized ED prediction tasks on the source EHR database and provides the acuity-prediction framing that the benchmark extends to interactive acquisition.","marker":"Xie et al., 2022"},{"why":"Defines and validates the five-level Emergency Severity Index used as the downstream prediction target.","marker":"Wuerz et al., 2000"}],"fun_headline_variants":["Triage conversations get EHR-grounded evidence audit","Benchmark tracks every triage disclosure to its EHR source","Conversational triage benchmark: measure evidence acquisition","EHR2Dial-Triage: linking dialogue to EHR for ESI prediction","New framework turns triage dialogue into auditable evidence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's core metrics rest on the assumption that its simulated patients answer like real patients and that its verifier maps utterances to EHR facts correctly, and the paper states that its validation is internal and does not establish clinical validity or deployment readiness.","fun_headline_variants_meta":{"raw":{"variants":["Triage conversations get EHR-grounded evidence audit","Benchmark tracks every triage disclosure to its EHR source","Conversational triage benchmark: measure evidence acquisition","EHR2Dial-Triage: linking dialogue to EHR for ESI prediction","New framework turns triage dialogue into auditable evidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000201,"raw_usage":{"total_tokens":1429,"prompt_tokens":1043,"completion_tokens":386,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":659,"completion_tokens_details":{"reasoning_tokens":306}},"tokens_in":659,"tokens_out":386,"duration_ms":3976,"temperature":1.0,"reasoning_tokens":306,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:17:05.089122+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the verifier on a set of real triage conversations recorded in an emergency department (with patient consent) and have clinicians annotate each disclosure's supporting EHR event and first relevant turn; if agreement between the verifier's lineage and clinicians' annotations is low, or if substituting real transcripts for the simulated patient responses changes fact-coverage and entropy trajectories materially, the acquisition metrics would not measure what they claim.","supporting_citations":[{"cited_title":"2025 , note =","cited_arxiv_id":null,"evidence_quote":"Demonstrates persona-conditioned patient simulation from EHR profiles, the method the paper extends with event-level provenance."},{"cited_title":"2026 , doi =","cited_arxiv_id":null,"evidence_quote":"Provides the note-to-dialogue EHR grounding paradigm that EHR2Dial-Triage extends with temporal and event-turn linkage."}],"review_version":1}