{"id":"b18616c6-1b8c-4a19-94b4-33ceb167d809","arxiv_id":"2608.03190","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A typed multi-agent protocol with a claim ledger, adversarial critic, and safety gate improves action F1 and safety over budget-matched baselines on a 360-case neuro-oncology benchmark.","lead":"TumorBoard is a multi-agent system for neuro-oncology that coordinates specialists through a shared patient timeline and an auditable claim-evidence ledger, with a safety governor that can defer advice to humans. On a 360-case benchmark it beat a strong baseline by 3.1 points in action F1 and reduced harmful recommendations, but the benchmark is self-built and the promised data and code are not yet public.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ablation token counts differ by 20% (Table 5), confounding the critic's measured contribution to the coordination advantage.","rationale":"The reader's primary concern is benchmark label bias, which threatens external validity. I agree that is important. However, the paper's most specific internal claim—that ablations isolate the contribution of the ledger, critic, and governor—is directly weakened by the token counts in Table 5. The 'No critic' variant uses 11,380 generated tokens versus 14,220 for full TumorBoard, a 20% difference, while the paper explicitly says gains are credited only after controlling for generated tokens (§1, §5.5). This means the measured effect of removing the critic is confounded with a large reduction in inference compute. If the critic improves results only by adding a reasoning loop (i.e., 'more thinking'), the mechanism attribution is wrong. The same concern applies to the 'No governor' ablation, though the F1 difference is small. A token-matched ablation would settle this. The reader's rationale did mention that the matched-budget control is not cleanly supported by reported token counts, so we partially agree; but the load-bearing assumption I focus on is the internal confound in the mechanism ablation, not the label-bias concern. This does not change the overall CONDITIONAL verdict: the paper needs both external validation and a token-matched ablation before the central claim is fully supported.","tokens_in":10829,"tokens_out":8842,"duration_ms":91807,"concrete_test":"Re-run the 'No critic' ablation with a token-matched filler: after removing the critic, add a non-adversarial 'summarization pass' that consumes the same number of tokens (e.g., re-formatting the ledger) and compare. If the F1 and contradiction-resolution drops persist, the critic mechanism is real; if they shrink, the effect is a compute artifact. Also report typed-council token counts for the primary comparison to confirm actual matching.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that ablations establish structured coordination as the source of the multi-agent gain (Abstract, §7) rests on comparing variants with unequal actual token use. §1 states that a gain is credited only after controlling for 'generated tokens,' and §5.5 says budget matching is implemented 'at the generated-token level,' but Table 5 reports 14,220 tokens for Full TumorBoard versus 11,380 for No critic—a 20% reduction—and 13,310 for No governor. The 'No critic' ablation therefore removes the critic and simultaneously reduces inference compute. The observed 0.039 F1 drop and 0.112 contradiction-resolution drop could reflect the extra 'thinking time' the critic provides rather than the challenge mechanism per se. No token-padding or compute-equalization procedure is described for ablations. Since the paper's own acceptance criterion is control for generated tokens, the mechanism attribution is not yet established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TumorBoard, a multi-agent decision-support system for longitudinal neuro-oncology. The architecture combines a timeline curator and shared case state, role-limited specialist agents (radiology, neuropathology, molecular, guidelines, therapy), a typed claim-evidence ledger, an adversarial critic, and a safety governor. The authors report a 360-case hidden benchmark with budget-matched comparisons against single-agent and multi-agent baselines. They claim action F1 of 0.772, evidence entailment of 0.914, a 3.1-point gain over the strongest baseline (adjusted p=0.0012), and a 7.8-point reduction in harmful release when the safety governor is enabled. They also present perturbation tests (evidence deletion, conflict, guideline shift, role corruption, prompt injection) and ablations of the ledger, critic, and governor.","tokens_in":11054,"tokens_out":4424,"duration_ms":50417,"significance":"If the empirical claims are sound, TumorBoard is a meaningful step toward auditable, evidence-grounded clinical decision support. The paper has real strengths: paired case-level statistics with bootstrap CIs and Holm correction, prespecified primary and safety endpoints, perturbation tests with parameterized severity, a hidden test set, and a commitment to retaining prompts, messages, ledger states, and outputs for audit. The typed claim-evidence ledger and explicit safety governor are sensible mechanisms for a high-stakes domain. The measured gains are modest but plausible. However, the benchmark is fully self-constructed with no external validation or public release, and the ablation token-budget mismatch documented below weakens the paper's central mechanistic attribution. The result is still worth serious consideration, but several load-bearing points need to be addressed before the claims can be accepted as stated.","major_comments":[{"comment":"The paper's own acceptance criterion for crediting coordination gains is control for generated tokens (§1, §5.5). Table 5 violates this criterion for the mechanism ablations: Full TumorBoard uses 14,220 tokens/case, No critic uses 11,380 (a 20% reduction), No governor uses 13,310, and No ledger uses 14,050. The reported No-critic degradation in Action F1 (0.749 vs 0.772) and contradiction resolution (-0.112) is therefore confounded by a 20% inference-compute reduction. The 'No critic' ablation cannot isolate the challenge mechanism from additional thinking time. The central attribution 'structured coordination as the source of the measured multi-agent advantage' (§7, Abstract) is not established by the reported comparisons. I ask for either token-equalized ablation runs (e.g., proportional critic iterations or filler reasoning) or an explicit covariate adjustment for generated tokens.","section":"§5.5 / Table 5"},{"comment":"The benchmark is self-constructed: gold action graphs and harm labels are authored and adjudicated by the research group, with no external validation or public release. The evaluation ontology (action graph nodes/edges, prerequisite and safety classes) closely mirrors the system's typed ledger schema and safety governor rules. This raises a circularity risk: the measured F1 and harmful-release reductions may partly reflect the benchmark encoding the same expectations the system is prompted to produce. Since the primary comparative claim rests on this benchmark, independent external adjudication (or at minimum a second institution's expert panel) and public release of case and label templates is needed. The paper should explicitly flag this limitation; it currently presents the benchmark as unproblematic.","section":"§5.1 / §7"},{"comment":"The No-governor row in Table 5 reports Action F1 0.776, numerically higher than Full TumorBoard's 0.772, despite using fewer tokens. The text in §6.3 says removing the governor 'increased harmful release by 0.078' but does not report that the decision-quality effect is negligible or slightly negative. The conclusion 'the ledger, critic, and safety governor jointly improve decision quality' (Conclusion) therefore overstates the evidence for the governor's effect on decision quality; at most the governor contributes to safety, not to F1. This should be stated precisely.","section":"Table 5 / §6.3"},{"comment":"The perturbation tests are a clear strength, but the safety outcomes are measured against the researchers' own harm labels, and the perturbation generator and the governor share the same prerequisite and risk rules. The 84.2% deferral under evidence deletion and the 7.8pp reduction in harmful release could thus be partly by construction: when the decisive fact is removed, a prerequisite check is expected to fail. This is not a trivial result, but it is less informative without an independent enumeration of harm types or a held-out clinician panel scoring potential harm with the governor's decision masked. Reporting such an external human harm review on a random subset would strengthen the safety claim considerably.","section":"§5.4 / §6.2"}],"minor_comments":[{"comment":"The 'Tokens' column reports values such as '14220.000' and '11380.000'; these should be integers. Also, the column header in Table 3, 'Deferral F1', appears to denote a deferral-related rate rather than an F1 score; please clarify the metric name.","section":"Table 5"},{"comment":"The appendix says every numeric statement maps to a locked CSV export, but the release 'will contain' prompts, retrieval snapshots, and logs rather than providing them now. For reproducibility, at least the prompts and retrieval collection procedure should be included as supplementary material.","section":"§5.5 / Appendix D"},{"comment":"Table 3 reports many point estimates without confidence intervals for baseline methods. The paired CI is given for the primary TumorBoard-vs-typed-council comparison, but CIs for other baselines would help assess the significance of the ordering in Table 3.","section":"§6.1 / Table 3"}],"recommendation":"major_revision","confidential_remarks":"I do not see evidence of deliberate distortion; the main concern is that the authors' own budget-matching criterion is not met in the ablation table, which is central to the mechanistic claim. The self-constructed benchmark is a substantial limitation for a clinical decision-support paper, and I would urge the editor to require external validation before publication. The paper's scope fits an AI conference/journal audience if the empirical gaps are addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The main head-to-head result is believable: TumorBoard beats the strongest budget-matched typed council by 3.1 F1 points (0.772 vs 0.741) with a tight confidence interval, and the safety governor's 7.8-point reduction in harmful release is a concrete, clinically meaningful outcome. The paper is careful about token control in the primary comparison, and the perturbation tests (evidence deletion, role corruption, guideline shift) are a genuine plus. The architecture—timeline curator, typed claim-evidence ledger, adversarial critic, safety governor—is a new integration that should be useful as a template for accountable multi-agent medical systems.\n\nThe internal consistency is actually quite good: paired statistics, confidence intervals, and appendix reconciliation all check out. And the ablation pattern is directionally right: removing the ledger raises unsupported consensus, removing the critic drops contradiction resolution, removing the governor raises harmful release.\n\nBut there are two real soft spots. First, the stress-test concern about the ablations is valid. Table 5 shows token counts varying across ablations: No critic uses 11,380 tokens versus 14,220 for the full system, a 20% drop. The paper says all variants receive the same total token budget, and that budget matching is implemented at the generated-token level, but no token-padding or compute-equalization is described for the ablations. So the critic's measured contribution is confounded with inference compute. The same issue affects No governor (13,310 tokens). This weakens the 'ordered failure pattern' claim in Section 7. The main result stands, but the mechanism attribution is not yet established.\n\nSecond, the benchmark is self-constructed. The gold action graphs and harm labels are authored and adjudicated within the research group, with no external validation and no released artifacts. The paper promises release but nothing is public yet. That creates a real circularity risk: if the labels encode the designers' expectations, the gains could partly reflect benchmark design. This is the main reason I would not yet treat the measured advantage as generalizable.\n\nThe authors are clearly thinking carefully, and the paper is transparent about its protocol, costs, and limitations. It deserves peer review, but the reviewers should push for (a) budget-matched ablations or an explicit acknowledgement of the token imbalance, and (b) release of the benchmark and scoring code. If those are addressed, this could be a solid contribution. My recommendation: send it out, with a request for revision.","headline":"Solid architecture paper with a credible matched-budget main result, but the ablation token imbalance and the self-built benchmark keep the mechanism claims provisional.","tokens_in":11508,"tokens_out":2804,"would_cite":false,"duration_ms":32562,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"TumorBoard shows that structured coordination, not extra tokens, drives multi-agent decision quality and safety in neuro-oncology.","keywords":["multi-agent systems","clinical decision support","neuro-oncology","claim-evidence ledger","safety governor","evidence grounding","longitudinal reasoning","action graphs"],"falsifier":"Have an independent neuro-oncology team re-annotate the 360 hidden cases (action graphs and harm labels) without seeing the system outputs, then rerun the budget-matched comparison on the same frozen prompts. If TumorBoard's action-F1 advantage over the typed council shrinks or reverses under those external labels, the measured coordination gain is an artifact of benchmark self-confirmation rather than a property of the architecture.","tokens_in":10774,"feed_emoji":"🧠","tokens_out":4747,"duration_ms":49363,"temperature":0.7,"pith_summary":"This paper tries to establish that multi-agent medical decision support improves decisions only when coordination is structured, not free-form. TumorBoard builds a typed claim-evidence ledger, an adversarial critic, and a safety governor so that specialist agents exchange auditable claims instead of prose. On a 360-case hidden neuro-oncology benchmark with matched token budgets, it reaches an action F1 of 0.772, beating the strongest budget-matched council by 3.1 points, and cutting harmful recommendation release from 0.117 to 0.039 when the governor is active. Removing any one mechanism produces the predicted failure pattern, which the authors take as evidence that the protocol itself, not extra compute, accounts for the gain. A sympathetic reader would care because this offers a concrete recipe for making clinical multi-agent systems safer and more auditable.","feed_headline":"Structured coordination lifts AI tumor-board decisions by 3.1 F1 points","feed_subtitle":"On a 360-case hidden benchmark, TumorBoard beats the strongest budget-matched council and cuts harmful releases.","key_machinery":"The load-bearing object is a typed claim-evidence ledger: a directed graph whose nodes are atomic claims carrying claim text, type, temporal scope, confidence, prerequisites, evidence pointers, and conflicts, and whose edges encode support, contradiction, supersession, and dependency. Evidence nodes are immutable, a claim cannot support its own ancestors, and a recommendation inherits the weakest confidence among its required upstream claims. The ledger makes coordination failures visible as graph defects; an adversarial critic then targets missing prerequisites, stale sources, temporal inconsistencies, and false consensus, while a safety governor releases, qualifies, or defers each recommen","core_discovery":"TumorBoard's central claim is that structured coordination—a typed claim-evidence ledger, adversarial critique, and a gated safety release—is the source of multi-agent decision quality in longitudinal neuro-oncology, and that this can be demonstrated by holding model, retrieval corpus, and total inference budget fixed while varying only the communication protocol. Under that constraint, full TumorBoard reaches an action F1 of 0.772 versus 0.741 for the strongest typed council, with evidence entailment of 0.914 and recommendation-to-evidence coverage of 0.927. Perturbation tests show the system defers 84.2% of cases when decisive evidence is deleted and caps harmful recommendations at 5.8%; e","pith_inferences":["Editorial inference: Because the gold action graphs and harm labels were authored and adjudicated by the research group with no external validation, the measured F1 gain over the typed council could partly reflect the designers' expectations rather than a generalizable property of the architecture.","Editorial inference: The 84.2% deferral under evidence deletion, while promising, may indicate an over-cautious bias; a testable extension would compare deferral decisions against human expert consensus on the same incomplete cases.","Editorial inference: The ledger records per-agent message cost, so the same data could be mined to identify which specialist role contributes unique evidence versus redundant prose, enabling a smaller council without loss of accuracy.","Editorial inference: The 4.3-point false-deferral cost is presented as an operational trade-off, but it lacks a clinical utility weighting; future work could link deferral decisions to downstream patient outcomes to set the threshold empirically."],"forward_implications":["Multi-agent medical systems should be evaluated with matched token budgets, because protocol structure, not compute, may explain reported gains.","The safety governor turns risk tolerance into an explicit release threshold, allowing deployments to choose how to trade harmful releases against false deferrals.","The claim-evidence ledger preserves the full reasoning dependency chain, making each recommendation traceable to specific evidence for human review and regulatory audit.","The same coordination protocol could transfer to other longitudinal, high-stakes decision domains where prerequisites and temporal validity matter, such as oncology treatment sequencing or chronic disease management."],"supporting_citations":[{"why":"Supplies the reasoning-and-acting primitive whose simple prompt baseline TumorBoard must beat.","marker":"[1]"},{"why":"Extends tool-use capabilities and informs the retrieval-capable RAG-agent baseline.","marker":"[2]"},{"why":"Provides the self-critique primitive behind the plan-and-critique baseline.","marker":"[3]"},{"why":"Supplies the conversational multi-agent orchestration pattern on which the free-chat council baseline is built.","marker":"[4]"},{"why":"Motivates clinical agent requirements for evidence acquisition, temporal synthesis, and accountable action boundaries.","marker":"[8]"},{"why":"Provides versioned neuro-oncology decision constraints that the guideline agent retrieves as evidence units.","marker":"[12]"},{"why":"Supplies current management and future-direction consensus clauses used in guideline-grounded decisions.","marker":"[15]"},{"why":"Provides clinical practice guideline clauses that the safety governor checks against prerequisites.","marker":"[24]"}],"fun_headline_variants":["Structured coordination beats solo AI in neuro-oncology decisions","AI tumor board gains 3.1 F1 via evidence ledger and critic","Safety governor cuts harmful AI tumor advice by 7.8 points","AI defers 84% of risky cases with evidence-grounding and critique","Multi-agent teamwork lifts tumor decisions by 3.1 F1"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the benchmark's gold action graphs and harm labels, authored and adjudicated by the research group with no external validation, do not encode the same expectations the system is prompted to produce.","fun_headline_variants_meta":{"raw":{"variants":["Structured coordination beats solo AI in neuro-oncology decisions","AI tumor board gains 3.1 F1 via evidence ledger and critic","Safety governor cuts harmful AI tumor advice by 7.8 points","AI defers 84% of risky cases with evidence-grounding and critique","Multi-agent teamwork lifts tumor decisions by 3.1 F1"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1382,"prompt_tokens":786,"completion_tokens":596,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":530,"completion_tokens_details":{"reasoning_tokens":503}},"tokens_in":530,"tokens_out":596,"duration_ms":7321,"temperature":1.0,"reasoning_tokens":503,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:55:57.623545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have an independent neuro-oncology team re-annotate the 360 hidden cases (action graphs and harm labels) without seeing the system outputs, then rerun the budget-matched comparison on the same frozen prompts. If TumorBoard's action-F1 advantage over the typed council shrinks or reverses under those external labels, the measured coordination gain is an artifact of benchmark self-confirmation rather than a property of the architecture.","supporting_citations":[{"cited_title":"Large Language Models as Agents in the Clinic","cited_arxiv_id":"2309.10895","evidence_quote":"Motivates clinical agent requirements for evidence acquisition, temporal synthesis, and accountable action boundaries."},{"cited_title":"van den Bent, Matthias Preusser, Émilie Le Rhun, Jörg C","cited_arxiv_id":null,"evidence_quote":"Provides versioned neuro-oncology decision constraints that the guideline agent retrieves as evidence units."},{"cited_title":"Wen, Michael Weller, Eudocia Q","cited_arxiv_id":null,"evidence_quote":"Supplies current management and future-direction consensus clauses used in guideline-grounded decisions."},{"cited_title":"Ahluwalia, Joachim M","cited_arxiv_id":null,"evidence_quote":"Provides clinical practice guideline clauses that the safety governor checks against prerequisites."}],"review_version":1}