{"id":"23d48291-fa44-4e76-ad9d-9fb468cc8b53","arxiv_id":"2607.28374","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A provenance-constrained ledger runtime improves multimodal agent accuracy and trajectory faithfulness by binding claims to tool evidence and restricting repair to typed, non-amplifying operators.","lead":"LedgerMind forces multimodal AI agents to store tool outputs in a structured evidence ledger and only reason from what that ledger actually contains. It raises both answer accuracy and intermediate-step faithfulness on visual QA benchmarks by blocking citation-backed hallucinations and unconstrained self-repair.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Faithfulness gains rest on a same-family LLM judge and deterministic ECC/NCC that may systematically favor ledger-shaped traces.","rationale":"The paper’s design is coherent and the empirical breadth (multiple vendors, ablations aligned with F1–F4, MC-Search chain metrics) makes accuracy improvements credible. The load-bearing soft spot is exactly the measurement of “trajectory faithfulness” that the strongest claim advertises alongside accuracy. The reader correctly flags judge–backbone family overlap and the limits of deterministic containment; that is the condition least secured by the manuscript, not the ledger abstraction or Proposition 1. A cross-family re-audit plus small human calibration is a decisive, feasible check. Until it lands cleanly, CONDITIONAL remains the right verdict—no upgrade to ACCEPT, and no grounds for REJECT given the supporting accuracy/ablation/chain evidence. Agreement with the reader is full on the weakest assumption; this pass only sharpens the concrete falsifier.","tokens_in":28402,"tokens_out":600,"duration_ms":12066,"concrete_test":"Re-run the full S-RFA on V*Bench and EMMA-160 with an independent non-Gemini judge (e.g., Claude-Opus or GPT-5.5) under the same ≤10-claim protocol, plus a human audit on a stratified 50-trace subset scoring entity/numeric support without seeing system identity. If LedgerMind’s GDR/R4R lift vs baseline shrinks by >50% or loses significance under the alternate judge/human labels while accuracy gaps hold, the faithfulness half of the strongest claim is overstated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim couples accuracy gains to trajectory-level faithfulness (UCR_reason, GDR, R4R, WDG; §3.4, Fig. 5). Faithfulness is operationalized by (i) deterministic ECC/NCC over a small alias table and token overlap (§3.2 Eqs. 3–5; App. C) and (ii) a fixed external auditor (Gemini-3.1-Pro) that decomposes both baseline CoT and LedgerMind traces into ≤10 atomic claims under an identical protocol (§4.1, §4.3). Gemini-family models are also evaluated backbones (Gemini-3-Flash / 3.1-Pro), so the judge can preferentially credit ledger-style citation structure and leaf-evidence phrasing even when paraphrase-level or cross-lingual unsupported content escapes the alias table (Limitation J). If that bias is material, the S-RFA polygons and R4R/WDG improvements overstate grounded reasoning relative to answer accuracy, weakening the joint claim even though ablations and MC-Search HPS/RD remain supportive. Proposition 1 is sound but nearly definitional and does not underwrite the faithfulness metrics.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes LedgerMind, a training-free runtime that treats multimodal agent trajectories as provenance-constrained state machines. Tool outputs are normalized into a Structured Evidence Ledger; reasoning and decision claims may cite only active entries; entity- and numeric-level containment (ECC/NCC) is checked; and repair is restricted to typed operators with a provenance non-amplification guarantee (Proposition 1). An Adaptive Dual-Path Dispatcher routes simple vs. complex queries, and an event-triggered verifier drives repair. The design targets four failure modes that final-answer accuracy obscures (unsupported intermediate claims, Phantom Grounding, over-reasoning, repair-time amplification). Experiments on VTC-Bench, MMStar, MMMU, MMMU-Pro, EMMA, MC-Search, and an in-house Hard-200 set, across six backbone MLLMs with matched tool budgets, report gains in answer accuracy and in trajectory-level metrics (UCR_reason, GDR, R4R, WDG), with ablations on MMMU-Pro and chain-alignment gains (HPS up, RD down) on MC-Search.","tokens_in":28762,"tokens_out":1737,"duration_ms":33699,"significance":"If the joint accuracy-and-faithfulness claim holds, the work is a useful systems contribution for multimodal agentic VQA: it makes provenance a structural runtime constraint rather than a prompting preference, names Phantom Grounding at trajectory level, and couples adaptive depth control with typed repair under a clear (if definitional) non-amplification guarantee. Strengths include multi-benchmark, multi-backbone evaluation under matched tool budgets; component ablations that isolate the ledger, typed repair, ECC/NCC, and the dispatcher; and MC-Search chain metrics that are harder to explain by answer-only post-hoc correction. Proposition 1 is correctly stated and proof-checked by operator enumeration. The main significance risk is that trajectory faithfulness is partly operationalized by a same-family external MLLM judge and by deterministic ECC/NCC over a limited alias table; if those instruments favor ledger-shaped traces, the faithfulness half of the claim is overstated even when accuracy rises.","major_comments":[{"comment":"§3.4 and §4.3 (S-RFA / Figure 5): Trajectory faithfulness (UCR_reason, GDR, R4R, WDG) is audited by a fixed Gemini-3.1-Pro judge that decomposes both baseline and LedgerMind traces into ≤10 atomic claims, while Gemini-3-Flash and Gemini-3.1-Pro are also evaluated backbones (§4.1). This creates a same-family auditor risk: the judge may preferentially credit ledger-style citation structure and leaf-evidence phrasing. The joint claim that accuracy gains come from grounded trajectories (not post-hoc rewriting) load-bears on this audit. Please add at least one of: (i) a second auditor from a different vendor family with agreement statistics, (ii) a human-labeled subset with inter-annotator agreement, or (iii) a blinded protocol that strips ledger IDs/formatting before judging. Without this, R4R/WDG and the enclosing polygons in Figure 5 remain only weakly identified.","section":"§4.3, Figure 5; §3.4; §4.1"},{"comment":"§3.2 Eqs. (3)–(5) and Appendix C: ECC/NCC operationalize claim–evidence containment via token overlap, a small alias table, and type-aware numeric tolerance. Limitation J already notes paraphrase, coreference, and cross-lingual gaps. Because Phantom Grounding (F2) and the Hard-split ablation drop for w/o ECC/NCC (Table 3, −5.19 overall, larger on Hard) are central to the paper’s diagnostic story, please quantify false-negative/false-positive rates of ECC/NCC on a labeled claim set (including paraphrases that are still licensed by evidence and entity substitutions that are not). Otherwise it is unclear whether ECC/NCC catches F2 or mainly enforces surface form that the ledger already encourages.","section":"§3.2 Eqs. (3)–(5); Appendix C; Table 3"},{"comment":"§4.2 / Appendix H (Hard-200): Hard-200 is committee-mined and partly self-constructed (RealCAR), then scored by a local LLM judge. Gains are large and uniform (Figure 4, no negative cell), which is encouraging, but the set is not a public fixed benchmark and selection uses cross-vendor failure rates that may correlate with the same failure modes LedgerMind is built to fix. For the stress-test claim, either release the full set with selection scripts and judge prompts, or demote Hard-200 to supplementary evidence and rest the main accuracy claims on the public suites (VTC-Bench, EMMA, MMMU-Pro, MC-Search), which already support a substantial part of the result.","section":"§4.2; Figure 4; Appendix H"}],"minor_comments":[{"comment":"Proposition 1 (§3.3) is correct but nearly definitional given the operator set R. In the main text, state explicitly that it guarantees provenance locality, not factual correctness of tools—this is in the proof paragraph and Limitations but should be adjacent to the proposition statement.","section":"§3.3 Proposition 1"},{"comment":"Appendix C: confidence demotion values (0.50/0.52/0.55) and σ_verify=0.6 are grid-searched on 50 held-out questions. Briefly report sensitivity of main metrics to these thresholds in the appendix so readers can see stability beyond the development set.","section":"Appendix C"},{"comment":"§4.4 Table 3: the dispatcher ablation shows Easy drop and stable Hard—good signature for F3—but absolute Easy/Medium/Hard definitions for MMMU-Pro should be stated in the table caption or appendix for reproducibility.","section":"§4.4 Table 3"},{"comment":"Figure 1 and Figure 2 are helpful; ensure vector text remains legible at single-column width (several labels are dense).","section":"Figure 1; Figure 2"},{"comment":"Related work (§2.1) correctly disclaims novelty of provenance tracing in isolation; a short pointer to how the 11-field schema differs from Open Provenance Model-style records would help systems readers.","section":"§2.1; Appendix B"},{"comment":"EMMA Coding subset shows a small regression (−0.18 vs Gemini 3.1 Pro, Table 1); the footnote explanation is fine—consider one sentence in the main text so readers do not treat it as a silent failure.","section":"Table 1; §4.2"},{"comment":"Typos/formatting: title casing inconsistency between running header (“A STRUCTURED EVIDENCE RUNTIME…”) and abstract name; occasional missing spaces in compound terms in the arXiv text dump. Clean for camera-ready.","section":"Title / running header"}],"recommendation":"major_revision","confidential_remarks":"The empirical package is stronger than average for agentic multimodal systems papers, and I would not reject on novelty grounds: the ledger-as-runtime-state plus typed repair package is a legitimate systems contribution even if individual pieces (citations, self-refine, adaptive depth) exist. The load-bearing weakness is identification of trajectory faithfulness under a same-family judge. If the authors add a cross-family or human audit on a subset and tighten Hard-200 disclosure, this could clear major revision at a solid venue. Scope fit for a serious ML journal is acceptable as a systems/evaluation paper; it is not a theory paper, and Proposition 1 should not be oversold in marketing."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a training-free runtime that turns multimodal agent traces into a provenance-constrained ledger, and the empirical sweep is wide enough that you should take the accuracy gains seriously. The faithfulness story is useful but softer than the answer numbers.\n\nWhat is actually new is the packaging, not any single primitive. They treat the trajectory as a state machine whose only writable evidence is tool-normalized ledger entries, force SC/DC claims to cite active entries, check entity and numeric containment (not just citation IDs), route simple vs complex queries with a dual-path dispatcher, and restrict repair to typed operators so free-form reflection cannot invent provenance. Phantom Grounding is a clean name for citation-backed entity/numeric hallucination at trajectory level. Proposition 1 is correct and almost definitional—it follows from the operator set—but it is still the right contract to write down.\n\nWhat they do well: matched tool budgets across six backbones and four vendors; ablations on MMMU-Pro where removing the ledger hurts most, then typed repair, then ECC/NCC, with the dispatcher showing the over-reasoning signature (easy drops, hard flat); MC-Search where HPS rises and RD falls, which is harder to fake with answer-only hacks; and a failure taxonomy (F1–F4) that actually organizes the design. Hard-200 is a stress set with some author-built RealCAR items—fine if labeled as such, not a substitute for the public benches.\n\nSoft spots, in proportion: the S-RFA faithfulness polygons rest on a Gemini-3.1-Pro judge while Gemini models are also backbones, plus deterministic ECC/NCC over a small alias table. That can favor ledger-shaped traces and miss paraphrase-level drift. Absolute faithfulness deltas may shrink under a different auditor; the accuracy and chain-alignment gains do not automatically vanish. Thresholds are small-set tuned; full repro needs paid APIs and supplementary code. None of that sinks the central engineering claim.\n\nWho it is for: people building tool-using VQA agents who care about auditability and repair safety. Cite the ledger interface, Phantom Grounding, and the typed-repair non-amplification idea. Send it to peer review—serious referee time is warranted; ask for judge-swap ablations and public artifacts, not a desk reject.","headline":"Solid systems paper: ledger-as-state plus typed repair is a real, usable pattern; faithfulness claims need a judge-robustness check but the accuracy and ablation story already stands.","tokens_in":29446,"tokens_out":586,"would_cite":true,"duration_ms":15840,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Multimodal agents become more accurate and faithful when every claim must cite a structured evidence ledger that repair cannot invent content into.","keywords":["multimodal agents","visual question answering","trajectory faithfulness","structured evidence ledger","provenance","phantom grounding","typed repair","adaptive inference"],"falsifier":"Run the same backbones with and without the ledger on matched trajectories: if answer accuracy rises while entity/numeric grounding rates, decision-grounding among correct answers, and chain-alignment metrics do not improve—or if typed repair still injects unsupported entities that the checks miss—the central claim fails.","tokens_in":29211,"feed_emoji":"📒","tokens_out":838,"duration_ms":18210,"temperature":0.7,"pith_summary":"Final-answer scores hide how multimodal agents actually reach answers: unsupported steps, citations that look valid while smuggling in new entities, needless deep reasoning on simple questions, and free-form repair that invents fresh unsupported content. LedgerMind treats the agent trajectory as a provenance-constrained state machine whose only shared state is a Structured Evidence Ledger. Tool outputs become ledger entries; reasoning and decisions may cite only active entries; grounding is checked at entity and numeric level; and repair is limited to typed transitions that cannot add content without tool provenance. Across several multimodal benchmarks and backbone models, that design raises both answer accuracy and trajectory-level faithfulness, with a formal guarantee that repair does not amplify provenance-less claims.","feed_headline":"Agent answers improve when claims must cite a locked evidence ledger","feed_subtitle":"Typed repair cannot invent content, and accuracy rises with grounded trajectories, not lucky finals","key_machinery":"The Structured Evidence Ledger: a normalized trajectory state in which every tool return becomes an entry with source, type, confidence, lifecycle status, and dependencies; claims may cite only active entries; a three-layer grounding protocol checks support coverage plus entity and numeric containment; and an event-triggered engine repairs only via typed operators that preserve provenance non-amplification.","core_discovery":"Treating a multimodal agent trajectory as a provenance-constrained state machine centered on a Structured Evidence Ledger—with entity- and numeric-level grounding and typed repair—improves both final-answer accuracy and trajectory-level faithfulness, while guaranteeing that repair cannot introduce ledger content without tool-produced provenance.","pith_inferences":["If provenance is structural state rather than a prompt preference, similar ledger contracts could transfer to tool-using text agents and retrieval pipelines where citation–content mismatch is already known.","Long-horizon or video agents would need the ledger’s lifecycle and time-to-live rules to become first-class memory, not just per-query scratch state.","A learned complexity router could replace the paper’s rule-based dispatcher once task mixtures grow more heterogeneous than the evaluated benchmarks."],"forward_implications":["Agent evaluation can report trajectory faithfulness (unsupported-claim rate, grounded decisions, right-for-right reasons) alongside accuracy instead of accuracy alone.","Citation-looking intermediate text is no longer treated as grounded unless conclusion-level entities and numbers are licensed by cited tool evidence.","Repair loops can be restricted to typed ledger/action transitions so free-form self-reflection cannot silently invent new provenance-less claims.","Simple and knowledge-heavy queries can be routed to a shallow path to avoid overwriting correct direct answers with noisy multi-step inference.","The same ledger state can later supply training signals that distinguish grounded from ungrounded claims at trajectory level."],"fun_headline_variants":["Claims must cite a locked evidence ledger—accuracy and faithfulness both rise","Provenance-constrained ledger stops phantom grounding and repair-time invention","Agent trajectories as state machines: grounded ledger, typed repair, no content freebies","Structured Evidence Ledger curbs unsupported reasoning and citation-backed hallucinations","Entity-level grounding plus non-amplifying repair beats lucky final-answer scores"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That checking whether claimed entities and numbers appear in cited tool evidence, scored by an external model that breaks traces into a handful of atomic claims, is enough to measure whether the trajectory was truly grounded.","fun_headline_variants_meta":{"raw":{"variants":["Claims must cite a locked evidence ledger—accuracy and faithfulness both rise","Provenance-constrained ledger stops phantom grounding and repair-time invention","Agent trajectories as state machines: grounded ledger, typed repair, no content freebies","Structured Evidence Ledger curbs unsupported reasoning and citation-backed hallucinations","Entity-level grounding plus non-amplifying repair beats lucky final-answer scores"]},"model":"grok-4.5","effort":"low","cost_usd":0.00193,"raw_usage":{"total_tokens":876,"prompt_tokens":777,"num_sources_used":0,"completion_tokens":79,"cost_in_usd_ticks":19304000,"prompt_tokens_details":{"text_tokens":777,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":20,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":777,"tokens_out":79,"duration_ms":2480,"temperature":1.0,"reasoning_tokens":20,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T09:22:32.741832+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same backbones with and without the ledger on matched trajectories: if answer accuracy rises while entity/numeric grounding rates, decision-grounding among correct answers, and chain-alignment metrics do not improve—or if typed repair still injects unsupported entities that the checks miss—the central claim fails.","supporting_citations":[],"review_version":1}