{"id":"54cb7238-7aac-4732-9bde-28418b52d3c5","arxiv_id":"2608.07943","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A controlled attribution study of multi-page document understanding finds that missing evidence and cross-page integration failures dominate over distractor noise and text extraction quality.","lead":"This paper tests where multi-page document question answering systems go wrong by changing one part at a time: how the page is encoded, which pages are selected, and how the model reasons over the evidence. It finds that missing pages hurt far more than extra pages, and that models struggle to combine evidence across pages even when all the needed pages are supplied.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Multi-page questions combine more evidence pieces than single-page ones; without a matched or page-merged control, the 25.9-point TLV deficit in Table 3 may be an evidence-combination difficulty effect, not a failure specific to cross-page integration.","rationale":"This is a serious, well-designed attribution study. The internal interventions — oracle gold pages, withholding and distractor sweeps, verdict transition analysis, document-level bootstrap, two rankers, and a scale sweep — give strong support for the representation and selection findings, and the reasoning-stage attribution is well motivated: with gold pages supplied at TLV, multi-page accuracy is far below single-page accuracy, and neither context length (Figure 3b) nor model scale (Table 18) explains the gap. I agree with the reader that the weakest assumption is pool comparability between single-page and multi-page questions; it is the load-bearing step for the claim that the deficit is specifically a cross-page integration failure. My sharpening is that this is not an abstract matching problem: the multi-page criterion is structurally correlated with the number of evidence pieces a question requires combining, a difficulty factor independent of page layout. The paper's controls (length via distractors, CoT, scale) all remain consistent with this confound. The decisive check is therefore a within-question intervention — presenting the same gold evidence as one synthetic page and measuring whether accuracy recovers. If it does, the cross-page mechanism is confirmed; if not, the contribution should be re-scoped to a multi-evidence reasoning bottleneck, which changes the recommended training target in Section 6.1. This test is cheap because the authors control the full pipeline. Lesser concerns (single LLM judge, Qwen3-family generality, page-level annotation granularity) are acknowledged in the Limitations and affect effect sizes rather than the central inference. The verdict stays CONDITIONAL, with the page-merge control (or an equivalent difficulty-matched comparison) as the condition to be satisfied.","tokens_in":29122,"tokens_out":14672,"duration_ms":154475,"concrete_test":"Page-merge control on the 352 multi-page questions from Table 3. Re-run the TLV condition feeding the same gold pages as a single merged page: concatenate the two page images vertically into one image at the same total pixel/token budget (token count unchanged), with no page-boundary marker; for the T and TL rungs, concatenate the two pages' text into one continuous block without a separator. If merged-page accuracy recovers more than half of the 25.9-point TLV deficit (moving substantially toward the 64.6 single-page level), the failure is specific to evidence spread across a page boundary and the cross-page integration claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim — that reasoners 'fail to integrate evidence across pages even when every required page is supplied' (Section 5.3, Table 3: TLV 64.6 single-page vs 38.6 two-page, deficit −25.9) — depends on the assumption that the single-page and multi-page question pools differ only in page spread. That comparability is not established, and the manuscript's Limitations section does not flag it. Multi-page questions are structurally more likely to require combining several evidence pieces (two dates, a chart value plus a sample-size table, two table cells), and multi-evidence combination is harder than single-evidence reading even within one page. The paper's controls do not remove this confound. The distractor experiment (Figure 3b) shows a single relevant page plus three distractors holds accuracy at 63.1, but that input still has only one locus of relevant evidence; it rules out length, not evidence-combination load. Chain-of-thought narrowing the gap (38.6 to 44.9, Table 3) is consistent with a difficulty confound, since CoT is known to help hardest questions most, and the 32B scale result (deficit −25.1) is a single comparison that neither closes nor clearly widens the gap. If the deficit survives when the same gold evidence is presented with no page boundary, cross-page specificity is genuine; if it disappears, the finding reduces to a general multi-evidence reasoning bottleneck, materially changing the claimed contribution and the training recommendation in Section 6.1.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an attribution framework for multi-page visually rich document understanding (MP-VRDU) that decomposes incorrect answers into three failure modes—representation, selection, and reasoning—each with two mechanisms, and instantiates the framework in a controlled single-pass retrieval-generation pipeline evaluated on MMLongBench-Doc. The empirical study intervenes on each locus while holding others fixed: representation is varied across text, layout, and vision encodings; selection is controlled by withholding gold pages or adding distractor pages; reasoning is probed by supplying complete gold evidence and varying prompting, scale, and model family. The main findings are that vision is necessary but does not replace text extraction, that missing evidence sharply bounds accuracy while distractor pages are largely tolerated, and that reasoners fail to integrate evidence across pages even when all required pages are supplied. The paper further reports that chain-of-thought prompting narrows the cross-page deficit more than increasing model scale does, and derives design guidance for building MP-VRDU systems under fixed compute budgets.","tokens_in":29360,"tokens_out":4156,"duration_ms":47056,"significance":"If the findings hold, the paper makes a useful contribution: it offers a common vocabulary for MP-VRDU failures, a controlled experimental design that separates selection from reasoning, and a concrete, falsifiable claim that cross-page evidence integration—not context length or evidence selection—is a major remaining bottleneck. The paper ships a careful experimental apparatus: oracle gold pages, per-question verdict transitions that distinguish systematic loss from churn, document-level bootstrap confidence intervals, and a rich set of worked examples. The practical guidance on routing, recall-oriented retrieval, quantization, and abstention is grounded in measured trade-offs. The central empirical claim, however, rests on comparing naturally occurring single-page and multi-page question pools, and that comparability is not established; this is the key risk to the paper's main conclusion and to its training recommendation.","major_comments":[{"comment":"The central claim that the reasoner 'fails to integrate evidence across pages' is supported by comparing accuracy on single-page questions (TLV: 64.6) with accuracy on multi-page questions (TLV: 38.6). This comparison assumes the two question pools are comparable in reasoning difficulty aside from the page spread of the evidence. No matching, difficulty rating, or single-page control with the same number of evidence pieces is reported. The distractor experiment in §5.2 (Figure 3b) rules out context length, but its single-gold-page condition does not reproduce the multi-evidence composition of multi-page questions. Chain-of-thought narrowing the gap (38.6 to 44.9) and the 32B result leaving the deficit at −25.1 are both consistent with multi-page questions being intrinsically harder (requiring combination of more evidence units), independent of page boundaries. The manuscript should provide a control that removes the page boundary while keeping the evidence set unchanged, or otherwise demonstrate that the single-page and multi-page pools are matched on evidence-unit count and type. Without this, the conclusion that the loss is specifically a cross-page reasoning failure is not uniquely supported.","section":"§5.3, Table 3"},{"comment":"The claim that a stronger representation 'enlarges the gap' and 'points to a visual reasoning failure specifically' relies on the same pool-comparability assumption, now made sharper: the deficit grows from −11.7 at T to −25.9 at TLV. Table 13 shows that the deficit varies wildly by domain, from +8.5 (Academic paper) to −54.1 (Financial report), and by evidence source, from −0.4 (Chart at T) to −34.5 (Table at TLV). This variation suggests that the composition of the single-page and multi-page pools strongly determines the measured gap. If multi-page questions draw disproportionately on evidence sources or domains where the image helps less, or where questions are intrinsically harder and hence less responsive to better inputs, the widening gap need not indicate a visual integration failure. The manuscript should demonstrate that the single-page and multi-page pools are balanced on evidence source, domain, and evidence-unit count before attributing the widening to a modality-specific integration deficit.","section":"§5.3, Table 3 and §D.1, Table 13"},{"comment":"The 'missing evidence bounds accuracy' result is presented as an empirical finding, but it follows largely by construction: the gold-page annotation defines the withheld evidence as required, so removing it removes information the question needs. The paper partially acknowledges this framing, but the practical recommendation to 'favour recall' would be more convincing if the coverage mechanism had a quantitative baseline—for example, an estimate of how often the remaining pages already contain enough evidence to answer the question, or a judged upper bound on accuracy given only the remaining pages. As reported, the near-collapse in accuracy when one gold page is withheld (Figure 3a) confirms the annotation, but does not by itself tell us how much of the drop is due to evidential necessity versus the reasoner's inability to use partial evidence. A control that presents the non-gold pages as the context, or scores the question with the withheld page's content paraphrased into another page, would strengthen the attribution without changing the framework.","section":"§5.2 and §4.1"}],"minor_comments":[{"comment":"The main text states that Gemini 2.5 Flash is used as the judge, while Appendix A.8 describes two judges (GPT-4o-mini and Gemini-2.5-flash) run at temperature 0. Please clarify which judge produced the main-text results and whether the two judges were combined or compared for agreement.","section":"§4.3 and Appendix A.8"},{"comment":"The paper states that 95% confidence intervals are computed from a document-level bootstrap, but the main-text tables do not display intervals. Adding intervals to the key tables, or explicitly referencing the appendix where they appear, would help readers assess the stability of the reported gaps, especially the 25.9-point deficit in Table 3.","section":"§4.3 and Tables 2–5"},{"comment":"The caption reports 'oracle 38.6, n=352' for the withholding panel and 'oracle 64.6, n=474' and 'oracle 38.6, n=241' for the distractor panel, but the relationship of these n values to the Table 1 counts (480 single-page, 246 two-page, 112 three-plus-page) is not explained on first reading. Please clarify whether the n values refer to the answerable subset used in each experiment and why the oracle for the two-page condition appears in both panels.","section":"Figure 3 caption"},{"comment":"The representation 'TV' is recommended in the main text ('TV is the cheaper choice on born-digital pages') but is not defined until Appendix A.2. Define TV and TLVi at first use in §4.2 so the discussion in §6 is self-contained.","section":"§6.1 and Appendix A.2"},{"comment":"The worked example M3 is labelled 'PROVISIONAL: partial G2 pool.' If any example is provisional, either complete it and remove the label or omit it from the final version; a provisional example undermines the otherwise strong illustration of the recall trade-off.","section":"Appendix F, worked example M3"}],"recommendation":"major_revision","confidential_remarks":"The paper has a strong experimental design and a clear, practically relevant message, but the central cross-page integration claim needs a control that separates page boundaries from evidence-combination difficulty. This is a fixable but load-bearing gap: a page-merging condition or a matched-difficulty analysis would either confirm or substantially qualify the main contribution. The framework itself and the selection/representation findings are publishable regardless. I would not reject on the current evidence, but the comparison in §5.3 should be reworked before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid empirical paper that earns its conditional accept, but the marquee finding—that reasoners fail at cross-page integration, not at capacity—is not as clean as the paper suggests. The three-locus framing (representation, selection, reasoning) is genuinely useful, and the intervention design is careful: oracle gold pages, paired perturbations, verdict transitions that separate churn from systematic loss, and document-level bootstrap CIs. The worked examples are very helpful, especially the one where the model counts a number off one page instead of counting objects across two. The representation and selection findings look robust, and the distractor-tolerance result is a nice, counterintuitive contribution.\n\nThe soft spot is exactly what your stress-test note says. The cross-page deficit in Table 3 compares naturally occurring single-page questions with multi-page questions. Those pools almost certainly differ in more than page spread—multi-page questions typically require combining several pieces of evidence, and assembling two dates or two table cells is harder than reading one value even within a single page. The paper doesn't match or counterbalance difficulty, and the Limitations section, which lists dataset, model family, granularity, and judge, doesn't flag this. The distractor experiment rules out length but not evidence-combination load, and the CoT narrowing is consistent with a difficulty effect. The fix is concrete and should be in the revision: take the same gold evidence and present it with and without the page boundary, or match questions on a difficulty rating. If the gap survives that control, the reasoning claim is solid; if it collapses, the finding is a more general multi-evidence bottleneck, which still matters but changes the headline.\n\nMinor: no code or data artifacts, and the 32B comparison is a single point that neither closes nor widens the gap. Both are addressable.\n\nBottom line: the framework and most of the empirical results will be useful regardless, and the paper deserves a serious referee. It should be accepted after the integration comparison is hardened, not before.","headline":"A careful, well-controlled attribution study whose headline reasoning claim rests on an unmatched question-pool comparison; the fix is achievable and the rest holds up.","tokens_in":29905,"tokens_out":2156,"would_cite":true,"duration_ms":24866,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper attributes multi-page document QA failures to three loci and shows that reasoners fail to integrate evidence across pages even when every required page is supplied.","keywords":["multi-page document understanding","visually rich documents","failure attribution","evidence selection","cross-page reasoning","chain-of-thought prompting","retrieval recall","document question answering"],"falsifier":"Take a set of questions whose evidence chain can be placed either on one page or split across two pages while keeping topic, answer type, and judged difficulty matched; if accuracy does not drop noticeably in the two-page condition under the combined encoding, the claimed cross-page integration failure is not confirmed.","tokens_in":28903,"feed_emoji":"📄","tokens_out":7656,"duration_ms":72853,"temperature":0.7,"pith_summary":"This paper tries to locate where multi-page visually rich document understanding fails by attributing incorrect answers to three stages: representation (does the encoding preserve the evidence?), selection (is the right evidence delivered?), and reasoning (can the model use it?). Using a controlled intervention design on the MMLongBench-Doc benchmark, it finds that vision is necessary but does not replace text extraction, that omitting one required page sharply caps accuracy while adding distractor pages costs little, and that reasoners fail to integrate evidence across pages even when every required page is supplied. The cross-page integration drop is not explained by context length, because chain-of-thought prompting narrows it while increasing model scale does not. The paper converts these findings into concrete design guidance for building MP-VRDU systems under a fixed compute budget.","feed_headline":"Document AI fails to join evidence across pages","feed_subtitle":"Even with every required page supplied, accuracy drops when evidence spans two pages; prompting helps, scale does not.","key_machinery":"The central object is the three-locus attribution framework, which decomposes an incorrect answer into representation, selection, and reasoning, each with two mechanisms: modality ceiling and conversion fidelity for representation; evidence coverage and distractor exposure for selection; evidence integration and response calibration for reasoning. The framework is made experimental by 'attribution by construction': each intervention changes only the condition a failure mode depends on while all other stages are held fixed. The load-bearing comparison is the single-page versus multi-page accuracy contrast at the same encoding, which isolates cross-page integration from mere input length.","core_discovery":"The paper's central claim is that incorrect answers in multi-page visually rich document understanding can be attributed to three hierarchical failure modes—representation, selection, and reasoning—and that once representation and selection are controlled, the dominant remaining failure is cross-page evidence integration. On MMLongBench-Doc, the default reasoner at the combined text-layout-vision encoding answers 64.6% of single-page questions but only 38.6% of two-page questions, a pooled multi-page deficit of 25.9 points despite the context being well within the model's window. Chain-of-thought prompting narrows the deficit to 16.5 points, whereas scaling the reasoner from 2B to 32B parameters leaves it essentially unchanged. The paper reads this as evidence that current reasoners fail to assemble evidence across pages even when every required page is fully supplied—a reasoning failure, not a capacity or selection failure.","pith_inferences":["If the cross-page deficit is truly a reasoning-stage failure, then training objectives that force explicit multi-page evidence assembly—for example, requiring the model to state which page contributes which fact—may close a gap that prompting and scale only partially address; the paper itself notes that multi-page integration is rarely a training target.","The deficit widening as representation improves (from 11.7 points on text to 25.9 on text-layout-vision) suggests a visual-integration bottleneck specifically; a testable extension would use interleaved per-page text-image ordering or visual grounding supervision to see whether the image-based gap shrinks.","Because the paper does not demonstrate that single-page and multi-page question pools are matched in difficulty, an exact attribution would require rewriting the same evidence chain into single-page and multi-page variants; such a control could revise the magnitude of the reasoning deficit.","The distractor-tolerance result is measured at modest context lengths; extending retrieval depth or moving to learned rankers could reverse the asymmetry if long enough contexts eventually impair evidence use, which the paper acknowledges as a limit."],"forward_implications":["Retrieval should be tuned for recall over precision: removing one required gold page drops multi-page accuracy from 38.6 to about 18.5 at the combined encoding, while adding three distractor pages costs only a few points.","The combined text-layout-vision encoding is the safe default, but routing representation by document class can match vision-only accuracy at less than half the input tokens, saving compute where questions do not need images.","Chain-of-thought prompting should be applied selectively to multi-page questions: it narrows the integration deficit by roughly 8–9 points while leaving single-page accuracy near baseline.","Under a fixed memory budget, quantization is the first thing to trade for parameters: four-bit weights cost about 0.6 accuracy points at the combined encoding while shrinking the reasoner footprint by 61%.","Abstention instructions should be enabled when a wrong answer costs more than no answer, because they raise correct refusal on unanswerable questions to about three-quarters while adding false refusals on answerable ones."],"supporting_citations":[{"why":"Supplies the MMLongBench-Doc benchmark, its gold-page annotations, and the question pools on which all interventions are run.","marker":"(Ma et al., 2024)"},{"why":"Defines the default Qwen3-VL reasoner and the model-family scale sweep from 2B to 32B.","marker":"(Bai et al., 2025)"},{"why":"Provides the 'lost in the middle' context-length degradation claim that the paper's reasoning-failure conclusion is contrasted against.","marker":"(Liu et al., 2024)"},{"why":"Motivates the conversion-fidelity mechanism by showing how parser errors cascade into downstream wrong answers.","marker":"(Zhang et al., 2025)"},{"why":"Supplies PaddleOCR-VL, the default parser used in the representation ladder and parser comparisons.","marker":"(Cui et al., 2025)"},{"why":"States the competing claim that irrelevant context dilutes attention, which the distractor-exposure experiments directly test.","marker":"(Shi et al., 2023)"},{"why":"Frames abstention as a reliable-answering strategy, grounding the response-calibration mechanism's hallucination/abstention tradeoff.","marker":"(Whitehead et al., 2022)"}],"fun_headline_variants":["Scaling won't fix multi-page doc AI, prompting might","Cross-page evidence stumps doc AI even at 32B","Prompting narrows doc AI's cross-page reasoning gap","Why doc AI fails to connect evidence across pages","Doc AI needs prompts, not bigger brains, for multi-page"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison between single-page and multi-page questions assumes the two question pools are otherwise comparable, so the accuracy gap measures integration ability rather than intrinsic difficulty; if multi-page questions are simply harder, the deficit is partly a task-difficulty effect.","fun_headline_variants_meta":{"raw":{"variants":["Scaling won't fix multi-page doc AI, prompting might","Cross-page evidence stumps doc AI even at 32B","Prompting narrows doc AI's cross-page reasoning gap","Why doc AI fails to connect evidence across pages","Doc AI needs prompts, not bigger brains, for multi-page"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1368,"prompt_tokens":857,"completion_tokens":511,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":473,"completion_tokens_details":{"reasoning_tokens":429}},"tokens_in":473,"tokens_out":511,"duration_ms":13343,"temperature":1.0,"reasoning_tokens":429,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:37:15.563004+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of questions whose evidence chain can be placed either on one page or split across two pages while keeping topic, answer type, and judged difficulty matched; if accuracy does not drop noticeably in the two-page condition under the combined encoding, the claimed cross-page integration failure is not confirmed.","supporting_citations":[],"review_version":1}