{"id":"2a4cb727-2154-44d3-bc17-f948578b1233","arxiv_id":"2604.03664","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"FinLongDocQA targets single- and cross-table numerical QA in long financial reports; FinLongDocAgent uses multi-round multi-agent RAG with intermediate calculation and verification.","lead":"The authors introduce FinLongDocQA, a benchmark for single- and cross-table numerical reasoning over long financial annual reports, and FinLongDocAgent, a multi-agent multi-round RAG system. They argue long context and multi-step arithmetic are the main LLM failure modes, and that iterative retrieve–calculate–verify helps.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Wrong full manuscript still leaves FinLongDocQA/Agent claims unauditable; evaluation fidelity and agent gains remain unchecked.","rationale":"The reader correctly treated this as abstract-only because the full manuscript body is a mismatched paper (Explainable PQC). I re-checked the cacheable context: it is still 2604.03665, not FinLongDocQA. No new experimental tables, ablations, or release artifacts for FinLongDocAgent are available. The weakest assumption the reader named—faithful question construction/evaluation and non-artifactual agent gains—is therefore still the single most load-bearing concern, and it remains untestable from the given materials. No internal contradiction in the abstract was found; the issue is absence of evidence, not a demonstrated error. Verdict should stay UNVERDICTED with low confidence until the correct PDF is supplied and the concrete checks above are run.","tokens_in":9365,"tokens_out":551,"duration_ms":11825,"concrete_test":"Retrieve the actual arXiv 2604.03664 PDF/source. Verify presence of: (i) dataset size, single/cross-table split, and annotation/agreement process; (ii) exact metrics and answer-normalization rules; (iii) ablations that remove iterative retrieval and verification; (iv) comparisons to strong long-context and multi-table RAG baselines. If any of (i)–(iv) is missing, or gains disappear under ablations/normalization variants, the strongest claim weakens and the verdict stays UNVERDICTED or moves toward REJECT.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that FinLongDocQA exposes two bottlenecks (context rot beyond ~129k tokens; multi-step numerical errors after evidence is found) and that FinLongDocAgent’s multi-agent multi-round RAG with intermediate calculation and verification improves reliable numerical QA. That claim is load-bearing on (1) the dataset’s questions and labels faithfully representing real cross-table financial analysis, and (2) reported agent gains not being artifacts of unstated baselines, retrieval setup, or answer normalization. The cacheable full text is a different paper (Explainable PQC, arXiv 2604.03665), not 2604.03664. From the abstract alone there is no N, single- vs cross-table split, annotation process, metrics, baselines, ablations, or normalization protocol. Without those, neither bottleneck nor agent improvement can be audited. This is the same gap the reader flagged; the supplied body does not close it.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission under review is titled and abstracted as FinLongDocQA / FinLongDocAgent: a dataset and multi-agent multi-round RAG system for single- and cross-table numerical reasoning over long financial annual reports, claiming two bottlenecks (context rot beyond ~129k tokens; multi-step arithmetic errors after evidence is found) and gains from iterative retrieval, intermediate calculation, and verification. The full manuscript body supplied with the submission is, however, an entirely different paper—“Explainable PQC,” a conceptual layered interpretive framework for lattice-based post-quantum cryptography (complexity vocabulary, exploratory combinatorial Hodge theory, and low-dimensional Julia LLL/BKZ experiments)—with no FinLongDocQA data, annotation protocol, baselines, ablations, or numerical-QA results. Consequently the claimed dataset, bottlenecks, and agent improvements cannot be audited from the materials provided.","tokens_in":9557,"tokens_out":915,"duration_ms":17437,"significance":"If the FinLongDocQA claims were supported by a matching manuscript (dataset statistics, annotation process, single- vs. cross-table splits, metrics, baselines, ablations, and answer-normalization protocol), the work would address a genuine gap in long-document financial numerical reasoning and could be a useful empirical contribution. As submitted, that significance cannot be assessed: the body is a different paper whose own contribution is conceptual/organizational rather than a new hardness result or attack, and it does not substantiate any of the abstract’s ML claims. No machine-checked proofs, reproducible FinLongDocQA code, or falsifiable numerical-QA results for the stated paper are present in the provided text.","major_comments":[{"comment":"Title/abstract vs. full text mismatch: the abstract and paper_id claim FinLongDocQA and FinLongDocAgent (cs.CL financial numerical QA), but the full manuscript is “Explainable PQC” (arXiv:2604.03665, lattice PQC interpretability). No section of the body defines FinLongDocQA, reports N, single/cross-table splits, annotation, metrics, baselines, or agent ablations. The central claims (two bottlenecks; multi-agent multi-round RAG gains) are therefore unauditable and load-bearing evidence is missing.","section":null},{"comment":"Even restricting attention to the abstract alone, the evaluation protocol is unspecified: no dataset size, construction/annotation process, answer-normalization rules, retrieval setup, or comparison baselines. Without these, neither the “context rot beyond ~129k tokens” bottleneck nor the claimed benefit of iterative retrieval/verification can be checked for artifacts—the weakest assumption flagged for this paper remains unchecked.","section":null},{"comment":"If the intended submission were instead the provided Explainable PQC body, its own central claim (a layered interpretive framework for PQC communication) is explicitly non-proof and non-parameter; Sections 3–4 remain exploratory/illustrative (low-dim 10–40 lattices far below NIST ranges; Hodge/local-generation link to hardness left open). That paper does not support the FinLongDocQA abstract and cannot substitute for it.","section":null}],"minor_comments":[{"comment":"Abstract of the claimed paper is clear on motivation but gives no quantitative preview (dataset size, accuracy deltas, model list), which would normally be expected even in a short abstract for an empirical systems paper.","section":null},{"comment":"In the mismatched body (Explainable PQC), Figure 1 and Tables 1–2 are useful for the interpretive triple (Sc, Sq, Sr), but the Julia toolkit results (Figure 3, 10D/40D) need clearer units and hardware details if that paper were under review separately.","section":null}],"recommendation":"reject","confidential_remarks":"The cacheable full text is a different arXiv paper (2604.03665 Explainable PQC) from the stated target (2604.03664 FinLongDocQA). This looks like a packaging/metadata error rather than a reviewable submission. I recommend desk-reject or return to authors for the correct PDF; I did not attempt to invent results for FinLongDocQA from the abstract alone. If the journal intended review of the PQC paper, that should be reassigned under the correct title and ID."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The package is broken in a simple way: the target is arXiv 2604.03664 (FinLongDocQA / FinLongDocAgent), but the full manuscript body is Explainable PQC (2604.03665). Treat this as abstract-only for the financial QA paper. Do not read the PQC text as evidence for Wang et al.\n\nFrom the abstract alone, the intended contribution is clear and useful for a narrow audience. Existing financial numerical QA suites are mostly single-table. They claim a long-report dataset covering single- and cross-table numerical reasoning, diagnose two bottlenecks (context rot past ~129k tokens when finding tables; multi-step arithmetic errors even after evidence is found), and propose a multi-agent multi-round RAG loop that retrieves, computes intermediate values, and verifies across rounds. That problem framing matches real analyst work better than single-table suites, and iterative retrieve–calculate–verify is a sensible engineering pattern for this setting.\n\nWhat is new is the setting and the dataset claim, not the agent architecture. Multi-round agentic RAG with tools and verification is an expected extension of current RAG/tool-use lines. The abstract’s “experiments highlight importance of iterative retrieval and verification” is the load-bearing result, and we cannot check it: no N, no single- vs cross-table split, no annotation protocol, no metrics, no baselines, no ablations, no answer normalization, no release artifacts. The weakest assumption is that FinLongDocQA questions and labels faithfully track real cross-table analysis and that reported gains are not setup artifacts. That is unchecked, not disproven.\n\nWho this is for: financial NLP and long-context RAG people, if and when the real paper, data, and code appear. From this package, there is nothing to run, cite, or stress-test. A serious editor would send a complete FinLongDocQA paper with proper evaluation to referees; they would desk-reject this mismatched submission. Recommendation: do not engage further until the correct full manuscript and artifacts are available. Then re-score from the experiments, not the abstract.","headline":"We only have the FinLongDocQA abstract; the attached full text is a different paper (Explainable PQC), so the dataset and agent claims cannot be audited.","tokens_in":10190,"tokens_out":532,"would_cite":false,"duration_ms":14134,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A multi-agent multi-round RAG system improves reliable numerical QA over long financial reports by iterating retrieval, intermediate calculation, and verification.","keywords":["document-level numerical reasoning","financial reports","cross-table QA","long-context LLMs","retrieval-augmented generation","multi-agent systems","FinLongDocQA","context rot"],"falsifier":"Re-run the same models and agent on a held-out set of real analyst-written cross-table questions from full annual reports (with independent human numerical ground truth) and check whether the iterative retrieve–calculate–verify loop still yields the claimed accuracy lift over single-pass RAG and raw long-context LLMs.","tokens_in":10208,"feed_emoji":"📊","tokens_out":943,"duration_ms":21226,"temperature":0.7,"pith_summary":"Large language models still fail at trustworthy numerical question answering over long, structured financial annual reports, where answers require arithmetic over evidence scattered across multiple tables and narrative text. Existing benchmarks mostly stay in single-table settings, so the authors introduce FinLongDocQA to cover both single-table and cross-table numerical reasoning inside full-length reports. Evaluating closed- and open-source LLMs on this dataset surfaces two bottlenecks: reports often exceed roughly 129k tokens, which worsens context rot when locating the right tables, and even after evidence is found models still err on multi-step arithmetic. FinLongDocAgent addresses both problems with a multi-agent multi-round retrieval-augmented generation loop that repeatedly retrieves, computes intermediate results, and verifies them. Experiments show that this iterative retrieve–calculate–verify pattern is essential for more reliable numerical QA on long financial documents.","feed_headline":"Iterative multi-agent RAG fixes numerical QA in long reports","feed_subtitle":"FinLongDocQA shows context rot and multi-step arithmetic errors; retrieve–calculate–verify helps.","key_machinery":"FinLongDocAgent: a Multi-Agent Multi-Round RAG pipeline that loops over evidence retrieval, intermediate numerical calculation, and cross-round verification rather than answering in a single pass.","core_discovery":"FinLongDocQA exposes that document-level financial numerical reasoning fails for two distinct reasons—context rot when locating relevant tables in reports longer than about 129k tokens, and multi-step arithmetic errors even after the right evidence is retrieved—and that a multi-agent multi-round RAG agent that iteratively retrieves evidence, performs intermediate calculations, and verifies results measurably improves reliability on both single-table and cross-table questions.","pith_inferences":["If context rot is the dominant retrieval failure, hybrid table-index or layout-aware retrieval may reduce the need for very long raw contexts more than larger context windows alone.","Separating calculation into an explicit, checkable intermediate step suggests numerical QA systems may benefit from tool-use or program-of-thought style execution rather than pure free-form generation.","Cross-table financial QA is a natural stress test for any long-context or agentic RAG claim, because wrong table selection and arithmetic drift are independently measurable.","Future work could ablate which agent role (retriever vs. calculator vs. verifier) contributes most of the gain to isolate whether the multi-agent design is necessary or whether multi-round single-agent verification would suffice."],"forward_implications":["Benchmarks that stop at single tables will systematically understate failure modes that appear only when evidence is scattered across a full report.","Simply stuffing an entire annual report into a long context window is insufficient; locating the right tables remains a first-order failure mode above roughly 129k tokens.","Even perfect table retrieval is not enough: multi-step arithmetic still needs intermediate calculation and verification steps.","Agent designs for financial QA should treat retrieval, calculation, and verification as separate iterative stages rather than a single generation pass.","The same iterative pattern is a concrete direction for other long, table-heavy document domains that demand exact numerical answers."],"fun_headline_variants":["Multi-agent multi-round RAG eases numerical QA in long financial reports","FinLongDocQA: context rot and arithmetic errors block document-level numbers","Iterative retrieve-calculate-verify lifts single- and cross-table financial QA","Reports over 129k tokens expose table-locating and multi-step math failures","FinLongDocAgent improves reliable numerical reasoning across financial tables"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The dataset’s questions and scoring truly reflect how analysts do cross-table numerical work, and the reported gains are not artifacts of how retrieval, baselines, or answer normalization were set up.","fun_headline_variants_meta":{"raw":{"variants":["Multi-agent multi-round RAG eases numerical QA in long financial reports","FinLongDocQA: context rot and arithmetic errors block document-level numbers","Iterative retrieve-calculate-verify lifts single- and cross-table financial QA","Reports over 129k tokens expose table-locating and multi-step math failures","FinLongDocAgent improves reliable numerical reasoning across financial tables"]},"model":"grok-4.5","effort":"low","cost_usd":0.004342,"raw_usage":{"total_tokens":1311,"prompt_tokens":785,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":43420000,"prompt_tokens_details":{"text_tokens":785,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":443,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":785,"tokens_out":83,"duration_ms":5040,"temperature":1.0,"reasoning_tokens":443,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T12:35:47.370969+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the same models and agent on a held-out set of real analyst-written cross-table questions from full annual reports (with independent human numerical ground truth) and check whether the iterative retrieve–calculate–verify loop still yields the claimed accuracy lift over single-pass RAG and raw long-context LLMs.","supporting_citations":[],"review_version":1}