Pith. sign in

REVIEW 4 major objections 8 minor 106 references

A holistic view of the whole document builds cleaner concept–chunk links and lets multimodal GraphRAG answer faster without dense graph walks.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-30 17:06 UTC pith:ZSMF6OTS

load-bearing objection Solid multimodal GraphRAG systems paper: conflict-aware concept indexing plus cheap online retrieval; big gains need code and tighter hyperparameter honesty. the 4 major comments →

arxiv 2607.24861 v1 pith:ZSMF6OTS submitted 2026-07-26 cs.IR cs.AI

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document

classification cs.IR cs.AI
keywords multimodal GraphRAGcomplex document QAholistic-view graph constructionconcept-level indexingcross-modal conflict resolutionmodality-aware evidence organizationretrieval efficiency
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Complex document QA fails when evidence is scattered across pages, sections, text, tables, and images. Graph-based retrieval is meant to fix that, but the paper argues two practical failures hold it back: knowledge is extracted chunk-by-chunk without checking the global graph, so cross-modal indices get noisy or contradictory; and online search then crawls large entity graphs, which is slow. HVM-GraphRAG’s answer is to keep a running holistic view while building the graph, resolve conflicts against supporting chunks, and collapse entities into a compact concept graph with direct concept-to-chunk indices. At query time it only anchors and lightly expands on that concept graph, pulls evidence straight from the index, optionally tops up with ordinary chunk retrieval, and groups the context by modality before answering. On three complex-document benchmarks the method leads most accuracy metrics and keeps online cost near ordinary RAG rather than heavy graph baselines.

Core claim

Guiding multimodal graph construction with a document-wide holistic view yields reliable concept-level indices to supporting chunks; searching that compact concept graph and reading evidence through the index—then organizing chunks by modality—improves answer quality while cutting the need for expensive entity-level traversal.

What carries the argument

Holistic-View-Guided Graph Construction (HGCM) plus Graph-Guided Holistic Retrieval (GHRM): a cross-modal holistic view M that detects and resolves conflicting facts during offline build, bridges entities to a compact concept graph with Icon concept-to-chunk indices, then retrieves by concept anchors, constrained relation propagation, direct index lookup, and modality-aware evidence organization.

Load-bearing premise

The whole claim rests on offline LLM and vision-language extractors plus an LLM conflict resolver producing concept–chunk links that are actually cleaner and more faithful than local extraction—not just differently tuned retrieval knobs or a shared answer model.

What would settle it

Hold the answer model and embeddings fixed, disable or reverse conflict resolution (or swap in independently extracted entity graphs), and check whether concept-to-chunk indices still reduce conflicting evidence and whether accuracy and online latency gains over BookRAG-style and flat baselines remain on the same three datasets, especially on the cross-modal question subsets.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • GraphRAG for long multimodal documents should prioritize reliable global indexing over denser entity graphs.
  • Online retrieval can stay near flat-RAG latency if concept nodes map directly to chunks instead of expanding entity neighborhoods.
  • Grouping retrieved text, tables, and images by modality before generation measurably helps heterogeneous evidence integration.
  • Gains should be largest where answers need cross-modal, multi-region evidence rather than single localized snippets.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same holistic conflict pass could be reused as a document-consistency cleaner even outside RAG, e.g., for building cleaner multimodal knowledge bases.
  • If concept assignment is too coarse, multi-hop scientific questions may still need a controlled entity fallback; hybrid concept/entity indices are a natural next test.
  • Dataset-specific thresholds and top-k settings suggest automatic per-document retrieval calibration may matter as much as the graph design itself.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 8 minor

Summary. The paper proposes HVM-GraphRAG, a multimodal GraphRAG framework for complex-document QA with two main components: (i) offline holistic-view-guided graph construction, in which an incrementally maintained "Cross-Modal Holistic View" of accepted facts is used by an LLM detector/resolver to filter conflicting or redundant extractions before building a concept-level indexing graph with direct concept-to-chunk evidence links; and (ii) online graph-guided retrieval over the compact concept graph (anchor selection + similarity-based one-hop propagation + direct index access), followed by modality-aware regrouping of evidence for the answer model. Experiments on MMLongBench, M3DocVQA, and Qasper against eleven baselines in four families report the best results on most metrics (e.g., MMLongBench EM 54.1 vs. BookRAG 43.8; Qasper Acc 66.1 vs. 55.2), online query time and token cost comparable to flat RAG, consistent ablation degradations, and improved performance when its graph constructor is swapped into other retrievers.

Significance. If the results hold, this is a useful empirical contribution to multimodal GraphRAG for complex-document QA. The paper is unusually complete on the empirical side for this venue type: three benchmarks reused without modification from prior processed splits, four baseline families re-run under a shared embedder (Qwen3-Embedding-0.6B) and shared generator (Qwen3-8B-AWQ) — a fairness step many GraphRAG papers omit — deterministic decoding, ablations of all three components, per-dataset hyperparameter sweeps, efficiency measurements, a constructor-swap generalization study (Table 5) showing the construction module transfers to other retrievers, and full prompt templates in Appendix D, which makes the pipeline reproducible in principle. The concept-level indexing design that trades offline construction effort for traversal-free online retrieval is a clean, falsifiable engineering idea, and the efficiency claim is scoped honestly to online cost. The gains over BookRAG on MMLongBench and Qasper are large enough to matter if they survive a clean tuning protocol.

major comments (4)
  1. [§4.5, Table 4, Figs. 6-7] Hyperparameters are tuned per dataset on what appears to be the evaluation set itself, and no development split is mentioned. Table 4 shows different θ, k_c, k_r, k_b per dataset, and Figs. 6/7/9/10 sweep these parameters directly on the benchmark questions (669/633/640 questions = the full test sets per Table 1), with the best cell apparently reported in Table 2. Meanwhile baselines are run at fixed defaults (k=10). This inflates the headline margins over BookRAG (e.g., MMLongBench EM 54.1 vs 43.8) by an unknown amount and is load-bearing for the RQ1 claim. Please either tune on a held-out dev portion and report test results at those settings, or report the performance of a single cross-dataset configuration alongside the tuned one.
  2. [§3.1.2, Eqs. (8)-(11)] The central mechanistic claim is that holistic-view conflict detection and resolution produce 'reliable evidence indexing' (Abstract; §3.1.2), but the manuscript provides no direct validation of the detector/resolver. Evidence is limited to end-task QA deltas (w/o CR ablation, Fig. 5) and a single qualitative case (Fig. 11). If the resolver systematically discards valid complementary facts or hallucinates merges, the indexing-reliability premise fails even when end-task scores improve for other reasons. A small human audit (e.g., 100-200 sampled conflict groups scored for detection precision/recall and resolution correctness) or a synthetic-conflict injection test would directly substantiate the claimed mechanism and is feasible given the prompts are already released in Appendix D.
  3. [§3.1.3, Eqs. (12)-(15); §2.1] The concept assignment function φ is never operationalized, and entity canonicalization across chunks is not described. Eq. (4) produces per-chunk concept/entity sets C_i, U_i, but the mapping from entity mentions to a global entity set V_ent and then to unique concepts via φ (Eqs. 12-15) is left unspecified. This matters mechanically: the conflict keys in Eqs. (6)-(7) (k_hr, k_ht, k_h) group facts by exact entity match, so without explicit entity resolution/normalization (e.g., 'Dusseldorf Stadtbahn' vs. 'Düsseldorf Stadbahn' — both spellings appear in Fig. 11), conflict groups and concept preimages φ^{-1}(c) are ill-defined. Please specify how entity strings are canonicalized and how φ is computed (LLM-assigned? clustered? extraction output?), since both conflict detection and the concept-level index depend on it.
  4. [§4.3, Fig. 4] The efficiency claim (RQ2, Fig. 4/Fig. 8) explicitly excludes offline construction cost, yet HVM-GraphRAG's offline stage is plausibly the most expensive among all compared systems: per-chunk modality-specific LLM/VLM extraction (Eq. 4), LLM conflict detection over candidate groups (Eq. 8), and LLM resolution with supporting-chunk retrieval (Eq. 11) for every chunk. Baselines like BM25/Vanilla RAG have near-zero offline cost. The online-only claim is honestly scoped in §4.3, but 'substantially improving online retrieval efficiency' (Abstract) is one of the paper's two headline claims, and readers cannot assess the total cost of ownership without offline wall-clock time and token consumption per document/dataset. Please report offline construction cost for HVM-GraphRAG and the graph-based baselines.
minor comments (8)
  1. [§B.3.2, Eq. (28)] The Accuracy metric (Eq. 28) is substring inclusion of the gold answer in the raw response. For short gold answers this can produce false positives (e.g., 'no' ⊂ a response containing 'not'). Qasper's official metric is token-level F1; consider reporting official-protocol numbers or at least discussing the false-positive risk.
  2. [Table 1 / Table 3] Table 1 and Table 3 are identical (both 'Statistics of the datasets'). Remove the duplication.
  3. [§4.2, Table 2] The M3DocVQA F1 result (66.0) is below BookRAG (66.2). The §4.2 wording 'best result on five of the six reported metrics' is ambiguous — there are six metrics per dataset and eighteen total. Please state precisely which dataset-metric cells are not best.
  4. [§4.1.4, Table 5] All results are single deterministic runs (temperature 0). No variance or significance information is given for gains as small as 0.4-1.4 points in Table 5. A brief note on run-to-run stability (or acknowledgment that determinism bounds this analysis) would strengthen the constructor-swap conclusions.
  5. [§4.4, Fig. 5] Ablation terminology is inconsistent: Fig. 5's 'w/o CG' is described as removing the 'Concept-Guided Graph', while §3.1.3/§3.2 use 'concept-level indexing graph G_con'. Also clarify whether w/o CG removes the concept graph (retrieval falls back to entity-level or flat?) — the replacement configuration for each ablation is not specified.
  6. [§3.1.2, Eq. (4)] Notation: Eq. (4) outputs (C_i, U_i, S_i, F_i) but U_i (entities) and S_i (schemas) are never referenced again; C_i's role in building V_con via φ is also unclear. Either use these symbols or trim the tuple.
  7. [§B.4] No code or data release is mentioned. Given that reproducibility rests on many prompts and per-dataset settings, a repository release (or a statement of intent) is strongly encouraged.
  8. [Global] Template artifacts remain: 'Conference acronym 'XX, June 03-05, 2018, Woodstock, NY' in the footer; title grammar 'on Complex Document' should be 'over Complex Documents' or similar.

Circularity Check

0 steps flagged

No derivation-chain circularity: empirical GraphRAG system evaluated on external benchmarks, not a self-defining prediction.

full rationale

HVM-GraphRAG is a systems/IR paper whose load-bearing claims are empirical (Table 2 QA metrics; Fig. 4 efficiency; Fig. 5 ablations), not closed-form predictions derived from fitted constants. Offline construction (document tree, modality-specific extraction, LLM conflict detect/resolve, concept–chunk index via Eqs. 1–15) and online retrieval (concept anchors, constrained propagation, modality-aware organization, Eqs. 16–26) are algorithmic designs; success is measured on held-out questions from prior processed splits of MMLongBench, M3DocVQA, and Qasper, against external baselines under a shared embedder/generator. Dataset-specific retrieval knobs (Table 4: θ, kc, kb, etc.) are ordinary hyperparameter choices, not parameters fitted to a target quantity that is then relabeled a “prediction.” Ablations remove CG/CR/ME rather than baking the claim into the definition of the metric. Author-overlapping citations (e.g., GraphRAG surveys) are background, not uniqueness theorems that force the method. No self-definitional loop, fitted-input-as-prediction, or load-bearing self-citation chain is present.

Axiom & Free-Parameter Ledger

4 free parameters · 5 axioms · 2 invented entities

Load-bearing content is engineering assumptions about LLM extraction/conflict resolution quality, layout parsing fidelity, and a large set of retrieval knobs chosen per dataset—not physical constants or formal axioms. Invented machinery is the Cross-Modal Holistic View and the concept-level evidence index built by entity bridging.

free parameters (4)
  • concept Top-k kc = 10 (MMLongBench, M3DocVQA); 15 (Qasper)
    Number of concept anchors retained; set differently across datasets (10/10/15) and shown to move EM/Acc in Fig. 7/10.
  • similarity threshold θ = 0.25 (MMLongBench, Qasper); 0.50 (M3DocVQA)
    Filters concept anchors; joint heatmaps (Fig. 6/9) show dataset-dependent optima used in main results.
  • supplementary RAG top-k kb = 2 / 4 / 5
    Extra dense-retrieved chunks merged with graph evidence; tuned per dataset in Table 4.
  • relation Top-k kr, graph chunk K, modality caps k_text/k_table/k_image = kr=10 or 15; K=10; 5 text, 3 table, 2 image
    Control propagation breadth and final context mix; fixed in main runs (e.g., K=10; 5/3/2 modality caps) after design choices that affect reported accuracy/efficiency.
axioms (5)
  • domain assumption Layout parsing (MinerU) plus LLM tree building yields chunk hierarchy PathT that is reliable enough structural context for extraction.
    Invoked in §3.1.1 Eqs. 1–3; errors here would mis-scope entities before conflict resolution.
  • domain assumption Modality-specific LLM/VLM extractors produce facts whose conflicts are detectable/resolvable by another LLM call using supporting chunks.
    Core of §3.1.2 Eqs. 4–11 and Appendix D prompts; not independently validated against gold KGs.
  • ad hoc to paper Each entity maps to exactly one concept via ϕ, and concept aggregation Icon(c)=∪ Ifac(e) preserves answer-relevant evidence without harmful over-merge.
    §2.1 and §3.1.3 Eqs. 12–15; uniqueness of concept assignment is a modeling choice that can collapse distinct senses.
  • domain assumption Cosine similarity in a shared embedding space is an adequate relevance signal for concepts, relations, and chunk reps (including VLM image summaries).
    §3.2 Eqs. 16–21; standard RAG assumption carried into graph propagation.
  • ad hoc to paper Grouping final evidence by modality and fixed order improves LLM integration of heterogeneous chunks versus interleaved context.
    §3.2.3 Eqs. 24–26; supported by ablation w/o ME but still a prompting/organization hypothesis.
invented entities (2)
  • Cross-Modal Holistic View M no independent evidence
    purpose: Running accepted-fact state used to detect and resolve conflicts while building the graph.
    Central constructed object in HGCM (§3.1.2); defined operationally via LLM resolver outputs, not an externally measured scientific entity.
  • Concept-level indexing graph G_con with evidence index I_con no independent evidence
    purpose: Compact retrieval substrate linking concepts directly to multimodal chunks without entity-level traversal.
    §3.1.3–3.2; standard graph abstraction specialized here via entity bridging after conflict resolution.

pith-pipeline@v1.2.0-grok45-kimik3 · 40129 in / 3877 out tokens · 84180 ms · 2026-07-30T17:06:08.694758+00:00 · methodology

0 comments
read the original abstract

Question answering (QA) over complex documents requires models to retrieve and integrate evidence distributed across distant document regions and modalities. Multimodal GraphRAG provides a promising direction by organizing document evidence with graph structures. However, existing methods often suffer from unreliable cross-modal evidence indexing and expensive graph traversal. To address these issues, we propose HVM-GraphRAG, a holistic-view multimodal GraphRAG framework on complex document. HVM-GraphRAG uses a holistic view to guide graph construction, thereby reducing noisy and conflicting graph updates and building reliable indices between concept-level graph nodes and supporting multimodal chunks. During retrieval, HVM-GraphRAG searches over a compact concept-level graph and directly accesses supporting evidence through the constructed index, avoiding costly traversal over dense entity-level graphs. After obtaining the retrieved evidence, HVM-GraphRAG further reorganizes chunks into modality-specific groups, enabling the answering model to better integrate heterogeneous evidence. Experiments on three datasets show that HVM-GraphRAG achieves the best answer performance in most evaluated settings while substantially improving online retrieval efficiency over representative graph-based baselines.

Figures

Figures reproduced from arXiv: 2607.24861 by Qinggang Zhang, Qing Li, Wenqi Fan, Xin He, Xin Wang, Yi Chang, Yili Wang.

Figure 1
Figure 1. Figure 1: Flat Multimodal RAG and Multimodal GraphRAG. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Comparison of answer accuracy and total online [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: The offline graph construction and online retrieval of HVM-GraphRAG. (1) A multimodal document tree provides [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Comparison of query efficiency. MMLongBench M3DocVQA Qasper 50 60 70 Performance (%) 54.1 63.0 66.1 49.0 57.1 59.1 50.7 59.9 64.4 50.8 60.3 63.4 Ours w/o CG w/o CR w/o ME [PITH_FULL_IMAGE:figures/full_fig_p007_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Ablation results on MMLongBench/M3DocVQA [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Effects of the similarity threshold 𝜃 and the num￾ber of supplementary chunks 𝑘𝑏 on MMLongBench and M3DocVQA. Darker colors indicate higher performance. 5 10 15 20 25 30 35 40 45 50 Concept Top-k kc 0.460 0.480 0.500 0.520 0.540 0.560 Exact Match MMLongBench 5 10 15 20 25 30 35 40 45 50 Concept Top-k kc 0.60 0.62 0.64 0.66 0.68 Accuracy Qasper [PITH_FULL_IMAGE:figures/full_fig_p008_6.png] view at source ↗
Figure 7
Figure 7. Figure 7: Effects of the number of retrieved concepts [PITH_FULL_IMAGE:figures/full_fig_p008_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: reports the total inference time and token consumption of all methods on MMLongBench, Qasper, and M3DocVQA. The re￾sults provide a more complete view of the efficiency differences among conventional RAG, graph-based RAG, layout-segmented RAG, multimodal GraphRAG, and HVM-GraphRAG. Graph-based methods introduce additional online retrieval overhead. Compared with conventional RAG methods such as BM25, Vanill… view at source ↗
Figure 9
Figure 9. Figure 9: Effects of the similarity threshold 𝜃 and the number of supplementary chunks 𝑘𝑏 on MMLongBench, M3DocVQA and Qasper. Darker colors indicate higher performance. 5 10 15 20 25 30 35 40 45 50 Concept Top-k kc 0.460 0.480 0.500 0.520 0.540 0.560 Exact Match MMLongBench 5 10 15 20 25 30 35 40 45 50 Concept Top-k kc 0.620 0.622 0.624 0.626 0.628 0.630 Exact Match M3DocVQA 5 10 15 20 25 30 35 40 45 50 Concept Top… view at source ↗
Figure 10
Figure 10. Figure 10: Effects of the number of retrieved concepts [PITH_FULL_IMAGE:figures/full_fig_p017_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Case Study of Conflict Resolution. Effects of 𝑘𝑐 [PITH_FULL_IMAGE:figures/full_fig_p018_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Case Study of Multimodal Evidence Integration. [PITH_FULL_IMAGE:figures/full_fig_p019_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: Prompt for extracting typed entities from textual chunks. [PITH_FULL_IMAGE:figures/full_fig_p021_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Prompt for extracting typed entities from table rows using table descriptions and column headers as context. [PITH_FULL_IMAGE:figures/full_fig_p022_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Prompt for extracting typed entities from images and accompanying textual descriptions. [PITH_FULL_IMAGE:figures/full_fig_p023_15.png] view at source ↗
Figure 16
Figure 16. Figure 16: Prompt for extracting high-confidence relations between given entities from textual chunks. [PITH_FULL_IMAGE:figures/full_fig_p024_16.png] view at source ↗
Figure 17
Figure 17. Figure 17: Prompt for detecting conflicts among candidate groups of extracted triples. [PITH_FULL_IMAGE:figures/full_fig_p025_17.png] view at source ↗
Figure 18
Figure 18. Figure 18: Prompt for resolving conflicting triples using their source passages. [PITH_FULL_IMAGE:figures/full_fig_p026_18.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

106 extracted references · 20 linked inside Pith

  1. [1]

    Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al . 2025. Qwen3-vl technical report.arXiv preprint arXiv:2511.21631(2025)

  2. [2]

    Ryan C Barron, Maksim E Eren, Olga M Serafimova, Cynthia Matuszek, and Boian S Alexandrov. 2025. Bridging legal knowledge and AI: retrieval-augmented generation with vector stores, Knowledge graphs, and hierarchical non-negative matrix factorization.arXiv preprint arXiv:2502.20364(2025)

  3. [3]

    Chenyang Bu, Guojie Chang, Zihao Chen, CunYuan Dang, Zhize Wu, Yi He, and Xindong Wu. 2025. Query-driven multimodal GraphRAG: Dynamic local knowledge graph construction for online reasoning. InFindings of the Association for Computational Linguistics: ACL 2025. 21360–21380

  4. [4]

    Wenhu Chen, Hexiang Hu, Xi Chen, Pat Verga, and William Cohen. 2022. Murag: Multimodal retrieval-augmented generator for open question answering over images and text. InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 5558–5570

  5. [5]

    Zerui Chen, Qinggang Zhang, Zhishang Xiang, Zhimin Wei, Linfeng Gao, Xiao Huang, Zhihong Zhang, and Jinsong Su. 2026. LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning. InProceed- ings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 37455–37484

  6. [6]

    Jaemin Cho, Debanjan Mahata, Ozan Irsoy, Yujie He, and Mohit Bansal. 2024. M3docrag: Multi-modal retrieval is what you need for multi-page multi-document understanding.arXiv preprint arXiv:2411.04952(2024)

  7. [7]

    Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira, Chaewoon Kim, Minjae Kim, Juneha Hwang, Jaeseon Ha, Hojun Choi, Suyeol Yun, Yongjin Kim, et al

  8. [8]

    Sijun Dai, Qiang Huang, Xiaoxing You, and Jun Yu. 2026. MG2-RAG: Multi- Granularity Graph for Multimodal Retrieval-Augmented Generation.arXiv preprint arXiv:2604.04969(2026)

  9. [9]

    Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner. 2021. A dataset of information-seeking questions and answers anchored in research papers. InProceedings of the 2021 Conference of the North Ameri- can Chapter of the Association for Computational Linguistics: Human Language Technologies. 4599–4610

  10. [10]

    Chao Deng, Jiale Yuan, Pi Bu, Peijie Wang, Zhong-Zhi Li, Jian Xu, Xiao-Hui Li, Yuan Gao, Jun Song, Bo Zheng, et al . 2025. Longdocurl: a comprehensive multimodal long document benchmark integrating understanding, reasoning, and locating. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1135–1159

  11. [11]

    Darren Edge, Ha Trinh, Joshua Bradley Newman Cheng, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. [n. d.]. From local to global: A graph rag approach to query-focused summarization, 2025.URL https://arxiv. org/abs/2404.16130([n. d.])

  12. [12]

    Xiang Fang, Wanlong Fang, and Changshuo Wang. 2026. Cogniverse: Revolu- tionizing multi-modal retrieval-augmented generation with cognitive reflection and geometric reasoning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7923–7935

  13. [13]

    Bernal J Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. Hipporag: Neurobiologically inspired long-term memory for large language models.Advances in neural information processing systems37 (2024), 59532– 59569

  14. [14]

    Giwon Hong, Jeonghwan Kim, Junmo Kang, Sung-Hyon Myaeng, and Joyce Jiy- oung Whang. 2024. Why so gullible? enhancing the robustness of retrieval- augmented models against counterfactual noise. InFindings of the Association for Computational Linguistics: NAACL 2024. 2474–2495

  15. [15]

    Chi-Hsiang Hsiao, Yi-Cheng Wang, Tzung-Sheng Lin, Yi-Ren Yeh, and Chu-Song Chen. 2026. MegaRAG: Multimodal Knowledge Graph-Based Retrieval Aug- mented Generation. InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 48031–48059

  16. [16]

    Yikuan Hu, Jifeng Zhu, Lanrui Tang, and Chen Huang. 2026. ReMindRAG: Low- Cost LLM-Guided Knowledge Graph Traversal for Efficient RAG.Advances in Neural Information Processing Systems38 (2026), 53757–53798

  17. [17]

    Haoyu Huang, Chong Chen, Zeang Sheng, Yang Li, and Wentao Zhang. 2025. Can LLMs be Good Graph Judge for Knowledge Graph Construction?. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 10940–10959

  18. [18]

    Gautier Izacard and Edouard Grave. 2021. Leveraging passage retrieval with generative models for open domain question answering. InProceedings of the 16th conference of the european chapter of the association for computational linguistics: main volume. 874–880

  19. [19]

    Cheng Jiayang, Chunkit Chan, Qianqian Zhuang, Lin Qiu, Tianhang Zhang, Tengxiao Liu, Yangqiu Song, Yue Zhang, Pengfei Liu, and Zheng Zhang. 2024. ECON: On the detection and resolution of evidence conflicts. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 7816–7844

  20. [20]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense passage retrieval for open- domain question answering. InProceedings of the 2020 conference on empirical methods in natural language processing (EMNLP). 6769–6781

  21. [21]

    Jungyeon Lee, Kangmin Lee, and Taeuk Kim. 2025. MAGIC: A Multi-Hop and Graph-Based Benchmark for Inter-Context Conflicts in Retrieval-Augmented Generation.arXiv preprint arXiv:2507.21544(2025)

  22. [22]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems33 (2020), 9459–9474

  23. [23]

    Huaxia Li, Haoyun Gao, Chengzhang Wu, and Miklos A Vasarhelyi. 2025. Ex- tracting financial data from unstructured sources: Leveraging large language models.Journal of Information Systems39, 1 (2025), 135–156

  24. [24]

    Longkun Li, Yuanben Zou, Jinghan Wu, Yuqing Wen, Jing Li, Hangwei Qian, and Ivor Tsang. 2026. SCOUT-RAG: Scalable and Cost-Efficient Unifying Traversal for Agentic Graph-RAG over Distributed Domains.arXiv preprint arXiv:2602.08400 (2026)

  25. [25]

    Pei Liu, Xin Liu, Ruoyu Yao, Junming Liu, Siyuan Meng, Ding Wang, and Jun Ma. 2025. Hm-rag: Hierarchical multi-agent multimodal retrieval augmented generation. InProceedings of the 33rd ACM international conference on multimedia. 2781–2790

  26. [26]

    Ziyu Liu, Yijing Liu, Jianfei Yuan, Minzhi Yan, Le Yue, Honghui Xiong, and Yi Yang. 2025. Graph-Guided Concept Selection for Efficient Retrieval-Augmented Generation.arXiv preprint arXiv:2510.24120(2025)

  27. [27]

    Haoran Luo, Guanting Chen, Yandan Zheng, Xiaobao Wu, Yikai Guo, Qika Lin, Yu Feng, Zemin Kuang, Meina Song, Yifan Zhu, et al. 2026. Hypergraphrag: Retrieval- augmented generation via hypergraph-structured knowledge representation. Advances in Neural Information Processing Systems38 (2026), 152206–152234

  28. [28]

    Yubo Ma, Yuhang Zang, Liangyu Chen, Meiqi Chen, Yizhu Jiao, Xinze Li, Xinyuan Lu, Ziyu Liu, Yan Ma, Xiaoyi Dong, et al. 2024. Mmlongbench-doc: Benchmarking long-context document understanding with visualizations.Advances in Neural Information Processing Systems37 (2024), 95963–96010

  29. [29]

    Aditya Nagori, Ricardo Accorsi Casonatto, Ayush Gautam, Abhinav Manikan- tha Sai Cheruvu, and Rishikesan Kamaleswaran. 2025. Open-Source Agen- tic Hybrid RAG Framework for Scientific Literature Review.arXiv preprint arXiv:2508.05660(2025)

  30. [30]

    Miaohe Niu, Lianlei Shan, Zhengtao Yu, Jingbo Zhu, and Tong Xiao. 2026. EfficientGraph-RAG: Structured Retrieval-State Management for Cross-Task Retrieval-Augmented Generation.arXiv preprint arXiv:2605.25379(2026)

  31. [31]

    Boci Peng, Yun Zhu, Yongchao Liu, Xiaohe Bo, Haizhou Shi, Chuntao Hong, Yan Zhang, and Siliang Tang. 2025. Graph retrieval-augmented generation: A survey. ACM Transactions on Information Systems44, 2 (2025), 1–52

  32. [32]

    Hongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao, Defu Lian, Zhicheng Dou, and Tiejun Huang. 2025. Memorag: Boosting long context processing with global memory-enhanced retrieval augmentation. InProceedings of the ACM on Web Conference 2025. 2366–2377. Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Xin He, Yili Wang, Wenqi Fan, Qing Li, Qinggan...

  33. [33]

    Derrick Quinn, Mohammad Nouri, Neel Patel, John Salihu, Alireza Salemi, Sukhan Lee, Hamed Zamani, and Mohammad Alian. 2025. Accelerating retrieval- augmented generation. InProceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume

  34. [34]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP). 3982–3992

  35. [35]

    Monica Riedler and Stefan Langer. 2024. Beyond text: Optimizing rag with multimodal inputs for industrial applications.arXiv preprint arXiv:2410.21943 (2024)

  36. [36]

    2009.The probabilistic relevance frame- work: BM25 and beyond

    Stephen Robertson and Hugo Zaragoza. 2009.The probabilistic relevance frame- work: BM25 and beyond. Vol. 4. Now Publishers Inc

  37. [37]

    Stephen E Robertson and Steve Walker. 1994. Some simple effective approxi- mations to the 2-poisson model for probabilistic weighted retrieval. InSIGIR’94: Proceedings of the Seventeenth Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval, organised by Dublin City University. Springer, 232–241

  38. [38]

    Albert Sadowski and Jaroslaw A Chudziak. 2025. On verifiable legal reasoning: A multi-agent framework with formalized knowledge representations. InPro- ceedings of the 34th ACM International Conference on Information and Knowledge Management. 2535–2545

  39. [39]

    Bhaskarjit Sarmah, Dhagash Mehta, Benika Hall, Rohan Rao, Sunil Patel, and Stefano Pasquali. 2024. Hybridrag: Integrating knowledge graphs and vector retrieval augmented generation for efficient information extraction. InProceedings of the 5th ACM International Conference on AI in Finance. 608–616

  40. [40]

    Parth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna, Anna Goldie, and Christopher Manning. 2024. Raptor: Recursive abstractive processing for tree- organized retrieval. InInternational Conference on Learning Representations, Vol. 2024. 32628–32649

  41. [41]

    Henri Scaffidi, Melinda Hodkiewicz, Caitlin Woods, and Nicole Roocke. 2025. GraphRAG on Technical Documents-Impact of Knowledge Graph Schema.Trans- actions on Graph Data and Knowledge3, 2 (2025), 3–1

  42. [42]

    Shreya Shankar, Tristan Chambers, Tarak Shah, Aditya G Parameswaran, and Eugene Wu. 2024. Docetl: Agentic query rewriting and evaluation for complex document processing.arXiv preprint arXiv:2410.12189(2024)

  43. [43]

    Joongmin Shin, Chanjun Park, Jeongbae Park, Jaehyung Seo, and Heui-Seok Lim. 2025. MultiDocFusion: Hierarchical and Multimodal Chunking Pipeline for Enhanced RAG on Long Industrial Documents. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 20996–21015

  44. [44]

    Joongmin Shin, Gyuho Shim, Jeongbae Park, Jaehyung Seo, and Heui-Seok Lim

  45. [45]

    Jong Keon Song, Dong Bin Youk, Hyery Kim, and Sang-Hyun Hwang. 2026. Multimodal Knowledge Graph–Guided RAG-LLM for Clinical Decision Support in Pediatric Leukemia.Cancer Research and Treatment(2026)

  46. [46]

    Manan Suri, Puneet Mathur, Franck Dernoncourt, Kanika Goswami, Ryan A Rossi, and Dinesh Manocha. 2025. Visdom: Multi-document qa with visually rich elements using multimodal retrieval-augmented generation. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologi...

  47. [47]

    Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. 2026. Ai-researcher: Autonomous scientific innovation.Advances in Neural Information Processing Systems38 (2026), 9481–9520

  48. [48]

    Xueyao Wan and Hang Yu. 2025. Mmgraphrag: Bridging vision and language with interpretable multimodal knowledge graphs.arXiv preprint arXiv:2507.20804 (2025)

  49. [49]

    Bin Wang, Chao Xu, Xiaomeng Zhao, Linke Ouyang, Fan Wu, Zhiyuan Zhao, Rui Xu, Kaiwen Liu, Yuan Qu, Fukai Shang, et al. 2024. Mineru: An open-source solution for precise document content extraction.arXiv preprint arXiv:2409.18839 (2024)

  50. [50]

    Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. 2025. Retrieval- augmented generation with conflicting evidence.arXiv preprint arXiv:2504.13079 (2025)

  51. [51]

    Jie Wang, Honghua Huang, Xi Ge, Jianhui Su, Wen Liu, and Shiguo Lian

  52. [52]

    Qiuchen Wang, Ruixue Ding, Zehui Chen, Weiqi Wu, Shihang Wang, Pengjun Xie, and Feng Zhao. 2025. Vidorag: Visual document retrieval-augmented generation via dynamic iterative reasoning agents. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 9124–9145

  53. [53]

    Shu Wang, Yingli Zhou, and Yixiang Fang. 2025. BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents.arXiv preprint arXiv:2512.03413(2025)

  54. [54]

    OMD-GraphRAG: Enhancing GraphRAG with Ontology-Guided Extrac- tion, Multi-Dimensional Clustering and Dual-Channel Fusion.arXiv preprint arXiv:2603.25152(2026)

  55. [55]

    Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng- ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, et al . 2025. Qwen-image technical report.arXiv preprint arXiv:2508.02324(2025)

  56. [56]

    Chuanjie Wu, Zhishang Xiang, Yunbo Tang, Zerui Chen, Qinggang Zhang, and Jinsong Su. 2026. MemGraphRAG: Memory-based Multi-Agent System for Graph Retrieval-Augmented Generation.arXiv preprint arXiv:2606.00610(2026)

  57. [57]

    Xihang Wang, Zihan Wang, Chengkai Huang, Quan Z Sheng, and Lina Yao. 2026. MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG.arXiv preprint arXiv:2604.24564(2026)

  58. [58]

    Peng Xia, Kangyu Zhu, Haoran Li, Tianze Wang, Weijia Shi, Sheng Wang, Linjun Zhang, James Y Zou, and Huaxiu Yao. 2025. Mmed-rag: Versatile multimodal rag system for medical vision language models. InInternational Conference on Learning Representations, Vol. 2025. 66188–66217

  59. [59]

    An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report.arXiv preprint arXiv:2505.09388(2025)

  60. [60]

    Xixi Wu, Yanchao Tan, Nan Hou, Ruiyang Zhang, and Hong Cheng. 2025. Molorag: Bootstrapping document understanding via multi-modal logic-aware retrieval. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 14035–14056

  61. [61]

    Chuanyue Yu, Kuo Zhao, Yuhan Li, Heng Chang, Mingjian Feng, Xiangzhe Jiang, Yufei Sun, Jia Li, Yuzhi Zhang, Qingyun Sun, et al . 2026. Graphrag-r1: Graph retrieval-augmented generation with process-constrained reinforcement learning. InProceedings of the ACM Web Conference 2026. 1398–1409

  62. [62]

    Junchi Yu, Yujie Liu, Jindong Gu, Philip Torr, and Dongzhan Zhou. 2026. Can Knowledge-Graph-based Retrieval Augmented Generation Really Retrieve What You Need?Advances in Neural Information Processing Systems38 (2026), 95653– 95682

  63. [63]

    R Yang, B Yang, A Feng, S Ouyang, M Blum, T She, Y Jiang, F Lecue, J Lu, and I Li. 2025. Graphusion: A RAG Framework for Knowledge Graph Construction with a Global Perspective (No. arXiv: 2410.17600). arXiv

  64. [64]

    Xiaohan Yu, Pu Jian, and Chong Chen. 2025. Tablerag: A retrieval augmented generation framework for heterogeneous document reasoning. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 14074–14093

  65. [65]

    Xu Yuan, Liangbo Ning, Wenqi Fan, and Qing Li. 2025. mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering.arXiv preprint arXiv:2508.05318(2025)

  66. [66]

    Wenwen Yu, Zhibo Yang, Yuliang Liu, and Xiang Bai. 2025. Docthinker: Explain- able multimodal large language models with rule-based reinforcement learning for document understanding. InProceedings of the IEEE/CVF International Confer- ence on Computer Vision. 837–847

  67. [67]

    Haoyun Zhang, Shihao Zhao, Zijun Zhou, Wenchao Zhang, and Yizhou Meng

  68. [68]

    Kepu Zhang, Weijie Yu, Zhongxiang Sun, and Jun Xu. 2025. Syler: A framework for explicit syllogistic legal reasoning in large language models. InProceedings of the 34th ACM International Conference on Information and Knowledge Management. 4117–4127

  69. [69]

    Haozhen Zhang, Tao Feng, and Jiaxuan You. 2025. Graph of records: Boosting retrieval augmented generation for long-context summarization with graphs. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 23780–23799

  70. [70]

    Qinggang Zhang, Shengyuan Chen, Yuanchen Bei, Zheng Yuan, Huachi Zhou, Zijin Hong, Hao Chen, Yilin Xiao, Chuang Zhou, Junnan Dong, et al . 2025. A survey of graph retrieval-augmented generation for customized large language models.arXiv preprint arXiv:2501.13958(2025)

  71. [71]

    In2025 7th International Conference on Internet of Things, Automation and Artificial Intelligence (IoTAAI)

    Domain-Specific RAG with Semantic Normalization and Contrastive Feed- back for Document Question Answering. In2025 7th International Conference on Internet of Things, Automation and Artificial Intelligence (IoTAAI). IEEE, 750–753

  72. [72]

    Yilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, and Arman Cohan. 2024. DocMath-eval: Evaluating math reasoning capabilities of LLMs in understanding long and spe- cialized documents. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers...

  73. [73]

    Mingtian Zhang, Yu Tang, and PageIndex Team. 2025. PageIndex: Next- Generation Vectorless, Reasoning-based RAG. PageIndex Blog (September 2025)

  74. [74]

    The answer is C

    Fengbin Zhu, Wenqiang Lei, Fuli Feng, Chao Wang, Haozhou Zhang, and Tat-Seng Chua. 2022. Towards complex document understanding by discrete reasoning. In Proceedings of the 30th ACM International Conference on Multimedia. 4857–4866. HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document Conference acronym ’XX, June...

  75. [75]

    Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, et al. 2025. Qwen3 embedding: Advancing text embedding and reranking through foundation models.arXiv preprint arXiv:2506.05176(2025)

  76. [77]

    Jiasen Zheng, Yumin Chen, Zijun Zhou, Chong Peng, Haozhang Deng, and Shihan Yin. 2025. Information-Constrained Retrieval for Scientific Literature via Large Language Model Agents. In2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering (ICBAIE). IEEE, 366–370

  77. [79]

    Retrieval Process

    Retrieval Query4. Retrieval Process

  78. [80]

    Was Essen Stadtbahn the system when Dusseldorf, Neuss was the city on routein the table of Component Systems and lines of Rhine-Ruhr Stadtbahn?

    Final Answer Traditional GraphRAGChunkAlists multiple component systems, includingDüsseldorf StadtbahnandEssen Stadtbahn.ChunkBfurther placesDüsseldorf Stadtbahn and Essen: Public Transport (incl. Essen Stadtbahn)in the same route-oriented neighborhood. Isolatedextraction:1(Essen Stadtbahn, part of, Rhine-Ruhr Stadtbahn)2(Düsseldorf Stadbahn, part of, Rhi...

  79. [81]

    Retrieval Process

    Retrieval Query3. Retrieval Process

  80. [82]

    Which player, in the scoring leaders of the 2000-01 AHL season, plays for the team with a pouncing wild cat on the jersey?

    Final Answer Traditional GraphRAGOur proposed Method Q: “Which player, in the scoring leaders of the 2000-01 AHL season, plays for the team with a pouncing wild cat on the jersey?” Uncertain/Incomplete Answer(Not grounded in the actual image or table content) Mikael Samuelsson(Correct & multimodally grounded)

Showing first 80 references.