Pith. sign in

Paper Citation Record · LEDGER

M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2411.04952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04952 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:49:54.263333Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:09:48.864251Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 77fc6296-a578-4773-9b75-fd9dc7431252 · inbound

Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks cites this paper.

Document Screenshot Retrievers are Vulnerable to Pixel Poisoning Attacks M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T05:49:54.263333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:49:54.263333Z digest=sha256:1a60ee66af9d8b8cc3a2f6e5a2b7ecd35df5bdbb5803300aabe282fc52403c95

Observation d32f2860-0616-4982-b3a8-d69e8cd7924b · inbound

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning cites this paper.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.927205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.927205Z digest=sha256:e39b2088fcd62d9c6bf0abcf16341f547a1cbe44eb52f7eda4faaed1642795b4

Observation b5544092-a869-4a6b-9b09-20613d430e27 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.298867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.298867Z digest=sha256:d13f5b0e8de8cf04b1bf3aa6cdf1ff3b38702c86ab7fe5e13af3ffff5dfff6bb

Observation 6f44c4a6-e6e8-44ea-b874-4bc444b6e429 · inbound

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers cites this paper.

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:50.994221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:50.994221Z digest=sha256:639429b00b55769bc18ecb0aeb9f15caf0bd4efcb3c7c6d498d464b3f8fb60d0

Observation 9cac0801-9dc0-44f0-8672-3fb25f9d59a4 · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.341207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.341207Z digest=sha256:2f9178739e5359c01d0de219484381fc1610e5c4d053abcefdaf2fcb0a1011cf

Observation a3903e34-0903-4e2f-93c1-04ecf26cc0d3 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:56:59.344346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:56:59.344346Z digest=sha256:a831f2fca2dd94a249c9307b00d0b1c8e6dde4d2fd3f73bd3ef6c2e61502e1a9

Observation c37fdab9-16b7-4324-ad2a-2c8dea816da8 · inbound

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs cites this paper.

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:20:35.791056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:20:35.791056Z digest=sha256:17ce461988a10f745f439058b16c4263c6e9e3102a24d6c7a8cf4a5347149723

Observation 87a50d81-abf0-42b3-a340-a0dc3d45a918 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:20.389650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:20.389650Z digest=sha256:959fd0d7471eda85613455fdd06a7076c198d79465dd4cf238fc8823a70be575

Observation 67e871a4-a0ed-4425-a552-41123f71c858 · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:28.060988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:28.060988Z digest=sha256:6825d5396b4798059df780ac583433a7d820673525c170260c37a1e944fb0027

Observation 5e9a7a58-70a6-4b00-b6a9-7c67a6af3e3c · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.309480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:e1dd5bb0d1884d0d7df5cae5a46a1c67cede6c60654ef182b9abcb6477e3c271

Observation 0dd7a4a8-5a2d-4165-94c7-d3bfe8dd94d2 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:45.508464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T10:46:17.411843Z digest=sha256:eb3e386d2074bf80d54a193320f29eb4505bbb53d4799b3b437a8ce2ddfd56d0

Observation 7091abbc-640a-41cf-a1f8-43fb994effe0 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:52.247742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:52.247742Z digest=sha256:e6fc13cbc926074ef408e4261ccaa2ff48072c2a2739d08be425cb00de72ca89

Observation ece0a592-947f-407e-aae9-53b47dab02d0 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.562485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:63167b8633fd91fa0887fd2cae5d158f2897f932df954b975e8335fd9fa1b758

Observation 5f01a93a-af31-4ff8-b618-a71912722062 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.561329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:3482625dc71d01fdd3f72aac5cac2bfb10c59801387e4b7732dec0d69c3d1ca4

Observation bb0b10d7-92e9-44ae-9928-9b6a349a1c2a · inbound

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework cites this paper.

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:28:13.938920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T20:27:57.033386Z digest=sha256:aa93af79d4cf031969d7b0ad2eb7b2393dc4f8be4754e57283c14fe75c340b02

Observation f84d7f04-5f93-48da-9d31-11c62b0d9b6a · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.210643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T09:29:32.250418Z digest=sha256:243faf9acd395e08e36595a2e7b2aae333229b27dba0782144afc7dae2264fe1

Observation 633e4897-0f0b-48eb-8a3a-729dd8243ac2 · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T06:07:40.071839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:07:40.071839Z digest=sha256:c5f943f00136ac40cb11ed983f8eeafe9b7f5520c280cb48f916d72c17c39250

Observation 36094f8a-e55c-40ca-8676-4de8322e971d · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.441239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:41918cba1ff532a58ac619d687c12fd1c074e0ca1a5e8507ee0a5a4cc876bde1

Observation 8773bdd0-5d75-45d0-bd6a-d352ee7ac2d4 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.537543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:bcd1c7e11d1bf123460a047185954c24e39b4010c47fd2c3488866b89e1e5146

Observation 1bbcf13c-0eae-4fc1-bb61-f7abe1ed8b08 · inbound

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs cites this paper.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.539181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:c568cd5d27941bd124df89c8f8228f2738e33603a64da25f05e100c5be4e0ad2

Observation b781defd-d22b-4f41-83f0-4d1187ad4eea · inbound

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning cites this paper.

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:09.400117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T12:42:39.231468Z digest=sha256:2b789be6431cc0ca957a8063ee082d83da3b824f624b6b4a144383a5e1783c0d

Observation 87c71117-c5f2-4adc-bcb2-f8428a09fa6a · inbound

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text cites this paper.

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.833254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:09:01.612988Z digest=sha256:e3c22142f8bf969b9eb45a979c0d2ce42d41b41ce51f561855b884ac3a3795e4

Observation c0c50346-4c21-4a50-b4b1-da2cec840893 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:37:58.177338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:7e11d4a3576383d9184ec658001cb75993a668caa69b4d93e4cc38f61a81598b

Observation e4023e49-b41e-4fcd-87cc-3d352f47e59c · inbound

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models cites this paper.

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.182184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:00:25.664841Z digest=sha256:ccca7cd1577db5498b82f391d6fed3a4b9f57a7854a085a406f6e6b797be6fd1

Observation d8dc2758-9de4-4ab9-8eea-c2ddd260589f · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.430081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:cfb48e48a1bafff4f15b3cd50d6cbe89efe85f9e1c5a08c331e9f682b1a0ce1e

Observation 3c1d3445-90a0-424e-9a15-34ecf41e108d · inbound

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1) cites this paper.

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1) M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:28.882201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:31:29.694131Z digest=sha256:70b753f3f4f93f09ce94a63811aa637a01c413889b1e1a5f54619f2da78c9e26

Observation 2616c404-a4d8-472f-92e1-5c410d18b5ec · inbound

Constrained Dominant Sets for Multimodal Document Question Answering cites this paper.

Constrained Dominant Sets for Multimodal Document Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:17:21.737164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T20:45:12.783718Z digest=sha256:c1b0081a3d881eafc83b7720e24c35b2824bf1d48e1fbbd675d03f9777b9bd1f

Observation b81377d0-739d-4711-9e80-6fcf20b96256 · inbound

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval cites this paper.

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:37:35.693848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T15:04:30.757297Z digest=sha256:bd40ffed922906ec8b0fd5ca15e84c8340935681808c1a9fff5be614f596aab1

Observation dba1760e-c9db-4743-af65-84f72f89ff38 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.504259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:b81de68171539c20cf56cec14fee5f642682219cc6dd1670a4ddec0cfd66631d

Observation 247aa86b-3d11-4ee9-ab1f-41c121830ab3 · inbound

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries cites this paper.

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.643693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T15:49:49.067128Z digest=sha256:3f1fbd0d6fe4bcea3b255febf5fe5dafcab99c0b32d168105994d02f2796e1b2

Observation 7fcdb0b4-12f7-4047-b3b4-50dbcb2b6dfd · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.865734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:baec13d455a571f3b15438ad9504f252c4e3ca66d0a21c4b7d1448121fda4d05

Observation f080281c-aff4-4c07-a7d9-95665d1e0f72 · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.580737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:35cf74a0b9be0078a437469c8309107ae8ff9befb48ba13005587610296741d4

Observation b4918ab3-ca04-4dc0-b40f-06ee414649de · inbound

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation cites this paper.

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:37.698496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T11:31:56.598221Z digest=sha256:08893841878a519ae28484f81c3bf15d71cc5296091c37871a4656f686939257

Observation c66f04c0-5fd5-4ec0-8cb9-0e742f609345 · inbound

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents cites this paper.

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:22.448947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T07:15:01.669896Z digest=sha256:d908719019762af92cdc7452ec9fba66a7d95b5302293c1159ec152bd20336a3

Observation c98a98fd-ce78-4469-9a8b-04c307e045cb · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.383295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.383295Z digest=sha256:b3608c4dd18bd0774653280af75205f24d07ca87baccd4c8e06a88bb8dcca540

Observation 2961fbf4-4c70-4e04-b49f-5794a096b1ba · inbound

When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering cites this paper.

When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:32:16.757013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:32:16.757013Z digest=sha256:bd266294b286f30b888818332b021003ef9c466baca6bc08b024294553395ee9

Observation 7c84fa5c-719b-4b6e-b92f-93ef095377da · inbound

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document cites this paper.

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T17:10:18.469354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T17:10:18.469354Z digest=sha256:8393db0a2274a469042683d56c0bd62790570499b3b86ff65aed202ae7a459e1

Observation 3cb8f18f-4f33-49d9-a1b6-0ed6457041ef · inbound

Evidence-Grounded Constraint Checking in Construction Documents cites this paper.

Evidence-Grounded Constraint Checking in Construction Documents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:28:21.243790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:28:21.243790Z digest=sha256:4c8324ddd557d1e2ff20bdbc002f05f6f75345024cece8831beb79834df7357e

Observation d772b5f9-1224-4359-b01f-e0895233a708 · inbound

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering cites this paper.

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:49.139003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:57:49.139003Z digest=sha256:8443e173c3cd4b0827893702bee8e74630111588415bf62e7f37d964f01fa27a

Observation 9a3ed520-b8da-4ed5-b808-563486dbab8f · inbound

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding cites this paper.

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:06.861120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:06.861120Z digest=sha256:b3cc6fdba06ba083ed306cbe388c3dd92be2c5e522a8741032136068c14ae319

Observation f8d507f5-ccbc-47d5-b24f-74c6a588ed79 · inbound

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval cites this paper.

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:09.444119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:09.444119Z digest=sha256:45883dbbdccf5e1adfb92d7ad061c20c55a91ce220c908a28430dde78973edd8

Observation 0f661bfb-1bc2-4751-84fb-d6d2dcbd2e5c · inbound

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning cites this paper.

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:42.455005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:42.455005Z digest=sha256:f5f0f390363cb210629e4a05d4b5d692f48e1af68c174d594f052a507d1b64e3