Pith. sign in

Paper Citation Record · LEDGER

M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2411.04952.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.04952 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:24:49.927205Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:09:48.864251Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d32f2860-0616-4982-b3a8-d69e8cd7924b · inbound

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning cites this paper.

VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:49.927205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:24:49.927205Z digest=sha256:e99c7a9816c04480eeb30895d3f5713dce3d3fcad9d46f90876bec96e0ec35f5

Observation b5544092-a869-4a6b-9b09-20613d430e27 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:48.298867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:48.298867Z digest=sha256:478d0dfe5ba63873b0e86127306fda52315699d41f98c8fbf702b79526df4b23

Observation 6f44c4a6-e6e8-44ea-b874-4bc444b6e429 · inbound

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers cites this paper.

The Next Phase of Scientific Fact-Checking: Advanced Evidence Retrieval from Complex Structured Academic Papers M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:50.994221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:45:50.994221Z digest=sha256:639429b00b55769bc18ecb0aeb9f15caf0bd4efcb3c7c6d498d464b3f8fb60d0

Observation 9cac0801-9dc0-44f0-8672-3fb25f9d59a4 · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.341207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.341207Z digest=sha256:e62c157206605f805907f24640056e8359b3be14e685c0560535ddc554e30a9b

Observation a3903e34-0903-4e2f-93c1-04ecf26cc0d3 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:56:59.344346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:56:59.344346Z digest=sha256:0d4f46b4ff9e52604ad3b6302a59db721448f706442410b8ce76dd27c6321c8e

Observation c37fdab9-16b7-4324-ad2a-2c8dea816da8 · inbound

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs cites this paper.

MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:20:35.791056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:20:35.791056Z digest=sha256:17ce461988a10f745f439058b16c4263c6e9e3102a24d6c7a8cf4a5347149723

Observation 87a50d81-abf0-42b3-a340-a0dc3d45a918 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:20.389650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:20.389650Z digest=sha256:959fd0d7471eda85613455fdd06a7076c198d79465dd4cf238fc8823a70be575

Observation 67e871a4-a0ed-4425-a552-41123f71c858 · inbound

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding cites this paper.

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T07:21:28.060988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:21:28.060988Z digest=sha256:6825d5396b4798059df780ac583433a7d820673525c170260c37a1e944fb0027

Observation 5e9a7a58-70a6-4b00-b6a9-7c67a6af3e3c · inbound

Attention Grounded Enhancement for Visual Document Retrieval cites this paper.

Attention Grounded Enhancement for Visual Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:55:15.309480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:53:15.920563Z digest=sha256:bf2d2391e12b5e8a5a881992225ebee9bcc14c3b4eb4443899777b8187a80b18

Observation 0dd7a4a8-5a2d-4165-94c7-d3bfe8dd94d2 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:47:45.508464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:46:17.411843Z digest=sha256:751a99db81a1b568f2be5d3959ab1ac87dcfb306ad0f1f0f3fe69294a92afb4a

Observation 7091abbc-640a-41cf-a1f8-43fb994effe0 · inbound

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation cites this paper.

R3G: A Reasoning-Retrieval-Reranking Framework for Vision-Centric Answer Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T08:13:52.247742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:13:52.247742Z digest=sha256:e6fc13cbc926074ef408e4261ccaa2ff48072c2a2739d08be425cb00de72ca89

Observation ece0a592-947f-407e-aae9-53b47dab02d0 · inbound

MoDora: Tree-Based Semi-Structured Document Analysis System cites this paper.

MoDora: Tree-Based Semi-Structured Document Analysis System M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:10:15.562485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T19:09:41.482713Z digest=sha256:26f0f75dcb6105c49883f6e7c071e08df62ab20f911730acb561c14d7f3cd303

Observation 5f01a93a-af31-4ff8-b618-a71912722062 · inbound

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training cites this paper.

Towards Real-World Document Parsing via Realistic Scene Synthesis and Document-Aware Training M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:18:26.561329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T01:15:26.757215Z digest=sha256:53d1a379d3aded326844ab57c265fa09c047cc014d223343f71bbf36f9bf5e72

Observation bb0b10d7-92e9-44ae-9928-9b6a349a1c2a · inbound

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework cites this paper.

Overcoming the "Impracticality" of RAG: Proposing a Real-World Benchmark and Multi-Dimensional Diagnostic Framework M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:28:13.938920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T20:27:57.033386Z digest=sha256:26cc6f389f2ee27e2930bb99a49636595f0d4a2b3ae2cee6d13b48fd001ac3b1

Observation f84d7f04-5f93-48da-9d31-11c62b0d9b6a · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:30:48.210643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:29:32.250418Z digest=sha256:77f1900c34144310db1d360a3efece5fb7e7f59772ba862120366fe0f133d86f

Observation 633e4897-0f0b-48eb-8a3a-729dd8243ac2 · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T06:07:40.071839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:07:40.071839Z digest=sha256:c5f943f00136ac40cb11ed983f8eeafe9b7f5520c280cb48f916d72c17c39250

Observation 36094f8a-e55c-40ca-8676-4de8322e971d · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.441239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:8ab7f7f0bd34328fc98be881da064da74d07764878282ff8a7f8d5b0d8f52653

Observation 8773bdd0-5d75-45d0-bd6a-d352ee7ac2d4 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.537543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:42299128bc37e2f59699c63c90386e54462dd1a0e74779c21a048d5366cdd12e

Observation 1bbcf13c-0eae-4fc1-bb61-f7abe1ed8b08 · inbound

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs cites this paper.

PaperMind: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:41:54.539181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T21:09:26.626179Z digest=sha256:d1654af44272a9552a8055dbcb768a5bb85192cea0678b1c75c2fcd1712e2194

Observation b781defd-d22b-4f41-83f0-4d1187ad4eea · inbound

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning cites this paper.

DocPrune:Efficient Document Question Answering via Background, Question, and Comprehension-aware Token Pruning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:09.400117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T12:42:39.231468Z digest=sha256:a4ac43c5f3c5cdbb244b3df58154ae46e04a5029a3241c23b4c2b9cf11ca6130

Observation 87c71117-c5f2-4adc-bcb2-f8428a09fa6a · inbound

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text cites this paper.

LensVLM: Selective Context Expansion for Compressed Visual Representation of Text M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:40:59.833254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:09:01.612988Z digest=sha256:cfb0f0a03d97aa119d17ea9ae5307a293e19af78eb7db002f511fc4e071c8ec4

Observation c0c50346-4c21-4a50-b4b1-da2cec840893 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:37:58.177338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:88e2ce0322e4f766c6387ddb45b6a37de4eb0fb50072b50cb9aece58e09fefdf

Observation e4023e49-b41e-4fcd-87cc-3d352f47e59c · inbound

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models cites this paper.

MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.182184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T21:00:25.664841Z digest=sha256:90da9f684f19f4a538b5704affeec5e098c70df7c8165a0b9a9d3f6a2890cf77

Observation d8dc2758-9de4-4ab9-8eea-c2ddd260589f · inbound

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval cites this paper.

Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:24:40.430081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:17:04.441743Z digest=sha256:640209c7b01a9e45ec2734f6ce4ec30d231a6e3dcab2af2a58ed24247341dac5

Observation 3c1d3445-90a0-424e-9a15-34ecf41e108d · inbound

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1) cites this paper.

Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1) M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:56:28.882201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T10:31:29.694131Z digest=sha256:26f2d432e6ce34af97ba5f09f3b2e2476db3a0afb8cad479878883b956d638c3

Observation 2616c404-a4d8-472f-92e1-5c410d18b5ec · inbound

Constrained Dominant Sets for Multimodal Document Question Answering cites this paper.

Constrained Dominant Sets for Multimodal Document Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:17:21.737164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T20:45:12.783718Z digest=sha256:181edde9a986f97777b5015a44a4addb0410b4dac36f8fcf506cf52d89e61155

Observation b81377d0-739d-4711-9e80-6fcf20b96256 · inbound

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval cites this paper.

EviProp: Seeded Relevance Diffusion on Chunk-Page Graphs for Long Multimodal Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:37:35.693848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T15:04:30.757297Z digest=sha256:45c6cd07bcdd61b492f68b63e83360e20731b4e1d1b57afe5e1f71dfdc8aa542

Observation dba1760e-c9db-4743-af65-84f72f89ff38 · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:57:41.504259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:5ecfddb9a90a39332c34b5b780ffc2f66db7eccf0012160316d313c6bbb9a2a4

Observation 247aa86b-3d11-4ee9-ab1f-41c121830ab3 · inbound

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries cites this paper.

Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:39:39.643693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T15:49:49.067128Z digest=sha256:b94b66b4d75da1397ff6f55f27637552ddf8797da19ebb657358f52a62bc0903

Observation 7fcdb0b4-12f7-4047-b3b4-50dbcb2b6dfd · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.865734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:da59e254a9144620c7ac14763075147050fad8675db3722822555cc09d3120b2

Observation f080281c-aff4-4c07-a7d9-95665d1e0f72 · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.580737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:962ae3781812b05bcebe835996ed6c29895323f0697469c83d4cd8c41a2aa92b

Observation b4918ab3-ca04-4dc0-b40f-06ee414649de · inbound

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation cites this paper.

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:37.698496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:31:56.598221Z digest=sha256:35e7437df6bf7acf510200bde860fb7c4f5ebfb450f66a937e87e39dc86dcce9

Observation c66f04c0-5fd5-4ec0-8cb9-0e742f609345 · inbound

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents cites this paper.

LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:24:22.448947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T07:15:01.669896Z digest=sha256:f08b79fce01e7776bd338571eb3dfc9801be4a108d7d067c876ba57804c44f9c

Observation c98a98fd-ce78-4469-9a8b-04c307e045cb · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 216

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.383295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.383295Z digest=sha256:b3608c4dd18bd0774653280af75205f24d07ca87baccd4c8e06a88bb8dcca540

Observation 2961fbf4-4c70-4e04-b49f-5794a096b1ba · inbound

When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering cites this paper.

When Do Multimodal and Graph-Augmented RAG Help? A Controlled Evaluation for Document Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T20:32:16.757013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:32:16.757013Z digest=sha256:bd266294b286f30b888818332b021003ef9c466baca6bc08b024294553395ee9

Observation 7c84fa5c-719b-4b6e-b92f-93ef095377da · inbound

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document cites this paper.

HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T17:10:18.469354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T17:10:18.469354Z digest=sha256:e02c55c4dc7f3f1e776ec5b93ddef82b80ea0a91f15fd970eb2fe14641f2ca01

Observation 3cb8f18f-4f33-49d9-a1b6-0ed6457041ef · inbound

Evidence-Grounded Constraint Checking in Construction Documents cites this paper.

Evidence-Grounded Constraint Checking in Construction Documents M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T14:28:21.243790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:28:21.243790Z digest=sha256:4c8324ddd557d1e2ff20bdbc002f05f6f75345024cece8831beb79834df7357e

Observation d772b5f9-1224-4359-b01f-e0895233a708 · inbound

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering cites this paper.

HierDoc: Hierarchical Page-to-Region Evidence Routing for Long-Document Visual Question Answering M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:49.139003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T02:57:49.139003Z digest=sha256:06d1feb28c21d17859561146449de5e860363092fce5bb0346f5e3b269affe39

Observation 9a3ed520-b8da-4ed5-b808-563486dbab8f · inbound

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding cites this paper.

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:06.861120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:06.861120Z digest=sha256:b3cc6fdba06ba083ed306cbe388c3dd92be2c5e522a8741032136068c14ae319

Observation f8d507f5-ccbc-47d5-b24f-74c6a588ed79 · inbound

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval cites this paper.

KoVRE: Training an Efficient Embedding Model for Korean Visual Document Retrieval M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:09.444119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:09.444119Z digest=sha256:45883dbbdccf5e1adfb92d7ad061c20c55a91ce220c908a28430dde78973edd8

Observation 0f661bfb-1bc2-4751-84fb-d6d2dcbd2e5c · inbound

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning cites this paper.

DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T21:35:42.455005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:35:42.455005Z digest=sha256:54962a84555a9c9e63763bb6c01ad21659f5651d47f8d1b5c2261b293e65275e