Pith. sign in

Paper Citation Record · LEDGER

Structured Attention Matters to Multimodal LLMs in Document Understanding

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2506.21600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21600 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:47:40.108730Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:19:43.621986Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T18:02:42.492424Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8acd3730-8961-4a10-b342-12ef1c3cf138 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Structured Attention Matters to Multimodal LLMs in Document Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:36.566350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:36.566350Z digest=sha256:2a8a3b5c5dc5faa9491713a814b01985c90857c3e566062c6025bab1b2afd688

Observation 30a5f94e-8089-4023-aafe-5992018c9785 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:41.547846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:36.657341Z digest=sha256:c870f711f2d89b6d2e01aa20630af85f1c44ff318494e0ebedcf498797e0e77c

Observation 0600b700-e1d4-4e72-8565-2a090f881c25 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Structured Attention Matters to Multimodal LLMs in Document Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:36.875706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:36.875706Z digest=sha256:df609852dba4e57db0493c67440b8b040277e9b32efab6687ace90ed1da0336c

Observation cc0b3ef1-c3a8-45bf-98d5-2e55cdffdfb5 · outbound

This paper cites LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models.

Structured Attention Matters to Multimodal LLMs in Document Understanding LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.029513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.029513Z digest=sha256:145738b5fa4cbb52447120c9e7207906de18d690b40730d210a7a4da366465e2

Observation 218b0545-9730-4ecd-87a5-7850e7bc3386 · outbound

This paper cites MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training.

Structured Attention Matters to Multimodal LLMs in Document Understanding MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.157188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.157188Z digest=sha256:fd1e15a793a3c7417fbfe4f01eb7faa595a8e8db25824fe649db828626bb358e

Observation b077fc46-2dbd-4faf-b1fd-3de8e4d6d3c6 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.248482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.248482Z digest=sha256:2f974bad3c51e8927d1ce7c848535cf3250261d645070947626cb6adebaefa38

Observation 9cac0801-9dc0-44f0-8672-3fb25f9d59a4 · outbound

This paper cites M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding.

Structured Attention Matters to Multimodal LLMs in Document Understanding M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.341207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.341207Z digest=sha256:690a18e7b94aa746341ac6da8999d0ddd29b6298100e0f4707ba4e01d42639e8

Observation 557e9e11-9923-40e7-ab49-7003a7901d62 · outbound

This paper cites LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating.

Structured Attention Matters to Multimodal LLMs in Document Understanding LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.435685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.435685Z digest=sha256:eddcc0fa3b3bafce8ab4fbfb06a3108a30abd9782aba1ab9381af75a92886e66

Observation c149758f-e1cf-48a9-93a7-df7ed2b5513e · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:41.365001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:37.571320Z digest=sha256:7f75da63ac2a13fecd34a7e63d618e3a4daa6788b88d8bb5ea4a77445dd285db

Observation 930e6e95-fd3f-40d4-aba1-44aa6cc5134f · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:41.188742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:37.737502Z digest=sha256:dbe452cc3f7df55eedeeabac566e59b3e3787cb01795c867846bb5739720d463

Observation bef3c24d-2df9-44be-96b1-744e84eb9ba4 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:40.966549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:37.851387Z digest=sha256:8875f75e80b245b5e30bdc7a31166a0660bf2a141756d2a585340cd02121ab99

Observation a3cc50d5-40d3-4842-8a6f-5866ae0646cb · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Structured Attention Matters to Multimodal LLMs in Document Understanding Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.917233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.917233Z digest=sha256:d1aefd55169ffa604175a5decf26440d0ce5299cb889f27ba7fadf3ae48a373b

Observation 412043ba-676f-4dd1-b790-7765516916aa · outbound

This paper cites MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding.

Structured Attention Matters to Multimodal LLMs in Document Understanding MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:37.986765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:37.986765Z digest=sha256:462e48ad781cfc05657dcc2142a379a306c20caf4ba85f5ba0b3f0d5119c34e4

Observation 57b3f667-02d7-41d7-844d-fc169ac4be3f · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.065585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.065585Z digest=sha256:e029ab10a68f46f0afeff5b3bab598e83c5895ce8375a1f63a8767e4ad49e78d

Observation bc40964e-1550-4d09-b463-3499dcb808a3 · outbound

This paper cites UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis.

Structured Attention Matters to Multimodal LLMs in Document Understanding UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.127865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.127865Z digest=sha256:af883cadda89aaa72e54a6406f1b9087f7a606e4c52cabcf02d76db08ea940d3

Observation 239b6cb1-c075-4c1e-bba7-b8492bc6a064 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.192562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.192562Z digest=sha256:539d9c1ab1211c3c9c6dace6461fa8249c1dce188545879aa3d200ffa5ea328d

Observation 2eec2257-a166-4c77-9288-4436926707e2 · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \.

Structured Attention Matters to Multimodal LLMs in Document Understanding u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.260351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.260351Z digest=sha256:eb27724a58210aa09cb8bac5ba32889289134f8feb7eca3a79f5801d4605b55b

Observation 9aa53227-272b-4a98-a7e1-e2a0188a9358 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Structured Attention Matters to Multimodal LLMs in Document Understanding MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.340852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.340852Z digest=sha256:519231c59295133a0686d011cbad100bef7187ec84ab16f164ea39b2a29d21af

Observation ebc8cab7-81b4-492d-88e4-331a0e42a767 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.403982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.403982Z digest=sha256:51dc6b6687b6eb7a845caa56d490272b2342f0e0c6c8cdc004d37436e0a73f36

Observation 2990d5ee-d5e0-4891-a994-6db0ccb3a5a5 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.468043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.468043Z digest=sha256:6f77a323a09ad345e0369c7657faa82fc9791483cdaaf35c17a72492cbcd4a6a

Observation 0a80537c-96f4-40f2-b0c5-5b88f29a12bf · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

Structured Attention Matters to Multimodal LLMs in Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.535627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.535627Z digest=sha256:f35ed3d5652bc91369bbf810a7596b454facd145daa3ca45545b497bb33f830d

Observation d6017804-7895-448c-a87a-29dde3b08fba · outbound

This paper cites From Text to Pixel: Advancing Long-Context Understanding in MLLMs.

Structured Attention Matters to Multimodal LLMs in Document Understanding From Text to Pixel: Advancing Long-Context Understanding in MLLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.551116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.551116Z digest=sha256:f7205981a3806bc94ecab33ce569b69ac0a13da520f0177cc1561849d9aa8698

Observation 79f003c5-613d-490b-919f-c1a4b29a166a · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

Structured Attention Matters to Multimodal LLMs in Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.698127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.698127Z digest=sha256:1f2192bfe3ed0edf8c43316877b9e1f8931aa759e2f724d1ce1197eace214c7c

Observation b1ca93ba-e2ea-40ef-a9fa-ae333345dcc7 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.746488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.746488Z digest=sha256:3e057518eff916af851f6227ba708aaad9293dfe8b6e9208980feb1f7e69cdfd

Observation d8faca11-f4e7-4fca-9394-9ae8d4b3c52c · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.889106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.889106Z digest=sha256:dc4f7e9f9241fca29f22308a11477b8eed6cce1ded2bb849259c8deaa0eeb811

Observation ff12f5d0-58bc-4e1d-b5f3-148668ad2fd0 · outbound

This paper cites VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation.

Structured Attention Matters to Multimodal LLMs in Document Understanding VisDoM: Multi-Document QA with Visually Rich Elements Using Multimodal Retrieval-Augmented Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.950010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.950010Z digest=sha256:41c891daa90ef5ec1782c87f1682067224dbfe463f36a540dac7d8cb5715a755

Observation 19f8b03b-c390-4246-bd4f-b382c59da653 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Structured Attention Matters to Multimodal LLMs in Document Understanding Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.053741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.053741Z digest=sha256:0ebb25265350b4343ea74e56cbffd7e6f675938991fb1a420ad365ac5fba03e5

Observation b047a5f4-3c5b-48f0-ad3a-3f8577e8d406 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.158696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.158696Z digest=sha256:c7ad04a4f8e90e14af6378132815fa1862643de9ce2bc6eff9e90947c8207b1e

Observation be76b27f-c18f-47c7-b5b5-5781ca506340 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.293516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.293516Z digest=sha256:47e44c32522e2da070fe5ea0357c9f8a00a97f1e4580c6feef59f271037cdd84

Observation 7c8db10b-9047-42fc-952b-08b260ff7073 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:40.770888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:39.392380Z digest=sha256:9a17dfe14559613a82d715f5f4e8488fd8852d831330ffcfbd4c6bdbee3cfab0

Observation ec82f0cf-2614-4cab-a515-ea9b2e0d0c18 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Structured Attention Matters to Multimodal LLMs in Document Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.500538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.500538Z digest=sha256:239ccba71293b9a72b78cd2bfa006ed61a14d9e2b0ab7016a10e9f9e1606cccc

Observation 82d1ece3-d31b-495e-8403-dbc485ef6800 · outbound

This paper cites an unresolved cited work.

Structured Attention Matters to Multimodal LLMs in Document Understanding Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:47:40.560631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-06T23:47:39.580634Z digest=sha256:394ee7d5ebab6476aac8f5980bcd4ae1227b9fbe59adbd11475b9b1fe9271852

Observation f6d2e2b8-8e95-45c1-8ff7-cd7cace1d5df · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Structured Attention Matters to Multimodal LLMs in Document Understanding mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.656467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.656467Z digest=sha256:79ecc16b70c2d83d858f32232039111ce5d3cf49585b795617571cfb79ea1ffb

Observation c30a6f23-c116-4a52-8b63-fd0d3e0629e1 · outbound

This paper cites OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation.

Structured Attention Matters to Multimodal LLMs in Document Understanding OCR Hinders RAG: Evaluating the Cascading Impact of OCR on Retrieval-Augmented Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.775273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.775273Z digest=sha256:ee7b897e8e836f54d6b28822b8ed6f5699ba995abbb622b248b7bfab941d30d1

Observation de3b0a21-a37f-471b-b485-d88fd6c34fe7 · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

Structured Attention Matters to Multimodal LLMs in Document Understanding Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.833498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.833498Z digest=sha256:03b69ef6df34dae3b79b0bd6fb34fc07188b170b174a465930e14f40b21ad3d6

Observation 7a4a81ef-f858-4a2b-89e4-0bbc5135869a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Structured Attention Matters to Multimodal LLMs in Document Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:39.913987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:39.913987Z digest=sha256:8a0ecbb6067ee72aa31548dca0585a6adabaef7e126cffcff7742fd6a2a9fa19

Observation a140df1b-6ba0-4536-8ff6-00458565147a · outbound

This paper cites URL: " 'urlintro :=.

Structured Attention Matters to Multimodal LLMs in Document Understanding URL: " 'urlintro :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:40.022629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:40.022629Z digest=sha256:54b9139d2414dc4bbe5ff557b98a37f81da81beca4667475244edb8713754182

Observation 822ab7c7-3dac-4a47-b228-fcb439828671 · outbound

This paper cites write newline.

Structured Attention Matters to Multimodal LLMs in Document Understanding write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:40.108730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:40.108730Z digest=sha256:642330d3ba18d816eb62c99400f3489a91e67164220810958a84722eac9cc9a2

Pith citing papers

Observation aacfdf11-dec1-4b6b-9b27-c0d10752de47 · inbound

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning cites this paper.

Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning Structured Attention Matters to Multimodal LLMs in Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T16:19:43.621986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:19:43.621986Z digest=sha256:9a174e99b33f8854c1106ae33cb16dcf7fb4b0d85dcdd3ae993875a3c8a1b082

Observation 7d9945e1-f706-42c4-ab82-c9d1b362c3f6 · inbound

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG cites this paper.

VideoStir: Understanding Long Videos via Spatio-Temporally Structured and Intent-Aware RAG Structured Attention Matters to Multimodal LLMs in Document Understanding

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:15:50.268399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T19:15:02.124035Z digest=sha256:6ce869d84ca4f3623c8225bfdeaa2410981d6bad53e7f56af500b449f940d5b6

Observation 20f42834-92cd-4a98-8b25-b80d443d2d51 · inbound

GenAI-Driven Approach to RISC-V Supply Chain Exploration cites this paper.

GenAI-Driven Approach to RISC-V Supply Chain Exploration Structured Attention Matters to Multimodal LLMs in Document Understanding

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T18:02:42.493996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T17:57:52.084110Z digest=sha256:376f7bcbd1579353a5fddf2b2f407deb8f183ff07d1dfcfdd7defa94a6d6d481