Pith. sign in

Paper Citation Record · LEDGER

Multi-Agent Interactive Question Generation Framework for Long Document Understanding

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.20145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20145 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:45:18.546931Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfcc0c4a-968c-4fb8-bb6e-00bee2d33549 · outbound

This paper cites GPT-4 Technical Report.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.393787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.393787Z digest=sha256:3f87d727cbf0457737f9c2f0f07046f860476e970d454d27a43782dc0f80ee56

Observation e484a1e3-89ae-4362-8052-87efc89c74fa · outbound

This paper cites Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Gemini 1.5: Unlocking multimodal un- derstanding across millions of tokens of context,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.148011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.400209Z digest=sha256:56886fcc383ae89afd226e0745d469a9f78ac68b0fe406d52ed12ed95c3efde8

Observation 4ebcdc18-ebf5-4418-a64b-d9259acee678 · outbound

This paper cites Claude 3 haiku: Our fastest model yet,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Claude 3 haiku: Our fastest model yet,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.129130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.405297Z digest=sha256:f07574d68de0289b20e4f96fbfcc3aa76a475c681becbe45edda8032fdce1f59

Observation af5ee380-4c5b-4f88-bb0d-b473a4b9828b · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.410519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.410519Z digest=sha256:afc7bf62c4031de1b06bf9e421054c9236584b24dadfa0ace9ce830531484763

Observation ca2994f1-3793-4b7a-9417-389b93fee3ce · outbound

This paper cites Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Llava-next: Stronger llms supercharge multi- modal capabilities in the wild,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.110730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.415715Z digest=sha256:ea98e5952d9bd028d2eb50205c650fc6474567a165eb58733d7f75e3daef3dcd

Observation 324f86a9-4301-49e3-ae5b-2611e8cd2e4d · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.420662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.420662Z digest=sha256:0975aede2c978d3703fc36644d67f94e0d9d2af138e63e00e51d48d35d74eea0

Observation 5acff4aa-9b05-4dfb-853b-431b69ccceef · outbound

This paper cites Docvqa: A dataset for vqa on document images,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docvqa: A dataset for vqa on document images,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.093513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.426258Z digest=sha256:8d1ee21e67f35b6777409a5c21b8051a6d18195b682a90f364053a1cf15449f3

Observation 27280278-a693-4ce7-8ca7-161302e5951f · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.431189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.431189Z digest=sha256:618d50fcf0e8bbccfe84b06fb5bd2c60cf9177674b7d5c8536ecd82f5be6d7c7

Observation 34bfc216-5520-4bd6-b9e3-8398fb41455f · outbound

This paper cites Infograph- icvqa,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Infograph- icvqa,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.076609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.436284Z digest=sha256:94fd592896d9fe2b1db5a6e03c63d47775793823e5eac938e1fc74e382d09ff2

Observation 64f418ca-3151-4c96-81d6-199f9909fefa · outbound

This paper cites Towards com- plex document understanding by discrete reasoning,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Towards com- plex document understanding by discrete reasoning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.060209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.441640Z digest=sha256:aeb077c93557e8f43e4bfefcaf9664184ffeff9a82bf4e0d476277bbf52ebb0f

Observation 27667108-53ab-483f-9c3c-38df8943549b · outbound

This paper cites Document understand- ing dataset and evaluation (dude),.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document understand- ing dataset and evaluation (dude),

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.041347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.446195Z digest=sha256:2a61fa09441068b56344c28c08c0d16c99c78e386dcd19d49a6144734039b372

Observation b6602d08-cec9-4ae5-8d63-dd8c648ee984 · outbound

This paper cites MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.451030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.451030Z digest=sha256:665579f1bc3e463b1472f50fdee161783c30da6e07742234a49bb6dda3e99cd6

Observation 63aeeae8-d904-4fca-aa40-385277a2c29e · outbound

This paper cites LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding LongDocURL: a Comprehensive Multimodal Long Document Benchmark Integrating Understanding, Reasoning, and Locating

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.456133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.456133Z digest=sha256:92a04f34bd289cb4d2506f174a20fa5a7e6377bee5e391076037f365b59c4d9b

Observation c931406a-2771-4c36-a718-3cb1e3f6d965 · outbound

This paper cites M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.461076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.461076Z digest=sha256:fd794085fb00e31bd4a38b7a0555505dcad0c77c8bf11d7d2e428f9192df9320

Observation 0547c844-27b8-40ba-a9c6-d9aa0da3ca25 · outbound

This paper cites CAMEL-Bench: A Comprehensive Arabic LMM Benchmark.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding CAMEL-Bench: A Comprehensive Arabic LMM Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.466100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.466100Z digest=sha256:3ccbbb79d7389083a7e06d38f85d39769eb460e53da8fd7efe512f90a752fa7b

Observation a370b885-cb9b-4f26-a4fc-50598c0175f0 · outbound

This paper cites Doclaynet: A large human-annotated dataset for document-layout segmen- tation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Doclaynet: A large human-annotated dataset for document-layout segmen- tation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.023515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.472174Z digest=sha256:adc2ad4ae93f331e32fc85eb47b6b585fed002e2e1911e9669620335a7116b97

Observation 654df614-46d7-4c05-b932-5b6438e5dccc · outbound

This paper cites A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding A Large-Scale Multi-Document Summarization Dataset from the Wikipedia Current Events Portal

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:45:18.705002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.477594Z digest=sha256:6353638b07601d30ed7ca1802a1a228715115c7e14e794ae50f63a6aa0d047fe

Observation 6c8fb9f5-c11f-45e1-9669-767c8bc8ad37 · outbound

This paper cites Docile benchmark for document information localiza- tion and extraction,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Docile benchmark for document information localiza- tion and extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:19.005952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.483478Z digest=sha256:55d7d4dfcfb6761dc9d7d0203e34a59dd417ddc61bf7f132c1f796029951228a

Observation 940ef01c-1f18-4173-b91b-80a37e052f3b · outbound

This paper cites Document Visual Question Answering Challenge 2020.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Document Visual Question Answering Challenge 2020

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.488425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.488425Z digest=sha256:a8078090e3d19f771396ef836f77fb7ba4e28c228e01e37fb87b8d80b290d2f8

Observation 5c6eff25-deef-4376-9f0c-648a7aac0fb1 · outbound

This paper cites No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding No- vachart: A large-scale dataset towards chart understand- ing and generation of multimodal large language mod- els,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.988127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.493435Z digest=sha256:e2c1cd09b8d0332209bef2ae067a7a8f8e033215593754fb8a52ea562cb852ff

Observation 91934601-72ae-42b7-92f3-3fbdcd9a8ec5 · outbound

This paper cites An overview of the tesseract ocr engine,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding An overview of the tesseract ocr engine,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.969588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.497787Z digest=sha256:956fd2acd9b6d595bdb5c47361c0f2dbb27695eb28e19a28b2e5b815651f5a4f

Observation 5fa76249-83fe-4102-9c5f-10abeb0d620f · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Learning Transferable Visual Models From Natural Language Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.502670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.502670Z digest=sha256:d02cca6ee536710aac5ce814df4d3c3c60a30ca574b4d34584e9c0eb89525815

Observation fc3a2fda-fdd8-421f-8669-fc293555ae44 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.507641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.507641Z digest=sha256:da6e364ae74ba48885f67a40c3ec9ed43857a2be12fb4485289136ab37f42525

Observation efa875fc-b398-4e67-abed-731a7b7b199a · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding BRAVE: Broadening the visual encoding of vision-language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.512376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.512376Z digest=sha256:801311bed757e4fd612844f766a4bd35425c5314412ee2d670028750fcf59fae

Observation 6fa8425a-0bf8-4c38-9f6e-1875c87d8dca · outbound

This paper cites Automated annotation with generative ai requires validation,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Automated annotation with generative ai requires validation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.950322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.517323Z digest=sha256:725f543b911d90f09f9f672c517680e886ab868877ea17e34802ac7bf75fadd0

Observation c80e8f12-6d94-420a-9b61-0cadd0596306 · outbound

This paper cites Labelvizier: Error profiling and interactive data annotation for long- document understanding,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Labelvizier: Error profiling and interactive data annotation for long- document understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.930246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.522182Z digest=sha256:f6584bc9ff48ea79221ac1523894e476093e08a4792ed8d912044229552b68ab

Observation c1b64abc-2170-483c-9c9b-e05c14af4a66 · outbound

This paper cites Meganno+: A human-llm collaborative an- notation system,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Meganno+: A human-llm collaborative an- notation system,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.911021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.526875Z digest=sha256:170fcfae185bc66929d11b6445ebd4c18222fdae51555019484818ae123c43ad

Observation 858801ef-4b1e-44e2-8afd-ee4c2fc5e5f9 · outbound

This paper cites pdf2image: A python library to convert pdf pages to images using poppler,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding pdf2image: A python library to convert pdf pages to images using poppler,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.888454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.532153Z digest=sha256:7c1c1eaf462e38885f7d36cb2a27ba32d0b3c2f5dbb0666a3c12154cfbd83379

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · outbound

This paper cites TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:470aa2ff90021713a8feb17e3d8c9de6645a93aca37a2c45accde0e748ed5e7e

Observation 06237e27-d010-4abb-9116-41f832c4d3e2 · outbound

This paper cites Ocean-OCR: Towards General OCR Application via a Vision-Language Model.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding Ocean-OCR: Towards General OCR Application via a Vision-Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.542039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.542039Z digest=sha256:bc39916088f7f4bced628f2f6446677cdc06a67789ebcb421d05dbd333e39d3e

Observation 977550ae-8cdf-4e78-a8db-8e7ddf9a203e · outbound

This paper cites The yolo framework: A comprehensive re- view of evolution and applications,.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding The yolo framework: A comprehensive re- view of evolution and applications,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:45:18.867108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T13:45:18.546931Z digest=sha256:f6dbf06597dfaa0ad5adf3a27d3e863558ebd614a6827bb97c3d9ab077f344ce

Pith citing papers

No inbound Pith citation observations are available.