Pith. sign in

Paper Citation Record · LEDGER

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2505.00063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00063 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.243457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:31:32.012726Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef23dea0-97f7-47c4-9b64-37c34d6f770a · outbound

This paper cites DeepSeek-V3 Technical Report.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.436465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.436465Z digest=sha256:3bcf61c7ec5b404ff9f24432b5af2cf4cdbdab62500c96db5a612f18f345abd9

Observation 0dfed56a-93db-49c8-9cfa-80589c13faca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.441586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.441586Z digest=sha256:d1e62699791ad6ce6966ad0b9e5c9b760b6661efda08f56321d7914fa96e99fc

Observation 7aae03d6-3165-4c7c-99a4-e49e9c0aa1a3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.447113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.447113Z digest=sha256:5a18b94fcf7fed5ebe291a1e8e8d5f802556903099113a9fc44ba52de92ceb37

Observation 8eee5f11-ab8e-41a7-911d-38b2212b53bc · outbound

This paper cites Gpt-4 technical report, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gpt-4 technical report, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.451755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.451755Z digest=sha256:1fdde35e257cd38a35fc12beeaed9daef3bdb048727930eef01d5962632d7511

Observation 95cb0b15-fffa-490b-9382-f9fcbe081f46 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.456539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.456539Z digest=sha256:5d3d6ab14132c2257829a6076b1d72a8446b7fe8dfd73bbfb06a66d280f98359

Observation 256032a7-155c-4392-8f4f-5144e4ffb11c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.461285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.461285Z digest=sha256:015ce17086e0ab6161254d372d6b821d64573f732b9e534fec341319a041f626

Observation e4037381-cace-4743-b36c-11cceef578b5 · outbound

This paper cites Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.458280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.465415Z digest=sha256:c13c7c3ac72784c762a255b7c7c852f562d5a43ceb13d3eca8eb20d0d761a343

Observation 48cd2450-3086-47aa-a10e-d035cfc5ca01 · outbound

This paper cites Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.445830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.469442Z digest=sha256:a2df5888f8bd5b47bdea5e04bb3ae97a0454283207bf2108c0159a0de61b1734

Observation 8e79c4fc-deb5-4c2f-a9a2-0ed26d45b394 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.473305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.473305Z digest=sha256:81360a1d44c7cf4a1e753756c75042520097c126aa0648e329df6495b7fa9096

Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.477616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.477616Z digest=sha256:f36c2ee898092bb7f877357cf966d81c01ddeddd3d4f5549c402992c55fec0c5

Observation 2df5237b-a74d-40ec-95c3-9d1bc6fc21a3 · outbound

This paper cites Tabpedia: Towards comprehensive visual table understanding with concept synergy.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Tabpedia: Towards comprehensive visual table understanding with concept synergy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.481951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.481951Z digest=sha256:8335bc9861ffc750622f2d831360352510217db1d3b35a96a205045a89a8d76f

Observation 2b21b38a-f6e9-47bc-bc22-727fa902dfad · outbound

This paper cites Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.423965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.485466Z digest=sha256:8731459c62ab6ce201e75b1233d234e9387018b22655396fbbde1526635111a5

Observation 26599e1e-d285-4baa-8b40-64cf9c1f40d5 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Overcoming catastrophic forgetting in neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.489725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.489725Z digest=sha256:c25df872ddebbe984de1c4f91c7fbbb17b96b633ec8ab68267ad463125f4ba5a

Observation 1d018a30-731c-4ad5-806b-4caf62dc6fec · outbound

This paper cites DuReadervis: A Chinese dataset for open-domain document visual question answering.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DuReadervis: A Chinese dataset for open-domain document visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.401634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.493483Z digest=sha256:bbd8ec5eeb944c3299a0d586594e7dc2430087f7b72b909cd4bfc68867f50432

Observation 0f36c01c-7d62-4a05-9aac-48fcb308fe91 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualmrc: Machine reading comprehension on document images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.388003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.497172Z digest=sha256:79045e04d135d425a9cfe0ea3b943c46af4679298b2b8add606101eac0766b3f

Observation 0de884fa-16c3-4040-b4a5-9b91c7146d85 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.374518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.500744Z digest=sha256:4b7cf0c2555befcd1ae81f5d9c8644ed991ffac5f341471a5cd1ca6582070c2c

Observation 33600cd0-fc4c-484c-9ec2-b8a8b6a34403 · outbound

This paper cites Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.360061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.504674Z digest=sha256:e2b01c0575cfa5ce34c013931fc8dd4610d5720435d7ab05e5f9a61e4c9227a0

Observation 63f34b33-3058-4466-ac11-71e449726465 · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.513802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.513802Z digest=sha256:ae2d435401dade40f968433f339efb2393dc183f7d8f6de928f7442bf2171137

Observation 10daf408-24a2-4a62-8ef9-73f593113801 · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.517859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.517859Z digest=sha256:5b12b0335c839aa042083f37fd1ace7cd81884d96a6914d6bf048dee11a1804d

Observation d75fb58a-1bd2-48e8-b3bc-ae8e248d874a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Nougat: Neural Optical Understanding for Academic Documents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.522410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.522410Z digest=sha256:1da693da1be415090cbe761b4096154244b3c294533bea016f2d73a7cc92b586

Observation 1712c746-288c-478f-ac47-77c4499f7029 · outbound

This paper cites PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.526740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.526740Z digest=sha256:b8822d3d61687d01e078211f8379017438cbd4771dbaccf5aef0cc7173bd9c0f

Observation 41a14685-2ba2-4da9-9981-83eadea739b0 · outbound

This paper cites Publaynet: largest dataset ever for document layout analysis.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Publaynet: largest dataset ever for document layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.337991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.531155Z digest=sha256:ce55d5780d141a1d86bbb37cb6fe0b3936f1d861d373ef67fcd0f7fe2dc315f5

Observation 8cd6ea47-e0cc-4c0b-95f2-c45de9725e72 · outbound

This paper cites Detecting text in natural image with connectionist text proposal network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Detecting text in natural image with connectionist text proposal network

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.324591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.535073Z digest=sha256:c2e6b79f8ed6ce2aa60d966f37eeb535f94d4974901384e66e0f0666e4a4cac2

Observation ffd37e88-9139-4325-a946-9233d44b18d3 · outbound

This paper cites Textboxes: A fast text detector with a single deep neural network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Textboxes: A fast text detector with a single deep neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.311428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.539314Z digest=sha256:329a58bbaa7e1001e2588d805a4203ef709ef8262343aff62dee52472514584e

Observation d82e9867-613e-4331-8d38-e1a3b6593bf2 · outbound

This paper cites East: An efficient and accurate scene text detector.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling East: An efficient and accurate scene text detector

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.298976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.543681Z digest=sha256:6b35d4d106e718c02c18db5de827f7cad12df9681417880787b35b45e05a58c8

Observation b9e2be46-9cc0-4de2-aab5-9ba6f40745cb · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Curved scene text detection via transverse and longitudinal sequence connection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.286625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.548442Z digest=sha256:911d08a2bcf5191e8d5dcf95962707d6aa3de6875808df8afca6814eccdc5982

Observation 9ec5d3f4-1911-47c7-82e1-7fd7155a6019 · outbound

This paper cites Gradient-based learning applied to document recognition.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient-based learning applied to document recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.553128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.553128Z digest=sha256:0501fd1245101eaab139659dd64ba0589ec26657e03a5c6d865e6500c971463a

Observation 2d3e4c22-bd05-493e-95c0-05731349d10b · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.265383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.557387Z digest=sha256:633ec2cd7af9a5dfc1a5d3cc70734bd1fdb7e0e4d43ba79df58ec47d46fc7dcd

Observation 8c6dc4d6-f1b3-4d83-8bad-d3bd2926a0e7 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Trocr: Transformer-based optical character recognition with pre-trained models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.561367Z digest=sha256:483aa43fb4d529cf2c9fdf14b4772088fdfa52d991c48a6e56a84c2b63566141

Observation c2e2f2d9-094c-4e41-b180-932d07ac1f6f · outbound

This paper cites General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.565654Z digest=sha256:cf4846cc5ad2dbaea326f31095f497f348e0d532fbb9bf45704a101e84417386

Observation 8996b0e3-6429-4831-a4a8-c789934f1974 · outbound

This paper cites Visual instruction tuning, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.569895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.569895Z digest=sha256:d7c004e1dbdc5fc38ad829de5cc1332c020e765b71d833ff1c4abc78e79985c6

Observation d0488b0b-c210-491b-993b-17f12a1606ed · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.574129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.574129Z digest=sha256:3a71e114695c7a7317ac1b2c5181a123b3acf2abe02e7e5c4e54901cd702c29d

Observation 922212e8-41a2-47e2-8247-2eb451015deb · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.578722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.578722Z digest=sha256:28e28c729bb1e932484a2ea5b029d759138f2599a1bcc26889f1ba4abadcf426

Observation e0997066-b011-4465-9296-49893f0d18e6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.583499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.583499Z digest=sha256:2dbf0c3f9a3532c8f53ec5296c832f42d5f6717effe34848755de8279250f932

Observation 1d82a257-955e-4ba6-9d4a-90268be311fe · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.588006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.588006Z digest=sha256:cf87d425f1320ab16e976381407a631c8b1798ba368a6f54cefbebfb81a485d2

Observation a886ceed-1c2f-4996-9721-b5f8b65c349d · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.592340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.592340Z digest=sha256:56ab56a623a625ceaadcf2ca277b3ac4cf938554b1a028735fb35bba28fc8119

Observation 5d83cfa6-9e42-4c5d-956a-efbde08c06dc · outbound

This paper cites Learning transferable visual models from natural language supervision.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.216221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.596918Z digest=sha256:d58f478371172f65cbb77d13061c52f0d5f571b85b2acc889d24d7719260a600

Observation 797929cd-e1e4-421a-ba37-31d011647556 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.600852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.600852Z digest=sha256:7c328cf0d0ce1732d57d29937f78321b3a85ccf2099cd7fee5f87d70e05970de

Observation e95e7813-7efa-4938-bf55-3a8a040aea4a · outbound

This paper cites Lamol: Language modeling for lifelong language learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lamol: Language modeling for lifelong language learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.203170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.606072Z digest=sha256:1b07888912e37c64dc10ec1dc6af2af79f6b68e631fdf32a81431008c18895be

Observation c54b0be1-acc4-4cef-9d30-0429b1f42896 · outbound

This paper cites Rational LAMOL: A rationale-based lifelong learning framework.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Rational LAMOL: A rationale-based lifelong learning framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.189763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.610475Z digest=sha256:d766d20c6ed8ece91befeb987f8f65caddc21d8c5c601c48d1bae175d9a6031f

Observation 82ee7177-343a-4e43-a0c8-b7b7b7adc093 · outbound

This paper cites Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.176505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.614501Z digest=sha256:c45b9fa5d479fdf14dafb537a4811585c82dddc303cb41861e9ad3c803ccdb11

Observation 81af4035-99a5-4568-aa02-a9478c71ba96 · outbound

This paper cites Progressive Prompts: Continual Learning for Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Progressive Prompts: Continual Learning for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.618680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.618680Z digest=sha256:376d1c1c4158d3e237786f2afb5a8bd5eb524f29a0f9fab58ee1101d1b525f86

Observation 173fb86d-f4b2-44c5-80fb-fd64ffbcf65b · outbound

This paper cites Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.622890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.622890Z digest=sha256:b6caa14a48a07d56851a2a65de605eabe69eb075ec4a46bfe2894c0f6c538354

Observation 68db0b5c-fd5e-40ed-87c2-ccfe3df8d7bd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lora: Low-rank adaptation of large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.626804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.626804Z digest=sha256:c9824f1a714a1868a729f4583bd1d5efc080cf60cf1efc7d273818c96d55022f

Observation 6d2a9d72-d308-4c3d-aeae-538066b61641 · outbound

This paper cites Continual sequence generation with adaptive compositional modules.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Continual sequence generation with adaptive compositional modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.154042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.631198Z digest=sha256:9a72e38c8435b21680cca992dea3124bbab4f7b2d71071b7649f2524ca4f1795

Observation 64ae4ae8-111b-4118-8b71-3ee8fb3e737a · outbound

This paper cites Preserving in-context learning ability in large language model fine-tuning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Preserving in-context learning ability in large language model fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.635375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.635375Z digest=sha256:81ffea344a815b819ec20994e0c75980c24cda1a4f27bcc69023b2a511ff44b7

Observation 03360ac3-5c97-4f2b-be02-6f313733a82d · outbound

This paper cites Editing models with task arithmetic.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Editing models with task arithmetic

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.130337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.639331Z digest=sha256:4e69967dc125285152d47f12b19f0618561679bbff3681358845c82e10d145d1

Observation d36672bf-a3f0-41d6-8f2a-af23e095b6ec · outbound

This paper cites Gradient projection memory for continual learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient projection memory for continual learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.116121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.643336Z digest=sha256:5f3c16761b1c8ed0818039e169353453e1afe13b34b7cdd049b84792653e1bc3

Observation 2c79a614-6c97-4c64-ba44-130819d5fd74 · outbound

This paper cites Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.102649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.647599Z digest=sha256:ecd572b42481d5d72b7cfe5669c90c176037e17107d268756b8673bb0f4cd9db

Observation bdc4997a-bc58-43a5-b345-6991e070fa4a · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Binary codes capable of correcting deletions, insertions, and reversals

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.652164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.652164Z digest=sha256:244b8fbebc1e8c10511aedaba42260b6567a44be83bf8d9fe047983bf132a4c7

Observation d4a37fe8-b0cd-4f5a-b8d4-e235d78f3865 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.656654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.656654Z digest=sha256:3b779ffd0be5d34fa562546ca6d6a37c250f7513d4f3dcb64874b17d8e690784

Observation 19630431-9e97-45b1-9486-d0f27d7a6590 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.661099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.661099Z digest=sha256:76871d6afb2d49f1f798e6597a2b6916df477ed7499ab9971ff07b77f7f44097

Observation a313a46e-512d-4f45-b274-2b209bb8c639 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docvqa: A dataset for vqa on document images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.665658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.665658Z digest=sha256:f765f1abdfef464d2e97f7d083ee5979a5e69fd62e1cf65202270f4a9e56589d

Observation 1235ce46-603f-4cfe-a02d-550ffdccf97b · outbound

This paper cites Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.069875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.669573Z digest=sha256:baf13ae4377362044fc4f27eade3827349cdfc825a77ce27a3812c9d6fd715fe

Observation 3f79ab19-d454-45af-825d-ef10ac8387e1 · outbound

This paper cites Towards vqa models that can read.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.674358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.674358Z digest=sha256:e90573d22f37192f55bd988e45334bd847f418395f7e3f3c3d8be2cf79df555e

Observation 0fd1fb34-22b1-47fb-8f26-ddfeff0de97c · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.678427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.678427Z digest=sha256:cf49b25e7e2b0990bd9f74e5838d5c01290be60380b5dabe1091ed50ee6f8d28

Observation 91f25037-2c36-4fc1-ac05-3854adee8f7c · outbound

This paper cites Infographicvqa.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Infographicvqa

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.682362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.682362Z digest=sha256:349ed6091858379ede77f95da294fa7c44e04952a59f61b8c4dad718c6a60e3c

Observation c0b50e32-06a3-45b4-9e96-a10264b63551 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.028544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.686867Z digest=sha256:52d08143cdfca0fa78a610fdf6d7d0915e231117adbf3fc64f59ef146d76e348

Observation 1bd3d5dc-95fc-4730-8710-cda898b85a6b · outbound

This paper cites GPT-4o System Card.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling GPT-4o System Card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.691092Z digest=sha256:968e1fac0512474534947a9d5f526872f137e6ca707a7536794f8d0dd9ee3bb9

Pith citing papers

Observation 8a091469-9a7d-47a4-8830-d8e7bd55df8e · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.243457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:13.243457Z digest=sha256:9b2afe8c93aeeecce0e151974d4d841d4846fb3a12c995751bd645a9d8f3f8c6

Observation 52333cf1-07aa-4464-b696-c4a155fd1775 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:32.084497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:20.388845Z digest=sha256:de41c67a0e3181a164ee70147a618524ae13bc6f6a3f8da3c74c3a18c92f004c