Pith. sign in

Paper Citation Record · LEDGER

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2505.00063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00063 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:01:15.691092Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:21:13.243457Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:31:32.012726Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ef23dea0-97f7-47c4-9b64-37c34d6f770a · outbound

This paper cites DeepSeek-V3 Technical Report.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-V3 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.436465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.436465Z digest=sha256:139e5a72db15c9525077c4856ead714f6a6c673f10fb02a3efb1263ca917ee63

Observation 0dfed56a-93db-49c8-9cfa-80589c13faca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.441586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.441586Z digest=sha256:03bc316b7fe0eac336c9df3c7a6762ed8618d9112fad2c12afe8a97c6b77f152

Observation 7aae03d6-3165-4c7c-99a4-e49e9c0aa1a3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.447113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.447113Z digest=sha256:d34ba6067ab55afc4daec25aa0ed349b7264e4a92eb39bd1501cd1efb33054c2

Observation 8eee5f11-ab8e-41a7-911d-38b2212b53bc · outbound

This paper cites Gpt-4 technical report, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gpt-4 technical report, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.451755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.451755Z digest=sha256:ee8a5c006d174e9e2d04fca225f9b227bde8f8afe3577c248bc5ade2f1ea69c4

Observation 95cb0b15-fffa-490b-9382-f9fcbe081f46 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gemini: A Family of Highly Capable Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.456539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.456539Z digest=sha256:55e7bbf50cc9d8f0e24fba938f2de365e2643295e9ecb4d3a4142329336c8c86

Observation 256032a7-155c-4392-8f4f-5144e4ffb11c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.461285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.461285Z digest=sha256:746305dbb907963cd9b0db7ba8247c0db87d208fa2b8a26c5d3a7eafc2dc550b

Observation e4037381-cace-4743-b36c-11cceef578b5 · outbound

This paper cites Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Autohallusion: Automatic generation of hallucination benchmarks for vision-language models, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.458280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.465415Z digest=sha256:cf9a62ae5d503ecfb0247f103cc660491ec6f52ab654a09bac185919c1c7e168

Observation 48cd2450-3086-47aa-a10e-d035cfc5ca01 · outbound

This paper cites Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Hallusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.445830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.469442Z digest=sha256:d71742a61e603339254256c46cb73f25d0ed7415135354cb69db551041d4a379

Observation 8e79c4fc-deb5-4c2f-a9a2-0ed26d45b394 · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.473305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.473305Z digest=sha256:f59ceb9eafd1141550b290acbb35acbb3573b2ea484f04a5452c98cef59725bc

Observation 1a1bcee3-8c27-4888-bb25-b0af90329a1c · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.477616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.477616Z digest=sha256:646a3abb964c3f82fdb1fe45557e937446a6346508071d3a28ceea7e29f1da68

Observation 2df5237b-a74d-40ec-95c3-9d1bc6fc21a3 · outbound

This paper cites Tabpedia: Towards comprehensive visual table understanding with concept synergy.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Tabpedia: Towards comprehensive visual table understanding with concept synergy

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.481951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.481951Z digest=sha256:68bbe5eb9eca5786c5b5bdea3562baee4e1661ee82d83100662d5baae0d0df20

Observation 2b21b38a-f6e9-47bc-bc22-727fa902dfad · outbound

This paper cites Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docpe- dia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.423965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.485466Z digest=sha256:64002b0b2066d16d3b74958ac1ba206cbab18931d9362b422cd4d0518a40e72f

Observation 26599e1e-d285-4baa-8b40-64cf9c1f40d5 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Overcoming catastrophic forgetting in neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.489725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.489725Z digest=sha256:59fbf264854216214eea5654dfb2372b513689b4f206bc9c8c7a1ebc72001d3d

Observation 1d018a30-731c-4ad5-806b-4caf62dc6fec · outbound

This paper cites DuReadervis: A Chinese dataset for open-domain document visual question answering.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling DuReadervis: A Chinese dataset for open-domain document visual question answering

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.401634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.493483Z digest=sha256:ee045467ef239b163c9db737171b36f845535b92db291ef50f8f98017eabf6bd

Observation 0f36c01c-7d62-4a05-9aac-48fcb308fe91 · outbound

This paper cites Visualmrc: Machine reading comprehension on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualmrc: Machine reading comprehension on document images

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.388003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.497172Z digest=sha256:18f647bb16c90d79a839f77d3933b0f712730de89c5ed4877abeac3f149c931b

Observation 0de884fa-16c3-4040-b4a5-9b91c7146d85 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.374518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.500744Z digest=sha256:d1124018872ab0418b250c6ffa741493e14a6c24fbf583b672e5271fd9479cbb

Observation 33600cd0-fc4c-484c-9ec2-b8a8b6a34403 · outbound

This paper cites Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench v2: An improved benchmark for evaluating large multimodal models on visual text localization and reasoning, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.360061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.504674Z digest=sha256:97a060894dfff5bebfa5c96b65da7cfd1dd8063fec86edaf810a5b6b467a8641

Observation 63f34b33-3058-4466-ac11-71e449726465 · outbound

This paper cites Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Omnidocbench: Benchmarking diverse pdf document parsing with comprehensive annotations, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.513802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.513802Z digest=sha256:8e5a953001c22be71395191301e3f090750a02fc892226415320495bb60112f8

Observation 10daf408-24a2-4a62-8ef9-73f593113801 · outbound

This paper cites MinerU: An Open-Source Solution for Precise Document Content Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling MinerU: An Open-Source Solution for Precise Document Content Extraction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.517859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.517859Z digest=sha256:e59a2d8e084db1da76d25284bbe7a03ce3612b31beda0c09feb8c98d2d9c0290

Observation d75fb58a-1bd2-48e8-b3bc-ae8e248d874a · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Nougat: Neural Optical Understanding for Academic Documents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.522410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.522410Z digest=sha256:b0c66cf0a535625e9f5abb2228033edd7b200c8f615b5e24cb00c1836ff6f620

Observation 1712c746-288c-478f-ac47-77c4499f7029 · outbound

This paper cites PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling PP-OCRv2: Bag of Tricks for Ultra Lightweight OCR System

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.526740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.526740Z digest=sha256:770e33704e9358d727175772c0ffcea3487bfd43a6e49a6ee9fcda120cc7c640

Observation 41a14685-2ba2-4da9-9981-83eadea739b0 · outbound

This paper cites Publaynet: largest dataset ever for document layout analysis.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Publaynet: largest dataset ever for document layout analysis

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.337991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.531155Z digest=sha256:c3b4c0a619f9de7754f89a566210aeb490648757defddfbc12d1c9f333dc77e1

Observation 8cd6ea47-e0cc-4c0b-95f2-c45de9725e72 · outbound

This paper cites Detecting text in natural image with connectionist text proposal network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Detecting text in natural image with connectionist text proposal network

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.324591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.535073Z digest=sha256:50fbe017e477989da333f2ddf998920cf26d6344813dc3c83ac7d351a8dc9308

Observation ffd37e88-9139-4325-a946-9233d44b18d3 · outbound

This paper cites Textboxes: A fast text detector with a single deep neural network.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Textboxes: A fast text detector with a single deep neural network

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.311428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.539314Z digest=sha256:42e1f47846d20b439ec9dd2128c09c79cafc19960541a7c9b86533759afd754c

Observation d82e9867-613e-4331-8d38-e1a3b6593bf2 · outbound

This paper cites East: An efficient and accurate scene text detector.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling East: An efficient and accurate scene text detector

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.298976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.543681Z digest=sha256:9b812c32ad4e3ab9857b6efedf16272e47886872e8e6d3b0b649234656a50503

Observation b9e2be46-9cc0-4de2-aab5-9ba6f40745cb · outbound

This paper cites Curved scene text detection via transverse and longitudinal sequence connection.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Curved scene text detection via transverse and longitudinal sequence connection

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.286625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.548442Z digest=sha256:54c5fa997e30cc6d05441c71e8f2c07201b2e12e06db260262ba12987d443478

Observation 9ec5d3f4-1911-47c7-82e1-7fd7155a6019 · outbound

This paper cites Gradient-based learning applied to document recognition.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient-based learning applied to document recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.553128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.553128Z digest=sha256:29daa1115df3096f7724bf7172b4955ac94366de5b9a5fdda0e7595773fab10f

Observation 2d3e4c22-bd05-493e-95c0-05731349d10b · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.265383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.557387Z digest=sha256:a2e236edbabb6d35b236ee87b29d817aa1b2d6d4897da966b309d55a31ec8966

Observation 8c6dc4d6-f1b3-4d83-8bad-d3bd2926a0e7 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Trocr: Transformer-based optical character recognition with pre-trained models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.251990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.561367Z digest=sha256:bb59f720845e3f000d0fd8ee046bf1f1275ca9ff761bddb2aece26f545d39f42

Observation c2e2f2d9-094c-4e41-b180-932d07ac1f6f · outbound

This paper cites General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling General ocr theory: Towards ocr-2.0 via a unified end-to-end model, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.565654Z digest=sha256:3395e2c7c8898f905c501df7fa1fbc37d591e3e35d0036ac7716e96c2cdd7754

Observation 8996b0e3-6429-4831-a4a8-c789934f1974 · outbound

This paper cites Visual instruction tuning, 2023.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.569895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.569895Z digest=sha256:5e1a5dfe637292da0877d4609cea6b66c082bfd49e3e88eea371dcad5a25dc0a

Observation d0488b0b-c210-491b-993b-17f12a1606ed · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.574129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.574129Z digest=sha256:c189c9090e036e627569cb8c2583e9f4b383cb9c254d05a0f44704dcf6b36ce3

Observation 922212e8-41a2-47e2-8247-2eb451015deb · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.578722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.578722Z digest=sha256:533160ccc0a71548889a777e7111365fccdedf637d3fcf6ce3e629d6f21470da

Observation e0997066-b011-4465-9296-49893f0d18e6 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.583499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.583499Z digest=sha256:fb75b70dc7a581886d43577a938365cb75e79dd489d9735cbd7591c4feda52bb

Observation 1d82a257-955e-4ba6-9d4a-90268be311fe · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.588006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.588006Z digest=sha256:cf81feb276087388554035da5eeb9372b7d0eaf075f2c3e4d7a22c58641e553b

Observation a886ceed-1c2f-4996-9721-b5f8b65c349d · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.592340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.592340Z digest=sha256:c731dfb889b62330515d5711005ac9bee9972e3c1e3b6eadf9928ed30d81789b

Observation 5d83cfa6-9e42-4c5d-956a-efbde08c06dc · outbound

This paper cites Learning transferable visual models from natural language supervision.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.216221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.596918Z digest=sha256:f937c254ece88d35db8320484ba01e40208aec72f2e6a3653d714ef8217e3373

Observation 797929cd-e1e4-421a-ba37-31d011647556 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.600852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.600852Z digest=sha256:354bccf9b881982f37dad8bf231ce2817727f0becefcf0e39010aca3448c6ee4

Observation e95e7813-7efa-4938-bf55-3a8a040aea4a · outbound

This paper cites Lamol: Language modeling for lifelong language learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lamol: Language modeling for lifelong language learning

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.203170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.606072Z digest=sha256:9aecb50e834acdfb0ac3b51590b87271e48628347fc9b674b6ed40cfe6a535d4

Observation c54b0be1-acc4-4cef-9d30-0429b1f42896 · outbound

This paper cites Rational LAMOL: A rationale-based lifelong learning framework.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Rational LAMOL: A rationale-based lifelong learning framework

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.189763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.610475Z digest=sha256:6e7cbedadf96de4864a4cc49d300601753f3161351502a5a19eea77a1049f842

Observation 82ee7177-343a-4e43-a0c8-b7b7b7adc093 · outbound

This paper cites Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Loramoe: Alleviating world knowledge forgetting in large language models via moe-style plugin

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.176505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.614501Z digest=sha256:40bc94ffa577cd0cdbed2819b50a9ed886673ad2d9a6f4f10b7d640e5132d2cd

Observation 81af4035-99a5-4568-aa02-a9478c71ba96 · outbound

This paper cites Progressive Prompts: Continual Learning for Language Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Progressive Prompts: Continual Learning for Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.618680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.618680Z digest=sha256:0a808eddb43ecf40aa8021b59953a8173731ff86b3404ef28a536e225f03be80

Observation 173fb86d-f4b2-44c5-80fb-fd64ffbcf65b · outbound

This paper cites Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Teamwork Is Not Always Good: An Empirical Study of Classifier Drift in Class-incremental Information Extraction

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.622890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.622890Z digest=sha256:bb02673e3c30df0c16324ac10360af5075dd3b57c0f2b8f746c6fdbc18a24b58

Observation 68db0b5c-fd5e-40ed-87c2-ccfe3df8d7bd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Lora: Low-rank adaptation of large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.626804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.626804Z digest=sha256:d5aae0b7b7496460d0f9285df3e9cae8f8abc4ad6a829d8a589aaa900dd85fde

Observation 6d2a9d72-d308-4c3d-aeae-538066b61641 · outbound

This paper cites Continual sequence generation with adaptive compositional modules.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Continual sequence generation with adaptive compositional modules

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.154042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.631198Z digest=sha256:d4a42ae6d57b71e84af11a1e06d396415be46add33281373e6fb429de5a20ad3

Observation 64ae4ae8-111b-4118-8b71-3ee8fb3e737a · outbound

This paper cites Preserving in-context learning ability in large language model fine-tuning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Preserving in-context learning ability in large language model fine-tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.635375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.635375Z digest=sha256:7b55ef780a62cd9f7906a4958f024156bc6aeef354711326b8847a48b3708e83

Observation 03360ac3-5c97-4f2b-be02-6f313733a82d · outbound

This paper cites Editing models with task arithmetic.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Editing models with task arithmetic

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.130337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.639331Z digest=sha256:f87ee41748c5c3d0a8c709f740ba44802d5575875bbaf63e3d3835f8831d8890

Observation d36672bf-a3f0-41d6-8f2a-af23e095b6ec · outbound

This paper cites Gradient projection memory for continual learning.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Gradient projection memory for continual learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.116121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.643336Z digest=sha256:266a9dc8873a795f79b683df1638060cc2a5285c575169e021d5fe49116637b7

Observation 2c79a614-6c97-4c64-ba44-130819d5fd74 · outbound

This paper cites Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Visualsimpleqa: A benchmark for decoupled evaluation of large vision-language models in fact-seeking question answering, 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.102649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.647599Z digest=sha256:20abbba17af4586732c829c7227df8815e172161943ac99a92fb3ef8a7288598

Observation bdc4997a-bc58-43a5-b345-6991e070fa4a · outbound

This paper cites Binary codes capable of correcting deletions, insertions, and reversals.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Binary codes capable of correcting deletions, insertions, and reversals

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.652164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.652164Z digest=sha256:e9a8a399aca17b4d77aa0603e40192cde0f30be6b8d96dd48d01f4a3c44451b2

Observation d4a37fe8-b0cd-4f5a-b8d4-e235d78f3865 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.656654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.656654Z digest=sha256:aefe6d77df0fbde603a0beaeabd143514f656a90c689f00ad34e3c63bad15337

Observation 19630431-9e97-45b1-9486-d0f27d7a6590 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.661099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.661099Z digest=sha256:746f0db2bc4dfb069b7180d33ed7aee5c527865c02d37904cd6687f9db61eda9

Observation a313a46e-512d-4f45-b274-2b209bb8c639 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Docvqa: A dataset for vqa on document images

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.665658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.665658Z digest=sha256:9a0a676bed385ed47852370dfab7576b14d05d0e8ba364926f8e9249d3527783

Observation 1235ce46-603f-4cfe-a02d-550ffdccf97b · outbound

This paper cites Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ai2d-rst: a multimodal corpus of 1000 primary school science diagrams

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.069875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.669573Z digest=sha256:488b379ded71cd2e2abff7d95dc133296ae41176601c9f3aa7d009bca02236e0

Observation 3f79ab19-d454-45af-825d-ef10ac8387e1 · outbound

This paper cites Towards vqa models that can read.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Towards vqa models that can read

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.674358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.674358Z digest=sha256:8f069117bf35756e7a2cc5bb73034fde58b4e12e124e77e48bc856232ec148ea

Observation 0fd1fb34-22b1-47fb-8f26-ddfeff0de97c · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.678427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.678427Z digest=sha256:808bb432ae490e0f344cf462cc1945645cbbce09967dc7ffcd150ddead044857

Observation 91f25037-2c36-4fc1-ac05-3854adee8f7c · outbound

This paper cites Infographicvqa.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Infographicvqa

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.682362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.682362Z digest=sha256:1cc57187283c944033e3fcb2303e404b313af27bca5ae2f2de571bda2da34223

Observation c0b50e32-06a3-45b4-9e96-a10264b63551 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:01:16.028544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T05:01:15.686867Z digest=sha256:7c70442be823ccc58960620879f55364e8a42b924535305064ad916daccf233e

Observation 1bd3d5dc-95fc-4730-8710-cda898b85a6b · outbound

This paper cites GPT-4o System Card.

GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling GPT-4o System Card

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:01:15.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:01:15.691092Z digest=sha256:da5c33bb3ff58f32a662e7682bd97753627786dc45f3f6246e000cc949d40a6a

Pith citing papers

Observation 8a091469-9a7d-47a4-8830-d8e7bd55df8e · inbound

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models cites this paper.

Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 255

Resolution
unresolved
no resolver link, observed 2026-08-15T23:21:13.243457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:21:13.243457Z digest=sha256:efaa33ad34c5fcf86a72b96b1a8615159c22b49d100fad74c38b3e08252bebe3

Observation 52333cf1-07aa-4464-b696-c4a155fd1775 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models GDI-Bench: A Benchmark for General Document Intelligence with Vision and Reasoning Decoupling

Reference 135

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:31:32.084497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T14:31:20.388845Z digest=sha256:590ae438aca824dbb019150c0436dff632ecb0b2b5bbb58be412a5824592566e