Pith. sign in

Paper Citation Record · LEDGER

Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2312.06109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06109 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:43:20.589681Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:15:47.123482Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ef0877b-8a2d-4ce0-bcb8-f8cb00d2e87e · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.746237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:bbfcc1fead7a428f5f2b2d61a58f17ac27d15b58df2df02fbbcde93af5280be6

Observation 3cb97303-4f3a-4cb8-9c4d-7f55523da5bb · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.114785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:fa08e1699aafc68d67428398e603b5cc854a358dda021192b9dfd432c4749a74

Observation df5d7acb-9b2d-4270-826e-e8951235dad4 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.251278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:eccac21de49339719bac2dd8dfabfb154aec098ce7d12beb51a7d0630f9468bc

Observation 8cc8cd89-b9b6-4665-bf4b-548ff47a3707 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.764945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:10cdda0ad2121e2ed748a26e957e00b4028f82b2099906f6908e7cb88868e480

Observation d04d7704-ece2-451a-9042-66889085de75 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.889473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:737678bace59dc7494900dc4b693cfd8f7dd7f96ae8e3c380fff2831f2b8da96

Observation 1df1c459-d0b3-480a-a4b6-574ad46b64a8 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:00:25.727604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:2f45dea6ece434283abd4adfada714c69970c8cf96b90dab96c526b208880bfc

Observation 58b1c4e7-4b60-4776-8ec3-6ad76ebe60e9 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.126496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:723c67d339a402266aeac571c9f2b1ed98202f0f1603d23145ea580b6599c84b

Observation 1019c037-da64-4eca-8f24-c4e4c910e79c · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.938795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:1417ad50dbd0f4f52d67edf54bfa63121a5d7c8be4ea20be21754dd90cc57983

Observation 367341b0-89e6-42f2-b246-9aedb38f19ec · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:20.589681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:20.589681Z digest=sha256:e58dced30d9c28c115e27b13e95f7e48e625804ba62fee985b7d9ff8e49089b7

Observation b199610d-4032-4a37-9d45-19e25c4ca3f1 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.951038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.951038Z digest=sha256:c987fa671e897679d066482db6a828374eccc999f1cfa42dc5566beade84e156

Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.124792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.124792Z digest=sha256:4deab0b17bf9268c4f51c283710b6fb3db53ce8d06a041e962f59d3622be649f

Observation ee790206-4216-4ee4-a93c-76a684ebae5d · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.552117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:7e8b49cf5820a2f8e81342a211b9d10dbdad1209d3b40f637ddc9047f4517305

Observation 521868e3-1820-4574-98d7-0385c01efb93 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.910094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:8374a0d6042e3d5521c794e74eea3a47aaa2af43dde97d970bf3c03546caadfb

Observation 721e2ad1-8cdd-4d8c-918e-d53071e49527 · inbound

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization cites this paper.

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.336351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:28:16.315767Z digest=sha256:56137e70fbfe0b99e915233439871a83c72a694c31dbad962b65f6bfbf5e66ba

Observation 7744dddc-a0cb-46dc-9637-e394843eb390 · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:15.855625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:67befcdfe941deac6dfea3dac072c731bd035741bb6a175cb85a514d23ffbcb5

Observation 7223f609-576a-4723-92e8-4aff541737be · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:4b568c3dde8c82779f145b39098ea519f30fef7149b2d7448cfabc8529abfcca