Pith. sign in

Paper Citation Record · LEDGER

Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.06109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.06109 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T06:01:13.400059Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:15:47.123482Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4ef0877b-8a2d-4ce0-bcb8-f8cb00d2e87e · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.746237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:8b0f12872daa56e38756859ee99341830bcb66c73b81ef0ac3ad1de1ebb0946a

Observation 3cb97303-4f3a-4cb8-9c4d-7f55523da5bb · inbound

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model cites this paper.

MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:27:52.114785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T15:27:51.839171Z digest=sha256:a6e7a84ad8dbe6ea2fd54c287e72ad70970f9c3d031de1a202271dfbe3b0815d

Observation df5d7acb-9b2d-4270-826e-e8951235dad4 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.251278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:3c3bb36e26222e5076e0462eeb217312178ffb5bd36f33653dbdf1f5b802d8d6

Observation 8cc8cd89-b9b6-4665-bf4b-548ff47a3707 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.764945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:2840abd8ae1e293a6f7294e41d6378750788553d4f2ac0915fe77b13c4708c6a

Observation d04d7704-ece2-451a-9042-66889085de75 · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.889473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:40d4ebff1dcf4f79b72f4f440768173db6a2cf09d59c6e541f589cc2291f1a87

Observation 1df1c459-d0b3-480a-a4b6-574ad46b64a8 · inbound

MinerU: An Open-Source Solution for Precise Document Content Extraction cites this paper.

MinerU: An Open-Source Solution for Precise Document Content Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:00:25.727604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T04:00:25.624430Z digest=sha256:477db50412d00e85030d26dc79b2e50fe9eb6521d8d920220928a10b52df51ea

Observation 58b1c4e7-4b60-4776-8ec3-6ad76ebe60e9 · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 254

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.126496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:a9e57abbd6a2f439a9bd4aa3b87b65d4a9b76a9ffb9b149be63c126576dcf488

Observation 1019c037-da64-4eca-8f24-c4e4c910e79c · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.938795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:0ed4c994b87c481f4e445578359d3f84a05c7e5dd504bdd7b7659f9c896630fa

Observation 3436fd0d-5181-4763-9b4c-c52b886c7a6f · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.400059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.400059Z digest=sha256:ae1104c9a3ba4eb5b406b1bcddc935674418ab7d7720b6ad4f740c33f1ef2b89

Observation 367341b0-89e6-42f2-b246-9aedb38f19ec · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:20.589681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:20.589681Z digest=sha256:95e3aa0068f8a49387145b9873a8dd97f66972d06c43c656a37eb5fe8a2cfde7

Observation b199610d-4032-4a37-9d45-19e25c4ca3f1 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:04.951038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:04.951038Z digest=sha256:e26d25bd3c1117cd73b2ccae65cdd0b453e2be722d23494979e4e948d3a8d43c

Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.124792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.124792Z digest=sha256:4deab0b17bf9268c4f51c283710b6fb3db53ce8d06a041e962f59d3622be649f

Observation ee790206-4216-4ee4-a93c-76a684ebae5d · inbound

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR cites this paper.

FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:20:11.552117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T20:19:52.701262Z digest=sha256:b9d0f037b7008dd212a00e5b31e09deb71c3a027a3e7416575ed2c70a7cc8a28

Observation 521868e3-1820-4574-98d7-0385c01efb93 · inbound

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models cites this paper.

Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:48:41.910094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T23:47:08.562575Z digest=sha256:8c38a67a0d70408fad645e46cfd1a9216df2eb48a734cb928a5ccbf0a77e134e

Observation 721e2ad1-8cdd-4d8c-918e-d53071e49527 · inbound

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization cites this paper.

Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.336351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:28:16.315767Z digest=sha256:258c8d512cad5c51ecf286d9b0e571d01f9b23adf6cf5fbf0174319539b79a9c

Observation 7744dddc-a0cb-46dc-9637-e394843eb390 · inbound

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction cites this paper.

ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:15.855625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:34:56.032634Z digest=sha256:79248da8394d331323e594ebcf5514fe42dfbf7229f6557b16d24aeecd96b01e

Observation 7223f609-576a-4723-92e8-4aff541737be · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:4b568c3dde8c82779f145b39098ea519f30fef7149b2d7448cfabc8529abfcca