Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2312.06109.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:43:20.589681Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-23T19:15:47.123482Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 4ef0877b-8a2d-4ce0-bcb8-f8cb00d2e87e · inbound
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3cb97303-4f3a-4cb8-9c4d-7f55523da5bb · inbound
MobileVLM V2: Faster and Stronger Baseline for Vision Language Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df5d7acb-9b2d-4270-826e-e8951235dad4 · inbound
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 121
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8cc8cd89-b9b6-4665-bf4b-548ff47a3707 · inbound
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d04d7704-ece2-451a-9042-66889085de75 · inbound
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1df1c459-d0b3-480a-a4b6-574ad46b64a8 · inbound
MinerU: An Open-Source Solution for Precise Document Content Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58b1c4e7-4b60-4776-8ec3-6ad76ebe60e9 · inbound
Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 254
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1019c037-da64-4eca-8f24-c4e4c910e79c · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 367341b0-89e6-42f2-b246-9aedb38f19ec · inbound
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b199610d-4032-4a37-9d45-19e25c4ca3f1 · inbound
Docopilot: Improving Multimodal Models for Document-Level Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 069b9a4b-b512-452b-ab20-7cb3fc479039 · inbound
Region-Level Context-Aware Multimodal Understanding Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee790206-4216-4ee4-a93c-76a684ebae5d · inbound
FinCriticalED: A Visual Benchmark for Financial Fact-Level OCR Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 521868e3-1820-4574-98d7-0385c01efb93 · inbound
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 721e2ad1-8cdd-4d8c-918e-d53071e49527 · inbound
Improving Layout Representation Learning Across Inconsistently Annotated Datasets via Agentic Harmonization Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7744dddc-a0cb-46dc-9637-e394843eb390 · inbound
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7223f609-576a-4723-92e8-4aff541737be · inbound
Mixture of Cognitive Experts in Large Vision-Language Models Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.