Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.17172.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:43.975806Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · inbound
BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5eecd7f3-3e5e-47de-9c20-db37a670b5de · inbound
SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5d5281de-347b-4e2b-b2d5-9b7b6ef518cb · inbound
Chameleon: Mixed-Modal Early-Fusion Foundation Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4e97cf76-c5ff-429f-a534-2f068825d26a · inbound
Learning Spatial-Preserving Hierarchical Representations for Digital Pathology Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4d64edaf-409b-4112-821c-a87bb58ef97e · inbound
MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 35db2fdd-d625-4dea-8c92-b57aef434d7b · inbound
PaliGemma: A versatile 3B VLM for transfer Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 13875ab8-b8a4-4864-9934-86a0932c7482 · inbound
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf89c0f2-b74a-44b5-9376-208df931765b · inbound
One Diffusion to Generate Them All Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · inbound
Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · inbound
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 237b7a46-057a-4fcd-8767-6c3bab50669e · inbound
MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce4e3727-b7e0-45e7-b3ad-69a5353cf09e · inbound
Next Patch Prediction for Autoregressive Visual Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31682340-93b8-4c4c-83fc-ac2dd32318cc · inbound
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 276
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · inbound
Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49637e00-5f91-43ca-929f-2b7bf399ef2e · inbound
Show-o2: Improved Native Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2e351de5-1f16-41ee-ba07-365d0832fabb · inbound
Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8b909db-3d64-4012-8b9b-39ba0cbf3a0f · inbound
Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b5181804-620d-4b7b-a402-dd4c8fec86f9 · inbound
Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e23884c3-5ce4-479e-8c1d-6ff67692997b · inbound
MentalThink: Shaping Thoughts in Mental SVG World Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.