Pith. sign in

Paper Citation Record · LEDGER

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.17172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17172 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:43.975806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.684545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:00bf141d827fb1b5bcff4779fed4b699e897c47b03e5173a613caa05305f07ef

Observation 5eecd7f3-3e5e-47de-9c20-db37a670b5de · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.134116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:cb43d6079d1fdac763638bd5e8470e20ca2a8aa24d725bfcb9208bff007b82f9

Observation 5d5281de-347b-4e2b-b2d5-9b7b6ef518cb · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:03:28.141894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:02a38f8a0273ac2c78ef2d28e84c11d6bdf4472709230776c1b7df659ba8b2fb

Observation 4e97cf76-c5ff-429f-a534-2f068825d26a · inbound

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology cites this paper.

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.798726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T00:20:31.729964Z digest=sha256:b88fabe4a22df57c027779e1c33e1d5813a7f272bfc0fab4be2118738c77fc20

Observation 4d64edaf-409b-4112-821c-a87bb58ef97e · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.405274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:04311550bcdb89172da691d8785b372f2ae75206428ac6a46983e1e6be50054c

Observation 35db2fdd-d625-4dea-8c92-b57aef434d7b · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.629194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:dc25760eaf6c784d31b633a33b67c056ed4dd1d669037d614b16dab4d2a52582

Observation 13875ab8-b8a4-4864-9934-86a0932c7482 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:43.975806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:43.975806Z digest=sha256:65248890b5cdeaed8f96124507afd5c240cd3dc905674b22d30d0614ef817ce8

Observation cf89c0f2-b74a-44b5-9376-208df931765b · inbound

One Diffusion to Generate Them All cites this paper.

One Diffusion to Generate Them All Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:54.679704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:54.679704Z digest=sha256:5c666b4c16a10a6b373af697131952aa2811ff08075279e09bcf3573e6049caf

Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.385354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.385354Z digest=sha256:53b639ae78a4eca7850f74efdb119a41583a227d63cfa2a633fe80e08fe8fe3a

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:01050ba303e289047a6dc66c2c5e2531db3354142635447046390e983871ea42

Observation 237b7a46-057a-4fcd-8767-6c3bab50669e · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.958991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.958991Z digest=sha256:bc9e68269c6a295e1b67c04d8e63f1cac6d857e4c5ced8bc7c17b29ecba4676a

Observation ce4e3727-b7e0-45e7-b3ad-69a5353cf09e · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.051512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.051512Z digest=sha256:5c1f0413c656e1090caa7fb48d2fc039d9863c04e28569f911e6a963b22eac3e

Observation 31682340-93b8-4c4c-83fc-ac2dd32318cc · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 276

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.354880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.354880Z digest=sha256:800ffc74cf3eadfdd4fcbb68a22a615b1e50d4789ee7e202a2f241ba6c2fb21f

Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.060895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.060895Z digest=sha256:01b7aa41a44e572be82c47bf5721cf554a696817ad5816c21a1ff5ffac775024

Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.170183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.170183Z digest=sha256:5d725796ced3e86c421c03041816a54b58926f90e9f431fcc961eedab01a8033

Observation 49637e00-5f91-43ca-929f-2b7bf399ef2e · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.073781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:199288d254c2008b0c0ca973c01f70fa47ce3eb08ae8ee764865c2b52a44b1b8

Observation 2e351de5-1f16-41ee-ba07-365d0832fabb · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.545721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.545721Z digest=sha256:005ebabe75468aea5fae79eb160c856e41ad22fcd5d80c81c11c8f792f7a6431

Observation d8b909db-3d64-4012-8b9b-39ba0cbf3a0f · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.304736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:1da2246cf4c8789e68613eaaaeaf0f45dfe3faae13d85ca5c6f754ff33e86e3d

Observation b5181804-620d-4b7b-a402-dd4c8fec86f9 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.290588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:174e63dbd8490c27998b4adecdf5c42ec8be436effaf625a08fda7401681c0a1

Observation e23884c3-5ce4-479e-8c1d-6ff67692997b · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:1189e23b5724eeb8cc97bbb098efd6fed6277ac306519db16de6a2796c7deb07