Pith. sign in

Paper Citation Record · LEDGER

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2206.08916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.08916 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:44:58.818749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:44.037310Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a1f6a5bc-6a6b-4eff-9270-32a7c46fc751 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.215627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:14d8ca7b0b41f3c3006f2f35864b163ee5a51303699debdb8b203c2006932dff

Observation 5af69b53-b45f-4708-8630-98950a2d779f · inbound

Objaverse-XL: A Universe of 10M+ 3D Objects cites this paper.

Objaverse-XL: A Universe of 10M+ 3D Objects Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:02:11.638947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T13:02:11.512409Z digest=sha256:9b345941e44afb5007e2189ad018fb3eb170d349e9f7e06b435c1c9858692828

Observation ae304f91-df2e-44eb-85c5-d36eacc2c498 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.723093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:b0665970ba109f93c4218ab07e884252985ede075797bb4860e675642100cb3c

Observation 57d8c058-f261-4e64-bae7-caa00386527f · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:09:16.941469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:92d52bb916651a14ed919732ae06796bdd3b8b326fddb2d63df70c96d3f13c35

Observation 18821cef-6e58-4eb4-af69-e4f3a8c919e3 · inbound

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living cites this paper.

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:44:58.818749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:44:58.818749Z digest=sha256:e0a0f4471e5c99d61d9047cda64bdfb1f7d206dd7e2ef935c74db9b810264748

Observation 8a54fbf0-55d9-4f9f-b408-f2d3202d4874 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.262482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.262482Z digest=sha256:8e08d32256a05335500d5405ec692614fc079ba39afe92130401a3d4eeec69b6

Observation a974054f-a8f6-4b30-9fba-5ae051fc1008 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.204865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.204865Z digest=sha256:6249210b8166931020797738e35580d2550017fa90f1ae321d6ba486afadb0b0

Observation e685d759-ba8f-47d8-b3a2-b88bacd1eddc · inbound

LlamaSeg: Image Segmentation via Autoregressive Mask Generation cites this paper.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.397888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.397888Z digest=sha256:7c0ded33c9df7064e029513fc5ca52269a1b2bcc7415b45b3801f5f74562b9a1

Observation 083b6baa-5d7f-478c-88d3-a32dcd1b8eb4 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:01.991113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:01.991113Z digest=sha256:c5d46e9d24c790d63ef4928a4b1178f00610c6ca563b4c1c90f8e61c13197118

Observation 529958fe-3f15-4769-8f18-a9bcdd368ede · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.825255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.825255Z digest=sha256:cd90265402e20c5c3d0363e22fca63d691d7d373333e3df19cba1620aaabdc15

Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · inbound

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens cites this paper.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.091512Z digest=sha256:d577f656bf911929a56cf203555c2b38f3b215daa193cc209ca6a56e71b65980

Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.505625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.505625Z digest=sha256:a15fe782a7f63a158a587b747e6a40c34a74d83f1748f18bf01135ab7d6964b6

Observation 446505ae-d8ab-4536-8a30-77cdf3a0a24c · inbound

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation cites this paper.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.310361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.310361Z digest=sha256:f01b289b02cfa212e2c28fb2bd1ffd9680b16d04c6484e4b81dedcb092bb78b1

Observation 71529ebe-e13d-4b3c-9d1c-94d7a2f340da · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:51.884460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:51.884460Z digest=sha256:88427baadb4a93b6854940b8550cb7d6fcaca61f5cef06c243df1af35fad3afa

Observation 40c626f2-5ee3-4f25-95ef-a2428b0320ff · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.615891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.615891Z digest=sha256:dfeb391ee7cde55490c3b07b11db4353d7d418dcd6b1790b6bf5f8d03709b8f1

Observation 5c500d54-2f44-47a8-bf42-e6aab6e177ed · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:05.931398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:05.931398Z digest=sha256:092896d13a5f771a7ba67f6b46c387d16eca691475f4fc3d13ab80439a60a50d

Observation e546804f-2ced-4318-aeb4-ad53f3e54008 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.623864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.623864Z digest=sha256:11fa82e6a756975262c790ad2c251463e412cb42fc610ba88b9888084c709f25

Observation 3d48944d-23b5-4973-ace7-88db099cecdf · inbound

Open-set Cross Modal Generalization via Multimodal Unified Representation cites this paper.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.595479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.595479Z digest=sha256:64a996e18bb8ecda4619c680934b72f38b921e2910759178fd8668d304227045

Observation 86f3f575-5e5f-445e-8920-db1c4aa906ec · inbound

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding cites this paper.

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:18:43.931214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:14:56.268778Z digest=sha256:3f9f70f0c2358d6fb282e7639ae3997af98a700ce7caf073dbff5f99ca583739

Observation 4414eaf0-8606-4f5f-90d1-088fd1508bb6 · inbound

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training cites this paper.

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-03T07:53:48.871339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:53:48.871339Z digest=sha256:172135191f219348499d52943ff2084852b5f247362f432e50e23a2a884c1cc4

Observation de1818f7-d813-403e-b2d3-770f40479d4f · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:02:19.639749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:a21f9379c2e3e1e83d8d0e4b494224f66a413488519a4bfe44102957d3052fb5

Observation 12a1438e-57bd-4ec8-b060-cf3f501a0141 · inbound

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models cites this paper.

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:09.717580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:12:14.321695Z digest=sha256:fd5a982c46d9c8a8eea2d541e55c26bf956572c2ed75610439c761ad8d8dfe65

Observation baf255d8-f9fd-46ca-8ed8-7e50cc722152 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.038647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:d7583a2093cc0c5e6dd0c0c5957c7a0cf7de3ec178fe806aa075f50349497ff1

Observation bde73f9d-46ea-42af-ba62-e49c096c564c · inbound

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation cites this paper.

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:38.368478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T16:22:31.982484Z digest=sha256:8bad0ebe7185641c8f647d03e5f15a9c7780cffab7034227ba5aa9de35b9151f

Observation 564e6f96-593b-4fac-a887-9ee827fb2ff0 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:93637c7213ce2f38899668dff3ad810b1cdb42fe79cfd6b747e08410c7669f86

Observation 44a4d8c8-6f4d-4fcb-8ebf-b5d8d4230cd6 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:1f150c6a5996b67b08222d3758b05e9ca297ed38832a974ab2acd3f8c88bb44c

Observation f98a7a04-3648-44f8-889e-ca36a41f7dc5 · inbound

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features cites this paper.

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:40:55.041093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:40:55.041093Z digest=sha256:ef65c8a135620238aa94d5f7ab5f52ae57e23ddc150eabc6d0c9d52365f246d6

Observation 5792015c-7b09-4844-ad46-4aa3ae0fcff8 · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.401226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.401226Z digest=sha256:526322cbdfb7a1ea631707ae9b179e768843dcc260893add246a4a7602e35ba6