Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2505.05422.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:51:05.377160Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.252935Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation cda058a8-54fd-4dc0-8247-e60690e21de7 · inbound
Show-o2: Improved Native Unified Multimodal Models TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3e3e0be-b8e7-4a3a-a3bd-34492836fef8 · inbound
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96348008-d968-42c8-8185-94a74ee37817 · inbound
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3a157326-a4d9-4ee4-a2bd-2f5483acdd24 · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 999349fb-97b6-42a4-8ecd-1f013600b7a1 · inbound
InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c48aaa9a-a0ea-4bea-ade8-504405aa3430 · inbound
WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c99b5ab4-0b59-49da-9c7c-ab35e12fd761 · inbound
ProductWebGen: Benchmarking Multimodal Product Webpage Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4af6f729-a158-4b93-9c62-6c1d8652f04c · inbound
Diffusing in the Right Space: A Systematic Study of Latent Diffusability TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6ee70d06-373a-4d63-aed5-41997f33e9c9 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 297
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6afbaf8d-4465-4986-90e9-94d5a2477fab · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 917a0248-2953-4160-a918-a61d19129d2a · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55211188-9943-41b8-963e-43475f5c3f7d · inbound
UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f247a7e-2267-4637-97d6-c2db6db2f4c5 · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4601cdfb-62c4-4bba-a418-da37901b710f · inbound
GroupVideo: Multi-Identity Customized Text-to-Video Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6527d101-63ed-4104-a6c3-74ef08f71b72 · inbound
dRAE: Representation Autoencoder with Hyper-Spherical Codes TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6858f2-cb07-4a73-bf3e-11d5b8722d22 · inbound
Twins: Learn to Predict Unified Representations with Focal Loss TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4112da53-a8be-491f-b95c-e48894b0801c · inbound
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.