Pith. sign in

Paper Citation Record · LEDGER

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.04423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04423 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.601708Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.732563Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31234611-dab3-4a1d-a3d0-a54352ddb84b · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601708Z digest=sha256:1de7674cc7986ccd7ded9dece7145f98f6eff5e30cb0faa6f4e5d54e47421ad4

Observation 4fb701a8-a7fa-4fd5-adfd-b33b36b63ab6 · inbound

R-Genie: Reasoning-Guided Generative Image Editing cites this paper.

R-Genie: Reasoning-Guided Generative Image Editing UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:35.925099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:35.925099Z digest=sha256:65f1af85ca5951ab4cd061ee6bd013a0f0ce6db7cfe85cf39390b777bb86a0e5

Observation 1ed1e2e4-6fd9-447c-89d2-e5d3ad74bb88 · inbound

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks cites this paper.

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:32.219643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:32.219643Z digest=sha256:9959e82958e92df75ed46e79d508cbc2b2d433bea0f35c84b4a7583ae86d4bd0

Observation b9e78131-bb3b-4edd-94c8-b901daf3346d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.736533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:9826640d92868fa80793eafd670369568a09858bbdb7e801db217f6a415d891d

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:05e6bdf2e87c807ba432686ca9dc503d587bc6bf7e19196f3b288ac0cd338524

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:ab267fd72a8275a18343b0c3775bcec6128632bf62cee7d506dcca470ada727e

Observation 5d1e0680-aea5-4e43-802b-bef60542ce56 · inbound

2D Gaussian Splatting with Semantic Alignment for Image Inpainting cites this paper.

2D Gaussian Splatting with Semantic Alignment for Image Inpainting UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:06:36.052885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:06:36.052885Z digest=sha256:4a0a94be1f5e9c95ef59e2186fda9afc4d7ceecaeedc8a9a375a37ac24b92f0b