Pith. sign in

Paper Citation Record · LEDGER

UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2504.04423.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.04423 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:53.601708Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.732563Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 31234611-dab3-4a1d-a3d0-a54352ddb84b · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601708Z digest=sha256:8ea47fd873125e420393b9bb97330931442705b6117bf33118c5b386e944ba84

Observation 4fb701a8-a7fa-4fd5-adfd-b33b36b63ab6 · inbound

R-Genie: Reasoning-Guided Generative Image Editing cites this paper.

R-Genie: Reasoning-Guided Generative Image Editing UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:35.925099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:35.925099Z digest=sha256:6e3bcdb9cb62f277560d775410df98cf8ff01cb11c83a0662dfce368c1422fc7

Observation 1ed1e2e4-6fd9-447c-89d2-e5d3ad74bb88 · inbound

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks cites this paper.

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:32.219643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:32.219643Z digest=sha256:2a411e16e973a7abf6e291af7f2b8055eb74aeadde4869c9d1e4f1eb53a69144

Observation b9e78131-bb3b-4edd-94c8-b901daf3346d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.736533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:ecb24c7e31bed3247c1f6cdb7a164cb3fa0b4df2079ce3007266648b392fb276

Observation 185a68ec-e17a-4264-a362-0c073d654f6d · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.667271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.667271Z digest=sha256:fc587cf7c064c5a6da297f15724dda471dedb166af9ded170fe114e297cb7ee5

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · inbound

NeoBabel: A Multilingual Open Tower for Visual Generation cites this paper.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:8ba6fd79b15792c9b8b306fc045fe79f2ba3d1f9ec68111316ad457a0fbb7627

Observation 5d1e0680-aea5-4e43-802b-bef60542ce56 · inbound

2D Gaussian Splatting with Semantic Alignment for Image Inpainting cites this paper.

2D Gaussian Splatting with Semantic Alignment for Image Inpainting UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:06:36.052885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:06:36.052885Z digest=sha256:2c75df9d607a5294dbc66969d649b7b6d1cfeb835fc58556f4117e9432a04668