Pith. sign in

Paper Citation Record · LEDGER

Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2503.21979.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.21979 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:04.231098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a80dd079-0f4a-4332-859b-5b1d5949f4c9 · inbound

DiSA: Diffusion Step Annealing in Autoregressive Image Generation cites this paper.

DiSA: Diffusion Step Annealing in Autoregressive Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.231098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.231098Z digest=sha256:1a45ef5c23c1f1fcec1c971d3cdfcd31151955d8a981622fc5a71707729f3a82

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:6035b673524d75d94bc59b7d1965bdc18286f0089851901e74c3cb670d5c6972

Observation 58a09def-c607-40c0-b40f-9ea28861d1b1 · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.859294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.859294Z digest=sha256:609e46007cb012229381994ea6e0ae15ad568f82b9460285e1430df3211abcd4

Observation e75b459d-78b3-443f-a894-9270dc10d264 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.248358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.248358Z digest=sha256:e62171a0b3fbb8546163eb41a2dffe00f408db25ffaaa6de855de62b5f9e3ce8

Observation 9f7a7664-068b-480c-a038-23757b08a911 · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:52:59.481610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T16:51:48.705876Z digest=sha256:887657df44f0fb3bffb81c1dbcfd0be00ed1a61ddc52c5dc40f6dd20fabc8e07

Observation 1149a9a4-a285-4b5d-b66c-f5aceed07d9c · inbound

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning cites this paper.

FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T12:10:53.720348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:10:53.720348Z digest=sha256:9d79a763c072efb83c0874842f20b8818fc909a566c67b8a71deb97202c178b4

Observation d00a1655-ebce-4abc-ba91-ce65d35f17d2 · inbound

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation cites this paper.

MAR-GRPO: Stabilized GRPO for AR-diffusion Hybrid Image Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:40:58.081884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:00:50.105629Z digest=sha256:3bc0f8ef74505f1e1384f33e1bc9ac551cc343f4cbdc51ddc00d5f13eb38fc01

Observation f4871b2c-0371-4777-8a3d-eedb6c0ba895 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.277057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:0b0972e7d82ef9a2919e33afe43edfefbd1385563d6eb7c8dab43adf017f821e

Observation 02848f04-40c5-4cc7-be1e-6f6c2432fe1b · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.458005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:2e78cd39f98fd57ef9d8fd77b768fc777a9dc60105e930d2bf3cf4add481c703

Observation 3af64175-1c8e-4cb9-9727-eba43ed6d6fd · inbound

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection cites this paper.

UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:44:15.287320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T22:21:10.133655Z digest=sha256:d28bb1d20ab8609cd0955b059d9331e42aca51a031d82c019f75aa1b7b8eb785

Observation b6af3b54-1388-4e76-b3c0-4b9ffa81aff5 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.714030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:f2e1f4b9bc0518b61a7ab6b328557a79c9ff8bf094c7107ca47d264a9d2ae460

Observation 40a9a53b-5dbf-41d4-bd34-b9cde05af524 · inbound

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens cites this paper.

WinTok: A Win-Win Hybrid Tokenizer via Decomposing Visual Understanding and Generation with Transferable Tokens Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:08:15.793520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T12:04:19.761430Z digest=sha256:27fd4d4c2de37f599ac9c562d9b151f81cee8a7aa32e50e562fb4a956914b46b

Observation 879ca5e2-fb28-429b-9c95-a44b0ce6226d · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.437930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:33c517834c020345c83b3444c4a93464f72abb9a72e46f5bbd0c9199731e458d

Observation f5913998-4216-4ce2-8925-878a650dcab1 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.881233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:ccce7fa5fe22997e1210065ca17ffbd9b4566adc0dedd1f1986dcb150bc7c53f

Observation f80ece7d-474a-4a7f-849e-0cf729c50f3b · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.677719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:8776773ecb8ef9b7872db00c5399c43c0c2191d97ee5d2c4394896009e862c9e

Observation 48829565-1a03-459b-a7f2-600152fa830a · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.477984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:fbd1195086459f70467243d944868bf703bb035c4da5c389c1864c506278aa58

Observation 02f63038-6940-4cd3-af38-483d52477528 · inbound

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation cites this paper.

UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:09:57.559037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T00:50:22.839005Z digest=sha256:c0f834581e2b32cdca528b0b8ff468062606934e2cb59ca64463dcbc5649b540

Observation 4784d571-1a55-4630-99cc-5d6dea84cfba · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 182

Resolution
unresolved
no resolver link, observed 2026-08-01T04:30:03.046647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:30:03.046647Z digest=sha256:ed23bc6b36609a7164ea76b7b0fbb90dcf1f215dcd48b672db71d9cf6965d6de

Observation e794fb98-3746-4095-a919-f93c121b53c6 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.791253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.791253Z digest=sha256:2eaa7b1deb51da6d62bf7dc51800995836dea09fb99726bdb6e494956e030de9