Pith. sign in

Paper Citation Record · LEDGER

DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2405.15232.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.15232 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:31.056352Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T18:35:00.270516Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08d5a66c-e4b3-448c-8329-64670617a31b · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T00:26:21.441735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:e9084ef3a153a308e4b80fee2461b8a1bb58a5251eb5117f6a1ad31a8e03b3f3

Observation 87e41b1d-d32c-4e1c-af7f-a153a7ead273 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:31.056352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:31.056352Z digest=sha256:75c92e3f9d5a282399db80b7317da9113a5e8ec719bd35cf43a823ac73453232

Observation f9361d6e-977f-4ec6-ac14-592518cc7e10 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.066201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.066201Z digest=sha256:ceab5c18f6c62f27f7532c8a19b4497b69ef1c07b9ff1abf8e3304a51b2d9ae0

Observation 9a9d91d1-ddca-4c88-ba94-e0a95ce3b47e · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.227597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:0f9dbc69d12f466b3a39f360780412ef2ca920695835fe6ddcfac7e4a8095c2d

Observation 1545c2a8-dd4f-4398-8250-d4b01290cf60 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.271961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:14b8ccabf4de1c355a69ecad900a6bca4002643cf0ba1d9a0fc751c4e2093089