Pith. sign in

Paper Citation Record · LEDGER

CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.10462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10462 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:02:19.589481Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T20:55:04.590393Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 37bfe57f-34a9-4846-81af-3f372d66cc03 · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:19.589481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:19.589481Z digest=sha256:e29b0617fc9f5ff6dcdeaa06b2e255a27cb33a9a086cb634b92c83e860d36506

Observation df1f0d85-c592-480c-b6ec-472cdc934e69 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.638404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:fc6f5371053a3681615024820fd878fe787c8c367aafc9f263c5cde9732f356f

Observation 5ae716e0-f550-4984-a7e7-c66ce3d745bf · inbound

CI-VID: A Coherent Interleaved Text-Video Dataset cites this paper.

CI-VID: A Coherent Interleaved Text-Video Dataset CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:54.222341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:54.222341Z digest=sha256:aae54406e0e639db3a39d9f322a007f0984133a57856c71ef2c0aac0c562baa5

Observation 10ef18b9-ce55-42a5-9e42-06e8c46ad0dc · inbound

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts cites this paper.

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:41:26.760116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T09:47:27.891288Z digest=sha256:34936f277d1fc7ff4a768cddce6b9285f8abd41b4956ac9428bb310b8b488188

Observation 916736f9-cbb1-4c1c-bd0e-abf0a03e2611 · inbound

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts cites this paper.

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:17:59.275253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:15:30.952247Z digest=sha256:cb1e22a9250da04a14c89eb24934f384d28c79b1e169280d29e1482eb0ea9bbe

Observation c47cd7de-b9ad-4a15-9df7-61f39115f6b1 · inbound

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA cites this paper.

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA CoMM: A Coherent Interleaved Image-Text Dataset for Multimodal Understanding and Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.591960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T20:46:54.667418Z digest=sha256:48747816bf8c9029c0bafba07769353d4fde608bc6fb1f6a85aec95de2661b53