Pith. sign in

Paper Citation Record · LEDGER

Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2304.06939.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.06939 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:19.293396Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:26:17.900939Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1a8c733-b4fb-417e-8c6c-509b6f930665 · inbound

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality cites this paper.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-24T09:04:15.092474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:c9f571c3892ac5418839a01d8d8ad79ed1b5332c8a59c9c396cbf0e47731ed4f

Observation ce814b00-ab6d-48e6-81f0-0f66a3411612 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.932640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:8b10b938836d9a4feb82a2c7c68d631aed056eeb00f509c41b46ccc4393f52b7

Observation 83655539-7804-4fc7-a7dd-c87582ab9c46 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.700766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:f440800d9e94ad4f513deb7d8f611d73db65debeecac8e0ddfca77cb1182229b

Observation 00094086-cbec-421b-98d8-166b7721447f · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:52:01.385739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:69823011c5a1f461a952cbc27d70d4fcac224be2a097c3910c6ff41fde54e0b7

Observation c1907c8b-6166-4e65-8e0b-6357b6777362 · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.745647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:b607de304be4fdd9f191db23042658e4593e72751ae3ad752fed2d385696aad4

Observation 8b267e48-0291-4722-8ad5-92c0ee025391 · inbound

Demystifying CLIP Data cites this paper.

Demystifying CLIP Data Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:20:20.241274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-16T09:20:20.143143Z digest=sha256:a3609e755c926017dfb592c5183bd92709e631483ceec5cd07ec2b5ade3ad3d0

Observation e5128dea-3a0d-4b63-83f8-af056b50585b · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.123983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:713ff928dda034d21ba8e45fb4002b16c3b59a15f2a42f0c85c19246367668c4

Observation 78b15180-5f55-4b42-9b42-fd794e3c1335 · inbound

ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval cites this paper.

ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.825940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-23T18:29:18.466941Z digest=sha256:adc141a69791ea46d09142db54a28a29b454dd3ece7b829dba5a638f360b26de

Observation 0489dc65-fdfd-472b-9299-b388e7b33607 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.293396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.293396Z digest=sha256:587a948b958d5d6535e67fadafa3dbbf99136f1963c21028f0fde969188735db

Observation 6ff885fa-82d4-455c-a507-07dc5c9787cd · inbound

Nature Language Model: Deciphering the Language of Nature for Scientific Discovery cites this paper.

Nature Language Model: Deciphering the Language of Nature for Scientific Discovery Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T12:31:52.057606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:31:52.057606Z digest=sha256:092be83171a8d8f93aa0a3c5143a32220f555626663e0e3aabf4a894d5f54d6d

Observation f7ca8001-b5f3-4738-bb41-372f06b969e9 · inbound

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling cites this paper.

StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:37:06.458067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:37:06.458067Z digest=sha256:f7ad1dd699cf0589304f4bab1037cfbd9dd894785fcec191291a88784fe04eff

Observation 2b5a7845-eb60-4c4a-a109-8d51fbb49255 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:54.830726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:54.830726Z digest=sha256:db16d76d8c86b33f744bf4c51df359337b03208e5cd51cd46ac9e2745ef834a1

Observation 922ea791-7568-45c2-a353-09ef6d0314ad · inbound

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding cites this paper.

ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:32.927393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T22:39:06.113655Z digest=sha256:2604b0d13294a0fd1e80d56568f81c11a59e9b0d72f825d2adfe03fc87d99f08

Observation 2a8080ea-7ead-49f4-951a-d2b9dba40426 · inbound

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence cites this paper.

CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 58

Resolution
malformed identifier
arxiv_id, observed 2026-05-14T20:37:58.324569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T20:37:36.144960Z digest=sha256:2b97ea6f4b01050b2679ef0a9d6241d5d2bc5e0ac6b1f1407e81af7ee90d650c

Observation 6453e69f-deaf-40f8-811f-7358abe13d2d · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention Multimodal C4: An Open, Billion-scale Corpus of Images Interleaved with Text

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.902355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:999ea6f6c2e844200700b30aff647671755f80769c925d93a51257eb99b3936a