Pith. sign in

Paper Citation Record · LEDGER

CM3: A Causal Masked Multimodal Model of the Internet

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2201.07520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.07520 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:54.303566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:08:24.801483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d129a54b-f7d7-4542-a8a0-6229153b06d7 · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:21:20.486086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:801754ccd3d8b30a7a1faa3402e6bcb9e91270302915840154e4fa4974b82b5d

Observation 564519e0-9d5d-4f56-8372-9c894ae2a97e · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.797952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:ffa6bedf15cf16f563e6f642af71146c235013fd7534e95f8cc84864dbf853f1

Observation 70778e2e-984c-40f1-9fb0-0c0b8d3eead3 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.133071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:3e339f4ca42e9306fc895aa024d6582d9d41cb6dee4ef30a69aa463be3c46930

Observation 42c24c9a-c0c3-4887-b7ae-353daa2905a0 · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle CM3: A Causal Masked Multimodal Model of the Internet

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:41.849849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:25517bc9dab9fb9d26d3d00f344c1e031a6bcfd7ff412c58e0aa54a908b96b01

Observation b95b0701-e95d-4948-8fa1-fc68dfabfa53 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:32:22.841609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:df85c9d44f364993de8543d71c60edbf11ec9f6218cceecb9c38573ddaba958f

Observation d24217c4-50ca-44bc-9ef2-85d3a12cc235 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:19:47.966474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:b54e26d6bab6572ee22084b3074f54955bd946071cdcf79bba7e2193c28eee4a

Observation afdeb8dc-7128-4bcd-adf4-67ee7f1c9560 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:52:01.207177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:fc8eb967607013e8ce2d8e89889a8651630b02669153995be299e341aab55097

Observation e6703626-86a0-4271-b486-0f31791b1518 · inbound

Finite Scalar Quantization: VQ-VAE Made Simple cites this paper.

Finite Scalar Quantization: VQ-VAE Made Simple CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:34:13.504879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T18:34:13.439993Z digest=sha256:013707790c170225a33eb97511468c6b486596c9f24f8aebd8a27ce6084d40b1

Observation a5d80f13-f7cd-4cb9-bf6a-28ab406f9b3e · inbound

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models cites this paper.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:21.394711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T23:00:20.720030Z digest=sha256:134b41790d9ae007db5973bcc3e74622258fdf91b8eddd51a8e7dee8f6632d55

Observation 7891e6dd-3675-4d72-a23a-b2c494a14419 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:03:28.160390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:e05f70136cd1dc89dd4e325d865ecf569d1a1415e7d448f18b303a519223d984

Observation d9e38b52-8ad4-4e11-b812-b781251217c3 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer CM3: A Causal Masked Multimodal Model of the Internet

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.375359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:9597e807aed895255d7564b0f747e800e2f495880d57e7fa6a6bd39ff23bb036

Observation 2b80459e-e98f-41cc-a8ce-8ffe572ed3b5 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:48:45.137735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:73d9d6ed8164f9e3d9e28682f49b9c0c4ce6fdb1ffafcfae7b1f5bc4c3b5714d

Observation 291d50e2-3c65-4e31-852d-384141d1ddeb · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning CM3: A Causal Masked Multimodal Model of the Internet

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.302427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:ffa5d4a60212789cf13bd2c8ade83a06831e32cd52033e37309e526206b0cb5d

Observation abc4ea03-17c9-4738-936f-47ed169a939d · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CM3: A Causal Masked Multimodal Model of the Internet

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.440871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.440871Z digest=sha256:1cc1ddf62c74c24ea93aadb060112b9502c092807bbb18e3d0f08cf295efef8f

Observation 9131fc63-07e3-43cb-bcc0-cad32f8bc346 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction CM3: A Causal Masked Multimodal Model of the Internet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.735481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.735481Z digest=sha256:4f8921d325c3a7224a256edf8d66ccdbc740765c311f0112ffb98c4ae7797224

Observation 3b459ec0-0e40-42a1-b02b-2887ec467367 · inbound

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning cites this paper.

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning CM3: A Causal Masked Multimodal Model of the Internet

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.803385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T15:06:53.377360Z digest=sha256:c6d6a1fad580a05bdf326ffb25adfaf99f92c2f4617a3ca5083cbd3e6958528e

Observation 869e90c3-232d-47c3-8cf0-af645cc7d399 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.802161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.802161Z digest=sha256:986162202809430c24a210b29cfb4e0d9af34b3c339f879c40f788259220b5db

Observation 63df0a23-6c2f-4179-a6b5-57144d3d1c24 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.303566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.303566Z digest=sha256:3eb47829293e125632f0cac430ee7a472869d622d096a7a598eebd60c94d9273