Pith. sign in

Paper Citation Record · LEDGER

CM3: A Causal Masked Multimodal Model of the Internet

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2201.07520.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2201.07520 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:54.303566Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T15:08:24.801483Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d129a54b-f7d7-4542-a8a0-6229153b06d7 · inbound

InCoder: A Generative Model for Code Infilling and Synthesis cites this paper.

InCoder: A Generative Model for Code Infilling and Synthesis CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T02:21:20.486086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:21:20.438666Z digest=sha256:3da714ee234f3b9ea33adb7a2e1125e9a11413bf1964ec2ee1a7c41339e19c96

Observation 564519e0-9d5d-4f56-8372-9c894ae2a97e · inbound

Hierarchical Text-Conditional Image Generation with CLIP Latents cites this paper.

Hierarchical Text-Conditional Image Generation with CLIP Latents CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T16:55:57.797952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:55:57.612364Z digest=sha256:67083cc8800422eec9d757746bae5889c6be40d4f72e49b1a09955f22ce6ff05

Observation 70778e2e-984c-40f1-9fb0-0c0b8d3eead3 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T04:22:30.133071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:4850b0c53a2082d21da05c9717e42ad65cc32d6c57e31bde4eca4347e579ac7d

Observation 42c24c9a-c0c3-4887-b7ae-353daa2905a0 · inbound

Efficient Training of Language Models to Fill in the Middle cites this paper.

Efficient Training of Language Models to Fill in the Middle CM3: A Causal Masked Multimodal Model of the Internet

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:41.849849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T00:40:41.647820Z digest=sha256:f18208666cc40f0e0f1ea80015127152b1b2dbce6caad41ac73c7e8df9cb3eaf

Observation b95b0701-e95d-4948-8fa1-fc68dfabfa53 · inbound

Language Is Not All You Need: Aligning Perception with Language Models cites this paper.

Language Is Not All You Need: Aligning Perception with Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T18:32:22.841609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T18:32:22.813668Z digest=sha256:21c77bcd2e31af651afd57ac42eb77ec6a962d5886b7ef5dfc615f51898d16c0

Observation d24217c4-50ca-44bc-9ef2-85d3a12cc235 · inbound

Kosmos-2: Grounding Multimodal Large Language Models to the World cites this paper.

Kosmos-2: Grounding Multimodal Large Language Models to the World CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:19:47.966474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:19:47.907355Z digest=sha256:606b2df2838cb3d189fd040ab0302c4824f8f044c1cbc46e849a96a3074d8aa8

Observation afdeb8dc-7128-4bcd-adf4-67ee7f1c9560 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:52:01.207177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:d0eab32310165ef8c3a06593e7202a0803935b77e6bc3c3d170fb71e1a5d5426

Observation e6703626-86a0-4271-b486-0f31791b1518 · inbound

Finite Scalar Quantization: VQ-VAE Made Simple cites this paper.

Finite Scalar Quantization: VQ-VAE Made Simple CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T18:34:13.504879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T18:34:13.439993Z digest=sha256:acaf7a4331972d3bdf4f4d43706c936a431e1b7a780901fbabcb016d15d93b1d

Observation a5d80f13-f7cd-4cb9-bf6a-28ab406f9b3e · inbound

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models cites this paper.

Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 290

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:00:21.394711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:00:20.720030Z digest=sha256:87ab2f50a2788813b92bb33df5a68773f2545356b3c8bbad536e2dbc38f0bec5

Observation 7891e6dd-3675-4d72-a23a-b2c494a14419 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:03:28.160390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:cfb6d2f90f350f77617d96ccb72ab5eafd6ad397f8e9e84c1b1766b76a7cac47

Observation d9e38b52-8ad4-4e11-b812-b781251217c3 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer CM3: A Causal Masked Multimodal Model of the Internet

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.375359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:81c16812287091931df53ef92d0b53cad521f311396f3fea3541af91187a5a99

Observation 2b80459e-e98f-41cc-a8ce-8ffe572ed3b5 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:48:45.137735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:efc314a38682259ef0777fa47816e5b36a4165950b31ac98e4021131ebaf2005

Observation 291d50e2-3c65-4e31-852d-384141d1ddeb · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning CM3: A Causal Masked Multimodal Model of the Internet

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.302427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:f8e4b833e1f4013349b48c473d03b6906aacc850abcde3c4f43fa0f52bb57249

Observation abc4ea03-17c9-4738-936f-47ed169a939d · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CM3: A Causal Masked Multimodal Model of the Internet

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.440871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.440871Z digest=sha256:1cc1ddf62c74c24ea93aadb060112b9502c092807bbb18e3d0f08cf295efef8f

Observation 9131fc63-07e3-43cb-bcc0-cad32f8bc346 · inbound

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction cites this paper.

Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction CM3: A Causal Masked Multimodal Model of the Internet

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:51:57.735481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:51:57.735481Z digest=sha256:4f8921d325c3a7224a256edf8d66ccdbc740765c311f0112ffb98c4ae7797224

Observation 3b459ec0-0e40-42a1-b02b-2887ec467367 · inbound

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning cites this paper.

Differentiable Optimization Layers for Guaranteed Fairness in Deep Learning CM3: A Causal Masked Multimodal Model of the Internet

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.803385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T15:06:53.377360Z digest=sha256:30dfb474d0c4565de9fbe0300ffa63cab51a96c7d96d4dca0fada3b4a569e377

Observation 869e90c3-232d-47c3-8cf0-af645cc7d399 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:27.802161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:27.802161Z digest=sha256:885689fa642bdcadbe3a6743d6414131a854d094f1c11aec54a602ef0e43c490

Observation 63df0a23-6c2f-4179-a6b5-57144d3d1c24 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes CM3: A Causal Masked Multimodal Model of the Internet

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:54.303566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:54.303566Z digest=sha256:5a7c626efd93f8dfcab94411f4f534ee8d7b35d8aa040d6474d7f88709e2c435