Pith. sign in

Paper Citation Record · LEDGER

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2309.02591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.02591 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:51:13.254591Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

27
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 739d5dad-b6b6-47d5-99ac-9a6729b8ee4d · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 214

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.615691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:57a7d44f71443169ae6dac90799cb835e825a1240f13dd8171cac13243dbde4a

Observation 2e63b695-acb3-402b-8270-6b1634298980 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.344214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:fcf1be6c048b477230e2327e4d0d9dd51216f85c17bd9e416cb00c3f4a3fdd72

Observation a157ff30-ed01-4622-8d76-af1ffb73c384 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:03:28.152026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:7b4c2698b89212fe67ea4f9d5112ed5da12909e0ebd8c7778ac362a98d8a0eb1

Observation d85a6cb4-c31f-454c-a2d5-09905940a6cd · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.609481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:c45625ae2c23dbdb5ebb2ed899b71388c305d6cd356e3a47a10adc391b8d2f75

Observation 3cdb2fec-7473-4b18-a944-37d7bf23c709 · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.997276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:e0505d3c790a4fbe8b09d4e655267bf1dd03600302eec1a4e89ae9a028839926

Observation 50606391-af01-487d-a526-855b56149c58 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.382770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:f2ef6d9190f21dc275a704675dfc15c4b109b9de8b2602f88ad1640677528278

Observation 7554d3b9-a55c-47e6-babf-3ee709a2ed4a · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.455699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:2f2311d6949c4a8c8889b52f13757de287563331be680d5b6a9b8bf981378713

Observation 5d99ca5f-b621-4cb0-8da2-388754be6df9 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:48:45.089761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:46b5e28881b6b269905f5fd6311b8c2dc85f30b03ae69cb5a35426be2edf4095

Observation 05b66629-7c06-48cb-b16d-cdbd42bfdf59 · inbound

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling cites this paper.

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:13.254591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:13.254591Z digest=sha256:c4c94b57f29cca19cc149946a1c367551d9545fd910712e44f8c1b10ddf64c67

Observation a911f41a-9533-4efb-ade8-38a6107d7050 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.729035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:7df05fccc97c17ee1d6b3e5ace8899511ab320dbfa873addcefbe789c6400dcf

Observation 1f6f407e-d188-444d-b36f-0b6e4a9b9b6d · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.730499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:f1ab802ecfedff174c79d9b32aeebdebfe5ecca550b141269fe383a721def78f

Observation 0349970a-a02d-4bfb-b826-ce4611bc89f9 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:00.168507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:00.168507Z digest=sha256:0653ea598d8049a50e01a2126ba776e10d2d94ad7aba002dd15b2474448405ee

Observation cddd270a-1f4d-403d-a095-4b938064210e · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.282820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.282820Z digest=sha256:f1111068c222b587942f92a2ba129b7158b3a4021eab19b6b330298e921577a4

Observation e9b4edd9-af5c-4184-8228-db135bf16a19 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.061508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.061508Z digest=sha256:d84471d59955a41562bc4cfa4217e3f1e1ad8c7995927cebe6829435d96ce495

Observation 3e051434-6923-46ac-b071-f1409063ba81 · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:30.191271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:30.191271Z digest=sha256:f28997fdb86b354b754d77fa56fa602fc03329160fc2b32303661055d2324fd8

Observation a3d45e71-80cd-47f3-bbb0-101445996b39 · inbound

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation cites this paper.

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:00:24.100992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T16:57:09.843691Z digest=sha256:854a753caf639b15de273723c993d7ed4f5d7c0582c318ca389b6a1b72cf320c

Observation c1c71beb-6d0d-4eec-b642-46710fa28895 · inbound

Mirai: Autoregressive Visual Generation Needs Foresight cites this paper.

Mirai: Autoregressive Visual Generation Needs Foresight Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:52.072386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:16:16.461488Z digest=sha256:092ed2f60a12d871f9c067598a384da1cfffd577921ec236e27c95e764c93c98

Observation a8c6eaa1-d52b-47f3-9230-edd866b78180 · inbound

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought cites this paper.

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T07:13:01.740855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:13:01.740855Z digest=sha256:0f6db318573fb73eec9b31c0cde43f630230df5e66b0f2b28bc3c5b841d4ddb6

Observation 6c18d0c1-30cf-445e-9666-814e7aa69a2a · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.450573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:96c583b34961107c5243bfa344fb8584f94e61ed2d5b67c20d06dd6cbb64df07

Observation eb19b34b-f87e-48c9-9918-9b17883db5ca · inbound

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation cites this paper.

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:44:03.082166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:40:36.206754Z digest=sha256:fc5c34a7c6ff76510b4d60ed3fa6e3f3c6638757db96f6d3413382464e97fed7

Observation 7533d770-7924-41f1-a666-d1af27ffe8ff · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.002018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:e87dcab2f2d4eeccd243e5df5168a8c021443ab496d76ab5e1193cd2065d1fde

Observation c5eac19c-a2da-47f4-9386-824529c178d0 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.661419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:5e93738863517b12c891af817ae19d37a62d28440bc332c8fb26b19ef7f929e7

Observation b62c772c-45f8-4b51-b1fc-f3b5fa752002 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.540114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:d15dddd9f77669f3739953336963dd4fd78590ac7bfa494ce7dff322452456cc

Observation 80d14c5e-5aca-41bd-bf26-dc472960ebc6 · inbound

Obliviate: Erasing Concepts from Autoregressive Image Generation Models cites this paper.

Obliviate: Erasing Concepts from Autoregressive Image Generation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.633519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:28:37.837090Z digest=sha256:fa390fdb867d075bd5e8c1d5a5838b455fb0079d4327155de5b0365a1a1fa9d3

Observation a6b91e9c-04c0-486f-8eba-a1f81998cbda · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:55:25.318583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:55:25.318583Z digest=sha256:ada6fead875d2720997ba8ba12774b35bd9a8400b5a1d938255844e19d039a02

Observation 6a7d5a16-0de4-403a-bb72-3732166dc2cb · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:13:03.762013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:13:03.762013Z digest=sha256:0e866f85ff27e14cc7a8dc4881ccff8547d6b77e6ef8d2a227ace0e380910e9a

Observation 18ce830d-80e2-4c80-87b4-e5957605ca13 · inbound

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications cites this paper.

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:51:50.147250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:51:50.147250Z digest=sha256:41c9653d9fd245d0793e529cc89e7e66e980de6dbfb8ca686f46598b2fa0fbdf