Pith. sign in

Paper Citation Record · LEDGER

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2309.02591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.02591 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:51:13.254591Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

27
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 739d5dad-b6b6-47d5-99ac-9a6729b8ee4d · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 214

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.615691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:44e96253bd39f030ad7cc03640d91ec219ff47c93fa7279f593f2d86accacfef

Observation 2e63b695-acb3-402b-8270-6b1634298980 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.344214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:1a5dc87573b6c53d3f3a78c125c6f98e91102b4a4b430514f5af2baaba4df91b

Observation a157ff30-ed01-4622-8d76-af1ffb73c384 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:03:28.152026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:cfb0c8e46b153088cb09ac3980a3b863c4e17bd79ff2218452de42755077d8f5

Observation d85a6cb4-c31f-454c-a2d5-09905940a6cd · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.609481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:e27e2a1b21278b72adc10e92d99338a2f3dd60165b4407f158c00350db794baa

Observation 3cdb2fec-7473-4b18-a944-37d7bf23c709 · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.997276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:bfff40b3d66f6f2c44242e3991caa9d8ed5d870c1f21e4417462c6f98f82af29

Observation 50606391-af01-487d-a526-855b56149c58 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.382770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:fc577691654bfaa3698ed65f3b8aeaff2705327083471e21bad27698d7301774

Observation 7554d3b9-a55c-47e6-babf-3ee709a2ed4a · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.455699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:ff4e91e0bb7b9915cad1ea43268952ca48b5fc1205aa086e37bd553ad0620e39

Observation 5d99ca5f-b621-4cb0-8da2-388754be6df9 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:48:45.089761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:74094ba6885204fe0b1a70ee0e77ba460d02488838ddc4b2e7400e5ba6d04e84

Observation 05b66629-7c06-48cb-b16d-cdbd42bfdf59 · inbound

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling cites this paper.

Next Block Prediction: Video Generation via Semi-Autoregressive Modeling Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T11:51:13.254591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:51:13.254591Z digest=sha256:c4c94b57f29cca19cc149946a1c367551d9545fd910712e44f8c1b10ddf64c67

Observation a911f41a-9533-4efb-ade8-38a6107d7050 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.729035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:2d4189285a71e85290657bfe4c1ca56bbef089dc763ee6280a414937c3f59490

Observation 1f6f407e-d188-444d-b36f-0b6e4a9b9b6d · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.730499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:d1015b1b724ef235c9c24c75217db1c19e78188da48ddcf8dcf2956e54282b94

Observation 0349970a-a02d-4bfb-b826-ce4611bc89f9 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:00.168507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:00.168507Z digest=sha256:93a5ac8fb1d1128eaacbec45793ea18eab02fe832c5c9a0a1772c23f433eac19

Observation cddd270a-1f4d-403d-a095-4b938064210e · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.282820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.282820Z digest=sha256:f1111068c222b587942f92a2ba129b7158b3a4021eab19b6b330298e921577a4

Observation e9b4edd9-af5c-4184-8228-db135bf16a19 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.061508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.061508Z digest=sha256:d84471d59955a41562bc4cfa4217e3f1e1ad8c7995927cebe6829435d96ce495

Observation 3e051434-6923-46ac-b071-f1409063ba81 · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:30.191271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:30.191271Z digest=sha256:f28997fdb86b354b754d77fa56fa602fc03329160fc2b32303661055d2324fd8

Observation a3d45e71-80cd-47f3-bbb0-101445996b39 · inbound

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation cites this paper.

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:00:24.100992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T16:57:09.843691Z digest=sha256:8ee75b19713578f443db0d0149537a9be9afc31b5957672282999d59198f77ee

Observation c1c71beb-6d0d-4eec-b642-46710fa28895 · inbound

Mirai: Autoregressive Visual Generation Needs Foresight cites this paper.

Mirai: Autoregressive Visual Generation Needs Foresight Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:52.072386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:16:16.461488Z digest=sha256:780582f1c82a1ba4c502935415f596e12decae702071173aa21eac5c88eebd9a

Observation a8c6eaa1-d52b-47f3-9230-edd866b78180 · inbound

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought cites this paper.

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T07:13:01.740855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:13:01.740855Z digest=sha256:bc9f664450308b121bf302265e507c8b1091b590929757c4737dcdb663875917

Observation 6c18d0c1-30cf-445e-9666-814e7aa69a2a · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.450573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:90db84a956b3cb3e3b478c83c4853cfb817b6becad0a94af0ed578514ce0d006

Observation eb19b34b-f87e-48c9-9918-9b17883db5ca · inbound

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation cites this paper.

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:44:03.082166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:40:36.206754Z digest=sha256:759a5a2dc2f9203266f612c885956337783a7fd79928ffa9ddecab11ceee8641

Observation 7533d770-7924-41f1-a666-d1af27ffe8ff · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.002018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:431da5cafbabb0fee289859e544ffa45f5b8262566d9562d44569edccdeb4d5a

Observation c5eac19c-a2da-47f4-9386-824529c178d0 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.661419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:48ad897ebeecdb20f13f0fde61b1fe3e528663cd3f187a21f5487f3211c0c0a5

Observation b62c772c-45f8-4b51-b1fc-f3b5fa752002 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.540114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:ddb3dc2097908ee6ecb9db8bb721e05bfd266bc225932d20b32a5128c44aa608

Observation 80d14c5e-5aca-41bd-bf26-dc472960ebc6 · inbound

Obliviate: Erasing Concepts from Autoregressive Image Generation Models cites this paper.

Obliviate: Erasing Concepts from Autoregressive Image Generation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.633519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T00:28:37.837090Z digest=sha256:478715fdcbf60f8e9c216e778e94793ead173633fe8092fdc1386c50330c6de1

Observation a6b91e9c-04c0-486f-8eba-a1f81998cbda · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:55:25.318583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:55:25.318583Z digest=sha256:ada6fead875d2720997ba8ba12774b35bd9a8400b5a1d938255844e19d039a02

Observation 6a7d5a16-0de4-403a-bb72-3732166dc2cb · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:13:03.762013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:13:03.762013Z digest=sha256:0e866f85ff27e14cc7a8dc4881ccff8547d6b77e6ef8d2a227ace0e380910e9a

Observation 18ce830d-80e2-4c80-87b4-e5957605ca13 · inbound

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications cites this paper.

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:51:50.147250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:51:50.147250Z digest=sha256:41c9653d9fd245d0793e529cc89e7e66e980de6dbfb8ca686f46598b2fa0fbdf