Pith. sign in

Paper Citation Record · LEDGER

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2506.07999.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07999 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:26:02.330264Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 544ef75f-082c-41c0-8402-9ca8daa351ab · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.141730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.141730Z digest=sha256:253d39c0ef5591d5f74b39b23e2a32f49ec0b6ff5bd53e6e51911378262a6891

Observation 494bde0b-a381-4826-bc7e-74cc21da2b3b · outbound

This paper cites Semantic-conditional diffusion networks for image captioning*.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Semantic-conditional diffusion networks for image captioning*

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.035305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.148374Z digest=sha256:c0aa08bea2c01b532a711f19b259bd710dc5a7033132cd340bd7922b553eb947

Observation 7210748d-ffab-4ce4-84b6-868e679ba753 · outbound

This paper cites Metaxas, and S.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Metaxas, and S

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.019678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.154536Z digest=sha256:efdc18dd580ecfd4a4833b2187dcd38e77d197eeebc86b6bcb051699841d841e

Observation 4fa29346-b78e-407d-ba57-8d494f21b8e0 · outbound

This paper cites Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.159306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.159306Z digest=sha256:0f8c19c76194c9208e11f5bbfc4fee1bf89278ee04b0f0c31111ec1d10972c5d

Observation 1acfa949-3647-486a-8f7a-3dcc41e3ac57 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Taming transformers for high-resolution image synthesis

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:03.003472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.164485Z digest=sha256:86a7ca6e1d090004e5edb09f7b9727312545283195769b0aa874a1c8780c5e91

Observation a9af8c98-d79e-47ad-a980-60bd0d5c3382 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.169516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.169516Z digest=sha256:4186b3494a360ef61e2855a895e10421a442d17a367bdc10f8ff5995956db2eb

Observation e296f37f-7262-4e36-9ab5-699e0e348ad8 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Cosmos World Foundation Model Platform for Physical AI

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.174857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.174857Z digest=sha256:51ad9b2a83a2d269836bdc2d548311e1fee3cd9481b7a654765b6c4de3b2faf4

Observation 861fa5a2-c893-4650-8695-d3812e7ed8e9 · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Blattmann, Dominik Lorenz, Patrick Esser, and Bj \"o rn Ommer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.987674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.179899Z digest=sha256:9734db9d96c3e213fa7fd60fa430d1a58c50cab0408db0e7b25529ccaf9d551c

Observation 40f1b01b-18b6-4358-a0fd-bcefc82a9227 · outbound

This paper cites Peebles and Saining Xie.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Peebles and Saining Xie

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.972644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.184466Z digest=sha256:b06fb8a0a7c7f8d5057df0f33cfc403b8b334f373e62a5cea57e6eb281a6e94d

Observation 9b1e849c-805a-457e-bf94-fd03f743576b · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.189060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.189060Z digest=sha256:cd57d58b938e8cf196eb6fbe7a7e61dabde978677453976818e27846e53fcb4a

Observation 68056b33-8d3e-40df-834f-f1d5b816e72a · outbound

This paper cites Announcing the flux pro finetuning api.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Announcing the flux pro finetuning api

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.956433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.195206Z digest=sha256:c635023c5006b9fb4e80ce1d386d81f2800abf52404546f9fd510f2d0169eb05

Observation 5f1e3e6b-773d-4c51-8945-3fe4e201cf36 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Autoregressive Image Generation without Vector Quantization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.200117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.200117Z digest=sha256:5438f4cb80343435fb1643f59196a3f518f32895e4e618554d4c7f689232893e

Observation c225a258-a4c0-40ad-ad33-589605391c9d · outbound

This paper cites Acdit: Interpolating autoregressive conditional modeling and diffusion transformer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Acdit: Interpolating autoregressive conditional modeling and diffusion transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.205671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.205671Z digest=sha256:448af80ad9aa4a1ba5a442690a2fc21d43cb764d57177679b2b1be5c9fe326b2

Observation 6fbc3b17-ad2e-44d7-8d58-a644b733a2a1 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.210366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.210366Z digest=sha256:a502f170838459bc23f25f0437604b94b4f0e9c446e1a1577e897330af86a767

Observation efe7341d-83ae-420a-bc88-319e07ac5896 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.215293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.215293Z digest=sha256:f8d744962281574db8394719e9f8f470fa529534a1a41be3e8db960eb423515e

Observation a6927dc3-ef34-4fbc-beff-b841f147b555 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Denoising Diffusion Probabilistic Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.219967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.219967Z digest=sha256:3bcbb5ecce75194550100f9a052645ae0fcae69b227576418468966f41431b58

Observation 592032d5-4885-454e-9b6a-b31df77da8b3 · outbound

This paper cites Neural discrete representation learning.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Neural discrete representation learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.938173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.224761Z digest=sha256:0568995a1bdb55d240e5f57530cd47ba5d8a8f343be9dc489412fdd22e54e5cc

Observation 52069c2e-0ba9-437b-8cac-ff9b27c1b071 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.229503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.229503Z digest=sha256:df01cff4ba4de891232e774c848a4fd28ed78a4d8009a8e2496370ecc0fa7c9c

Observation fbaad28b-6672-4741-97a4-0bfbedee047a · outbound

This paper cites Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Minnen, Yong Cheng, Agrim Gupta, Xiuye Gu, Alexander G

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.921578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.235441Z digest=sha256:ab17917b6a82b345cf1d06b5f8c2f56126008a404dbeb26090aa2d4491ccac29

Observation 1d12df17-efb9-41c9-8cea-fa22c4d26da0 · outbound

This paper cites GIVT: Generative Infinite-Vocabulary Transformers.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation GIVT: Generative Infinite-Vocabulary Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.240085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.240085Z digest=sha256:c8f0db3a72fe7992d7236befe04e4c99e378d36ccc70511eef4a302c52d86a6c

Observation 58e5e446-5684-49ed-a9fb-a57566674917 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.245217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.245217Z digest=sha256:54f6d91d77cb4b491b900fc9464054b0522d49b8e9d95925d261dd201bda50b7

Observation bd83759b-6a84-4ff7-9ff3-95f419ceca7e · outbound

This paper cites The Llama 3 Herd of Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation The Llama 3 Herd of Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.250162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.250162Z digest=sha256:4cf73f3eac271493e61de8c0ea2fc1fd164f70003cd63a474f00b95f8fe5ad19

Observation f8e9cff5-7b83-49e2-8c18-1caf4883be3c · outbound

This paper cites Improved Denoising Diffusion Probabilistic Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Improved Denoising Diffusion Probabilistic Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.255039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.255039Z digest=sha256:5ec93dad48525232335730a46193219f00c84e92884701226de5ef64a510079b

Observation 4fba9174-df27-40ef-a29b-901eeec878d1 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.259935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.259935Z digest=sha256:24c8b36869df16228a6eec30090fe03dd57e93ae54139f39878d2a25fe4c69f5

Observation b1c20c57-3a1e-4fa7-ad98-4618190b6905 · outbound

This paper cites Flex Attention: A Programming Model for Generating Optimized Attention Kernels.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Flex Attention: A Programming Model for Generating Optimized Attention Kernels

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.265059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.265059Z digest=sha256:87fd68e9ae98a64b4ddbfa8f70c1e6393ec1784ee5eca7ee2ad1fa44e7b0d5fd

Observation 4db2e591-13d1-4e93-93a7-dde297427fd4 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Progressive Distillation for Fast Sampling of Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.270330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.270330Z digest=sha256:e4e26f6f3c836e8fe7ed2cfc3cca097a9e455a82a724dc9b7256ee89c5c1f0b2

Observation 62d869dd-57ac-48be-966f-c3e3d4421037 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation A style-based generator architecture for generative adversarial networks

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.905182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.275413Z digest=sha256:07e0e05daa75a006e0c3ce398aef46653ffd47a4b459d73aaaf2a2b7c94fa8c7

Observation dec8887e-146a-497a-9fcb-531037e4a453 · outbound

This paper cites Berg, and Li Fei-Fei.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Berg, and Li Fei-Fei

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.280007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.280007Z digest=sha256:e7f823b1832fbd404cc718fa15031c8eb479108f0094fb2e1b99433117ab43dc

Observation e2b7d0af-685f-409a-9696-75ee733aa485 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.284927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.284927Z digest=sha256:860b30c8ba47e590713ad7f1af4260bd64b8ec491ab6a5d9c8af8baef43f3abf

Observation b359263d-7854-4482-bff0-19015f64197b · outbound

This paper cites Denoising Diffusion Implicit Models.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Denoising Diffusion Implicit Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.290423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.290423Z digest=sha256:45fa7d7a4386b20a1f95a6efca932850b14da792f3635fafa8113aa1a5f6e75b

Observation 87f949b3-56af-4d57-b1b8-c37f3b2ffbad · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.295862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.295862Z digest=sha256:a832697aebe5cdb43797b233944922beb558c67c47f34355cbc128e46190c3cd

Observation 69d98c81-9564-48fe-8a07-aeefee81b03d · outbound

This paper cites MonoFormer: One Transformer for Both Diffusion and Autoregression.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation MonoFormer: One Transformer for Both Diffusion and Autoregression

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.300549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.300549Z digest=sha256:0826b3cd9192b6319ca5d2136ace51598e7fc33cfd2630161e4f5c66ec6cfc9a

Observation c019fe1d-5e97-499b-8135-b37cb455203b · outbound

This paper cites Multimodal Latent Language Modeling with Next-Token Diffusion.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Multimodal Latent Language Modeling with Next-Token Diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.305545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.305545Z digest=sha256:d4b3142af1aec3d3c779f2f2b4fd58ae67e5706f5a9eb4ea0b0a324ae205a3e9

Observation 446505ae-d8ab-4536-8a30-77cdf3a0a24c · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.310361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.310361Z digest=sha256:b02b3c3a45a799c1048495e82a55066bb75d1b5be9f7e7dce742396fca523f23

Observation 736137f4-926f-4033-9fb4-1590dd309969 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-io 2: Scaling autoregressive multimodal models with vision, language, audio, and action

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.887324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.315361Z digest=sha256:7591daddc52e672a446871ef83a704eb28908263c4963b79a2c255e49b7db738

Observation d9adef13-b6d4-4578-a126-2b3c99bbc6c7 · outbound

This paper cites Addendum to gpt-4o system card: Native image generation, 2025.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Addendum to gpt-4o system card: Native image generation, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:26:02.867367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:26:02.320633Z digest=sha256:d7d431bb320e2d81dbe76672f09544198b845e8384739523809fab35cabfdea3

Observation 5a82c5ce-1d22-45c3-b8f0-cae910510371 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.325132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.325132Z digest=sha256:4dc32ea16c8e2337c040f82b4f148bcc7bc106b6380656ad4a3608dd2feb6673

Observation 7afd44aa-38ce-4c14-996a-5b623d8e2a32 · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.330264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.330264Z digest=sha256:379359182ca623ebe56cfe73df3bba293a5bf722d5a98c2e7fae9d9b098b411f

Pith citing papers

No inbound Pith citation observations are available.