Pith. sign in

Paper Citation Record · LEDGER

Next Patch Prediction for Autoregressive Visual Generation

As of 12 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 5 inbound Pith citation observations for arXiv:2412.15321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15321 v3

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:37:42.215299Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:12:05.860105Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:17:07.173716Z

Reference resolution

100 of 120 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved92
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2678b4c-a8a7-472e-bd8d-4f7c35dade05 · outbound

This paper cites Revisiting neural scaling laws in language and vi- sion.

Next Patch Prediction for Autoregressive Visual Generation Revisiting neural scaling laws in language and vi- sion

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.867328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.867328Z digest=sha256:3013f6e4e18995ab552e340a455e1e87348e3655f6d0f00982cb6e2b67582ebe

Observation c6e1ecb5-c01d-4f59-95d1-a8187bc02fc8 · outbound

This paper cites PaLM 2 Technical Report.

Next Patch Prediction for Autoregressive Visual Generation PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.871583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.871583Z digest=sha256:220d2d64030c45af6d74c5e00ec7fe7eb0bc3b7315911a0e15ccca530c0c20ef

Observation 04f60c00-a63e-4c40-aaf7-247fbdc99b95 · outbound

This paper cites an unresolved cited work.

Next Patch Prediction for Autoregressive Visual Generation Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.875792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.875792Z digest=sha256:a00d57b44f77fc661236ba6c458f521c87d886046fac11dca33bb5b860722277

Observation e3d8c713-e14b-4b1c-9628-d7c70cc5060e · outbound

This paper cites Qwen Technical Report.

Next Patch Prediction for Autoregressive Visual Generation Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.879419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.879419Z digest=sha256:6f0171e405606db75fd0b676f2716d18a8f0307698d7dd3878ae1af117df1d9a

Observation 2f64a344-5a30-4338-ab0a-b66c14123c28 · outbound

This paper cites Sequential Modeling Enables Scalable Learning for Large Vision Models.

Next Patch Prediction for Autoregressive Visual Generation Sequential Modeling Enables Scalable Learning for Large Vision Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.882975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.882975Z digest=sha256:52e639975a7e87f7730a75fde2e266be051ce8bf8658a2710c786a87206c09bd

Observation 89581c38-6b10-45d2-8fc5-8ca4948c2250 · outbound

This paper cites Improving image generation with better captions.

Next Patch Prediction for Autoregressive Visual Generation Improving image generation with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.886704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.886704Z digest=sha256:5b3afbd63362c071ebdbbbdbc1dfa7292f9bdd8f1749e1f5c8bbf95f3d830665

Observation dfcf8ced-fa11-4fc5-95f2-c0a211c7eaae · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Next Patch Prediction for Autoregressive Visual Generation DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.890237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.890237Z digest=sha256:85a0594aa37a9856ff430a37ae71561725d96d7d8be4870999117ded269ec9a8

Observation 90f1afab-de10-40bc-847a-7494ccf00d03 · outbound

This paper cites Large Scale GAN Training for High Fidelity Natural Image Synthesis.

Next Patch Prediction for Autoregressive Visual Generation Large Scale GAN Training for High Fidelity Natural Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.894050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.894050Z digest=sha256:828e9d7128e8608c3dbdc761f3ead53a5d86f8c8bc76572bf09bc57b4e0f9aa8

Observation 5f44bb7f-536d-4a2f-a0d7-131cb81d77ad · outbound

This paper cites Language models are few-shot learners.

Next Patch Prediction for Autoregressive Visual Generation Language models are few-shot learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.897862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.897862Z digest=sha256:d903edb035935297948f26cfdb6ffafa37501305febb2bb20081e27175ba8828

Observation c4b1364c-ebe0-47d4-91d6-c760ef16ec34 · outbound

This paper cites Maskgit: Masked generative image transformer.

Next Patch Prediction for Autoregressive Visual Generation Maskgit: Masked generative image transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.901097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.901097Z digest=sha256:b60812eac53690c71d2d4803391ab15446c0441223535449085e20fd65c93e7a

Observation d99c398d-87eb-4f0b-adf8-dd4f3625554c · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Next Patch Prediction for Autoregressive Visual Generation Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.904410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.904410Z digest=sha256:dfad4a744938cff20f5870476f6e862ab9aa84db885dd80adb1dbcde9d199979

Observation 7e9c81ff-e777-491c-b4f5-50c16d550645 · outbound

This paper cites Softvq-vae: Efficient 1-dimensional con- tinuous tokenizer, 2024.

Next Patch Prediction for Autoregressive Visual Generation Softvq-vae: Efficient 1-dimensional con- tinuous tokenizer, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.908140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.908140Z digest=sha256:413e0900e8904783b49f558ddb4c26e0a5bf258f8bf91e322baa6230a7c686f5

Observation 8bf2de5f-4eef-437a-aa5e-36ad1a70eddc · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Next Patch Prediction for Autoregressive Visual Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.911439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.911439Z digest=sha256:bb1e9edee68265559f046abcfe0484adf736a83b9b49ccf346cc7f73264cacba

Observation 220ef9b3-00dc-4421-824e-efdf971ba888 · outbound

This paper cites GenTron: Diffusion Transformers for Image and Video Generation.

Next Patch Prediction for Autoregressive Visual Generation GenTron: Diffusion Transformers for Image and Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.915325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.915325Z digest=sha256:8b66ffacddc5ec05783ef5afcd67ab1c64b70d5c8464fdcb21a4ebfa1e9ef9b7

Observation c9d3ace6-f9f3-434a-9176-17a4333537e4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Next Patch Prediction for Autoregressive Visual Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.918863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.918863Z digest=sha256:8c5afd3b5f6a458d7e736ba6fe5ed7c80fcbf386b8d67d887d34eb53c39a2f92

Observation 1fb8c0bb-85e2-4cdb-910d-2b0c8c946382 · outbound

This paper cites Palm: Scaling language modeling with pathways.

Next Patch Prediction for Autoregressive Visual Generation Palm: Scaling language modeling with pathways

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.922411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.922411Z digest=sha256:231536cf09d8d5c029f0f1a86c2a76caeb1f7d7b4101fd1e1382039d421ee838

Observation 2ad967ac-6d14-4e0e-9ddf-0902be8561a9 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

Next Patch Prediction for Autoregressive Visual Generation Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.925388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.925388Z digest=sha256:78e114879198d2ea2e7167b0bcf4dc0661c2623525a5a28786d7ea725fcf7b8a

Observation f5c6e865-33db-44b1-8925-007933e685e3 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025.

Next Patch Prediction for Autoregressive Visual Generation Deepseek-r1: Incentivizing reasoning capa- bility in llms via reinforcement learning, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.928364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.928364Z digest=sha256:fa970dd07d51ef50e900aa7011ed8cc5602a0507e65e20e6a0cdd84d4c01a2be

Observation b8798375-3b51-41d9-8eea-3106105cdfe3 · outbound

This paper cites Autoregressive video generation with- out vector quantization, 2024.

Next Patch Prediction for Autoregressive Visual Generation Autoregressive video generation with- out vector quantization, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.931046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.931046Z digest=sha256:12ea0fb842cbe8a1dda2e97a16bfc2aa1ab250eac0d01704d31bbc1438dbc5df

Observation 33125c8e-daec-4ff4-9ecf-514bcd806083 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Next Patch Prediction for Autoregressive Visual Generation Imagenet: A large-scale hierarchical im- age database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.933725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.933725Z digest=sha256:5e9e6dfaf38adb88965da8ef77ae6a801cd54790a6e2ef995210f64b15904af5

Observation 2f67ab22-9600-4b34-b43e-ebdd56750ba5 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Next Patch Prediction for Autoregressive Visual Generation BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.936489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.936489Z digest=sha256:67b95f58e3c8f0500d83200aff2d7b9adbb105d126bfd8335d08f6c345b92204

Observation e14eb8f3-a7ff-461d-aa05-69f6fa20ba06 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Next Patch Prediction for Autoregressive Visual Generation Diffusion models beat gans on image synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.939360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.939360Z digest=sha256:45b866da0d840dca1a66224516c94fe577df823870673cc6a0446ac07d87ce45

Observation b15cd339-15e3-4bc5-ad53-6ce19fcbc589 · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

Next Patch Prediction for Autoregressive Visual Generation DreamLLM: Synergistic multimodal com- prehension and creation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.941969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.941969Z digest=sha256:cc2f2a71d21e06f4b18f483d00beaa9dd0b9a05b59e04baacef9689c1f0681e0

Observation 501a85bf-e889-47f6-a35a-fd197efe3114 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Next Patch Prediction for Autoregressive Visual Generation Taming transformers for high-resolution image synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.945072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.945072Z digest=sha256:82491b1758d1cf1501d4b921fa717a451493aa01b5d4a1282daf4cba57fd5bd2

Observation 35823c8f-67a9-4f11-ab41-2a32498df1f2 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

Next Patch Prediction for Autoregressive Visual Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.948693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.948693Z digest=sha256:2a4781dcaf34689e81b1d6bee85393359f359af923a51c852e2fb1c44d6a7aae

Observation 50ed3788-2b69-4a11-939a-294476e8b714 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

Next Patch Prediction for Autoregressive Visual Generation Making LLaMA SEE and Draw with SEED Tokenizer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.952024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.952024Z digest=sha256:8890aed7deff8178cf44cf608b4a68671f99e3763f86164f840bbbac5205c832

Observation 8214e03a-a089-41b2-b337-074ea15b8a8b · outbound

This paper cites Generative adversarial nets.

Next Patch Prediction for Autoregressive Visual Generation Generative adversarial nets

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.955591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.955591Z digest=sha256:354fd229d7d66b43895ffd39f446c7d695bbdcdf71314fb9000190532359b436

Observation e9b80d46-a249-4253-8782-73b8fc4d698a · outbound

This paper cites an unresolved cited work.

Next Patch Prediction for Autoregressive Visual Generation Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.958864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.958864Z digest=sha256:8e1eb26ae385e6af2399a2bc77fe6faae8fc7f82a945c59ae131e31bdcbf94f0

Observation 1d92170d-9fd9-41d1-8f30-38b7edebabe2 · outbound

This paper cites DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation.

Next Patch Prediction for Autoregressive Visual Generation DART: Denoising Autoregressive Transformer for Scalable Text-to-Image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.962012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.962012Z digest=sha256:dcfb057ea2084c41e547dff7db85c301b877197710a9364aed6be976da78b124

Observation 5f7a98de-699f-4d50-b3d1-1e38eb6e50d6 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024.

Next Patch Prediction for Autoregressive Visual Generation Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.965659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.965659Z digest=sha256:ec1afef712438b3ed844fb4a0573e74c3badde6a6aba2330e36d861434359c05

Observation 5c62801d-ef20-4771-a732-61e78b761156 · outbound

This paper cites ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality.

Next Patch Prediction for Autoregressive Visual Generation ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.968876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.968876Z digest=sha256:c45292d707048e0dc665b77010fa1f999d9f7ebc31895e883aad577fcdca600d

Observation e6493f08-ffd8-49d4-8262-3177b447e512 · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Next Patch Prediction for Autoregressive Visual Generation Scaling Laws for Autoregressive Generative Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.972477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.972477Z digest=sha256:1054a0a69702251f120389b1bd6f87b3eca7d1f14d8fc5c65ae7e0caad575c32

Observation 9cf56f31-d9af-46d3-bd69-4bbf2a8b6c9c · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Next Patch Prediction for Autoregressive Visual Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.976043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.976043Z digest=sha256:aff93d6ce31302651919f7f1b351d03a2d02f7de37205d71e82c55c0554e2fce

Observation 800ae06a-9a0d-4fb2-9142-753bc256bf3b · outbound

This paper cites Classifier-Free Diffusion Guidance.

Next Patch Prediction for Autoregressive Visual Generation Classifier-Free Diffusion Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.979203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.979203Z digest=sha256:59b6e86057d1832bd80533c7b56acd74d15d2c0cd5700f1e5f5d2f69f247f3b5

Observation 535a6ee1-8ef3-4c14-b499-eb2a4d5b0045 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Next Patch Prediction for Autoregressive Visual Generation Denoising dif- fusion probabilistic models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.982421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.982421Z digest=sha256:2b23acc65a52cfad8e7371d857fae0a09c3b624eb47423b48720e1f098aba4b0

Observation b905402e-2b15-4010-889e-781cb79b73a1 · outbound

This paper cites Cascaded diffusion models for high fidelity image generation.

Next Patch Prediction for Autoregressive Visual Generation Cascaded diffusion models for high fidelity image generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.989225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.989225Z digest=sha256:67f8c1e501bc7b78b2dd0c1ce6cc16f40e64da2c2e7a9bf3f5267ad8677ca51a

Observation 8eae2842-301c-4d4e-8b59-368800c09ca9 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Next Patch Prediction for Autoregressive Visual Generation Training Compute-Optimal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.992403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.992403Z digest=sha256:3f4a4d1e040d679f2859555804d98eafff4249810c4a5e1a755380cd65bb92d7

Observation 5f6d48e9-ffdc-450c-8651-02ffab45ed0e · outbound

This paper cites ARFlow: Autoregressive Flow with Hybrid Linear Attention.

Next Patch Prediction for Autoregressive Visual Generation ARFlow: Autoregressive Flow with Hybrid Linear Attention

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:37:42.787730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:41.995809Z digest=sha256:90e6f810211404daa150a22b80c48bbe07f268785bd40dcbcb0006f291f24ea2

Observation 6ea85648-9296-4445-a02c-77370a3e7108 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

Next Patch Prediction for Autoregressive Visual Generation Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:41.999372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:41.999372Z digest=sha256:596a1847e2415d766c863cf777ad197fbeb89cc62eeb40bba5978fbb85756fbd

Observation 4c86337c-43ae-4bc6-8a3b-b74af73c6735 · outbound

This paper cites Scal- ing up gans for text-to-image synthesis.

Next Patch Prediction for Autoregressive Visual Generation Scal- ing up gans for text-to-image synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.002716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.002716Z digest=sha256:a4e787e6ffaf57fec83b1170a2a73a5d89fd03c4f2114aad12fa643d478bdaa0

Observation 876a8aef-bea5-4f1b-a1be-9877573cfc1a · outbound

This paper cites Scaling Laws for Neural Language Models.

Next Patch Prediction for Autoregressive Visual Generation Scaling Laws for Neural Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.009440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.009440Z digest=sha256:5edc6b009ae9205a1af852b376ba57e654b0954bc73be381ee2d8633bfb9fa94

Observation cbdf7a2e-51f6-4154-9921-ab86926f9011 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Next Patch Prediction for Autoregressive Visual Generation A style-based generator architecture for generative adversarial networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.012633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.012633Z digest=sha256:c8ffc47834d9527dbbd5c50ccd653984999be9503ee508b4c191ddc2c3e4b36a

Observation 63dde569-3f17-425e-89e3-6f47c39b9c42 · outbound

This paper cites Auto-Encoding Variational Bayes.

Next Patch Prediction for Autoregressive Visual Generation Auto-Encoding Variational Bayes

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.015717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.015717Z digest=sha256:3bdcd5ea834e08fae5357addcf3c6b60b41d295c68b683fd0c8709105885c89e

Observation 421e1f81-4312-4385-97fc-6c3f7f722b44 · outbound

This paper cites Autoregressive image generation using residual quantization.

Next Patch Prediction for Autoregressive Visual Generation Autoregressive image generation using residual quantization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.019036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.019036Z digest=sha256:4060e6b45faeb5e3b4916466787e45e2fec7aaac9d6a671480f48f51dbd04c7b

Observation cc82f279-1b58-4284-ac40-d8741e717245 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Next Patch Prediction for Autoregressive Visual Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.022394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.022394Z digest=sha256:a636aaf0c2aeb2f0148b36027240559f1e544f0df8503c5b4aee12ed1560a3f9

Observation ee3dcc74-d745-4c7e-864c-0a92866d13ff · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Next Patch Prediction for Autoregressive Visual Generation Autoregressive Image Generation without Vector Quantization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.025724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.025724Z digest=sha256:6171f461bc6002a8fe4bfe0c39a69a5a60a4e1657083784fcca66bfe790dd325

Observation f2321931-1577-4967-8bc1-521a7293ab18 · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

Next Patch Prediction for Autoregressive Visual Generation ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.028863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.028863Z digest=sha256:36d68704e7dbd1dc6349731f7aa1aaade174d46667ea002d923183ace58d9279

Observation 26180138-cf7f-4607-876f-22251b42a286 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Next Patch Prediction for Autoregressive Visual Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.031898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.031898Z digest=sha256:21b31c2163ddf52e98b736aba12d7e1515aa8ac216ab0b70c6f3912d147c3ec0

Observation 2be723d2-c6a8-4e83-8fda-1b05bc489594 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Next Patch Prediction for Autoregressive Visual Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.034923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.034923Z digest=sha256:88178388112311908f6387afb91eb7aaa94404cbd6be5b55c4e8f9b148bea413

Observation 97a851dc-b0f4-443f-b02f-9eefac836bd3 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Next Patch Prediction for Autoregressive Visual Generation MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.038070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.038070Z digest=sha256:d9051e82863823320fe4d0ee13b4d5bea80c9fb1eb713f0f08db988657b7dff7

Observation 22e999eb-ce76-47a9-bdbf-cf7890f7e178 · outbound

This paper cites Visual instruction tuning.

Next Patch Prediction for Autoregressive Visual Generation Visual instruction tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.041346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.041346Z digest=sha256:ae593040d440cfe4e1e59592ed8faa7b3344d57f3382abb86d75019b3ce1664a

Observation 222481ce-06b5-4bd8-b2a5-209fde848ae4 · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.

Next Patch Prediction for Autoregressive Visual Generation Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.044781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.044781Z digest=sha256:2a8e814059d73a64d921e6b310a82104f8ecf196edf34fcc6015851207ce043c

Observation 1c079d4e-525b-46a4-bf3a-e520f7620d1d · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Next Patch Prediction for Autoregressive Visual Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.047969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.047969Z digest=sha256:a0ef421dd6a99af7dbb1551012d7e16c4260fcf0cad0debb15c8e35fce985dfc

Observation ce4e3727-b7e0-45e7-b3ad-69a5353cf09e · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

Next Patch Prediction for Autoregressive Visual Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.051512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.051512Z digest=sha256:5c1f0413c656e1090caa7fb48d2fc039d9863c04e28569f911e6a963b22eac3e

Observation d461c9b0-df8f-42f6-9589-d06f33176927 · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

Next Patch Prediction for Autoregressive Visual Generation Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.055300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.055300Z digest=sha256:10e76d7e2751ca377b1e5b10c7e6ac45dbe1e101ec7729f593bd3f0bf0d9842c

Observation 344f0e95-2292-46ac-8815-10bac8f8b852 · outbound

This paper cites Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models.

Next Patch Prediction for Autoregressive Visual Generation Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.058992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.058992Z digest=sha256:62d6e709be1e06a047a468e71083f31355b6424f0494eafdcba5a8ad26ee6350

Observation 47e94fa9-10c3-407e-8be6-86889bdea2b1 · outbound

This paper cites Janusflow: Harmonizing au- toregression and rectified flow for unified multimodal un- derstanding and generation, 2024.

Next Patch Prediction for Autoregressive Visual Generation Janusflow: Harmonizing au- toregression and rectified flow for unified multimodal un- derstanding and generation, 2024

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.062549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.062549Z digest=sha256:694f4a1e08acb00f3776963de86598c2d872a863564f9edf3fbd96a64e72b70a

Observation 747b278c-b834-4575-a2d4-2400aa666b9e · outbound

This paper cites an unresolved cited work.

Next Patch Prediction for Autoregressive Visual Generation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.065960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.065960Z digest=sha256:c46bce622860c5fb30895295eb8b1409f85f025cd1f01c952b5c75478484f418

Observation c5743155-79c4-4e45-9599-a06ecd2315ff · outbound

This paper cites GPT-4 Technical Report.

Next Patch Prediction for Autoregressive Visual Generation GPT-4 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.069358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.069358Z digest=sha256:a4742cdf20e456d82c027a098fcc0b9072fba2d2268e1434d54911ed9f1e9173

Observation 84bab1b5-fc09-4168-83d9-87a0911dd6ec · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

Next Patch Prediction for Autoregressive Visual Generation Training lan- guage models to follow instructions with human feedback

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.072935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.072935Z digest=sha256:d3d8bb5c88cfea0c1aefbe5de94ba3578aa24175e7349f3267c1f49640fa2a4f

Observation be3b8a50-5726-48a0-aff9-e1c5cb954ada · outbound

This paper cites Byte la- tent transformer: Patches scale better than tokens, 2024.

Next Patch Prediction for Autoregressive Visual Generation Byte la- tent transformer: Patches scale better than tokens, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.076202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.076202Z digest=sha256:513f54a4f54e3ca648ef7e669269c29c6a8d0f22b43ce6d200e8943cd0bff850

Observation 010df9a7-dbe2-4bbc-b694-a159ce69c75c · outbound

This paper cites RandAR: Decoder-only Autoregressive Visual Generation in Random Orders.

Next Patch Prediction for Autoregressive Visual Generation RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.079502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.079502Z digest=sha256:65e3bb31c2dcec3efe0b9be67f170543eff62e5ce7cec447426f708db32dd6d9

Observation 84c08f94-ac3f-4a52-88ea-04485cccb29a · outbound

This paper cites Scalable diffusion mod- els with transformers.

Next Patch Prediction for Autoregressive Visual Generation Scalable diffusion mod- els with transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.083216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.083216Z digest=sha256:b78307f374131d741b4ec2888aea63a6034df97be3678ea4354f45c3b07aa715

Observation 352da051-374e-461e-9da2-fc27bd87d23e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Next Patch Prediction for Autoregressive Visual Generation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.086422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.086422Z digest=sha256:1867725e2c618a74c7a1c3af1802fe1b7309f80d7ce9ec245389ce6ca1f4653f

Observation 119419f8-3985-4d79-ba27-2728fd90e623 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Next Patch Prediction for Autoregressive Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.089803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.089803Z digest=sha256:ab2b8fdbc4cc5eea35084feeca1013866de203e30633360b32d8612c39aa88b3

Observation 43d0d774-8cf4-4c06-babb-9e4c11199368 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Next Patch Prediction for Autoregressive Visual Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.093550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.093550Z digest=sha256:8f11cb484d5d0f65180ee9d5e99d136e005b297acdc3197cbe659b06960da02d

Observation 28c8210c-6b4f-46d9-80e9-58dba463fb9d · outbound

This paper cites Improving language understanding by gen- erative pre-training.

Next Patch Prediction for Autoregressive Visual Generation Improving language understanding by gen- erative pre-training

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.097252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.097252Z digest=sha256:f26283f15782efcc77b6bdb7d6b9e7c24e8c1de4569a0b4df4ccb07c9d1f4d7a

Observation 91d201b5-9dbd-4f3d-8109-145311e58c45 · outbound

This paper cites Language models are unsupervised multitask learners.

Next Patch Prediction for Autoregressive Visual Generation Language models are unsupervised multitask learners

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.100596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.100596Z digest=sha256:b5b521198c503811d51ee8df192246249eebb135e52b55597812c1277b0e6fed

Observation 7d49ddf5-f181-468e-b9e1-08e9c62d1fd4 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Next Patch Prediction for Autoregressive Visual Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.103831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.103831Z digest=sha256:0220b5d0d4f86c0ce98d2877b251bc3f01d77402bd5c8f7db7120b76bf7c116d

Observation 95437200-fb32-4168-ad0a-d1eee3cff138 · outbound

This paper cites Zero-shot text-to-image generation.

Next Patch Prediction for Autoregressive Visual Generation Zero-shot text-to-image generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.107311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.107311Z digest=sha256:31fb3426b355b8b2ecac9139189c9ebf9e7ee676d6a153799d44011f5580e6f8

Observation 8e7fb964-0b2c-4f56-80c8-37e24d39b611 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Next Patch Prediction for Autoregressive Visual Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.110942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.110942Z digest=sha256:c9fd89a28969b0facbeb1f7dc47ea79e842bcc1ddd49b41a2092b0693eaa49e8

Observation 230d4f1f-3548-4492-92cc-51894f39cb84 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Next Patch Prediction for Autoregressive Visual Generation Gener- ating diverse high-fidelity images with vq-vae-2

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.114916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.114916Z digest=sha256:4aed6a29f5d354d3e5454f9efc268c2d57afdad3afb4c7c5a7a28dbfe56111ec

Observation d1b46f05-cda0-43fb-a284-4eaa9070c5df · outbound

This paper cites FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching.

Next Patch Prediction for Autoregressive Visual Generation FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.118172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.118172Z digest=sha256:2c4616f9dd0197d87e79f3c29592a07156a97db557e8db64d0df43653ebed65f

Observation c4cfca07-c8e9-4a6c-8ca5-064760d7897a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Next Patch Prediction for Autoregressive Visual Generation High-resolution image synthesis with latent diffusion models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.121492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.121492Z digest=sha256:f658125808f661a59db3f1ec7928d2eca8dff13d767d426ed429bee2d0366da7

Observation 39790062-eb6a-404d-acbc-20f8bd60ae56 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Next Patch Prediction for Autoregressive Visual Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.124504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.124504Z digest=sha256:8cfafddd891e455d462c328bae6b545e274f5380d2ccbbfbfb9ce52a2b0e22b7

Observation 9ea46b6f-7b70-435b-8c3a-d23e0c548708 · outbound

This paper cites Improved techniques for training gans.

Next Patch Prediction for Autoregressive Visual Generation Improved techniques for training gans

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.127401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.127401Z digest=sha256:b641fd7e95a644c79290e94e33d1fff0a73addb107fad820bbfff5a0cc963199

Observation e75e3496-e8f0-4581-87dd-217dc28f34d4 · outbound

This paper cites Stylegan- xl: Scaling stylegan to large diverse datasets.

Next Patch Prediction for Autoregressive Visual Generation Stylegan- xl: Scaling stylegan to large diverse datasets

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.074724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.130384Z digest=sha256:ad8387f8171f7b83069878af8181f4c74e8e96b240fbe1068efded19ed699a06

Observation d7c9fc83-589a-4bb6-9318-afcfa41bd12e · outbound

This paper cites Beyond Next Token Prediction: Patch-Level Training for Large Language Models.

Next Patch Prediction for Autoregressive Visual Generation Beyond Next Token Prediction: Patch-Level Training for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.133398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.133398Z digest=sha256:fb6212323837d05321a068f9bd45b6eb08bb605012363d4b678ddbf087933c4d

Observation e81aac83-706a-4d85-957a-9aaccec93474 · outbound

This paper cites Scalable Image Tokenization with Index Backpropagation Quantization.

Next Patch Prediction for Autoregressive Visual Generation Scalable Image Tokenization with Index Backpropagation Quantization

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.137388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.137388Z digest=sha256:42926121df6c331ee22e4969753be46b6eff1a97114e69c4dca553f495e0d9bf

Observation b17f3eb0-7adb-44d8-9d91-04b679d818ce · outbound

This paper cites Denoising Diffusion Implicit Models.

Next Patch Prediction for Autoregressive Visual Generation Denoising Diffusion Implicit Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.141135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.141135Z digest=sha256:61cb5c87cd0d8802da90da76e1dcca33fe6e9989f53b2c26cc842f43b26c4411

Observation e3732777-565b-4398-96e2-6b4adfa0abb9 · outbound

This paper cites Generative modeling by estimating gradients of the data distribution.

Next Patch Prediction for Autoregressive Visual Generation Generative modeling by estimating gradients of the data distribution

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.064149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.144971Z digest=sha256:46175dd9a9fb4e0f5c1d802837c05c1eaca11b98d1561582ff497e162f7cf532

Observation 60ee9139-fb02-49c9-a5e0-55456bcc2eb3 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Next Patch Prediction for Autoregressive Visual Generation Roformer: Enhanced transformer with rotary position embedding

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.053252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.148225Z digest=sha256:c5ae9b172fea4713e309fb591a5df0bd9c6401a86b8e7706246c26d1c40d5651

Observation 3c7429a4-bfe4-4359-81a3-67a4f62ccdfe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Next Patch Prediction for Autoregressive Visual Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.151652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.151652Z digest=sha256:bbdea281f8e0e26dd47be4a88f5b774322e51f9dfecaf913d10946a742b6d618

Observation 25e286b3-7cab-42b4-85d8-9b50c1fdb019 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Next Patch Prediction for Autoregressive Visual Generation Emu: Generative Pretraining in Multimodality

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.158805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.158805Z digest=sha256:6824e2f06fdcaed68cf89f94fa0e57ce31d17817c291ba199a74683a38dbf4b7

Observation 6556fe5e-a01b-4071-b3ea-089c1b9dca15 · outbound

This paper cites HART: Efficient Visual Generation with Hybrid Autoregressive Transformer.

Next Patch Prediction for Autoregressive Visual Generation HART: Efficient Visual Generation with Hybrid Autoregressive Transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.162071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.162071Z digest=sha256:5ab0bf1af6a2fa45b464c26b976ca65b103ead77963db5703ace7936f14fffa8

Observation b67454dc-65b8-4ba6-a8bf-d428cfc0dbe1 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Next Patch Prediction for Autoregressive Visual Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.166114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.166114Z digest=sha256:cd5d888c859bb46af42fdf644aac5bcd0efb93e3f50199792a42e49a3f36295b

Observation e72a95e6-e806-4b7e-a280-142087eefdc2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Next Patch Prediction for Autoregressive Visual Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.169547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.169547Z digest=sha256:b15715dfbc0fc44c80dc7b2b044809845792e1cdf01a246b728d476dfecaed19

Observation 8a862b20-96f0-46d4-8832-7862484ffe40 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities, 2023.

Next Patch Prediction for Autoregressive Visual Generation Internlm: A multilingual language model with progressively enhanced capabilities, 2023

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.042849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.173333Z digest=sha256:4792f19a0255f6758332408b0a523ee0a0b9a264151c011c475529b1a094c2e6

Observation efce938b-7501-4164-9820-ccc8cea5c48f · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

Next Patch Prediction for Autoregressive Visual Generation Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.176631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.176631Z digest=sha256:2d94b1409d8e0b83966b1b648a04adb0be6925a61fb1ef9ce233f57366657d0b

Observation c1be90a2-340c-42e3-95ac-1d02a93b629d · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Next Patch Prediction for Autoregressive Visual Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.180110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.180110Z digest=sha256:6464b3a1c65aa75385652033d985f79d48a226702019fa2f215d46c056506db8

Observation 01591285-72fa-49cd-9248-628e800370d4 · outbound

This paper cites Metamorph: Multimodal understanding and generation via instruction tuning, 2024.

Next Patch Prediction for Autoregressive Visual Generation Metamorph: Multimodal understanding and generation via instruction tuning, 2024

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.032784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.183850Z digest=sha256:8552be4b3858bff5931696843d7bccd8f9aae90663f27bf8bc429ee0a7e63540

Observation 9c3b8d69-2aa4-42ac-b238-7b0ef9964ab1 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Next Patch Prediction for Autoregressive Visual Generation LLaMA: Open and Efficient Foundation Language Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.187211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.187211Z digest=sha256:2ee2994633c4e5186a6a60f88a0174bf4e9e3448664adcac43bc3ee464b4c0c7

Observation 16b3cb20-640d-4926-b4f0-19045500349b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Next Patch Prediction for Autoregressive Visual Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.190894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.190894Z digest=sha256:58055dd83dcd428545776f0dbc6bec03985cac9cc5bda412522eb4195775606f

Observation 706ce238-2539-4469-970a-39fc0e5693aa · outbound

This paper cites Neural discrete representation learning.

Next Patch Prediction for Autoregressive Visual Generation Neural discrete representation learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.021678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.194543Z digest=sha256:a3049d1d6d3d601aa40bb9ccea9ff210b1e20133100b61ddb289b932760e1a78

Observation 073051cf-b76e-409d-b11d-cb1567c6d097 · outbound

This paper cites Attention is all you need.

Next Patch Prediction for Autoregressive Visual Generation Attention is all you need

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:37:43.010880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T11:37:42.197801Z digest=sha256:645e8b78ecedd8662d262b94091c98b7f34a4fb8e6cd28c1d46512ab8ba8fc69

Observation b7dcabfc-5d01-435e-b835-f2bb60a781fd · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Next Patch Prediction for Autoregressive Visual Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.201149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.201149Z digest=sha256:a3bd516428fc369fce4b1a598e1bb89210b0a9bb81d0db17b96e730d74bd7b45

Observation 78316c24-1984-437a-87a8-5911c83fe11d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Next Patch Prediction for Autoregressive Visual Generation Emu3: Next-Token Prediction is All You Need

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.204639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.204639Z digest=sha256:3abb1a37f38814efdca5c8acae58f0b530d95a2917344f1489c72770963175dd

Observation 1f8065e1-da55-4e18-932f-76ea0599a604 · outbound

This paper cites Parallelized Autoregressive Visual Generation.

Next Patch Prediction for Autoregressive Visual Generation Parallelized Autoregressive Visual Generation

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.208403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.208403Z digest=sha256:26e487475387b8170a5b1097304be37613438b862002365bc758fca702c8dc6c

Observation 7afc6367-d954-4d82-9032-49ae7226e8f9 · outbound

This paper cites MaskBit: Embedding-free Image Generation via Bit Tokens.

Next Patch Prediction for Autoregressive Visual Generation MaskBit: Embedding-free Image Generation via Bit Tokens

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.211813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.211813Z digest=sha256:e4fc09d7e4f2bec8797062e1a52b7697fb3f5a289544fe1b53064a7a9f057041

Observation 4541546c-a0c4-45cd-97ed-49ef68583e5d · outbound

This paper cites Emergent Abilities of Large Language Models.

Next Patch Prediction for Autoregressive Visual Generation Emergent Abilities of Large Language Models

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.215299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.215299Z digest=sha256:e905784a6de685f1ad3e65ace2c88d048c8a1cd23abd95ab20c24627f8cca157

Pith citing papers

Observation 1457285d-4191-414b-bd55-2cec95ab0033 · inbound

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning cites this paper.

Hierarchical Banzhaf Interaction for General Video-Language Representation Learning Next Patch Prediction for Autoregressive Visual Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T23:12:05.860105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:12:05.860105Z digest=sha256:7e82ae17577dcf5f64301c5c7884edad42710767a7d2259ad2723aa7ffe5203a

Observation d2fb46c3-45f2-4412-b0c8-2329aa1008c0 · inbound

AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene cites this paper.

AE-NeRF: Augmenting Event-Based Neural Radiance Fields for Non-ideal Conditions and Larger Scene Next Patch Prediction for Autoregressive Visual Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:11:01.792501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:11:01.792501Z digest=sha256:1dabfbd78de18f3699fad786a2f8911a03954cd287eccde8291ba0d1101b29e8

Observation b1206060-f4c2-4603-a6da-2288fef146a0 · inbound

Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective cites this paper.

Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective Next Patch Prediction for Autoregressive Visual Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:50.885357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:50.885357Z digest=sha256:c66a8a9ac4a7dcab8f419050a7ebc2723fa8269a2031ef9991267368920d76fc

Observation e3225084-d36a-40d3-8b5c-273ad0ed1248 · inbound

E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras cites this paper.

E-4DGS: High-Fidelity Dynamic Reconstruction from the Multi-view Event Cameras Next Patch Prediction for Autoregressive Visual Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:51:42.414004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:51:42.414004Z digest=sha256:ecd561a5bcec84dd392751303ce3f08ae81b63623b4f3266edc6a3347c9eca4e

Observation babc43c3-2653-46cd-b273-e9fa7cb8192a · inbound

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts cites this paper.

MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts Next Patch Prediction for Autoregressive Visual Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.176042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-02T15:14:36.946247Z digest=sha256:0ef994023e97588bf2a6efb2f997d8ce536718839ab08e905303a57508affde2