Pith. sign in

Paper Citation Record · LEDGER

High-Resolution Image Synthesis via Next-Token Prediction

As of 13 August 2026, this Paper Citation Record lists 100 of 137 outbound references and 0 inbound Pith citation observations for arXiv:2411.14808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14808 v2

Coverage vector

measured 100 of 137 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:55:42.363118Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 137 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved91
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e86bbf5c-659a-45af-9892-d3d006f66bf2 · outbound

This paper cites GPT-4 Technical Report.

High-Resolution Image Synthesis via Next-Token Prediction GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.595842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.595842Z digest=sha256:2932c6ab8834c09af46c631e39ef814224bd091112adfe6833aff0f6f464b320

Observation 36c78d02-b30e-4b77-8b99-d1e2e4921d57 · outbound

This paper cites A window with raindrops trickling down, overlooking a blurry city.

High-Resolution Image Synthesis via Next-Token Prediction A window with raindrops trickling down, overlooking a blurry city

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.600514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.600514Z digest=sha256:9a1534c1283e41ba38334c6166212d87767ce050208727988e57dee34c347284

Observation fbd11f71-a9fe-416f-8e13-5582dbed13c3 · outbound

This paper cites Stochastic Interpolants: A Unifying Framework for Flows and Diffusions.

High-Resolution Image Synthesis via Next-Token Prediction Stochastic Interpolants: A Unifying Framework for Flows and Diffusions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.604928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.604928Z digest=sha256:9591a5721e2903d1de331657125c460ca80830f8e413a95d728ae195230a8633

Observation 6f2730ed-997d-449d-b13d-cead22f0f1cb · outbound

This paper cites PaLM 2 Technical Report.

High-Resolution Image Synthesis via Next-Token Prediction PaLM 2 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.609639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.609639Z digest=sha256:bb2ae9fc69ac570e3183688108a4522ccf6ddcfde2f22fa1376a5d9e1d85fc68

Observation 733d3135-d48b-4284-a61b-d1ec8c9b7038 · outbound

This paper cites Qwen Technical Report.

High-Resolution Image Synthesis via Next-Token Prediction Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.614007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.614007Z digest=sha256:6f065c082b78c387ffc1c9a3a2815c1c3c76980afe9501cfdb27a855d2450d44

Observation 63b347ca-1b64-432e-9fdd-2d8d81eb5b1b · outbound

This paper cites Layout control by relative positional offset b.

High-Resolution Image Synthesis via Next-Token Prediction Layout control by relative positional offset b

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.618699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.618699Z digest=sha256:7de4515b9dc5eecd964f1e0583954f6202dc9fefa817f58699f40e8b7ba15b2e

Observation 8ba0d1be-cd99-4371-883b-613c11083623 · outbound

This paper cites Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models.

High-Resolution Image Synthesis via Next-Token Prediction Analytic-DPM: an Analytic Estimate of the Optimal Reverse Variance in Diffusion Probabilistic Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.622653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.622653Z digest=sha256:e5ea04e6e87c640a5df9d37843d627aeefc470283027fab07a2199c0897de0e7

Observation 5c7df0cb-1da1-4c9c-b3ae-3ca5368a4a87 · outbound

This paper cites All are worth words: A vit backbone for diffusion models.

High-Resolution Image Synthesis via Next-Token Prediction All are worth words: A vit backbone for diffusion models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.627308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.627308Z digest=sha256:4cfb3f1dfab92751215174ede457468ba2b451eae2646234ce1dd3c77e10e79d

Observation d83fc5a1-b65c-4edc-b103-26e11cd5cd2f · outbound

This paper cites Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models.

High-Resolution Image Synthesis via Next-Token Prediction Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.631468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.631468Z digest=sha256:8b0ae25e38a433de7ada540a971945c756c2e85d5896454dc5b7ff63c9b88638

Observation 68c33c2b-2b2e-4da4-8653-9487330c8c10 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

High-Resolution Image Synthesis via Next-Token Prediction BEiT: BERT Pre-Training of Image Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.636000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.636000Z digest=sha256:bb4c6cabd217c8bd5a3016e2173affdb51c865d505352e0a9b0d79cb16259a56

Observation fccd1052-9e83-4755-b61b-8e5563fdd234 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

High-Resolution Image Synthesis via Next-Token Prediction Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.640035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.640035Z digest=sha256:0a2600913c7cf0ad0433756b23590d68e74b9139dd484dc60f29918ec520e3a5

Observation a80c798c-485e-4546-b80d-00e02854cf87 · outbound

This paper cites Improving image generation with better captions.

High-Resolution Image Synthesis via Next-Token Prediction Improving image generation with better captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.644734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.644734Z digest=sha256:e4e92c380544e3ea10357c17a5c99d1dd401aa46cc50cd4e327014064ac9ecbb

Observation 63de849b-7c06-4ff4-8926-169e3a8dc27f · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

High-Resolution Image Synthesis via Next-Token Prediction Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.648696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.648696Z digest=sha256:79f24e0dadf1123418baebafab5eae5e357503e4a34e4211dde9a2ade3cae6de

Observation 561b0ad8-7ebf-4995-b073-95a1898313d6 · outbound

This paper cites Align your latents: High-resolution video synthe- sis with latent diffusion models.

High-Resolution Image Synthesis via Next-Token Prediction Align your latents: High-resolution video synthe- sis with latent diffusion models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.653484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.653484Z digest=sha256:0687024085b234c914f477d223f26d7eaec295b7592e605679bf4b1adfbc2269

Observation bcb9320e-472c-4ff1-ad17-1b6f7723abbc · outbound

This paper cites Brooks, B.

High-Resolution Image Synthesis via Next-Token Prediction Brooks, B

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.657660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.657660Z digest=sha256:224292e07fa4cc4bd90bebc3f5ec73046982ceeacba037d7ee3b95364766b8d5

Observation b86ac1be-a17d-4194-acaf-6937846a4fc7 · outbound

This paper cites Language Models are Few-Shot Learners.

High-Resolution Image Synthesis via Next-Token Prediction Language Models are Few-Shot Learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.661652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.661652Z digest=sha256:ccf8e011ac4491175580a16864052c0be08599a740a6ed5c1bb43c3bcb8057f5

Observation ffc01eaa-85a7-410f-8097-bedc24b2c961 · outbound

This paper cites Maskgit: Masked generative image transformer.

High-Resolution Image Synthesis via Next-Token Prediction Maskgit: Masked generative image transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.667292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.667292Z digest=sha256:646702048c4df4620e4b52e44032d2ba446fd195bffcd5bfec40a4c7e9c86036

Observation d2503e56-de50-4299-ad80-7a716e804acb · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

High-Resolution Image Synthesis via Next-Token Prediction Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.671777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.671777Z digest=sha256:207aaf20d24fcf2cfc76e15348436ff4299b24fad53e5264b1fda7665af8851d

Observation 9f240309-bf83-45f3-be72-bdd4545240d9 · outbound

This paper cites Bfloat16: The secret to high performance on cloud tpus, 2019.

High-Resolution Image Synthesis via Next-Token Prediction Bfloat16: The secret to high performance on cloud tpus, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.676078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.676078Z digest=sha256:2b6a6d5f79ccddeb6f8d5fdfecba5bdfc9ac3aed2d0c5c7690c6fdba01412d47

Observation dfb66dac-56dc-4fbf-b749-928f2e56722b · outbound

This paper cites Denoising with a Joint-Embedding Predictive Architecture.

High-Resolution Image Synthesis via Next-Token Prediction Denoising with a Joint-Embedding Predictive Architecture

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.684165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.684165Z digest=sha256:817d5fa2b1280177f08ed016be19555fdfc4bd4d2544cc9fe1c5d6cf7b000104

Observation ccd1b0b6-e81d-4394-9c20-c61dac582042 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

High-Resolution Image Synthesis via Next-Token Prediction PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.690352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.690352Z digest=sha256:fac9e2764a7014c41ec20f5c92cccf2f5e7c187e46d3bc65175e4d86d3aba7c2

Observation 6c3ca0b3-6f17-4278-91ea-8bdc770e6ab8 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

High-Resolution Image Synthesis via Next-Token Prediction PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.696009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.696009Z digest=sha256:573f277e1a4c9cbc786b47b3871a6de92ff2b1a46f52efd6bf325328c7c0e120

Observation 176a03e3-3585-4682-b716-9361c632876a · outbound

This paper cites Neural ordinary differential equa- tions.

High-Resolution Image Synthesis via Next-Token Prediction Neural ordinary differential equa- tions

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.701937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.701937Z digest=sha256:1aafc2def065a9ecbdf9ebf5081610ebcd6f0d6891fa96c2ac4af39db5d52d96

Observation d175ca6c-0151-41b9-931d-d88d3b7e5aa5 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

High-Resolution Image Synthesis via Next-Token Prediction InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.708051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.708051Z digest=sha256:f836b26b58abfa723206802d3e6878c538e5a492c0fe7fd5f0f445ede82161db

Observation 25191846-4385-45d2-9b3c-2d089c505330 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

High-Resolution Image Synthesis via Next-Token Prediction How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.716145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.716145Z digest=sha256:320c844a431714982942fcb85901aef32af107233c6b1d79061a236737ee0fa0

Observation 501a4058-965e-4527-9e11-ea3145a7bef8 · outbound

This paper cites Palm: Scaling language modeling with pathways.

High-Resolution Image Synthesis via Next-Token Prediction Palm: Scaling language modeling with pathways

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.722362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.722362Z digest=sha256:824cfb95326200caa7e4aaca8d7dc35001828c5ec18d8eec30d7986e30bc5e03

Observation 5b726e95-19f8-45f0-8b58-35bc3250a702 · outbound

This paper cites Scaling instruction- finetuned language models.

High-Resolution Image Synthesis via Next-Token Prediction Scaling instruction- finetuned language models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.728465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.728465Z digest=sha256:71d909b1d811ff3ea2e4d0cf9421d3b8a03bcc0833776377b630ba05e9657fe9

Observation eba7e90d-336e-4732-8b2c-0d22fe4c8c15 · outbound

This paper cites Emu: Enhancing image generation models using photogenic nee- dles in a haystack, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Emu: Enhancing image generation models using photogenic nee- dles in a haystack, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.732712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.732712Z digest=sha256:2ac60cf6e281a67e9b4159fae4142219099999d6717b3b0cb2c1079cb77a0d95

Observation d8c1c698-32e5-456b-8cd3-8567002c9c10 · outbound

This paper cites Flow matching in latent space, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Flow matching in latent space, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.737254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.737254Z digest=sha256:0c8bc53fb2190ad443353b2ec42f3c64a9f55a9bb174565223bcb790319369e2

Observation 30160b63-1f7e-4bd1-b454-e6220f846029 · outbound

This paper cites an unresolved cited work.

High-Resolution Image Synthesis via Next-Token Prediction Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.741239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.741239Z digest=sha256:093fec2b9b891cf3268c5920dc1d59b53423a78d93ec09f20830568d36ab911f

Observation 25e9dd81-7d42-468a-a8f1-1e7c95279ef6 · outbound

This paper cites Diffusion mod- els beat gans on image synthesis.

High-Resolution Image Synthesis via Next-Token Prediction Diffusion mod- els beat gans on image synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.745818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.745818Z digest=sha256:9d6ef43dce367d2c97e37b6e7e5428a270e82b9fc99266fb4d6f158f31070414

Observation 3170d992-da09-4b55-b246-f300abb86650 · outbound

This paper cites Score-Based Generative Modeling with Critically-Damped Langevin Diffusion.

High-Resolution Image Synthesis via Next-Token Prediction Score-Based Generative Modeling with Critically-Damped Langevin Diffusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.750016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.750016Z digest=sha256:08aba6db006d70787e6df91005228131758ebf0120f8695e5c0d4accbca06e82

Observation 75ed7e6a-011b-47e4-bc52-d0c44e83d0bf · outbound

This paper cites Genie: Higher-order denoising diffusion solvers, 2022.

High-Resolution Image Synthesis via Next-Token Prediction Genie: Higher-order denoising diffusion solvers, 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.916893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.916893Z digest=sha256:16814c6cd1140b83a82c5391792dbc6269fc7397764c066f2b1e9907d8490794

Observation f4ddbcc5-6fb1-4c1f-baa9-c945f0f687fc · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

High-Resolution Image Synthesis via Next-Token Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.926416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.926416Z digest=sha256:5ccbb36c191b9612e0bd843777a1601ac0a8bf48fa450950f3236349dc58cbe6

Observation 1ccdc236-c7d0-4296-997e-de66917651fc · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

High-Resolution Image Synthesis via Next-Token Prediction Structure and content-guided video synthesis with diffusion models

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-12T14:55:41.941194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.941194Z digest=sha256:8baed2fd1d4b63d44ce16db097d30cdea5154d5cd573ef6c3db9734dd4cd289a

Observation 775a4ac7-8ab4-49e9-a83f-433612101e38 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

High-Resolution Image Synthesis via Next-Token Prediction Scaling rectified flow transformers for high-resolution image synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.948763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.948763Z digest=sha256:b803a1f4abd19dc5eed7f38f6e6fb8e8f05d7eb18c5421dbb9f1c648f513afbb

Observation 123275f2-91bb-431b-84be-8620261977bc · outbound

This paper cites Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens.

High-Resolution Image Synthesis via Next-Token Prediction Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.953638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.953638Z digest=sha256:734b5a7e8ece67ba7ea9a6f2e5e6147e348d4ae0a466f1dc6fdeaf388ff68af4

Observation 827f6d5a-a70b-4b57-a91e-8f14fd744b5c · outbound

This paper cites Ernie-vilg 2.0: Improving text-to- image diffusion model with knowledge-enhanced mixture- of-denoising-experts.

High-Resolution Image Synthesis via Next-Token Prediction Ernie-vilg 2.0: Improving text-to- image diffusion model with knowledge-enhanced mixture- of-denoising-experts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.961852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.961852Z digest=sha256:9c88aa31a606a14d9f63cfd6a8550058397d36c5d34a240686ea9f1412272be3

Observation 30e130dd-b2b9-4a9b-804b-43137a2e8500 · outbound

This paper cites Boosting Latent Diffusion with Flow Matching.

High-Resolution Image Synthesis via Next-Token Prediction Boosting Latent Diffusion with Flow Matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.966445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.966445Z digest=sha256:904034aefaecac30e9f56af799707ce9f5abe16de32e0086bc13966294153eb0

Observation a64af342-f6fd-43cc-a7ee-a4bf09b8903a · outbound

This paper cites If: A github repository.

High-Resolution Image Synthesis via Next-Token Prediction If: A github repository

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.971810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.971810Z digest=sha256:72a27392aa15bd4a888ac9f4a712fe53cd0efa2b510bba3071d5a24e0fc48e9f

Observation 9e7b1660-f2e0-45cf-9b58-eb7aebe52921 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

High-Resolution Image Synthesis via Next-Token Prediction Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.976670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.976670Z digest=sha256:9bedb5f43affd1716cd3fdc4f3b1c72981db45ae58fe10da16b7447336a274fb

Observation 493dd55e-efbd-4bb5-9e0a-b3d3d770b1bd · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

High-Resolution Image Synthesis via Next-Token Prediction Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.982687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.982687Z digest=sha256:7631fbe1cb480b59cbe057324ccac8942fabd2c83cf9a3b44a8d84d90ba24bb1

Observation 171b0ba8-928f-4bda-a73f-14e0cc7d0887 · outbound

This paper cites Photorealistic video generation with diffusion models,.

High-Resolution Image Synthesis via Next-Token Prediction Photorealistic video generation with diffusion models,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.989159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.989159Z digest=sha256:75b1990059740c567ad3b9f7aa7545824bc42a5412ab2fc17330ca742b2ffc71

Observation 8e41062b-ce9d-4c5b-b924-3a182ad37fb4 · outbound

This paper cites Deep residual learning for image recognition.

High-Resolution Image Synthesis via Next-Token Prediction Deep residual learning for image recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.994325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.994325Z digest=sha256:5b48efef136ed6a82069cf77b2a68f14293830967dd32ff4e40a289b141c06ba

Observation f604fd9f-0ae3-48bd-a25a-59023889cfdf · outbound

This paper cites Masked autoencoders are scal- able vision learners.

High-Resolution Image Synthesis via Next-Token Prediction Masked autoencoders are scal- able vision learners

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:41.999554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:41.999554Z digest=sha256:33cea733de49ac2378fa2bc35b0866c6fb4ce320f1b8b6147dd584b77afc9296

Observation 01123c0b-380a-44ab-96da-ae7b8e6f2770 · outbound

This paper cites Rethinking image aesthetics assessment: Models, datasets and benchmarks.

High-Resolution Image Synthesis via Next-Token Prediction Rethinking image aesthetics assessment: Models, datasets and benchmarks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.009907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.009907Z digest=sha256:1257d3ad98df57d75b9de9ac698afc201870547c4d1476fba71f278da06d8912

Observation 346414f4-f99a-4b80-8222-d80be90dbfc4 · outbound

This paper cites Classifier-free diffusion guidance.

High-Resolution Image Synthesis via Next-Token Prediction Classifier-free diffusion guidance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.015663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.015663Z digest=sha256:83ed62f338007c98fe7fbf5e6ee6bd29f49633299aa1b83f303cd52d4f0f93cb

Observation ad773235-83c2-40c7-afe0-53d6bed67ec8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

High-Resolution Image Synthesis via Next-Token Prediction Classifier-Free Diffusion Guidance

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.022584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.022584Z digest=sha256:c207367f0b045756ca2cb9b4c390dd909f128d94a74e8fc30a3312c551c5e815

Observation 8ab4d461-7a3a-4ec3-8b19-39ec996c897b · outbound

This paper cites Denoising dif- fusion probabilistic models.

High-Resolution Image Synthesis via Next-Token Prediction Denoising dif- fusion probabilistic models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.027402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.027402Z digest=sha256:54f37fdaf7a3f08484170d5ca7dc075757051623fc8d8afd8fec3d17e1384f63

Observation 0c4c436e-0262-4875-bac0-a25a22202396 · outbound

This paper cites Visual comparison between HunyunDit [65] and D-JEPA·T2I.

High-Resolution Image Synthesis via Next-Token Prediction Visual comparison between HunyunDit [65] and D-JEPA·T2I

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.035731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.035731Z digest=sha256:a4ccc2c4be11375acd9cdc1f44cdb68440a607a0675cba68e09be3672041944f

Observation 625453d7-db29-4d20-98de-85cc8b11e3ec · outbound

This paper cites Cascaded diffusion models for high fidelity image generation.

High-Resolution Image Synthesis via Next-Token Prediction Cascaded diffusion models for high fidelity image generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.041829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.041829Z digest=sha256:e7b3d2369807213179d1be00652888f62aa31ab5467f503ec16222d741393102

Observation 46007515-6fb6-4363-ac6b-3047160199fe · outbound

This paper cites Training Compute-Optimal Large Language Models.

High-Resolution Image Synthesis via Next-Token Prediction Training Compute-Optimal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.047079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.047079Z digest=sha256:dcb6a7c8b965cbc6c8f0ed78c5708416993222a750fd554ce576940391f864af

Observation c55bc952-1cd6-4db5-9232-4e47da42a719 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

High-Resolution Image Synthesis via Next-Token Prediction T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.051812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.051812Z digest=sha256:2a1506f6aad167fb5e1f5266be1018aa33cd84e074828b132ce4170755381c14

Observation f5051507-7291-4b4d-9bd4-7c150936dc17 · outbound

This paper cites Estimation of non- normalized statistical models by score matching.

High-Resolution Image Synthesis via Next-Token Prediction Estimation of non- normalized statistical models by score matching

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.056336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.056336Z digest=sha256:d1377f1607d9415139c5137a9cc324cd719e13c51abbcbf9f8b70cf3de619c7b

Observation 797ad938-a9aa-4dc9-ac36-b168be063a4d · outbound

This paper cites an unresolved cited work.

High-Resolution Image Synthesis via Next-Token Prediction Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.061298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.061298Z digest=sha256:76f9d39fff7a20ce7f6708c7d4ced34b28e4f85c329990a5d6550ee805b45ae8

Observation 03962d96-9ce8-42eb-a06a-a7bb6ff0cd4c · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

High-Resolution Image Synthesis via Next-Token Prediction Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.065590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.065590Z digest=sha256:bbee80b4f8b03db1688eb1af48995b401061cc16a5d02ced9c2027a2f4deffd5

Observation 3674d533-0937-4fd2-b872-58f55359acb1 · outbound

This paper cites Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction.

High-Resolution Image Synthesis via Next-Token Prediction Computational Tradeoffs in Image Synthesis: Diffusion, Masked-Token, and Next-Token Prediction

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.069501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.069501Z digest=sha256:844dcfc64efe9471b2d0911940ed779ba738b4ffc7b8d68781be3437fec0d31e

Observation b43e82ad-9cde-4715-bc95-04c89d18399c · outbound

This paper cites Understanding diffu- sion objectives as the elbo with simple data augmentation.

High-Resolution Image Synthesis via Next-Token Prediction Understanding diffu- sion objectives as the elbo with simple data augmentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.074264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.074264Z digest=sha256:c18ea5de24ff9226e2756ea3a40eacbd54e4e845cc51231fae28dca55f886103

Observation 93249733-0401-44bf-96ea-109e8e4bc64c · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

High-Resolution Image Synthesis via Next-Token Prediction Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.078521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.078521Z digest=sha256:95c456255ec95fafd2f1e86e945c8d3c4bdfc2707f53fd299d8d1ba00a528a15

Observation f94f2d03-08da-4764-93be-0635b5ba7ea5 · outbound

This paper cites Flux.1: An open-source image genera- tion model.

High-Resolution Image Synthesis via Next-Token Prediction Flux.1: An open-source image genera- tion model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.083271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.083271Z digest=sha256:ae55d0dcd923ea0dbb7da4cd62240f28e79b9cd11b8f589b437287444b2aaa71

Observation 32bd2fd6-2c0d-4cf2-8970-78b683854cb0 · outbound

This paper cites Bloom: A 176b-parameter open-access multilingual language model.

High-Resolution Image Synthesis via Next-Token Prediction Bloom: A 176b-parameter open-access multilingual language model

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.088512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.088512Z digest=sha256:8e7862b8f387035810ba0a559706398100ddd21374a18f7671b3c62d46f963fb

Observation 22183a6d-747b-4455-ab08-2e92c642dd15 · outbound

This paper cites Minimizing trajectory curvature of ode-based generative models, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Minimizing trajectory curvature of ode-based generative models, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.093959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.093959Z digest=sha256:b3363e1147a30a77fec5e959e6b48e04e051421db86786dce9a3028402f543f8

Observation c00f906e-fc29-4576-8777-748cefb19812 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

High-Resolution Image Synthesis via Next-Token Prediction GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.098059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.098059Z digest=sha256:4e44f82875af53479d32659786c5bd5eec5d4f71e93480eb70a2ca02dcfd75f7

Observation 5b8e5a99-7f69-462f-928b-6f66e0c95504 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

High-Resolution Image Synthesis via Next-Token Prediction Autoregressive Image Generation without Vector Quantization

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.102956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.102956Z digest=sha256:bc18a4ae216c887ce201f81bc1338ca326ea6b02d06ee0f66c0a4d877652affa

Observation 981345a0-6025-41c6-a7b9-3e4e260114b9 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

High-Resolution Image Synthesis via Next-Token Prediction Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.108525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.108525Z digest=sha256:af08e999d4f644e103d9f878caa47079cce246c4f4ca6d2cc996f5f213362ce0

Observation ce07c99d-54cc-412d-86a2-fed3b7a3d65d · outbound

This paper cites Rich human feedback for text-to-image generation.

High-Resolution Image Synthesis via Next-Token Prediction Rich human feedback for text-to-image generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.114847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.114847Z digest=sha256:7f3cb5ad1604761d96c6f13c3ed6cfc23606122c37c2339c2528e1d6d9adc2ad

Observation 973cf82a-65dd-4955-a322-fedb30a20778 · outbound

This paper cites Flow Matching for Generative Modeling.

High-Resolution Image Synthesis via Next-Token Prediction Flow Matching for Generative Modeling

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.119342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.119342Z digest=sha256:c105e97f6a07cc8e244fb5a2b430c67957f67cdc7cf457f266bf6edbdc5ad31b

Observation a146a521-5fcc-44d5-bccb-3d0e092cd65c · outbound

This paper cites an unresolved cited work.

High-Resolution Image Synthesis via Next-Token Prediction Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.124347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.124347Z digest=sha256:5c4218112b514990371ea1a6a8727c47d5b4f7e5f48c040c8aa16df0068578c2

Observation 3ab9b0db-78f6-4386-8037-753973592291 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

High-Resolution Image Synthesis via Next-Token Prediction Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.128419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.128419Z digest=sha256:a2c1072333cbf84795b226cc65a308c0310f83bf460ed5e6b3f711a15bfe0c40

Observation 5cfc6750-c564-4847-9fff-7e049a86656e · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

High-Resolution Image Synthesis via Next-Token Prediction Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.133163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.133163Z digest=sha256:d3e41ce359ca508cc60a42dec77cb069e1af7bf44d5159174516ba63011274b2

Observation dad1a6cd-6496-4255-a290-0da1f7e2bcda · outbound

This paper cites Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Instaflow: One step is enough for high-quality diffusion-based text-to-image generation, 2023

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.138637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.138637Z digest=sha256:a64302ddf2b4dda3892a87cb94f0f2a4017085dd223d60d02ea1608fef1069d1

Observation 7501e36b-52e3-400d-b252-ca053b142071 · outbound

This paper cites Decoupled Weight Decay Regularization.

High-Resolution Image Synthesis via Next-Token Prediction Decoupled Weight Decay Regularization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.144742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.144742Z digest=sha256:c9e15816457649dbebf82bd58e0f428c6317c49e3bcc0143ef308331c8fbb944

Observation f6d2ce02-bb0e-4e7e-9d77-96ec4d540d00 · outbound

This paper cites Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models, 2023

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.154155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.154155Z digest=sha256:3b397d1d45df148f92b3c97d61222f1e9279eb481b91c16a6323f85d1b163c7f

Observation a67f0147-cfae-4eb5-95b4-1c9a075bec2c · outbound

This paper cites Albergo, Nicholas M.

High-Resolution Image Synthesis via Next-Token Prediction Albergo, Nicholas M

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.184771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.184771Z digest=sha256:583a4d5003d21b43c2d81d05b24e044bdcfc4fe412605384c6532d5062080026

Observation 184e538d-03d6-4500-8c57-e860fa9ce292 · outbound

This paper cites Midjourney v6 - ai art generator,.

High-Resolution Image Synthesis via Next-Token Prediction Midjourney v6 - ai art generator,

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.190069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.190069Z digest=sha256:91c632ad26da833947cdf5105cb30f4b3217ba56752b86a443ef5fced9a8c02e

Observation 720488a1-b541-4823-94a0-3abf6b0ae351 · outbound

This paper cites Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition.

High-Resolution Image Synthesis via Next-Token Prediction Freecontrol: Training-free spatial control of any text-to-image diffusion model with any condition

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.207634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.207634Z digest=sha256:6a17af912bb872b537cf4704ec0acaf5fa7586fff206712e965a6b986dc0f755

Observation eb74614f-2570-48b6-8470-889085f45230 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

High-Resolution Image Synthesis via Next-Token Prediction GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.211836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.211836Z digest=sha256:8181d636e6cfb8d0b062899d2551ace62a991510bfdaa7a154f7a437601cd9d6

Observation d9703236-9be4-4609-b8b7-04899cf26c66 · outbound

This paper cites Training lan- guage models to follow instructions with human feedback.

High-Resolution Image Synthesis via Next-Token Prediction Training lan- guage models to follow instructions with human feedback

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.219595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.219595Z digest=sha256:2069ca2609c43dcc6f42ef80e37ebc0baef01cfa8a0be48eff1711bb80d70081

Observation 9bbf1d3b-efe2-4016-95ab-d5620993449f · outbound

This paper cites Pavlov, A.

High-Resolution Image Synthesis via Next-Token Prediction Pavlov, A

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.225299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.225299Z digest=sha256:54e1649cf22eeb3f4362224ea70a38aa348b7fe67f991f5b8795093b2b83d79e

Observation b235576b-496f-4815-b535-6671cefa6d87 · outbound

This paper cites Scalable diffusion mod- els with transformers.

High-Resolution Image Synthesis via Next-Token Prediction Scalable diffusion mod- els with transformers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.235756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.235756Z digest=sha256:b981454915c3eb4b3a25e4a609d934cdafa59c79244e13f1bb2d995efde43907

Observation 532bff5e-111b-47a4-af38-cdae587a83db · outbound

This paper cites Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degrada- tion, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Ntk-aware scaled rope allows llama models to have extended (8k+) context size without any fine-tuning and minimal perplexity degrada- tion, 2023

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.245981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.245981Z digest=sha256:89041484d033ccbb3133e57e750926ef17457cbc8d1232ed112d7ce6e6300493

Observation 0d26dbf9-3e13-46bc-9647-05942d5ac412 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

High-Resolution Image Synthesis via Next-Token Prediction Film: Visual reasoning with a general conditioning layer

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.251520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.251520Z digest=sha256:233cf0afdf8de53ab2a398b3a89d727dd7a3129acd0489455cb2cda6c89f92b8

Observation b3ebb336-6d74-48fa-8eb7-5e418d594e38 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

High-Resolution Image Synthesis via Next-Token Prediction Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.258081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.258081Z digest=sha256:38ae4e7b5d981689a42edb09573d30bffdd539d3ced3d8aec6cb1102196b4304

Observation b55c53be-783d-43ae-a606-839ec604809d · outbound

This paper cites an unresolved cited work.

High-Resolution Image Synthesis via Next-Token Prediction Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.265203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.265203Z digest=sha256:cc49afdded45f0171522382d510283a9b4ab3546d45947643c41fef76b1fd6d0

Observation dd30535c-b4b3-4531-adb1-884d277402de · outbound

This paper cites Improving language understanding by gen- erative pre-training.

High-Resolution Image Synthesis via Next-Token Prediction Improving language understanding by gen- erative pre-training

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.271730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.271730Z digest=sha256:35e0a73999f823e9199e2d2941b12a68d29d541020cefce1aa1fdeb280c0963d

Observation 523f7d07-4038-472a-894c-d77794d8fcc9 · outbound

This paper cites Language models are unsupervised multitask learners.

High-Resolution Image Synthesis via Next-Token Prediction Language models are unsupervised multitask learners

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.279705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.279705Z digest=sha256:280a071665a52d6d1d76ce84323e228b59d7191e6985bf3000fb176870ceecc6

Observation aef469e7-8091-4bd4-a64c-d2b2e709b7b1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

High-Resolution Image Synthesis via Next-Token Prediction Direct preference optimization: Your language model is secretly a reward model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.284217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.284217Z digest=sha256:41ae4cc355c76729ddd0a3d15d045811824c63e3f52118e6a42d2e02567bd4f6

Observation d9550ff4-1773-4d70-a7f1-68a102e830d0 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

High-Resolution Image Synthesis via Next-Token Prediction Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.289957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.289957Z digest=sha256:d01317cb767ae9e407a94f2433d3ce0054fc1c585a9a8d8a7829449b74158f6e

Observation 7c6fbb83-ca38-49dc-97f0-dd0e5f929ba5 · outbound

This paper cites Zero-shot text-to-image generation.

High-Resolution Image Synthesis via Next-Token Prediction Zero-shot text-to-image generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.306062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.306062Z digest=sha256:a995ff0a12428af9b31dff39aa6b0ee429b572a038bd2bbbae81e051ab8df8bf

Observation acf0ed78-5932-4687-8357-83f0fe47cf56 · outbound

This paper cites Hierarchical text-conditional image generation with clip latents, 2022.

High-Resolution Image Synthesis via Next-Token Prediction Hierarchical text-conditional image generation with clip latents, 2022

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:44.054627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.313742Z digest=sha256:b8b39f412ae3b94ef2e4803e2a8b6484ececa09654927359fb2260d655c332c6

Observation fb5d152b-6ea2-4863-a1d3-ba8902b12b47 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters.

High-Resolution Image Synthesis via Next-Token Prediction Deepspeed: System optimizations enable training deep learning models with over 100 billion param- eters

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:44.041121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.318387Z digest=sha256:4327b0fd6072fd3eee281a0b9fa6f864e7572201c104a934f3cde30654398bf8

Observation 66626cc8-837f-4fed-9111-ebfdb6106ec1 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

High-Resolution Image Synthesis via Next-Token Prediction Gener- ating diverse high-fidelity images with vq-vae-2

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:44.026994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.322486Z digest=sha256:f8c10e975037a22aeb57a8b9e4b46c3bce5cb20dec5d79dd2ffe4e7636952909

Observation 0ee604bb-cd28-485a-8613-4341e0824535 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

High-Resolution Image Synthesis via Next-Token Prediction High-resolution image synthesis with latent diffusion models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:44.014565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.326308Z digest=sha256:571ccc122ef5f5cc7600aa36d0d48a1c251db4c24a83bfaf6dd6ca23b0ee9603

Observation a421d31d-c8fe-47f5-ab1f-e1b1294a76ef · outbound

This paper cites Im- agenet large scale visual recognition challenge.

High-Resolution Image Synthesis via Next-Token Prediction Im- agenet large scale visual recognition challenge

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:44.000541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.331271Z digest=sha256:cb118eaed7a4012dac8323c29c08c36a53af2b736ed1df084fce4acbcf2c840f

Observation bdfbfe9f-f59b-4095-bf98-a26a3d2d8d47 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

High-Resolution Image Synthesis via Next-Token Prediction Photorealistic text-to-image diffusion models with deep language understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.335352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.335352Z digest=sha256:3b3a35c5d5d3a542d0dcd40f2c6604cf6332e484118e141fee772fa095df440a

Observation 37572a9e-eeac-4e01-bb1f-7f1eae695253 · outbound

This paper cites Adversarial diffusion distillation.

High-Resolution Image Synthesis via Next-Token Prediction Adversarial diffusion distillation

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:43.977762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.339216Z digest=sha256:8aa4de524e501a10fc76072247107f8f33ff6ce2922146188fb987c19de8c5cc

Observation 3275ad53-baff-4a89-bfe3-fbaf859699c4 · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

High-Resolution Image Synthesis via Next-Token Prediction Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:43.936623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.345329Z digest=sha256:9cdc3718a3bd8ac26c8262b5716e36238aeb6f84fc0b79ef5d2293a90990b6ab

Observation 2f352fd4-3af4-4ace-ac93-4ab8fbb9c95f · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data, 2022.

High-Resolution Image Synthesis via Next-Token Prediction Make-a-video: Text-to-video generation without text-video data, 2022

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:55:43.921928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T14:55:42.349972Z digest=sha256:c65cb31ca23472a40aa5f189d70d144079461bfe6f5aa6dff48608e00c6228a4

Observation 688b3afa-b894-4c42-828a-57d96c244586 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

High-Resolution Image Synthesis via Next-Token Prediction Deep unsupervised learning using nonequilibrium thermodynamics

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.357262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.357262Z digest=sha256:a1c14953c6c682a68af9fae5975cd8b3345db23377a17563cb402ba5c0fea5d7

Observation 489d3f92-fdf3-4947-8d14-095fd99fb28d · outbound

This paper cites Denoising Diffusion Implicit Models.

High-Resolution Image Synthesis via Next-Token Prediction Denoising Diffusion Implicit Models

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T14:55:42.363118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:55:42.363118Z digest=sha256:f688121f3dc27e9ece8ba4f0d5df0588b52ce8919083bf05e5feef0fa6e1666b

Pith citing papers

No inbound Pith citation observations are available.