Pith. sign in

Paper Citation Record · LEDGER

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 2 inbound Pith citation observations for arXiv:2505.10046.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10046 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:21:22.544386Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T20:44:20.190064Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T20:52:37.899434Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 84166abf-e733-4b08-8131-92ecb78713f0 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.301918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.301918Z digest=sha256:f61bf3cde0bc1a73e501a0b755a36070c63f54618bedc97499108bbb2e82fabf

Observation 38107a9c-d4ce-48ae-80e2-1ac0682d6741 · outbound

This paper cites Improving image generation with bet- ter captions.https://cdn.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Improving image generation with bet- ter captions.https://cdn

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.411407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.307045Z digest=sha256:17e3939374d659de219e102279cfddcdfb05a0c0a878daffb2d63ed338677574

Observation 099dd9c6-6957-48f1-8e80-d41b57328739 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis PaliGemma: A versatile 3B VLM for transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.312956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.312956Z digest=sha256:d52f0b87f18dcd143e3860a53f83d1110a356b73a73a0cc6c3cd7211cd3f3ec8

Observation 6924cf41-b96d-4d04-b941-de9e9bf73817 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.395722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.317960Z digest=sha256:f2b4483424379a8087c5c3ece3fd9852f29f8fa16b8191a570ba1d20b34005f9

Observation 1d980851-702e-4225-8b0b-a89b2b9e4e75 · outbound

This paper cites Pixart-σ: Weak-to-strong training of diffu- sion transformer for 4k text-to-image generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Pixart-σ: Weak-to-strong training of diffu- sion transformer for 4k text-to-image generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.322726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.322726Z digest=sha256:1145852a588df4c4399f2548ce3a4da1148b7b5b31f94517bb188ba6415ac9e9

Observation e5d32a4a-2b8f-4010-8297-357af72623d9 · outbound

This paper cites Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.327511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.327511Z digest=sha256:922b047eb857bc6786e20b5b1b53b5fe89364df414da414fa3f0e04c4968dc35

Observation 7d1e53b3-7ec3-4164-a779-abcef857e1f3 · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Dreamllm: Synergistic multimodal com- prehension and creation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.360640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.332279Z digest=sha256:1901c7a601c94a9a9135417aa97aeccdb69f2bbbe28713473dc8d646c0ed010a

Observation 18600dbf-4d7b-4c32-9c3e-50e178a4a01c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis An image is worth 16x16 words: Transformers for image recognition at scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.337102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.337102Z digest=sha256:49d3ae41a632f35922914edeb7a6115d35faf230b8165d7ce82344347436e4a2

Observation ea9feaa5-d32c-4a32-97ab-86555854e8b1 · outbound

This paper cites conceptual-captions-cc12m- llavanext.https://huggingface.co/datasets/ CaptionEmporium / conceptual - captions - cc12m-llavanext, 2024.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis conceptual-captions-cc12m- llavanext.https://huggingface.co/datasets/ CaptionEmporium / conceptual - captions - cc12m-llavanext, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.329413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.342342Z digest=sha256:0b1e536637f9c8c35f6363a80b6fdab6c213678101a0448d7e7b63f22afcd66b

Observation a2d01e66-b909-4997-8f7d-94bdd26dd061 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.310435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.347166Z digest=sha256:bfb873a7b0613a89597bebbd138b4d43ef1cf2996d24012283006a5009479fba

Observation 5d519a93-288e-437c-85e1-b0880a5f1302 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.351916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.351916Z digest=sha256:4185878a8b142bfe1922e6ae944f2d3e6584f66186ce6b496519d921eb737c04

Observation 1c1b25db-b937-4edf-a7e1-d762c5b8e6bb · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.357600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.357600Z digest=sha256:3a8fe9743e064819ec5eca30d9186f1543e7ec5496d5c9c87df4830f43297e61

Observation 87813c34-a662-4e5e-935a-201f7f17690f · outbound

This paper cites The Llama 3 Herd of Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.362449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.362449Z digest=sha256:77735341eca67c378de26e9fa954dece5784834482832d1e53de1b1ceb959a95

Observation 9daa17b9-ba47-4df7-8437-507b94b4de11 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.293896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.366760Z digest=sha256:3c6833b3f844459c7d8e8d34e9f9ec95ae7ed5a7e131a842b01e0ebe51bd0f08

Observation cba3a9ee-932d-4840-977a-abba7820078c · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.370956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.370956Z digest=sha256:6404d0b86fb9f0cb434604361e515bf444b097085bb40c5a14212e4c3c00fea6

Observation e4628444-0840-4eb3-8b39-5cfed74b87b6 · outbound

This paper cites Segment any- thing.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Segment any- thing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.276634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.375713Z digest=sha256:6414e4313d9b724a6382d68ae422c440e37ac5d5c92feec3863390606fda38ef

Observation d5f0aabc-8081-4ed7-8bc5-be63e2764acb · outbound

This paper cites Announcing black forest labs.https: / / blackforestlabs.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Announcing black forest labs.https: / / blackforestlabs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.259190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.379767Z digest=sha256:9eb9e4e01a3e6d76e66afc95c6d388945fd5863e96fe67103c11c3cc21599f28

Observation ed48c7a0-a327-49ce-ad36-75df7b37c629 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.384379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.384379Z digest=sha256:8fb304dfc22068df5d9f919d2fbc576ccdadaf91d4a77335290266d7f6c275f5

Observation 9b08194e-ee6d-4bc4-b43b-b3e109e38284 · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.389546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.389546Z digest=sha256:18eeec115dcc5c3bc7aa68964a3786d68abcc2d39c2de8e82717de9f7142c6b6

Observation 0bef1506-ddd8-453f-ae31-edce49ed054c · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.394751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.394751Z digest=sha256:1dcd01e813d7fefa3447eff525c43d47f971fd94db3558f2742f98cb2eae73b7

Observation d548729c-5054-4608-bfba-5b440b84d94b · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.400326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.400326Z digest=sha256:a2329027576e47542a22f776c7d0d93c2c1a4a23f55304bdfe71b8899bd4db0c

Observation e9d11945-1a16-4623-beba-1900f437cd1b · outbound

This paper cites LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LLM4GEN: Leveraging Semantic Representation of LLMs for Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.405865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.405865Z digest=sha256:2a011402bbfe247adc8f45f848ab2f05b716b7e1e8d8eac69abb1323552f410f

Observation 9a13fcdd-d4c0-417e-abf1-64ce1ff25263 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.410575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.410575Z digest=sha256:b262ce845b2383e1dca53da9b488bcb36d3f4d309f859f77539923ca36dbc640

Observation 67fc0407-e9f6-4173-a7b4-8f0d5506609f · outbound

This paper cites Decoupled weight decay regularization.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Decoupled weight decay regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.415252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.415252Z digest=sha256:629d64cd4ac7c89649687ea3df5acb8fe19b5c7d9f6b4b2d6a717e586f436940

Observation a07767c5-ac48-40b4-a783-d8c948b9a28e · outbound

This paper cites Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.419884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.419884Z digest=sha256:d96e56d5c4c94aa379263ab505e9d26c5ba85b3f96e315e535127264c9866833

Observation a79a9207-d8ca-48fc-9664-19cfb0e9705f · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.425364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.425364Z digest=sha256:11768c43c3ffb48cbd480c03302850cc674573695577fd0623d0c83ff2389de7

Observation 13c44ca0-43f1-49fb-a9ff-afb09bda4419 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gemma: Open Models Based on Gemini Research and Technology

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.430243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.430243Z digest=sha256:a308370d731998bcfe868b43ac7040ec10d85b083b2f08a0eeaf2f18441aead4

Observation 74ab8f56-896d-436d-b5a8-245951ea52ae · outbound

This paper cites Kosmos-g: Generating images in context with multimodal large language models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Kosmos-g: Generating images in context with multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.231608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.434395Z digest=sha256:5c582e59e221509f1a15a0c91768f20d57cf84d10ba30c77ece32050a55fc3cf

Observation 12de4571-f47f-4484-a08e-f83f1b68a468 · outbound

This paper cites Scalable diffusion models with transformers.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Scalable diffusion models with transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.216837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.438510Z digest=sha256:82564a27eb89a9a8f18525e58b3c4c596189303ceaf901c3d641d57d548b946e

Observation 41e8791c-670a-47a2-a30a-b81753ade952 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.442388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.442388Z digest=sha256:c75d20d2c38a2900836ecfe15382aa8687d7ffcbecb2503670c511b7da04d319

Observation aaae57d3-25d5-4324-9ede-01b31b3f3f7f · outbound

This paper cites Learning transferable visual models from natural language supervision.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Learning transferable visual models from natural language supervision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.200891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.447075Z digest=sha256:0a230cd3b0d2976e5c08ff6b5a14c66195a4b8e3e062c40647076bfb06ada853

Observation fdd8e30c-961c-4114-8a27-3cbf787658ba · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.185200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.451755Z digest=sha256:639966e0038634fc9a7172a29f0a17a48f446303c83aedd727d33bab6c5c414f

Observation 1c29957c-fcc7-458d-97c1-17ac02b26bdf · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.456347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.456347Z digest=sha256:5dea5896a7af1cd1c2dc2d81a38c3dd88e2da5aa9439305845fb42f96e93b314

Observation 766d0d3c-8087-4f3b-a042-c97d9b6bed46 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gemma 2: Improving Open Language Models at a Practical Size

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.461406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.461406Z digest=sha256:2431ef5fc8df6ddd661b540f7860998ac8616db4aa093b5775726f6e7844391f

Observation c6d5c7ef-ac60-426d-a695-ce10885587f4 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis High-resolution image syn- thesis with latent diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.170014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.466130Z digest=sha256:9112a3a14dcee11a34423406ac87b5ff71d32ef56b412e8172f19890be247005

Observation fce47d88-2f5d-4a7a-a3f7-620ede3bba3a · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis U-net: Convolutional networks for biomedical image segmentation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.471083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.471083Z digest=sha256:3b25efb16b04c758130fb7a0a5137e37578dec82bd9f136c6f75b497b0a08814

Observation 770ff9e0-87c8-4cb8-9e05-40f25708c8ea · outbound

This paper cites Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Stretching Each Dollar: Diffusion Training from Scratch on a Micro-Budget

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.475887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.475887Z digest=sha256:438e80131fc960689a1b32b8049bc3cfc684efddd7d1ddae9a2606f58c78af67

Observation 04671a1f-5f16-4c92-8f25-e16a2ac660be · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.481319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.481319Z digest=sha256:28d58f9e619b83dcdd7d769072323b7f0cdab877930b5314406c63f1988a18c5

Observation 32912298-69e7-4213-9abf-079c51bc419b · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.487821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.487821Z digest=sha256:219f76279ba8c806b76fa80a49f6ea4988d067ab5c7e1e59e8232cff564b8ec5

Observation 0726554d-c66b-4672-a230-a6fb5c664a3c · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.145142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.492712Z digest=sha256:780d4e0c1a0d193833d4ae286f21bdd3cc0f38b502d54cc1bc48c77251cc6227

Observation 4e2894ff-197f-47f8-84d4-4fca02de2027 · outbound

This paper cites Journeydb: A benchmark for generative im- age understanding.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Journeydb: A benchmark for generative im- age understanding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.130660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.496996Z digest=sha256:d800d5eb3780929da91a28966671d6faa7e90a096e49ca8723398c148d41daf7

Observation 3b2636fd-e03b-4e50-8807-cb94f2c15be9 · outbound

This paper cites Is noise conditioning necessary for denoising genera- tive models?arXiv:2502.13129, 2025.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Is noise conditioning necessary for denoising genera- tive models?arXiv:2502.13129, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.501812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.501812Z digest=sha256:d9609852b415449deeff58aa53c1f864056b8e34f2caeb669c006d960922b84c

Observation 7ab2af98-7c52-410b-9a57-5542dc66eef5 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.506037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.506037Z digest=sha256:e5d196a8aacbe39bba30a052fd605fea930404ff547f7b2a6288c20f5b9acd6a

Observation 45c80d04-10c2-482d-b653-eea19c028557 · outbound

This paper cites Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Divide and Conquer: Language Models can Plan and Self-Correct for Compositional Text-to-Image Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.509934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.509934Z digest=sha256:db89ead8369e1a90b24894f9a9f97a587986741a32d8e8d2d76607c94ede658e

Observation 4f92cb59-87ca-4751-bb31-ecba7051fa41 · outbound

This paper cites Self-correcting llm-controlled diffusion models.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Self-correcting llm-controlled diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.102824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.514397Z digest=sha256:f8e8f1aa25d32332a68308f256bc1bb9f4d071fc1ee42e1319ead528a3a82e64

Observation 349b1ff8-2fd8-4e95-b63e-8894acf49083 · outbound

This paper cites OmniGen: Unified Image Generation.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis OmniGen: Unified Image Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.519750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.519750Z digest=sha256:86447c0be5e3d5f4623ec61a15641b8bc84d58347d6e9f415c6ef192ec3a1c0c

Observation 1d3a0de2-db54-4969-b304-7f068e3226fa · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.524502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.524502Z digest=sha256:d249430af19f8193299a6c47e9a35a611ac8be55809eeca29678cc4b937cc687

Observation ccb77f9a-1191-4d2b-b06b-5e21026ca24c · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.529564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.529564Z digest=sha256:15444308a851e6f7c2d98a9cc917e8a860416e4915f802ee3c85696807d30eea

Observation 9ccefb93-3e28-44bd-934f-c88b21e75d7b · outbound

This paper cites Mastering text-to-image diffu- sion: Recaptioning, planning, and generating with multi- modal llms.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Mastering text-to-image diffu- sion: Recaptioning, planning, and generating with multi- modal llms

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:21:23.086820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:21:22.534197Z digest=sha256:21abcf41e209e7f403ea97077699ba9099e84341c2aa58caa788bd15a45419c1

Observation 24a5858c-0b98-4d62-94c8-41a92f6c5dd5 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.538806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.538806Z digest=sha256:223c2f1fbbf9a49de172615d4a1651a4fe35f8b1585abb5d4c15f8ba60e5af6f

Observation 7c23bef9-2e55-45cc-88e1-922db9ce5157 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T21:21:22.544386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:21:22.544386Z digest=sha256:7b33e8e5dc0e7ce85f1fd512bce172ea3b660c8f412b881a45e9ea09a241d97a

Pith citing papers

Observation 20536b70-f7f7-486f-aba4-ca14b5b728e9 · inbound

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds cites this paper.

Exploring Time Conditioning in Diffusion Generative Models from Disjoint Noisy Data Manifolds Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.483764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T16:52:29.681159Z digest=sha256:2f53c474bcbe2da4b3163531e0279833fa1656621cb1b9c48bcb23445e3f287f

Observation 1453aca1-62fc-4390-8b98-9d0544cdf4a0 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Exploring the Deep Fusion of Large Language Models and Diffusion Transformers for Text-to-Image Synthesis

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.900963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:d96f05bf49bf913d2cf3f3f86a97823c40bd05d49d4cdeeee22a957c2426a54f