Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

As of 17 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2506.08210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08210 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:51.244779Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 398c47e8-fb10-4e8a-a0e9-f818b192f0c9 · outbound

This paper cites Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.977218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.977218Z digest=sha256:4bf4743b42cdc81dad57879700cc455a028993a2a6f1b0e30e85d899d8dfdf2c

Observation aa460d2e-26ff-4f00-8a19-b5c09a34d2db · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.982334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.982334Z digest=sha256:7a7453c132f90be26cd38513ec652b0b23894111df06b69b3d01b88d7dd611a3

Observation 61d7c6d3-68f9-41d1-a5cb-65495b3ba30f · outbound

This paper cites Imagen 3.arXiv preprint arXiv:2408.07009, 2024.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Imagen 3.arXiv preprint arXiv:2408.07009, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.986447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.986447Z digest=sha256:c43165770cc46aab62353d29664ce8ee17eefa31c4765d17f390c9439a835744

Observation 1634a2a4-269c-4570-bf60-f9799e1fad0e · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.990041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.990041Z digest=sha256:3d42327a5582e45994c34de13ec983e4e9d03474bf06e273bb1b47177465861f

Observation e697754a-d932-4d75-9ef7-6b3d7692d37a · outbound

This paper cites Scalable Performance Analysis for Vision-Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable Performance Analysis for Vision-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:20:51.597266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:50.994543Z digest=sha256:ac6b29195e483af1832145431e1ca74b81c9749e4b89fd54e66dc6a2c8f55b7a

Observation 9b84a85a-28d1-424a-ae5d-98817c55b024 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.998688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.998688Z digest=sha256:5500eda1cae8ddbcbb56b5f5d0d31d125f8bb42983f10c390ffee630d34d9b77

Observation c575c6f6-6984-4190-8a71-a26d056f684d · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.027527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.006114Z digest=sha256:98d25994bc6ec26d78f7446e740489ba1a43625e136823ef4e49954dac96ce95

Observation 853b019a-3335-47a9-8701-ce431034a2e2 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.010211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.010211Z digest=sha256:5dfd25caabbe69ddb79c070ca3c5f92f21c7699ac1fdaffa78dc8fbfafb466cb

Observation c497bdec-1a28-48ba-876e-9ebb948eda0c · outbound

This paper cites Visual pro- gramming for step-by-step text-to-image generation and evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual pro- gramming for step-by-step text-to-image generation and evaluation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.015758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.014921Z digest=sha256:066f4964c1ef025ef4a8b43d031fbf2523f7309cf2a1670aa03bcb20c188b58c

Observation 23195d1b-b58d-4436-8aed-e94a525398a1 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.004536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.018690Z digest=sha256:e32850518e8fd8dfe63998eebc9f4539fbcbfa570c9c46f208924acccbb822f8

Observation 9924718c-6b1b-4cff-a924-0ff9529cfff8 · outbound

This paper cites Analyzing Transformers in Embedding Space.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing Transformers in Embedding Space

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.022450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.022450Z digest=sha256:6b983c0b728ad20cbaebbd0840a7f60438d2bb2db83e3ced25a0c6c3c4be1cad

Observation e5739900-abb2-4b2c-a8d6-cef0fd7ec774 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.993821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.027375Z digest=sha256:7b0fa92513a5a8f480920030e52b2fb5d315f4dbcae5b317c161c02768ac8930

Observation becaf3ad-2e97-42fc-9820-e757bec27070 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.030846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.030846Z digest=sha256:5135ee3ddf354d6d395bf325fa8fff0f4cd9c1d2f2b422f152812f40e4fa4218

Observation 941aed29-e666-4cc1-ae4c-391ecdcb3f49 · outbound

This paper cites Visual fact checker: en- abling high-fidelity detailed caption generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual fact checker: en- abling high-fidelity detailed caption generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.982389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.034740Z digest=sha256:795278d50ae247c1617ecc50e2f89de1eca63fa67abc5afb2a113e543fe8b38c

Observation c4abb752-e600-45c3-8439-85cbfcb3b568 · outbound

This paper cites The Llama 3 Herd of Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.038814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.038814Z digest=sha256:7da448bbcec14357d1a4fbcb76ca98c7c9a76dc74fc7a99af19aec6c75c04062

Observation fbc237b3-b337-479e-8e9a-c9f999690795 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.042386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.042386Z digest=sha256:3b5f68b754c31a954d46f428da28a613617ebe3d6a0e8ffc785553ee480829a8

Observation ea91496f-857e-48f3-9b51-e200fe3e3686 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.046257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.046257Z digest=sha256:9f0c4263f48455e2c9f161f2d8b20d1ba0127c740a8b48c3fd6f000cfe7c25c9

Observation 6d8bb99b-0573-4499-ae53-3369c5a78347 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.972571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.051035Z digest=sha256:9dd35a4638c780cb7bad4436ce09d6b544d692b8a696793677992606e0ce95b8

Observation 039cbba8-4e40-41f8-b62d-971cdaf3469e · outbound

This paper cites Classifier-Free Diffusion Guidance.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.054665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.054665Z digest=sha256:a044a3546a723da254483add64162583060115b310a96674cf9d7a16ba8b14c1

Observation 927c9c3c-c01b-44c8-bdbb-8c2e2d661b09 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.058500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.058500Z digest=sha256:c2294ca0a1ac4cea13844918b81898331fe44634113a9ce81219d6128a9ded64

Observation 6f26b31b-c794-4861-bf2b-2305a5e532db · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.961439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.062379Z digest=sha256:a3396609c3de51a80b5dd79dfbecf4eff8a3739de5f6a2f3fe94f7dc512876ac

Observation 87b63063-c52e-4bc7-bce0-91687dd6dc4d · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.950410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.066724Z digest=sha256:8526cdfd72962794874fe8c44620f2642cc51d1cc2ded8cbf7c75543231cf309

Observation 23a02f78-89ab-48d2-941a-1acf24767390 · outbound

This paper cites GPT-4o System Card.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.070312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.070312Z digest=sha256:d5a4ebe0eb3517543a2048a2f00bbff9d0d73112f671518894846cc0ab0088ed

Observation 5f9e4d79-9629-4745-bedf-a538a28a0496 · outbound

This paper cites What does bert learn about the structure of language? InACL,.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What does bert learn about the structure of language? InACL,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.938958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.074586Z digest=sha256:779cb5e1dde6c091fc0270b5e5114553ea8cb451022b6ae1f09deb830da1f25e

Observation a744e85c-dbf0-40ea-8e1b-74831bdc6daa · outbound

This paper cites Mistral 7B.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.078207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.078207Z digest=sha256:419fac7111315b1fe2a3f58206d9da6a155edde9e71d27200fdc36cf78e6330b

Observation d6902ad0-3b3b-41c4-8c12-dd1929b6dc40 · outbound

This paper cites Analyzing the Role of Semantic Representations in the Era of Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing the Role of Semantic Representations in the Era of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.082207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.082207Z digest=sha256:a6ed88faaba08ecfbb716dc74a4c60c30d44b44be9f614206443a0c9226863ee

Observation 79b400d3-f8ac-4ef6-9dff-60a433fc1530 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.927459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.086581Z digest=sha256:2eeeaaea4528532847364cf3323b7464409e8090d2c603f173c94f8a3b919d21

Observation b44455d0-a6cc-4098-bd00-f884dce11158 · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing and improving the training dynamics of diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.916207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.089829Z digest=sha256:96c792421bfe442cdfea26466c218b12f5fae44c1f169a5bed67f803d57bc4ff

Observation 8f166b0a-9b9b-4618-8442-4e80d4ff018d · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.093052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.093052Z digest=sha256:fa86a9ded7b576f33d147820e3c8476421d29b2adb11e3ffe968c113ed0fd970

Observation 4093adbf-4e7f-4333-9802-e5abbf15e686 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.096656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.096656Z digest=sha256:ceab700a13883f01facb202c6ba74ea4c656c7af1f45811e16b4545878d52752

Observation ce625a39-7946-45c8-88db-e4bd5b49427b · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.100614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.100614Z digest=sha256:6179ea949f6db532b9497f2036f3092a162d4d12ec1b04ac0459091fd5575d31

Observation 6fe2513e-05dc-4728-b10f-6418a9231af6 · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.903817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.104499Z digest=sha256:704c1f3e6d91907c63b37e9af5ccb668f1b8cefb7353a99227f73cf2d5f1a74e

Observation e25b3597-2306-45ac-b073-ecdae4731aad · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Common diffusion noise schedules and sample steps are flawed

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.892766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.108275Z digest=sha256:5c6515f4046c1be50b8d62152993c557f5d7629e9a229de29e305b63e019b215

Observation 04ead64d-cb04-459d-bb57-c2d51b0d4bf9 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.882601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.112353Z digest=sha256:b8564e12794e4b167e28762156b2bd56e37f9c1819ba824f70292c43d830c386

Observation 78c10748-bbdd-4373-b842-1ba1ad57c327 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.115963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.115963Z digest=sha256:41317fd527d493e7bcc6a377b1cfcbac4eb9d42f9f916bbadd7db6a063c512e4

Observation 46685bd4-1272-4c9a-a0e6-f226a9a9bce3 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Character-Aware Models Improve Visual Text Rendering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.120173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.120173Z digest=sha256:a217a168af1f8627d4c38e38275362da8e630b3fc9645daeff06cc896a1f807a

Observation 9494912b-8c09-407f-8ccd-0cbc9de6a8ff · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.123785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.123785Z digest=sha256:311abf06db0690f9158d18a30ad309919f4f2969627288e7dddce4261b7bcb51

Observation f5a33c4b-94c4-443c-99bd-e73523f50273 · outbound

This paper cites Decoupled weight decay regularization, 2017.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Decoupled weight decay regularization, 2017

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.870995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.127292Z digest=sha256:33691414f8f1fb3371b27d58587b189338d7e3fec8bdf57e86e6d4ee05aa7603

Observation 5e1143d1-e17e-46dc-b0ad-588a8094bacf · outbound

This paper cites Salesforce AI Research’s SFR-embedding, the top performing text-embedding model.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Salesforce AI Research’s SFR-embedding, the top performing text-embedding model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.860167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.130898Z digest=sha256:eb26ca9577c6edb72ab1181b6c8decaa39d724dd8b08db9163a91be3ec454c54

Observation 08adad73-ee57-430b-936f-99f7ffdda3e4 · outbound

This paper cites MTEB: Massive Text Embedding Benchmark.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MTEB: Massive Text Embedding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.134528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.134528Z digest=sha256:cd7209086c1c3137199a84bf3562149dae0f51870c7c1f5a92b7f1767459316c

Observation b6d0ab25-1849-45fd-b079-deb6eeff7e55 · outbound

This paper cites Scalable diffusion models with transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.139116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.139116Z digest=sha256:02f83a3a1f21ed8b5dde36f3c62984f9680c0cac4a23820f5467c03f07cf9865

Observation 08aee82c-0c41-472e-9dcb-815af9af02ab · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.143069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.143069Z digest=sha256:59ac0f18b4b20ea8cdf5e47d32e55efe41a24629f5b3e095d358a67e570752c6

Observation be6f771f-b28b-44e4-b94d-fd8c5e1e1042 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Learn- ing transferable visual models from natural language super- vision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.842843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.147265Z digest=sha256:c5ca90d9f8f4d59665d961b06da43870b0d019b3d99354b16b6402edaaec2593

Observation 30bb8684-6000-4e2c-a8f7-3ee077efd3ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.832409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.150903Z digest=sha256:cc4447bae84e51ae680a28cba69429c673e628784b5372a160d8d6aaf12fdf84

Observation 8f990867-8353-45e4-9157-07fab43f9504 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.154397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.154397Z digest=sha256:9094b26ebee433c39cff92fb61687d6038571f50556346188b0f162d1c6ca0e3

Observation d242fceb-18b2-42b1-b1f2-e59bf4045b48 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.822625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.158439Z digest=sha256:26c2f332087258f529480cfd338957d9805aa7a10b099aa3ab7f7bd2d8ffad4d

Observation cf875356-8584-4dfa-b89e-7565d2e257a6 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation U-net: Convolutional networks for biomedical image segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.811734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.161994Z digest=sha256:8bb0a5a6a763766e4be95b027b023a78c7c29c8530b0c1446904961a9ea39473

Observation 1d1708d4-fdd0-4ca0-8d8e-d1275221830b · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.800992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.165570Z digest=sha256:a86421ff9276f3450b3b9518c99bc21d6acbaf8a2193e09b904861ad6d7062c1

Observation c2cf3cc2-a1d7-4f07-9be7-d6131df0fd71 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.790009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.169459Z digest=sha256:afdb0116df80a734d4d2ec7174d61224ca9094636085a75c7971aba8086ffc35

Observation 554f8853-d4a0-49a4-9075-33be8ee69c87 · outbound

This paper cites Repetition Improves Language Model Embeddings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Repetition Improves Language Model Embeddings

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.173283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.173283Z digest=sha256:87548630244e21bc27406204fcd741fa4cbe07955047ebbd9f7bfcba07536308

Observation 8f9e3754-0106-4a25-8f2d-4963b637f538 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.181220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.181220Z digest=sha256:f23047e91112d2283d8e9161fec1cb90ffc5089d64cd29b9a10d9be2a2990ff2

Observation ce1863b8-9a5e-46e5-a0ea-e282955f68fd · outbound

This paper cites Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.779633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.184771Z digest=sha256:7ce20d85b6671c295844353979ce466194e348ebadea20b4651d0ddf172ea8ce

Observation 650e3c80-0662-460e-a644-f33a1de5cf4b · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.188280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.188280Z digest=sha256:ae8277c2274263322d7856abf2f95e2e69a7e494d6a793f8be7b061944317020

Observation 8256558b-81e8-4136-ab26-a8ed29f1c08f · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.768198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.192566Z digest=sha256:5832a6e8d5d11a99ffb3b170cf85c00d2741fbc6f7e9d0ca7f7421abc0e16679

Observation df1e8bc0-cccc-4034-90a9-ac8b7fbfef8c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.196243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.196243Z digest=sha256:7a6f3999ca201a784f3a23e7e3ccf673c8a184150c4c9905992f5b48e0209a6f

Observation c1827cbc-35e5-458e-96d8-66c9a8fe0679 · outbound

This paper cites Diffusers: State-of-the-art diffusion models.GitHub repository, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Diffusers: State-of-the-art diffusion models.GitHub repository, 2022

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.756013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.199902Z digest=sha256:baaccf46b0c509bd1e637d9676dd7c021c909e2688e0b45430796ae9955eb442

Observation 5727adb9-dd26-4d48-b90c-70e3e7357347 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.203423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.203423Z digest=sha256:78f216c0d09858338e3720d05a945a67296a329a25123236ba89d741ae193159

Observation 3189fe56-8ea4-442d-b5f1-2e8540898742 · outbound

This paper cites Improving text embeddings with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Improving text embeddings with large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.744839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.206990Z digest=sha256:76eab5ea8bd0c54c087d6e70eaa9d441bac768ba73fd2e9f2ea9169be1649f2f

Observation f715e3cf-606d-4af0-979d-491ffc7800f5 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.210596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.210596Z digest=sha256:3424e1ecf42328b3cbf1e0a422bc5ca8794e001fdcd536f506f94784c9a3cd93

Observation fd00a9b6-0f84-43ba-831f-ebc5df508ea0 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.214409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.214409Z digest=sha256:2542e93e2e3e2828dd4380bc7467f4477173cd1efc59f90b63bfb799ce6d50f0

Observation eca2e838-1853-4e5e-aa53-4efa5424ae8d · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.733137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.218461Z digest=sha256:62c624787044e1afc6549b472f6a8f2429900df0c055d66247bd67ff14a1abb1

Observation 1f715680-9d55-44af-af67-261e85e2dc28 · outbound

This paper cites Qwen2 Technical Report.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.222983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.222983Z digest=sha256:2877b34c09d27cabe0892b9cb212f3309f378b82712d24eb3caf8b7a29ec7ff9

Observation 8ca66ea2-ab10-4a4d-8f5a-92033c24c6b7 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What you see is what you read? improving text- image alignment evaluation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.721893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.226960Z digest=sha256:6b65b6e2cbbd4929c9136182b8e677da238ab680c7a137675fbb78bec98b2324

Observation c5623f59-da6b-48f4-81dd-6f7bd2954a2d · outbound

This paper cites Investigating Layer Importance in Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Investigating Layer Importance in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.230669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.230669Z digest=sha256:68667dee7f7f2315cb12ec032211c0a1b2471ad398d87174ccd2c8259d448c81

Observation d7afd82b-89f1-4642-9c27-61a1bdf32373 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.235247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.235247Z digest=sha256:75f638e52489cf57865ac51f3a4cb48bbf341ca41303ad56f469015efc144a77

Observation ae66d45d-f3f5-45a0-91f6-994ca9ad6a26 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.240632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.240632Z digest=sha256:bdc33e09e412c81d1e0781520c86069d5536175916f4c1742efaed3e2007ce54

Observation 275b4bab-9df2-45e8-a30a-4291d5d7f937 · outbound

This paper cites a beautiful morning in the woods with the sun peaking through the trees.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation a beautiful morning in the woods with the sun peaking through the trees

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.711029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T05:20:51.244779Z digest=sha256:b87d960245504ed582db2235443c0fe1950a3f1cf5d86970dd6d8bbb3e88dc30

Pith citing papers

No inbound Pith citation observations are available.