Pith. sign in

Paper Citation Record · LEDGER

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2506.08210.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08210 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:20:51.244779Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 398c47e8-fb10-4e8a-a0e9-f818b192f0c9 · outbound

This paper cites Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.977218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.977218Z digest=sha256:5f3383bc13c118c781856e539e4756c16a830d1ace0570daf88c193aa9d2f8d9

Observation aa460d2e-26ff-4f00-8a19-b5c09a34d2db · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.982334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.982334Z digest=sha256:1e4aeb7522b91471afe8ebdcab69ce0a80f1bfdb32a188df6d39c31a08b5e92c

Observation 61d7c6d3-68f9-41d1-a5cb-65495b3ba30f · outbound

This paper cites Imagen 3.arXiv preprint arXiv:2408.07009, 2024.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Imagen 3.arXiv preprint arXiv:2408.07009, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.986447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.986447Z digest=sha256:c43165770cc46aab62353d29664ce8ee17eefa31c4765d17f390c9439a835744

Observation 1634a2a4-269c-4570-bf60-f9799e1fad0e · outbound

This paper cites LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.990041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.990041Z digest=sha256:010b6466b80a05acc5296ad9e6a7ffbd462da32caf9db6b1d57ca46e02c0496e

Observation e697754a-d932-4d75-9ef7-6b3d7692d37a · outbound

This paper cites Scalable Performance Analysis for Vision-Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable Performance Analysis for Vision-Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:20:51.597266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:50.994543Z digest=sha256:b6844577095f998e8d64fecc40eeccbda5ecead50c024c8b401217f7b0cebd9a

Observation 9b84a85a-28d1-424a-ae5d-98817c55b024 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:50.998688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:50.998688Z digest=sha256:5500eda1cae8ddbcbb56b5f5d0d31d125f8bb42983f10c390ffee630d34d9b77

Observation c575c6f6-6984-4190-8a71-a26d056f684d · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.027527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.006114Z digest=sha256:68158f53a75f9debbf2111d67af1bb1afe29b440bde15ea06ef4de8ee87537f2

Observation 853b019a-3335-47a9-8701-ce431034a2e2 · outbound

This paper cites M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.010211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.010211Z digest=sha256:5dfd25caabbe69ddb79c070ca3c5f92f21c7699ac1fdaffa78dc8fbfafb466cb

Observation c497bdec-1a28-48ba-876e-9ebb948eda0c · outbound

This paper cites Visual pro- gramming for step-by-step text-to-image generation and evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual pro- gramming for step-by-step text-to-image generation and evaluation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.015758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.014921Z digest=sha256:c7fd072907b4428a5a9135080015bef944900b6cf4935200425b12c4dfc3cd1e

Observation 23195d1b-b58d-4436-8aed-e94a525398a1 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:52.004536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.018690Z digest=sha256:28afff9cd5673aba022ce7e6e06d2ea3e52d12c8eca6a328e7883b29903a561e

Observation 9924718c-6b1b-4cff-a924-0ff9529cfff8 · outbound

This paper cites Analyzing Transformers in Embedding Space.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing Transformers in Embedding Space

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.022450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.022450Z digest=sha256:0b29284ae6ca72d56de21f81bcaa383916ff8eb75aab6403f4ba9e870c8fb236

Observation e5739900-abb2-4b2c-a8d6-cef0fd7ec774 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.993821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.027375Z digest=sha256:247318b4f490199642a9906cea4360f3103163806597b30425ffe2eb3469e8f5

Observation becaf3ad-2e97-42fc-9820-e757bec27070 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.030846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.030846Z digest=sha256:81599fc987a3431d7005ebf484c8fce432bbfbf51004a3cecc8bf0e1f07eb98b

Observation 941aed29-e666-4cc1-ae4c-391ecdcb3f49 · outbound

This paper cites Visual fact checker: en- abling high-fidelity detailed caption generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Visual fact checker: en- abling high-fidelity detailed caption generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.982389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.034740Z digest=sha256:4e83f1b2c0e6db1f7b42527872dbb3c1c3a21c307c95fc385c7cae4c04e57eff

Observation c4abb752-e600-45c3-8439-85cbfcb3b568 · outbound

This paper cites The Llama 3 Herd of Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.038814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.038814Z digest=sha256:7da448bbcec14357d1a4fbcb76ca98c7c9a76dc74fc7a99af19aec6c75c04062

Observation fbc237b3-b337-479e-8e9a-c9f999690795 · outbound

This paper cites MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MARS: Mixture of Auto-Regressive Models for Fine-grained Text-to-image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.042386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.042386Z digest=sha256:6ac80915bf8ebd7bec7f5757f5d2f86ec6bc0bce9b8afc8e80bd01d8b0d727a4

Observation ea91496f-857e-48f3-9b51-e200fe3e3686 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.046257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.046257Z digest=sha256:9f0c4263f48455e2c9f161f2d8b20d1ba0127c740a8b48c3fd6f000cfe7c25c9

Observation 6d8bb99b-0573-4499-ae53-3369c5a78347 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.972571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.051035Z digest=sha256:0702e184628eae70f06271157bde1e1e7831dc36ae79f53567391a0113e47961

Observation 039cbba8-4e40-41f8-b62d-971cdaf3469e · outbound

This paper cites Classifier-Free Diffusion Guidance.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Classifier-Free Diffusion Guidance

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.054665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.054665Z digest=sha256:a044a3546a723da254483add64162583060115b310a96674cf9d7a16ba8b14c1

Observation 927c9c3c-c01b-44c8-bdbb-8c2e2d661b09 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.058500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.058500Z digest=sha256:c2294ca0a1ac4cea13844918b81898331fe44634113a9ce81219d6128a9ded64

Observation 6f26b31b-c794-4861-bf2b-2305a5e532db · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.961439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.062379Z digest=sha256:f72fd158d0ed8db1a1cb2279df7f0c417eed85b5a06ee0c1a5eee71e513f0321

Observation 87b63063-c52e-4bc7-bce0-91687dd6dc4d · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.950410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.066724Z digest=sha256:42d7fc1a97852ebdf8100a0613f7560e34022881495917012a085b1ed5e5c667

Observation 23a02f78-89ab-48d2-941a-1acf24767390 · outbound

This paper cites GPT-4o System Card.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.070312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.070312Z digest=sha256:f0eb81664bd61d2847bb74543f6ee3ba85d085bece21de8dea299c89680e1328

Observation 5f9e4d79-9629-4745-bedf-a538a28a0496 · outbound

This paper cites What does bert learn about the structure of language? InACL,.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What does bert learn about the structure of language? InACL,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.938958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.074586Z digest=sha256:26a427f981d5b89500a0c41a22e7c68ba444185445e2564fbd3eca59ba5bf8d9

Observation a744e85c-dbf0-40ea-8e1b-74831bdc6daa · outbound

This paper cites Mistral 7B.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Mistral 7B

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.078207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.078207Z digest=sha256:394a1a38066e42e84c4bec13c06ee06a798149e70b7671d8ad65702e97509955

Observation d6902ad0-3b3b-41c4-8c12-dd1929b6dc40 · outbound

This paper cites Analyzing the Role of Semantic Representations in the Era of Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing the Role of Semantic Representations in the Era of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.082207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.082207Z digest=sha256:1f85a04f865acbdc9519e2cdb58b2e027d7289e6114068477e7db4ebe28f7424

Observation 79b400d3-f8ac-4ef6-9dff-60a433fc1530 · outbound

This paper cites Elucidating the design space of diffusion-based generative models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Elucidating the design space of diffusion-based generative models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.927459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.086581Z digest=sha256:c3436539735a16d44ba1a5d951b59b08d9939fdd258f9ff6c837931554bf481c

Observation b44455d0-a6cc-4098-bd00-f884dce11158 · outbound

This paper cites Analyzing and improving the training dynamics of diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Analyzing and improving the training dynamics of diffusion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.916207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.089829Z digest=sha256:57b6d5af2b385d88bd62b43d6e713c42c8e7e49ae8059ef4f032b87db4bc6927

Observation 8f166b0a-9b9b-4618-8442-4e80d4ff018d · outbound

This paper cites NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation NV-Embed: Improved Techniques for Training LLMs as Generalist Embedding Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.093052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.093052Z digest=sha256:fa86a9ded7b576f33d147820e3c8476421d29b2adb11e3ffe968c113ed0fd970

Observation 4093adbf-4e7f-4333-9802-e5abbf15e686 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.096656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.096656Z digest=sha256:93fae6fc014f7a289c0c25890ba9b5df46265ca7f94f6c5fcf9f5d3bc08aaae5

Observation ce625a39-7946-45c8-88db-e4bd5b49427b · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.100614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.100614Z digest=sha256:71f27e1c9a475b13bcb48c1c0dc67fbc806a6432332a7c4cbfe136b89165d3a9

Observation 6fe2513e-05dc-4728-b10f-6418a9231af6 · outbound

This paper cites Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Llm- grounded diffusion: Enhancing prompt understanding of text-to-image diffusion models with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.903817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.104499Z digest=sha256:f254b2133ce15fa2faba005a1233d9cb1064f602882bbfba5e04361377874cd1

Observation e25b3597-2306-45ac-b073-ecdae4731aad · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Common diffusion noise schedules and sample steps are flawed

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.892766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.108275Z digest=sha256:0ab841f353dc2a7b376132bdde75a3d92e08003ecf9a74e439e13f4d7ec47975

Observation 04ead64d-cb04-459d-bb57-c2d51b0d4bf9 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text gen- eration.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Evaluating text-to-visual generation with image-to-text gen- eration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.882601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.112353Z digest=sha256:250ba83f7c60a9ea4b960fa21c7fec6036b0fad2127a21bad02a766c1c9682d7

Observation 78c10748-bbdd-4373-b842-1ba1ad57c327 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.115963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.115963Z digest=sha256:ae656691343bc9d6a41eda2a836a62b1288893dc671b2e406751bc863110e609

Observation 46685bd4-1272-4c9a-a0e6-f226a9a9bce3 · outbound

This paper cites Character-Aware Models Improve Visual Text Rendering.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Character-Aware Models Improve Visual Text Rendering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.120173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.120173Z digest=sha256:54dfad9ee1dcebd3873cbd4eef13b3e9470b63e44268d12c08601b47cc3ed69b

Observation 9494912b-8c09-407f-8ccd-0cbc9de6a8ff · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.123785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.123785Z digest=sha256:806fbafee7b59b8178ad23acb6678774e71f2ff628ec5ffd1348ae67e5f18514

Observation f5a33c4b-94c4-443c-99bd-e73523f50273 · outbound

This paper cites Decoupled weight decay regularization, 2017.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Decoupled weight decay regularization, 2017

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.870995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.127292Z digest=sha256:b27a1469370d33dbc6457dcad6cf1ab772c1b8873ab38547eff240536ad7b677

Observation 5e1143d1-e17e-46dc-b0ad-588a8094bacf · outbound

This paper cites Salesforce AI Research’s SFR-embedding, the top performing text-embedding model.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Salesforce AI Research’s SFR-embedding, the top performing text-embedding model

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.860167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.130898Z digest=sha256:c98be7c848df84cd6000f5cc5ff037162738f347e108d6b4d23cd9efc60cdc0b

Observation 08adad73-ee57-430b-936f-99f7ffdda3e4 · outbound

This paper cites MTEB: Massive Text Embedding Benchmark.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation MTEB: Massive Text Embedding Benchmark

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.134528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.134528Z digest=sha256:b12bd0772c5580823c96770dc7a4239732a4ffd709d53c7df9d575f3779849ec

Observation b6d0ab25-1849-45fd-b079-deb6eeff7e55 · outbound

This paper cites Scalable diffusion models with transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.139116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.139116Z digest=sha256:02f83a3a1f21ed8b5dde36f3c62984f9680c0cac4a23820f5467c03f07cf9865

Observation 08aee82c-0c41-472e-9dcb-815af9af02ab · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.143069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.143069Z digest=sha256:63e0ffcb164f0beca8c45f15737e5b33e98814426edb75fcaf57974f71b490a2

Observation be6f771f-b28b-44e4-b94d-fd8c5e1e1042 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Learn- ing transferable visual models from natural language super- vision

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.842843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.147265Z digest=sha256:cfe886bdcdca90c33f4448eec2cfb09a923f3bcfc9e846dac5fc60c1155ec0a3

Observation 30bb8684-6000-4e2c-a8f7-3ee077efd3ab · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Exploring the limits of transfer learning with a unified text-to-text transformer.JMLR, 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.832409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.150903Z digest=sha256:f1576a1c6f621dac1bff112270077b05397a08c39fb30c6e9eb659b034b91665

Observation 8f990867-8353-45e4-9157-07fab43f9504 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.154397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.154397Z digest=sha256:997c40e3ca441635501480de43845b5c6404ac1ee1754ff655e4c146db591b12

Observation d242fceb-18b2-42b1-b1f2-e59bf4045b48 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.822625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.158439Z digest=sha256:472d53d52503b1e167db275b16a31e62fc1b0111360077755b1031ac2189c2ac

Observation cf875356-8584-4dfa-b89e-7565d2e257a6 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation U-net: Convolutional networks for biomedical image segmentation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.811734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.161994Z digest=sha256:cefd6c1a280f1e6be185d19e0a9fe06300a8c6f7a129a6205a91481317763c5e

Observation 1d1708d4-fdd0-4ca0-8d8e-d1275221830b · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.800992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.165570Z digest=sha256:f96af54f4f8b3a71b81d815a660f75545f3fdb2ebc5d43106c80b4c1644e6cf1

Observation c2cf3cc2-a1d7-4f07-9be7-d6131df0fd71 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.790009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.169459Z digest=sha256:c97ddd349a1a87951748fa9d72a6206d42d22bf85cdfd2754f99f13055631292

Observation 554f8853-d4a0-49a4-9075-33be8ee69c87 · outbound

This paper cites Repetition Improves Language Model Embeddings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Repetition Improves Language Model Embeddings

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.173283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.173283Z digest=sha256:631b640a6f08414d0fd1c5270023be92f8c46e8f9ebf78b2efff68a945596cfe

Observation 8f9e3754-0106-4a25-8f2d-4963b637f538 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.181220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.181220Z digest=sha256:f23047e91112d2283d8e9161fec1cb90ffc5089d64cd29b9a10d9be2a2990ff2

Observation ce1863b8-9a5e-46e5-a0ea-e282955f68fd · outbound

This paper cites Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Stable diffusion training with mo- saicml.Mosaic Research Blog, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.779633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.184771Z digest=sha256:d493260d52ec473b6feb8cb32e733e728e8ae9e12c13a0bb63a177e157f2e03b

Observation 650e3c80-0662-460e-a644-f33a1de5cf4b · outbound

This paper cites What do you learn from context? Probing for sentence structure in contextualized word representations.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What do you learn from context? Probing for sentence structure in contextualized word representations

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.188280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.188280Z digest=sha256:3f27253e0ee26331cc2b76121feeba7f58a0d63eb6ed3493fee696d9cfa0bb95

Observation 8256558b-81e8-4136-ab26-a8ed29f1c08f · outbound

This paper cites Winoground: Probing vision and language models for visio- linguistic compositionality.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Winoground: Probing vision and language models for visio- linguistic compositionality

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.768198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.192566Z digest=sha256:b8797869ff82d7424e5c7f48f67eb783b74543981c879ae05dbce3bddbdb78f9

Observation df1e8bc0-cccc-4034-90a9-ac8b7fbfef8c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.196243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.196243Z digest=sha256:7a6f3999ca201a784f3a23e7e3ccf673c8a184150c4c9905992f5b48e0209a6f

Observation c1827cbc-35e5-458e-96d8-66c9a8fe0679 · outbound

This paper cites Diffusers: State-of-the-art diffusion models.GitHub repository, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Diffusers: State-of-the-art diffusion models.GitHub repository, 2022

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.756013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.199902Z digest=sha256:aaee20a0783d6d25087f0f2d13c937ae6e88c8eef24b7193def06791d3669144

Observation 5727adb9-dd26-4d48-b90c-70e3e7357347 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.203423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.203423Z digest=sha256:78f216c0d09858338e3720d05a945a67296a329a25123236ba89d741ae193159

Observation 3189fe56-8ea4-442d-b5f1-2e8540898742 · outbound

This paper cites Improving text embeddings with large language models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Improving text embeddings with large language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.744839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.206990Z digest=sha256:bbcc8a537795397b4375dccf2c204067b414e33356bef79ce5114125de241aeb

Observation f715e3cf-606d-4af0-979d-491ffc7800f5 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.210596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.210596Z digest=sha256:497cdae5c38b877c250aa354c3b35806f1190b87c4c26f98f343c6ca462a543d

Observation fd00a9b6-0f84-43ba-831f-ebc5df508ea0 · outbound

This paper cites SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.214409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.214409Z digest=sha256:cceb9428d28d92b6123694a290ee68082032043b66c71cc0f95fd6f207bd2bff

Observation eca2e838-1853-4e5e-aa53-4efa5424ae8d · outbound

This paper cites ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation ByT5: Towards a token-free future with pre-trained byte- to-byte models.TACL, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.733137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.218461Z digest=sha256:36bdb2fdc9adf3bc49dee76d560573606dd8189ee21a54c006e6f5033c54d90f

Observation 1f715680-9d55-44af-af67-261e85e2dc28 · outbound

This paper cites Qwen2 Technical Report.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.222983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.222983Z digest=sha256:cf0e12bc21d876e76e116b1b1c1ef598b5a7b9e4bc439963ecd003d421a87dc4

Observation 8ca66ea2-ab10-4a4d-8f5a-92033c24c6b7 · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation What you see is what you read? improving text- image alignment evaluation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.721893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.226960Z digest=sha256:7d6c4efa52809f657da22757720406fed820132a5ed026d2111c32f649345f6c

Observation c5623f59-da6b-48f4-81dd-6f7bd2954a2d · outbound

This paper cites Investigating Layer Importance in Large Language Models.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Investigating Layer Importance in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.230669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.230669Z digest=sha256:32882ad3e7d3a0ea78a075e539c827a4f4bb654eec3229b7cb021f79fb1fce54

Observation d7afd82b-89f1-4642-9c27-61a1bdf32373 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.235247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.235247Z digest=sha256:75f638e52489cf57865ac51f3a4cb48bbf341ca41303ad56f469015efc144a77

Observation ae66d45d-f3f5-45a0-91f6-994ca9ad6a26 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:51.240632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:51.240632Z digest=sha256:94489d13f5892635d5f4b8e2b356196fc6e949c8cb41425a92582d8bfac4bd1a

Observation 275b4bab-9df2-45e8-a30a-4291d5d7f937 · outbound

This paper cites a beautiful morning in the woods with the sun peaking through the trees.

A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation a beautiful morning in the woods with the sun peaking through the trees

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:20:51.711029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:20:51.244779Z digest=sha256:a4c1d4d3db07ef69151cad71ed1c40183df62a216facaf0dcf971a31c933b233

Pith citing papers

No inbound Pith citation observations are available.