Pith. sign in

Paper Citation Record · LEDGER

Test-time Prompt Refinement for Text-to-Image Models

As of 9 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 3 inbound Pith citation observations for arXiv:2507.22076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.22076 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:01:52.609484Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T02:13:10.574530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:29.278949Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy31
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 43569dd6-df28-472c-b09b-7dd3ef991d24 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Test-time Prompt Refinement for Text-to-Image Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.176538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.176538Z digest=sha256:9c62e77655f43417fd861e8a716e22039275b0811bd35c65b44c003073e1111f

Observation d086ba3e-d8bd-4d9a-b8bc-3c593c135bf0 · outbound

This paper cites Blended latent diffusion.

Test-time Prompt Refinement for Text-to-Image Models Blended latent diffusion

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.721694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.182180Z digest=sha256:ca5fd4fcd72e30d2e74a284de1a5d289a9fd92a4ec6f85b10186a31a9a4a77f6

Observation 2e92cf0d-7eeb-4d4c-9d31-50ebba10271f · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Test-time Prompt Refinement for Text-to-Image Models eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.186385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.186385Z digest=sha256:aa1946bc39c727e18fdeb99b28ec3c96458538a69f931f481ded875cbfddf204

Observation 7f4b6d50-f2df-4c29-b53a-be447b2216c3 · outbound

This paper cites Multidiffusion: Fusing diffusion paths for controlled image generation.

Test-time Prompt Refinement for Text-to-Image Models Multidiffusion: Fusing diffusion paths for controlled image generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.700213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.191212Z digest=sha256:27aa1486d5750b8ac10dfbca87d5e83f21fa83ea2f8dfc53c73fd055d4cc78a1

Observation 9c19095c-5274-4f7e-8dbd-ea1d0be6be35 · outbound

This paper cites Improving image generation with better captions.

Test-time Prompt Refinement for Text-to-Image Models Improving image generation with better captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.679210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.195527Z digest=sha256:60034fdb613ab40a105ae09569f3330ec8f2a971fc3b0c3cd80aa60a33797457

Observation f1065118-1e2e-437d-8dbc-4419629261b4 · outbound

This paper cites Training-free layout control with cross-attention guidance.

Test-time Prompt Refinement for Text-to-Image Models Training-free layout control with cross-attention guidance

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.663088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.199854Z digest=sha256:cf33d2c58c9455103b9b7f0b432d08afc3674373df3f603d798e50758db58e05

Observation e5531e3a-38e1-423a-8b42-ad44ed35c59a · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Test-time Prompt Refinement for Text-to-Image Models Masked-attention mask transformer for universal image segmentation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.645208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.204947Z digest=sha256:dd1c3e646c6ae6d0c14819316f8070abb791574e5f7b0a86e262c119332c0c94

Observation 7a840b82-e008-4775-abc9-d7e62f1a7d72 · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Test-time Prompt Refinement for Text-to-Image Models Cogview: Mastering text-to-image generation via transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.627074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.210094Z digest=sha256:b31811f696d0897cd31344207268b44ab075e9ea5ad0b55ba11f63b2a18ab3bb

Observation 8cad3c74-3d26-436f-ac96-f036b47e3f99 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Test-time Prompt Refinement for Text-to-Image Models Taming transformers for high-resolution image synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.215333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.215333Z digest=sha256:38fa15128db8218d5e87c83cd8d983e12e12c912c1be1dac71c80a3d5409cd01

Observation 41cb1387-aa9d-4c76-ae4e-f2046d7fd1be · outbound

This paper cites DPOK: Reinforcement learning for fine-tuning text-to-image diffu- sion models.

Test-time Prompt Refinement for Text-to-Image Models DPOK: Reinforcement learning for fine-tuning text-to-image diffu- sion models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.598235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.220132Z digest=sha256:96dcb03fd85fed51b3780ce9f8c94958639c89d2a117390b4a315767cf9844b1

Observation 2ec0318e-7b3d-4610-9ee5-19cc19ca7bb2 · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

Test-time Prompt Refinement for Text-to-Image Models Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.224431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.224431Z digest=sha256:0893fa8ad8690127b7612f19cab9759cc41f25999eb61b0fdfc13e547d2710a1

Observation b668eda1-6ab9-4385-b393-4e9b5541dd2f · outbound

This paper cites Layoutgpt: Compositional visual plan- ning and generation with large language models.

Test-time Prompt Refinement for Text-to-Image Models Layoutgpt: Compositional visual plan- ning and generation with large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.570232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.228582Z digest=sha256:a123a4154a6d843d5acf62d0cefc9f071489387b9d343e660925258916f8d803

Observation d2feb80e-38fb-46c3-bf8f-41a51c693c73 · outbound

This paper cites LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts.

Test-time Prompt Refinement for Text-to-Image Models LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.232968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.232968Z digest=sha256:2a762bfbfeadbf596846a793b792529790aac20ffe84dde74e6576b5edb417b9

Observation 7fb07d2c-bc98-4ba5-a989-8c9a404d06ba · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

Test-time Prompt Refinement for Text-to-Image Models Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.551291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.237769Z digest=sha256:2ce35ad4ae873fe42933ad91365a8b8456f0155a556e26006686b3f9167e7d69

Observation 607496b9-0335-4ba5-8b1e-70b831dbe442 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Test-time Prompt Refinement for Text-to-Image Models Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.241657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.241657Z digest=sha256:171eddf059813eaf438e2d34d934e9c488e3fced940df8855fd3a91873540fd9

Observation 5d66487a-27c2-485b-a521-03b2ffe11767 · outbound

This paper cites GPT-4o System Card.

Test-time Prompt Refinement for Text-to-Image Models GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.246022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.246022Z digest=sha256:2a1853c558798b093037ded866587a31813bb6ef262221b912e110f451194e60

Observation 7b7040ec-03fc-4003-bbd9-b6762caf11d4 · outbound

This paper cites Few-shot classification and anatomical localization of tissues in spect imaging.

Test-time Prompt Refinement for Text-to-Image Models Few-shot classification and anatomical localization of tissues in spect imaging

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.529471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.250084Z digest=sha256:fe2bf8b0e316ed5915838fd87327a3ecade26ca8a811ff28a7a7b179067545ec

Observation f334e179-6c0b-4836-805a-a3a84e487e3a · outbound

This paper cites Clas- sification of microstructure images of metals using transfer learning.

Test-time Prompt Refinement for Text-to-Image Models Clas- sification of microstructure images of metals using transfer learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.510571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.254159Z digest=sha256:a48ceb4941a9664209b7ff86f5fcd9ea26bbffb45d97aacfd6c50d317d1d4f80

Observation ac1bfb0e-b880-4f05-a352-0d20f21c9884 · outbound

This paper cites Alina: Advanced line identification and notation algorithm.

Test-time Prompt Refinement for Text-to-Image Models Alina: Advanced line identification and notation algorithm

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.492598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.258108Z digest=sha256:ac1d434235a4c818b877b6ca32cb883bc6a9ca02c81a2bc3f606d9b1ed6a8993

Observation 4fbd46a4-5c4d-46b5-96cd-ec8e50de5cb6 · outbound

This paper cites Gen- erating images with multimodal language models.

Test-time Prompt Refinement for Text-to-Image Models Gen- erating images with multimodal language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.476091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.262537Z digest=sha256:fe9f707c755dfa8b44d40b935ff23ae0b56104c1a3c7df0203ddebe6adc9a459

Observation 1e8729aa-64a5-400f-8ff6-9fa8c1c86155 · outbound

This paper cites Zero-shot Text-guided Infinite Image Synthesis with LLM guidance.

Test-time Prompt Refinement for Text-to-Image Models Zero-shot Text-guided Infinite Image Synthesis with LLM guidance

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:01:52.927761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.267531Z digest=sha256:4b9ef4d2e83f4fa5898b9729a41040831398e2396bfe3bfad5397c4e304a4fe3

Observation 0b04754e-ca52-4c10-965f-b0fc759b9aa5 · outbound

This paper cites an unresolved cited work.

Test-time Prompt Refinement for Text-to-Image Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:53.458307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.273106Z digest=sha256:2284b6a1352871295e48c13853fa08bad198568adafe8c8fd4eb24eb81b26184

Observation c458aae0-ddd2-42cd-8521-ce422abd5c84 · outbound

This paper cites Aligning Text-to-Image Models using Human Feedback.

Test-time Prompt Refinement for Text-to-Image Models Aligning Text-to-Image Models using Human Feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.277213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.277213Z digest=sha256:8d63a883584883abcef05efb940ec9b3ff888020863ff8345d28a05ef0188a25

Observation ed207dd5-73fc-477d-b770-76d4b4edfa32 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Test-time Prompt Refinement for Text-to-Image Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.281822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.281822Z digest=sha256:2a4b292ddc9fd817896ba14d3b72cf67928535f1d30c0b5994f72127f129ae47

Observation b870a995-e2de-4ba2-9621-f8a92b33a46a · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

Test-time Prompt Refinement for Text-to-Image Models Gligen: Open-set grounded text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.429341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.285748Z digest=sha256:dcfb3bf96c6d9e171ef76b85af79431e051f82d9ead1e8028d6eb38e9b3ed83d

Observation 0f586462-6019-4bfd-a6d6-84efb249685e · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

Test-time Prompt Refinement for Text-to-Image Models LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.290435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.290435Z digest=sha256:4e663857263836a1c1703787cdac263d39c07ca1a96e53dd75b7e545eb214b2d

Observation 666964f2-506d-4ca3-860c-6dda68be8baa · outbound

This paper cites VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning.

Test-time Prompt Refinement for Text-to-Image Models VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided Planning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.294824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.294824Z digest=sha256:c1df04a4cbe711f8eaa78bf901bdbe2822ba3d9c0fb1384445a7e4f14cf51464

Observation d2b90bad-7946-40f5-a4b1-644a8afd93a2 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Test-time Prompt Refinement for Text-to-Image Models Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.299646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.299646Z digest=sha256:175f8056e6df4f60f58666907d2c0766b58f96fcb302077d40e5464cb0fcb0e9

Observation c69a07c1-9aec-482b-b49f-5660feabae24 · outbound

This paper cites PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.

Test-time Prompt Refinement for Text-to-Image Models PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.303932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.303932Z digest=sha256:9a3fb3253f29deea472f8e40e3d629dbe62b6d7add286e78ef89ba300a489c7f

Observation 0827f241-f3f2-4252-abbf-8283292521de · outbound

This paper cites Simple open-vocabulary object detection.

Test-time Prompt Refinement for Text-to-Image Models Simple open-vocabulary object detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.409155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.308732Z digest=sha256:ae17f4b22bdbc6fe181ffe580b8fef54bfe5cce921d08a957220238ae8a939d3

Observation 7e5399fb-d84e-42d3-b10c-bd914fde2f4b · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Test-time Prompt Refinement for Text-to-Image Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.313124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.313124Z digest=sha256:cf0d62af7cec946beef198552bec138fd81c4ed8f05f2ae397978b0d0623e826

Observation 3edf1190-2cbd-4ab6-9b5a-52d72c946208 · outbound

This paper cites Localizing object-level shape variations with text-to-image diffusion models.

Test-time Prompt Refinement for Text-to-Image Models Localizing object-level shape variations with text-to-image diffusion models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.390860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.318058Z digest=sha256:571036f68659ed5d11410089b908f92f1ff41ad6a4bd367aa87f5494af3f5407

Observation 7e66b17b-116f-467a-a772-e431decfacdc · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Test-time Prompt Refinement for Text-to-Image Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.322258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.322258Z digest=sha256:7c6aee1129a938516fff7d7e1bab2cde363ae2a0b899852a09385cce1dfcc538

Observation 8fc2749c-911d-4d41-a400-bcca2f8dbe4f · outbound

This paper cites Diffusiongpt: Llm-driven text-to-image generation system.

Test-time Prompt Refinement for Text-to-Image Models Diffusiongpt: Llm-driven text-to-image generation system

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.327594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.327594Z digest=sha256:6d3d922158cc11369adde0ce1655182d91c88b9dd672ba960f5b08df587fb105

Observation 4be411d5-7a06-439d-8ffb-7f5fd95a4a65 · outbound

This paper cites Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation.

Test-time Prompt Refinement for Text-to-Image Models Layoutllm-t2i: Eliciting layout guidance from llm for text-to-image generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.373242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.331685Z digest=sha256:023ca6f7afc56455e83ed397b452743216502b2c89be6112a4e6d7152bc3c428

Observation 95ae111d-5157-41a2-bd76-2f672cfde516 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Test-time Prompt Refinement for Text-to-Image Models Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.356271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.335654Z digest=sha256:84a4608afa12327208786344f025b2ce85f6a80e931a0eb446ea9ca9e86851bc

Observation dbe6e7c9-470f-4e2d-94b5-667c65a5c52f · outbound

This paper cites Zero-shot text-to-image generation.

Test-time Prompt Refinement for Text-to-Image Models Zero-shot text-to-image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.335735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.339583Z digest=sha256:1f49b0a171e261c669807bd75d91379ad19d38147d4a94802455d771a1f76d27

Observation f24677d5-041d-4729-8ad2-7a39540dd740 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Test-time Prompt Refinement for Text-to-Image Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.343453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.343453Z digest=sha256:7faecf36ecf195bd4dd3ff1421169fbff410d77cc22cc3d72b62721db00553ef

Observation 5b6bbfd6-dd56-4e85-bb72-145d4820cb02 · outbound

This paper cites Generative ad- versarial text to image synthesis.

Test-time Prompt Refinement for Text-to-Image Models Generative ad- versarial text to image synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.347957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.347957Z digest=sha256:9f36cd5ab064872476d95b58bd5ae10579f3a464adc61bb5437b0f7fc53c4e11

Observation 054a60bb-6615-44f5-bcb2-ca5f51cd2a55 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Test-time Prompt Refinement for Text-to-Image Models High-resolution image synthesis with latent diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.307768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.352213Z digest=sha256:e30bb5e85474cc7ccdcd6be4e3145730221b7cf2dc99445a028b54ed0b04e049

Observation 34456057-989b-4c33-95f7-f18177b0478c · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Test-time Prompt Refinement for Text-to-Image Models Photorealistic text-to-image diffusion models with deep language understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.357931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.357931Z digest=sha256:754b991a468add3991c96817691ed32b054f83f6b68b1900269bf671c9ff8f2e

Observation a3b624fe-b3e8-4325-9b8a-77b0efa9309a · outbound

This paper cites Benchmarking awesome diffusion mod- els.

Test-time Prompt Refinement for Text-to-Image Models Benchmarking awesome diffusion mod- els

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.281385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.362885Z digest=sha256:10eaab967d3098d3f9396ad89e35018460fb19e6261fe217598195c8ad007777

Observation 6f5a8638-aeef-488c-9eb8-5dd93c618c76 · outbound

This paper cites Df-gan: A simple and effec- tive baseline for text-to-image synthesis.

Test-time Prompt Refinement for Text-to-Image Models Df-gan: A simple and effec- tive baseline for text-to-image synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.249741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.371093Z digest=sha256:d566db33093d225d16f1a897e073ef413355958a649ba9297cd7ccf50e83ea7b

Observation 52b951c6-ecf2-45a9-ab21-7763cce1a45b · outbound

This paper cites Qwen2.5: A party of foundation models, 2024.

Test-time Prompt Refinement for Text-to-Image Models Qwen2.5: A party of foundation models, 2024

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.233186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.375107Z digest=sha256:0e2968eedbe9a0c0d159021ca4188552e35cf999a6be2f689bad13b779dc307f

Observation 8757f061-212b-41f5-9a4d-2d7b2673e23a · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Test-time Prompt Refinement for Text-to-Image Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.379789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.379789Z digest=sha256:1ec008a0caaff5873fcefee70ee63f3da8ee0a33a93b56151c658aa5ed98d5a7

Observation 306d0874-42d8-4f2c-8762-a23c7dabc5b3 · outbound

This paper cites Self-correcting llm-controlled diffu- sion models.

Test-time Prompt Refinement for Text-to-Image Models Self-correcting llm-controlled diffu- sion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.218936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.384252Z digest=sha256:55ba27711987741ba439470e2e1261b6f1cb50b050e630930f1440ff51b4d2d1

Observation c40d7134-6a2b-4b31-99d2-0c92c2e7e2f6 · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

Test-time Prompt Refinement for Text-to-Image Models Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.201848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.388650Z digest=sha256:a8b92a4c72618133d33a9b72529bd38acbbecf0b88d84147a77593de787db7c4

Observation 585525c1-dfa8-4590-8882-11ae613bb283 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Test-time Prompt Refinement for Text-to-Image Models Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.187417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.393591Z digest=sha256:25dc8b1dde3d05d940cdd69f23436ee2000d892abab89c4c38fc56bdfb16823a

Observation e1077ab1-cc15-4f07-adb7-24b287ce6577 · outbound

This paper cites Qwen2 Technical Report.

Test-time Prompt Refinement for Text-to-Image Models Qwen2 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.540099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.540099Z digest=sha256:3245ee8f926b989b39a7d6e9d9faf060bc92409181beddb55821fc7002de3db2

Observation de47594c-35f6-4206-8575-02af20cd8924 · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion mod- els.

Test-time Prompt Refinement for Text-to-Image Models Paint by example: Exemplar-based image editing with diffusion mod- els

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.172630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.547403Z digest=sha256:a14fc4400d8dd0b26cdac2ecdbab7fa6d069f3d0e64fa6f27516a02961da507f

Observation f64aab34-d6b6-48cc-8bc1-1d97b6cceb17 · outbound

This paper cites Reco: Region-controlled text-to-image genera- tion.

Test-time Prompt Refinement for Text-to-Image Models Reco: Region-controlled text-to-image genera- tion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.554275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.554275Z digest=sha256:f520f9ae0fb9a0e837d4eaa1799dfce5f157c28ccc92d7e4d0b8295766a6fea7

Observation 958ad8bd-cfe6-412a-8dfe-7f255a1b844e · outbound

This paper cites Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks.

Test-time Prompt Refinement for Text-to-Image Models Stack- gan: Text to photo-realistic image synthesis with stacked generative adversarial networks

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.148697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.562603Z digest=sha256:257d6f249f0c83b783cfa9a1ab9477449ae89b5d8234581eee19da13987e9c05

Observation 85a3587b-fc2c-4b79-928e-7d51e7409aa8 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Test-time Prompt Refinement for Text-to-Image Models Adding conditional control to text-to-image diffusion models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.569855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.569855Z digest=sha256:a53f8958475f7acc38780b48b726201fb89113dd4906d437bae229b09fee5fed

Observation 2513d08e-c38d-43c6-8d7f-5e9f15a9e864 · outbound

This paper cites Controllable text-to-image generation with gpt-.

Test-time Prompt Refinement for Text-to-Image Models Controllable text-to-image generation with gpt-

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.123489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.576873Z digest=sha256:bea310ed779b3758aed17e02aeff223ee16685dad0611a075d2a02a6e4742efd

Observation e7e3f439-9ce0-49a1-891c-c8ddd560bf47 · outbound

This paper cites Dm-gan: Dynamic memory generative adversarial networks for text- to-image synthesis.

Test-time Prompt Refinement for Text-to-Image Models Dm-gan: Dynamic memory generative adversarial networks for text- to-image synthesis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.108208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.593493Z digest=sha256:eff32c798ccf5d5bda93d86a0e2ee96429704726aad818b2770ed8a9d1995d74

Observation 6c7405c6-5644-493e-8975-3742ad478835 · outbound

This paper cites Controllable Text-to-Image Generation with GPT-4.

Test-time Prompt Refinement for Text-to-Image Models Controllable Text-to-Image Generation with GPT-4

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T15:01:52.585344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:01:52.585344Z digest=sha256:afc4554822ef96ceab0a5cd2367ddc8f629974f628d8ebf85d227264131116eb

Observation b152ec5c-59c1-4354-b5b8-6dd4a6c28df3 · outbound

This paper cites Benchmark Datasets We use three benchmark datasets to assess compositional fidelity, prompt comprehension, and generalization: 9.1.1.

Test-time Prompt Refinement for Text-to-Image Models Benchmark Datasets We use three benchmark datasets to assess compositional fidelity, prompt comprehension, and generalization: 9.1.1

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:01:53.093043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.601255Z digest=sha256:4c67959650f6bc7255506259baf2f5fc044ce19274e20cb0088303a61d6e79e2

Observation 0c053d58-a99d-4125-a188-9f923a76db6d · outbound

This paper cites an unresolved cited work.

Test-time Prompt Refinement for Text-to-Image Models Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:53.078443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.609484Z digest=sha256:c562779254a46ffc40e9bdd84183eba387018af080f253de97c32430a91b7e8f

Observation 7bb962e7-2f81-48b0-a76f-fcc33561f370 · outbound

This paper cites an unresolved cited work.

Test-time Prompt Refinement for Text-to-Image Models Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:01:53.266861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:01:52.367049Z digest=sha256:389504e55d8fd66bb8e10e947f9b7dc030d57915383ea881c0902068b0bda7b1

Pith citing papers

Observation 5af3e4d7-3f17-45fc-acab-134cacffe0e0 · inbound

Evolutionary Token-Level Prompt Optimization for Diffusion Models cites this paper.

Evolutionary Token-Level Prompt Optimization for Diffusion Models Test-time Prompt Refinement for Text-to-Image Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:40:57.856590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:05:00.723170Z digest=sha256:a64aa4dfb4b1677c026b725fb283313172066e4bbb74f1211a703f3e25b1b73a

Observation 5a7d0d4a-ab9d-4063-99a1-367dc1aa8057 · inbound

DuET: Dual Expert Trajectories for Diffusion Image Editing cites this paper.

DuET: Dual Expert Trajectories for Diffusion Image Editing Test-time Prompt Refinement for Text-to-Image Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:29.280305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T06:58:39.254427Z digest=sha256:47e964661d46ab04ec64a2afb3af978ff0434e04035c5f5957bebeb4f15c15c0

Observation ee1caf41-9d5b-4243-9317-9ad86911a2f3 · inbound

DuET: Dual Expert Trajectories for Diffusion Image Editing cites this paper.

DuET: Dual Expert Trajectories for Diffusion Image Editing Test-time Prompt Refinement for Text-to-Image Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T02:13:10.574530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:13:10.574530Z digest=sha256:958f95d43e7f8517ca163927dc87668e1f2965ad57b8101edad2d198b4dfa2d5