Pith. sign in

Paper Citation Record · LEDGER

Text-to-Image Alignment in Denoising-Based Models through Step Selection

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2504.17525.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17525 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:42:25.555555Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:42:04.111093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5ab7b4a0-1b24-4c1b-b8a4-c0f6bc18cad6 · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Text-to-Image Alignment in Denoising-Based Models through Step Selection A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.124943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.383107Z digest=sha256:6d8c8356940ce41b79ec34272458b2b2dac78978d6c16c4522bf919cb8070418

Observation 6d3002d6-6c8b-4138-85b4-7f7aa780f4ab · outbound

This paper cites Attend-and-excite: Attention-based semantic guidance for text-to-image diffu- sion models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Attend-and-excite: Attention-based semantic guidance for text-to-image diffu- sion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.114275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.390996Z digest=sha256:3f2f696c81d0de89ea922cb9941b09a774b271ad124238f9c1e6f726e9d88fb3

Observation bd363c7f-2116-419c-90ee-98d4eba56b0f · outbound

This paper cites Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Pixart- α: Fast training of diffusion transformer for photorealistic text-to-image synthesis,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.105328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.394885Z digest=sha256:ed22e01406616e966889bca7171e57ae123aa3373204a0e246141124ea7c997d

Observation 3f668409-2eaa-4f7d-a0b7-8edffb70c318 · outbound

This paper cites While improvements may be possible, further research is required to identify optimal hyperparameters.

Text-to-Image Alignment in Denoising-Based Models through Step Selection While improvements may be possible, further research is required to identify optimal hyperparameters

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.700422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.541197Z digest=sha256:c950a2140d3053698fafec2daa7675cb6c12d8cb254f8e2a914042f47c6e975b

Observation c6767a6d-ec20-4069-b967-d140af4f1572 · outbound

This paper cites Ilvr: Condi- tioning method for denoising diffusion probabilistic models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Ilvr: Condi- tioning method for denoising diffusion probabilistic models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.087420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.401761Z digest=sha256:4ef14cbbe6015f0ae548f537194031e885f59d92c115b8db8862c9b4f28b6b06

Observation a597db35-672c-4596-b701-4a1e350c9a2e · outbound

This paper cites Exploiting the Signal-Leak Bias in Diffusion Models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Exploiting the Signal-Leak Bias in Diffusion Models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.058602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.411653Z digest=sha256:d3d93818b909b6f44f2ab43763caeb7d46ddeef16cb3d107f5c5d95916d377ef

Observation 79bd7936-4a12-42ab-a1f1-08463d373761 · outbound

This paper cites Training- free structured diffusion guidance for compositional text- to-image synthesis,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Training- free structured diffusion guidance for compositional text- to-image synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.049287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.414791Z digest=sha256:15f8926175dfbbbb15c5ddd8f36ce6f91f10de2b2991d06e5c23c7a7934404df

Observation 71e7588a-7604-449c-9198-7e38619dd751 · outbound

This paper cites Fleiss et al.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Fleiss et al

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.039438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.417918Z digest=sha256:bbde076c758c12cac2c8bac8da65336c2aa433766b9b2efe7568f6bae00e8287

Observation ec31a6ff-98c4-4574-8344-b58b39a45f54 · outbound

This paper cites an unresolved cited work.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:42:25.657349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.555555Z digest=sha256:d0b45d8c360a5f49d9fd31e190ba6050555a0db46f5acca55c37cf0a103d1a31

Observation d818f417-804a-48f6-a5b7-c2369db70acf · outbound

This paper cites Signal dynamics in diffusion models: En- hancing text-to-image alignment through step selection,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Signal dynamics in diffusion models: En- hancing text-to-image alignment through step selection,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.019765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.424620Z digest=sha256:f7f1c839e008a3098419fc2f49811b58a2517fad9d52cc8569268c7aa1317a21

Observation bb017783-2f27-4bc0-b5f9-12fb4589e7d1 · outbound

This paper cites Prompt-to-prompt image editing with cross attention con- trol,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Prompt-to-prompt image editing with cross attention con- trol,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.999321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.431180Z digest=sha256:4d36d53ce6db038ce142064e1789006a545b4d9521d333cc2e8415fadfb01990

Observation f7bc0201-0b48-4398-8bb5-1f983e2c0405 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Classifier-Free Diffusion Guidance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.434297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.434297Z digest=sha256:59abad63da8bdfdb6b06afa040e438be138463ddaae03f3b8ba5735e17954eb4

Observation 18498b87-1b4e-4f84-8970-d37f4e95bc1a · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Denoising Diffusion Probabilistic Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.437937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.437937Z digest=sha256:7252a0e6837c01ee4e5bc252b0a75599e137d4b8c1407748149f2e68d9a36106

Observation 602e508f-663b-438d-a2f3-1e82a78fd3d9 · outbound

This paper cites Kingma and Jimmy Ba.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Kingma and Jimmy Ba

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.444893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.444893Z digest=sha256:d63f605c21bfe95bf673ad2dc47bbebafc657368aa245cea0e92741a01be8d89

Observation ae27a167-8571-4d5c-9401-a79483548cdd · outbound

This paper cites Richard Landis and Gary G.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Richard Landis and Gary G

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.963524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.451210Z digest=sha256:f3f6ff319e0f2c8bd8a395af13069e3a4b4e4081c0bd895c4d0a8067e89bd42c

Observation 26f4833c-41be-46a5-b605-2568f49da848 · outbound

This paper cites Divide & bind your attention for im- proved generative semantic nursing.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Divide & bind your attention for im- proved generative semantic nursing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.943981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.457302Z digest=sha256:c5bd37cdc6d68a11476c7c74ad9146f2ec282a8cefcb5e4c5896366d29bb4920

Observation 0c4cde37-8d20-49d2-9bb5-06f24e87f7cf · outbound

This paper cites Lawrence Zitnick.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Lawrence Zitnick

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.934469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.460342Z digest=sha256:d31530df7edfaea093af5cac949494fb5784e39128466e61d5cda2a720408ba5

Observation f7dab736-88ac-4bfc-ab07-331cdceedbb3 · outbound

This paper cites an unresolved cited work.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T10:42:25.914901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.466582Z digest=sha256:ddefd8d02c58499779333d2d23ecf43edcc0dfdd7257dd0963282e89d9b94445

Observation 4dba8fa5-9330-4db4-957b-0d2514d3430a · outbound

This paper cites SDEdit: Guided image synthesis and editing with stochas- tic differential equations.

Text-to-Image Alignment in Denoising-Based Models through Step Selection SDEdit: Guided image synthesis and editing with stochas- tic differential equations

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.903671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.469762Z digest=sha256:b813088a7d33e4ef2e7fc8f1971c1a8428dbe9b8d15218eb3e4f96d708194f94

Observation 1c34649a-57df-48d1-a1bc-aae00c996cb7 · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.473075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.473075Z digest=sha256:5c36a67635a1ef76c4f44e21f8e0ed53b1a6094cd8dcfc05b563b948f4d4ddfc

Observation bec9d93c-f3bf-41f2-a533-ddc047fb8b70 · outbound

This paper cites Teaching clip to count to ten.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Teaching clip to count to ten

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.893243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.476551Z digest=sha256:6851393a48bbca54c557e9d1d553c15eefb0f09698c4148f6439b9f0776f531d

Observation 04aaf966-5db7-481b-8a45-1da8bf4bc239 · outbound

This paper cites Understanding the latent space of diffusion models through the lens of riemannian geometry.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Understanding the latent space of diffusion models through the lens of riemannian geometry

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.882361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.479770Z digest=sha256:fd71779643afd624aec32faed7f1cee652e9331f9eedc3a90028b73d7691cec5

Observation c3a51c59-69a3-44c1-b7bd-11cbd0289dc0 · outbound

This paper cites Scalable Diffusion Models with Transformers.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Scalable Diffusion Models with Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.482789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.482789Z digest=sha256:4dad6b7d4c1178a68bb8250d55723c31f4ce5a29757996708003d51b8de8e6d2

Observation c8b67286-38d2-48b9-a30b-d06d3409c9f6 · outbound

This paper cites Sdxl: Improving latent diffu- sion models for high-resolution image synthesis,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Sdxl: Improving latent diffu- sion models for high-resolution image synthesis,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.871729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.486115Z digest=sha256:ddd6d60b45cae0e797c13ac9aec469a3c4ab9791751f8eb8c115b5e082415a30

Observation d9a6e444-5863-4e4a-b9a6-c5b3c7df82f4 · outbound

This paper cites Learning trans- ferable visual models from natural language supervision,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Learning trans- ferable visual models from natural language supervision,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.861062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.488872Z digest=sha256:18ba71f602a02b8c80a799bf167c43e7c5a938782eef27be3f5940b24ddf89d9

Observation 8b25cda6-12f7-4c50-8e6f-31aeb9a269b8 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.491701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.491701Z digest=sha256:81c672f97ea0f3a95971f23d109dda212ed94b428ec5fcfb5d7bacfd218e8b7a

Observation ee57f05f-7803-40bb-a3ca-0ad23ee0f322 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.494927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.494927Z digest=sha256:59461473af847822196d87d4c07ef3ac3c5f28edb1c6f48633ed215bc1b4e60d

Observation 304bfb3f-15a6-49ba-b1ba-ecfc6fdf6909 · outbound

This paper cites Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.849661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.498150Z digest=sha256:5b95ab3304ea3809975cafc588030fdd63c34f9577a7b21baca48fe4950d79aa

Observation 3bbac828-79bc-4ad8-bad6-9b42eb70d892 · outbound

This paper cites Generative modelling with inverse heat dissipation.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Generative modelling with inverse heat dissipation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.838843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.501121Z digest=sha256:6f69bffc28ab89602321f5007faaa7aa4a9cbf22f682765e0332c133d7fa56d8

Observation 37d6201a-6b1b-4c8f-8765-7e1cb507fc2a · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection High- resolution image synthesis with latent diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.827082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.504022Z digest=sha256:3155e57aab27fde6894d0fc10cc155ac833da7f53fc83fc6c41a233fe1a98a5c

Observation 8170c227-ccae-4352-940e-da8dfb364fd8 · outbound

This paper cites Fleet, and Mohammad Norouzi.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Fleet, and Mohammad Norouzi

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.815661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.507181Z digest=sha256:bf49d74d76b7e5e70886357436a56b531d0684f468418bfd265fcad0bdd75e99

Observation fa7de1bb-53b8-4b71-b7a1-b8ab68d229ca · outbound

This paper cites Laion-5b: An open large- scale dataset for training next generation image-text models,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Laion-5b: An open large- scale dataset for training next generation image-text models,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.805071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.510193Z digest=sha256:eeb36ecdb0f01b5c1b42da170c9237bccf39f1ba00d653a56e480e32dd35f3e9

Observation 5116c87f-7fd3-4b18-84fe-cfea6873e777 · outbound

This paper cites A picture is worth a thousand words: Principled recaptioning improves image generation,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection A picture is worth a thousand words: Principled recaptioning improves image generation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.794019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.513243Z digest=sha256:a442fc74bc17dc2d4bdba6ffa9d6dca94ac0b8c3149c93a0b8c377aa6b880c90

Observation f83b8276-509e-4a1b-93ce-b76421b246f8 · outbound

This paper cites Spatial-aware latent initialization for control- lable image generation,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Spatial-aware latent initialization for control- lable image generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.783269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.515956Z digest=sha256:4c095a78283d238a58fbdc2bf6f1807c49881eaf67e352186796bca391d6abac

Observation cf224006-bea0-4305-8ce4-803898fdbb18 · outbound

This paper cites What the DAAM: Interpreting stable diffusion using cross attention.

Text-to-Image Alignment in Denoising-Based Models through Step Selection What the DAAM: Interpreting stable diffusion using cross attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.773700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.518884Z digest=sha256:5744435d641434d82a500ca81eb12923d5d5458b06e5521d0dba4ffc41444563

Observation 83be4b21-5aac-4747-b806-32f9f653d2a1 · outbound

This paper cites zero-shot.

Text-to-Image Alignment in Denoising-Based Models through Step Selection zero-shot

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.764464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.522530Z digest=sha256:a472823192cd1c941dce92eb77218cea58ed460afb3c0013eb2892a785823e08

Observation 0ab3065c-10de-4fcd-865f-c6dcd033156e · outbound

This paper cites Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Boxdiff: Text-to-image synthesis with training-free box-constrained diffusion

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.755289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.525528Z digest=sha256:e8614643fdec76782a97988c1c4bc5215976fa2048ff26757fd7f2f577c50873

Observation 5417e8dc-26b6-43e6-a546-caf9964d2117 · outbound

This paper cites Freedom: Training-free energy-guided conditional diffusion model.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Freedom: Training-free energy-guided conditional diffusion model

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.745194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.528579Z digest=sha256:a2c92c4089e1176a897b852f970470c907051ec92fe49316770b060014a6247d

Observation 6e24ea95-4b68-4524-9e8a-118f701ce33d · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations ,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations ,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.735050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.531470Z digest=sha256:95df85c4f08c7b280671b1ab943f4b194691e652cd2b99b260cd6cf6dc2f4c55

Observation a71674ce-e733-4510-8620-03c523164ad7 · outbound

This paper cites Adding conditional control to text-to-image diffusion models,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Adding conditional control to text-to-image diffusion models,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.724812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.534465Z digest=sha256:8a175d471ebc7df57076521dcea77426871371f6a52ef82c5417a086352f6528

Observation 91517407-7cae-4d7f-9d7a-caf327e34bb4 · outbound

This paper cites 11 A.1 Text-to-Image Methods Setup For Stable Diffusion 1.4 We provide here some implementation details about Stable Diffusion 1.4.

Text-to-Image Alignment in Denoising-Based Models through Step Selection 11 A.1 Text-to-Image Methods Setup For Stable Diffusion 1.4 We provide here some implementation details about Stable Diffusion 1.4

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.712406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.537799Z digest=sha256:5ea487ffdc3ceb799e171760e8736a43f21c3d1659b752c82610d585702324db

Observation 81727da3-926d-4e4c-ad7e-12af6d30a3da · outbound

This paper cites The model learns to transport points from one distribution to another.

Text-to-Image Alignment in Denoising-Based Models through Step Selection The model learns to transport points from one distribution to another

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.690359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.544601Z digest=sha256:7ba4fbbfdc084d84b68a3311a9eaf702ee2db9cf52d4583a8ecd485dc02f5022

Observation 7e875081-d86f-4022-a227-3b76e4ec0d0a · outbound

This paper cites For subject tokens that span multiple tokens (e.g., due to subword tokenization), we average their respective attention maps.

Text-to-Image Alignment in Denoising-Based Models through Step Selection For subject tokens that span multiple tokens (e.g., due to subword tokenization), we average their respective attention maps

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.678498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.548531Z digest=sha256:18b1b40fe4039f3a8eb091865687ecb5425fe01d000856afb4c127bd3e99cfdd

Observation 43f9a52d-446d-444f-9d20-f93d402dc3db · outbound

This paper cites a photo of det(o1)o1 anddet(o2)o2.

Text-to-Image Alignment in Denoising-Based Models through Step Selection a photo of det(o1)o1 anddet(o2)o2

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.667690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.552073Z digest=sha256:81759ddf62f277089a51bfccb6984881ead9684749afbeb193e0747255a78579

Observation 4147eca8-9dc8-4df9-9dd8-e7a605ce6a8f · outbound

This paper cites Tiam - a metric for evaluating alignment in text-to-image generation.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Tiam - a metric for evaluating alignment in text-to-image generation

Reference 1971

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.029434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.421379Z digest=sha256:a6036a0d3ddace60b52e88f3fd06634604851f848766e3a4a619928e97376bbd

Observation 5e8a9ac7-3e01-4a53-8036-1da17a62a3b7 · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and gen- eration.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Blip: Bootstrapping language-image pre- training for unified vision-language understanding and gen- eration

Reference 1977

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.953898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.454263Z digest=sha256:bbc8536480e3e12acf83dfd7ed5d61be197a6e1a828e0266486f01b61c808bbf

Observation 02360b9b-dfdb-4cd3-b952-4f351516d716 · outbound

This paper cites [Lin et al., 2024] Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang.

Text-to-Image Alignment in Denoising-Based Models through Step Selection [Lin et al., 2024] Shanchuan Lin, Bingchen Liu, Jiashi Li, and Xiao Yang

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.925082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.463345Z digest=sha256:73554d6bc08a1cdc6a6aabaa0e24aca4921fa3e8efb8cba8c7f6a7a52f5b7833

Observation d2e1ad9f-3a6e-4a75-9ceb-726d72f879c4 · outbound

This paper cites Diffusion models already have a semantic latent space,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Diffusion models already have a semantic latent space,

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.973310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.448153Z digest=sha256:824523b325b5c71076ba42a712b33b8d8c13e64b430c1b865c57f88ea019e00b

Observation a0060848-6bb8-46c5-b548-cc41aeb7be03 · outbound

This paper cites Yolo by ultralytics,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Yolo by ultralytics,

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:25.989450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.441786Z digest=sha256:97544e3ae7914d2edba805a3de8ae9aaf9d7b0b36af3956c680b21adcacd63ad

Observation 603a38b1-d893-470d-9e3d-a974a1c1d5c7 · outbound

This paper cites Perception prioritized training of diffusion models.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Perception prioritized training of diffusion models

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.077524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.405112Z digest=sha256:e3aa7b3d3a8ea23ccab0c772964f93a6e17900ff22e1e4d8f18abe0f91ca44a3

Observation fd10c6d9-dda6-4aef-a527-95402133f937 · outbound

This paper cites Scal- ing rectified flow transformers for high-resolution image synthesis,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Scal- ing rectified flow transformers for high-resolution image synthesis,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.068045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.408698Z digest=sha256:caa3a3119dbd9f0d98b485fb90b803069914290b8d11a201082a253417976e06

Observation 6eca2714-fe4c-4b3d-ad36-811ae536e729 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Text-to-Image Alignment in Denoising-Based Models through Step Selection eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:42:25.387277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:42:25.387277Z digest=sha256:de6db112a434cc8977fb2b54ce70af859392e3c5caadb883849bd6646ed53db4

Observation 1e8c1210-fe91-450d-9048-0082b2bdd389 · outbound

This paper cites On the importance of noise schedul- ing for diffusion models,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection On the importance of noise schedul- ing for diffusion models,

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.096168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.398236Z digest=sha256:7d8f574e8c0df5a1bb3de00576a34ee5e31a7aaefed4856c21c34ccfce6a7198

Observation 7f6c983b-374e-4e53-802d-0cefae0e9d4c · outbound

This paper cites Initno: Boosting text-to-image diffusion models via initial noise optimiza- tion,.

Text-to-Image Alignment in Denoising-Based Models through Step Selection Initno: Boosting text-to-image diffusion models via initial noise optimiza- tion,

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:42:26.009188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:42:25.427971Z digest=sha256:5eed6b93da487892a0a59d3164d832a2e78f27d21a1bc2916b3d8fe1e75be1aa

Pith citing papers

Observation c660c21f-e36c-4b9e-b4e7-a67efcf78104 · inbound

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time cites this paper.

Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time Text-to-Image Alignment in Denoising-Based Models through Step Selection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:42:04.111093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:42:04.111093Z digest=sha256:fd27dab03fc8d2b8c72438a93cf3e18e600864cb03891c34219b38c8d82059a1