Pith. sign in

Paper Citation Record · LEDGER

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2411.15236.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15236 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:08:54.669287Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:53.297589Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:58:37.263440Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 860654df-32b1-470a-a7be-77ce9b923569 · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.587696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.220813Z digest=sha256:ea05dba38c58c47fcb949439bfdc77a2ea33bfd1f0672d8a8ad6379f5d2ad39f

Observation 00464a84-ec3f-481f-9cba-3fb45053ed2d · outbound

This paper cites AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps AlignIT: Enhancing Prompt Alignment in Customization of Text-to-Image Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.655150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.416975Z digest=sha256:25738e52d0cf6fd58a6dfb1bde84bfadf005bf34f599057aedcf6c9bad25fa92

Observation f3a26ef5-c45c-459a-a889-f967ed3a35fd · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.510085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.538102Z digest=sha256:2a6b52ef043937da81bae06dbb8fd3dc49910bf7336d6481c016f814deda657e

Observation 86b56c1a-ed34-4e6f-a549-f9508e3e5072 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.586669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.586669Z digest=sha256:d4fbd231bda520105078ada734f63160729173b7e39f4b4c3ab1a90ff70ccba9

Observation c1746190-6bc2-4565-81e2-59ffb069b285 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.596085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.596085Z digest=sha256:20d422245db2fd8dcecd73f41aa15e7844fe252683aca04d3d5ae6b315ccea2d

Observation 338f1936-24c5-4ed3-94a4-fa7ee20f7b0b · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.459247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.629726Z digest=sha256:f6056e8e4cca49d9fa88fd3844279655aa33ab2b92f8b68b4c52bea4ea297a0b

Observation a614f833-0c21-4618-8e76-d663731cda2e · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.675405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.675405Z digest=sha256:0cf55500f5e1e88de8e49353d3398c8817aa42c977dd6aaee65172a3ba27a87a

Observation 38e3d833-4cdc-45bd-8fd6-117feeac26d1 · outbound

This paper cites Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Training-Free Structured Diffusion Guidance for Compositional Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.683502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.683502Z digest=sha256:9f34cefbd7c62e0b2861d08584db2dd9c03209c6628fb25f982e74ff47456695

Observation 735bb36e-7c43-4f19-bb83-85f56785282f · outbound

This paper cites When Attention Sink Emerges in Language Models: An Empirical View.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When Attention Sink Emerges in Language Models: An Empirical View

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.730797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.730797Z digest=sha256:a2ae709c4ca7e32f8f00166a518e82240d2bb14ba7932106c29c6a35782e9f71

Observation 9f994c93-3122-47c8-bd26-2ca7ec81dc13 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.768421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.768421Z digest=sha256:b4611d355c158511c741a1710a99eba7cd8664233afc111d468ac9628079ce11

Observation e0a854e5-19ca-4449-8159-8bf6b8d2fed6 · outbound

This paper cites spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps spacy 2: Natural lan- guage understanding with bloom embeddings, convolutional neural networks and incremental parsing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.362672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.783216Z digest=sha256:2deb4f499411a607addb9c7f0539b43e8eb82bf5fd6ee9d26db3aeb59c65495e

Observation ce474741-af47-465b-bc87-cbc72ca61bd7 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.821243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.821243Z digest=sha256:705c630dee190067d7af262a6c46a6b25ef934a1d21a9c3a4a7c5fade7670a87

Observation 7ffb49e3-94ef-49ad-a930-5195b5240431 · outbound

This paper cites Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Tifa: Accu- rate and interpretable text-to-image faithfulness evaluation with question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.288975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:53.842397Z digest=sha256:7835c9494a951e0c997d1dd4ae25a69721646415eb1a5f9e4c9844d785729b68

Observation 4139fcd9-6cf9-4359-9bca-762418452722 · outbound

This paper cites MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps MC$^2$: Multi-concept Guidance for Customized Multi-concept Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.854867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.854867Z digest=sha256:7f194d458fa3b9a83491001d897c9ea271caff7d9ecb29776f93d698791a492b

Observation 166d9dcb-4ec2-453d-b425-f61bd04fea16 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dense text-to-image generation with attention modulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.879934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.879934Z digest=sha256:870ae95d0645b1e8c667559d631a5d9720000a987194a239d3880333d5b042e9

Observation 9be8d80a-c345-45a0-a1d6-eb7b3402a058 · outbound

This paper cites mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connections

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.954810Z digest=sha256:49ee9fc8532ca5c307daa10d8c7bf077d96b423aee2af92aa031da15c644cc16

Observation 978e7c52-b3e2-417e-9378-8a35852e42b8 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.982683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.982683Z digest=sha256:4e3aea02380263b789dacd2b3bfd24db5e31ecd23eb68121d93b0c4c48e73f25

Observation cd0d00d4-ea19-4ec4-9c31-f439499692a8 · outbound

This paper cites Microsoft coco: Common objects in context.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Microsoft coco: Common objects in context

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:53.993512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:53.993512Z digest=sha256:feebbbc9f26941520c38d576880a2786d8eb7f67a95703498c1ed60f158ad468

Observation 37b1587b-5b1e-4954-a3ad-5feb3b425a7a · outbound

This paper cites Improving Text-to-Image Consistency via Automatic Prompt Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Improving Text-to-Image Consistency via Automatic Prompt Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.008711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.008711Z digest=sha256:7e23896d5cbd9daccafe0d80fcbb17d6d41c5acdbf20c26ac12dc073e4200f16

Observation c62ebd91-f847-4e46-a36a-22593e22af9e · outbound

This paper cites Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:08:55.214378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.067179Z digest=sha256:608820a31dfcb92039192f1e987289383e85008616ab5afe0d356afcca12127f

Observation 22456177-2924-4e5a-80e3-d21454ed916a · outbound

This paper cites Conform: Contrast is all you need for high- fidelity text-to-image diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Conform: Contrast is all you need for high- fidelity text-to-image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.127872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.125524Z digest=sha256:a6d182bec3764b5d58215716774543b48b927ddb4804ec8328fa15bb3dc28681

Observation 70120ce3-9e5c-4932-97f2-7d689dec5a89 · outbound

This paper cites Openai gpt-3 api [gpt-3.5-turbo], 2024.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Openai gpt-3 api [gpt-3.5-turbo], 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:56.003417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.138455Z digest=sha256:34b45e0e159fda9f59805844fb4df00f85f54878f1fb1da06acafe7034a66c9f

Observation 1b51419f-9f9c-4f71-bb8d-f1216a926d55 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.150477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.150477Z digest=sha256:bf371cd565accc5a3f8bd2d9d5577f3e64a077e26ba9a35bd07b3fa567684353

Observation f7c116de-be8d-4031-88e7-7e733ee9434a · outbound

This paper cites Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Not All Noises Are Created Equally:Diffusion Noise Selection and Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.173401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.173401Z digest=sha256:5f2e6796a8269616c0ad4cf0c87faf1140dfa9cd9b08388ed27689e726f1538b

Observation 1ade333e-63e3-43c6-9714-1cd9299d83c7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.188458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.188458Z digest=sha256:98694bd81bf10938a97605fc2a2893b1b156c03907d0e164dd104a657302be0c

Observation d350a184-03ee-4255-a4d6-34504c349c54 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.235980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.235980Z digest=sha256:62598d68ff9bae7078bb9bd9f1b5289ae535875552e16261c0292836b9d8a3ad

Observation 0b906bca-8342-4e94-a1ca-9173e430b137 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.256485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.256485Z digest=sha256:dadaa9c7b1e823adfaac3fb0d47f50cc2161c58207575b70f2c8ebefad68f00b

Observation 11e5687e-4de8-4665-8bde-81fe11a06ab5 · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.870591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.270378Z digest=sha256:91f83d7bbd825d22371f0ce9b4d539cc5a1bba6ce1dd80800d18e8b1dd9adc51

Observation b3976570-40f6-4017-8709-f07fdd194195 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps High-resolution image synthesis with latent diffusion models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.331546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.331546Z digest=sha256:de6d665555b245431d95289021f9d50100c73bca74b82aa419536987b278d7a7

Observation e67aa0c6-9e1a-4956-a64a-a3190cd26d30 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Photorealistic text-to-image diffusion models with deep language understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.791771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.376477Z digest=sha256:602b2346fe7f30d23a3cb7165afc937c4216267ebf38816faf1057348103b2fc

Observation bfd31bc5-6cb1-49fd-be3d-344a8282fc4d · outbound

This paper cites Rethinking the spatial inconsistency in classifier- free diffusion guidance.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Rethinking the spatial inconsistency in classifier- free diffusion guidance

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.762366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.386550Z digest=sha256:c5c9eea9780336a3ebc8911530e2fe16a32271d77cad178e83960923c1fa3a0c

Observation fd48f046-aa25-41ac-bf2b-ac475226fc33 · outbound

This paper cites Massive Activations in Large Language Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Massive Activations in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.393582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.393582Z digest=sha256:4848df33571ffb4b3e1a07e11a89457c27cf36406d2c2393300d7641d7140e1d

Observation f95914e2-e112-4fba-9254-4cc7a3ab33f7 · outbound

This paper cites Attention is all you need.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.448164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.448164Z digest=sha256:ffc92da856fa93819de401bd912ee349238bd810476602d9d8905c9550c31aec

Observation 61e611f5-978f-42c8-af57-4d3effba7709 · outbound

This paper cites TokenCompose: Text-to-Image Diffusion with Token-level Supervision.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps TokenCompose: Text-to-Image Diffusion with Token-level Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.481001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.481001Z digest=sha256:788d15a2fd4b46ff8f255f0eda60a9c5d9a85d34d7f08822bfbc94496780259f

Observation 7ea65f26-5a4c-4102-a902-2ab30b770d93 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Efficient Streaming Language Models with Attention Sinks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.493758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.493758Z digest=sha256:21dd796ca91ab463574bfe8d8c6d0978750fac443331bb0c4b4bf06e43399c34

Observation 68feadc4-3997-4a2a-89c4-757edabb1c2d · outbound

This paper cites Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Dynamic prompt learning: Addressing cross- attention leakage for text-based image editing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:08:55.688429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:08:54.508264Z digest=sha256:aebd44a675ab1a123b1d80cb32dc32d0446bf1fe72be209122f12df77b8e41b5

Observation f24ac3fb-769c-4092-b47d-a0ea293bd3a1 · outbound

This paper cites Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Towards Understanding the Working Mechanism of Text-to-Image Diffusion Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.570225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.570225Z digest=sha256:0b96a27f216df138bf105f335c4eeaa69f7c9e5125aad2eafb1af65af4c2c79c

Observation 5a92f757-1784-42dd-8ef2-10fe272ceac7 · outbound

This paper cites Uncovering the Text Embedding in Text-to-Image Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Uncovering the Text Embedding in Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.605688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.605688Z digest=sha256:4eb9ee7d11b6480dda57d277687134b4d17e42bdccf6da4ad2c4cbbbd74167b2

Observation 988c4283-3a19-43a4-aa41-6e97888d8372 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.617831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.617831Z digest=sha256:64fead92dbe113bdab4ef06c3344f18e73c2360e00fd413af36f513203501cf2

Observation 1c5deeee-8208-4051-8fa1-6835cb49f8d4 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.627473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.627473Z digest=sha256:46bb8fdc21af56c42f526403f964fe5845587b5a159903de587d256d5af71f2e

Observation 2d139588-cb0e-453b-8fe0-eca0c44a2ca8 · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:08:54.669287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:08:54.669287Z digest=sha256:a8da4c739fa86cabf1a33d8120889721452db4f9753f9624b5e75fa58efea6d0

Pith citing papers

Observation fdbcc39c-bc22-4799-9ead-0a82d02e5fc4 · inbound

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models cites this paper.

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:53.297589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:53.297589Z digest=sha256:cfdfc47c811e383eef66ff37812a12bc18ce27ebfbffcffd15371ff2a502564e

Observation 46a62a12-f3ce-40f9-a558-0886d5c4ec17 · inbound

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation cites this paper.

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation Text Embedding is Not All You Need: Attention Control for Text-to-Image Semantic Alignment with Text Self-Attention Maps

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:58:37.265067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-07-03T15:56:38.037304Z digest=sha256:0f6bf4182ecbe3bda03bb362add42f9436d44cd162ea5a69933fb79c25d1e0eb