Pith. sign in

Paper Citation Record · LEDGER

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

As of 16 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 6 inbound Pith citation observations for arXiv:2411.18301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18301 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:23:39.194911Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:55.020093Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T22:02:50.398472Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b7ac967-ad2d-469f-b8f1-0833d46bd194 · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.043236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.043236Z digest=sha256:9c9de2d134dbf3015cd07358355293899ef8553dc6d9f6d168e4704b9a82387f

Observation b6b36b74-55fa-402f-a2f7-786da14c8a48 · outbound

This paper cites Separate-and-enhance: Composi- tional finetuning for text-to-image diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Separate-and-enhance: Composi- tional finetuning for text-to-image diffusion models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.116696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.116696Z digest=sha256:ea7b7b12ebdd1422e9075aaa4e7841214281a79cef20e20b24c26d53702d99f2

Observation fa70afa9-7a6c-4d04-81f7-b2c69bd1fbbb · outbound

This paper cites Make It Count: Text-to-Image Generation with an Accurate Number of Objects.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Make It Count: Text-to-Image Generation with an Accurate Number of Objects

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.197831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.197831Z digest=sha256:e516f0621e61a9df8a7251c5e6bc8b4646ec35029ea3327a2d5d88fd72e5a7e1

Observation 40e1ba90-e595-4ba4-9e11-527794ca7a5d · outbound

This paper cites Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Attend-and-excite: Attention-based se- mantic guidance for text-to-image diffusion models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.544596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.201989Z digest=sha256:5f18b744db00ad350ffbedb9cd4bbe013229eaf7072b6782a26668fa2c77e145

Observation d6dea554-32b1-4bff-afb4-d1a4bf206666 · outbound

This paper cites Training-free layout control with cross-attention guidance.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Training-free layout control with cross-attention guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.205808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.205808Z digest=sha256:4bff5b07d444291d8a4d47151fa4919235b5b2685a3684707c0d7fb5f18dfe1e

Observation 83f0ef26-6420-4f2f-8cac-dd06f889bf94 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Reproducible scal- ing laws for contrastive language-image learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.209892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.209892Z digest=sha256:1877d909f131a73aec0e1cebb48b1b1d375f945db534f057ba08294f070973db

Observation 7cf3d7aa-6c68-418a-8afe-ab3ca7c4052b · outbound

This paper cites Be yourself: Bounded attention for multi-subject text-to-image generation.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Be yourself: Bounded attention for multi-subject text-to-image generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.518462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.214371Z digest=sha256:709fea81ebf1c25d6fcdc722076e5f1e40a3287e7101a87a1ef46d992e739b8f

Observation 1b5e9962-00ee-4125-b24f-dd661b0b1715 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.505675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.218347Z digest=sha256:fad26ed68911c26b5ae2ac841ea7429b2a94de5226adf31afbcea7d4cd934086

Observation 1daa7a5c-3cf4-46b9-9a32-f40123d61b80 · outbound

This paper cites Training- free structured diffusion guidance for compositional text-to- image synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Training- free structured diffusion guidance for compositional text-to- image synthesis

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.401492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.221834Z digest=sha256:bfd49c2f37072a49639d32545e94ab1081853f38859a373b4a079524afc213a5

Observation bdccbab0-6763-4c00-b01c-b515802fe92d · outbound

This paper cites Initno: Boosting text-to-image diffu- sion models via initial noise optimization.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Initno: Boosting text-to-image diffu- sion models via initial noise optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.225772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.225772Z digest=sha256:45f9e066124897d60fef88010e1f3b161b20b6b6a717ed2320cf9974f8ba82f7

Observation 8d9f2fa3-cece-4ba7-aef5-a1ea67106a9a · outbound

This paper cites Optimizing prompts for text-to-image generation.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Optimizing prompts for text-to-image generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.381571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.338635Z digest=sha256:f235b654e9fe3fa746a09e86e7cc53c8136a02be394829258b2b84f066847279

Observation 3f75a776-ea8c-411a-a646-47dcc18c425a · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.315509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.342137Z digest=sha256:797d4e1a8b77884c68785d26ba065299050016e8f90bb8cbdefd57b097945ad9

Observation 49d280ce-a727-4b10-8bd7-f62a4de6b2e1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Classifier-Free Diffusion Guidance

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.346111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.346111Z digest=sha256:20f93a171c2f0948e0b3134084510e22e4ee3d8ea6b61f0212f0457bd4a2a64a

Observation bdd953a8-8959-458d-a57c-51bd1f9222d3 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Denoising dif- fusion probabilistic models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.350184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.350184Z digest=sha256:1070c464e8ad1aaaf94f1920aaf95e2afb0bb39ba81369daa1154755ba537cd2

Observation 1a7572d9-e64a-4f5d-afbc-e07231b7c4a4 · outbound

This paper cites Scal- ing up gans for text-to-image synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Scal- ing up gans for text-to-image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.353310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.353310Z digest=sha256:024c2b3e6ca1decca46016f65487a565733dcbc124fa003ad103e1398da20857

Observation 579048d3-32b6-456e-9780-b08236d76551 · outbound

This paper cites Counting Guidance for High Fidelity Text-to-Image Synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Counting Guidance for High Fidelity Text-to-Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.356572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.356572Z digest=sha256:5bd91bca5dac0cf953e5b0f42cbf76b2c8571f4948fe0f08115eac3ae0d71d5b

Observation a996a307-abfb-422b-ba7c-1e7d880a0cf2 · outbound

This paper cites Compositional visual generation with composable diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Compositional visual generation with composable diffusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.236070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.360200Z digest=sha256:b2ce01e3544b4994815a56919f816b7f69cfafd6f24a2259a87ba1947f200dd8

Observation e2196b6c-4a59-4244-8517-5ccbd57801d6 · outbound

This paper cites Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Grounding dino: Mar- rying dino with grounded pre-training for open-set object detection

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.222124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.363438Z digest=sha256:681d6c964ac065a1c8a21e50224f237f4c0a8f8c721a3801d83a771049037f2a

Observation edeb82ba-38fb-4ce5-a7a5-c462eabb3979 · outbound

This paper cites Design guidelines for prompt engineering text-to-image generative models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Design guidelines for prompt engineering text-to-image generative models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.208524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.366729Z digest=sha256:005a5bbb9af3ffd90efbb1a6c6dbf3f9a4b8d98725f1e0ff7219a87efcf08137

Observation 7118c30c-170f-499b-9d9c-6a3c05c1c738 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.195227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.370565Z digest=sha256:99afc801becea92f0fadf13c3d0e2442d9eaca594978f46102cc7bc170a4ea49

Observation 990cc766-901a-481e-8f4e-b0dac7a31e8f · outbound

This paper cites Conform: Contrast is all you need for high- fidelity text-to-image diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Conform: Contrast is all you need for high- fidelity text-to-image diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.156326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.374575Z digest=sha256:8bc3e3a211bfb8530cfd13cdfa8525339453c4236578f1892cc79d351f171675

Observation 1e700a80-f3a4-4760-b714-24a165f5a6cd · outbound

This paper cites Glide: Towards photorealis- tic image generation and editing with text-guided diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Glide: Towards photorealis- tic image generation and editing with text-guided diffusion models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.378170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.378170Z digest=sha256:8ac4a1568baa3703fbefef5b81b4f27119570f5d7cdd6d52f6e35446be3951d0

Observation 228509c5-4103-470a-9d7f-963f73be38c6 · outbound

This paper cites Gpt-4o mini: Advancing cost-efficient intelligence.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Gpt-4o mini: Advancing cost-efficient intelligence

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.449916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.449916Z digest=sha256:1ecdd3848256045c229ab4257a0bccb7cd3c9590ba8741baacfc5f9781012bb3

Observation d882c672-e587-4f36-bbd5-b90109b6f892 · outbound

This paper cites an unresolved cited work.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:23:40.024964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.745755Z digest=sha256:9f1ff79b1084a01f619d931a1e89ed648a9a3f38c6f1fdaf006f1957b4ba3991

Observation 71b2bb50-a706-4842-8904-281f3d02f709 · outbound

This paper cites Scalable diffusion models with transformers.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Scalable diffusion models with transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.800435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.800435Z digest=sha256:5796431ea1877757577bd12c61cdb11f3425a72ca90d2e8782a65629bc456025

Observation 017b59e2-16c9-4f77-9d47-34dc4eba6bc1 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.804926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.804926Z digest=sha256:95666ae8d715d1f137061be2aa5d2b5246409698c00a1c7746e729d6028ced7d

Observation b4ca5c12-b16d-42e1-93e7-ad7b07fa1124 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Learning transferable visual models from natural language supervi- sion

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.809870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.809870Z digest=sha256:3d98738ac55da238e16d9b573f1c08a97ac7e59bb911c1cdff3a0178d633dc0b

Observation 33675c2e-1086-4084-9b9f-8670baea0a5d · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.992127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.813882Z digest=sha256:f63de649c644ec7724c3b19c14cb73a72fa62c11e77389a953c73c0ca1a31ab2

Observation 4d9d0ce6-8b5b-413d-a542-f780ee614bf0 · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Hierarchical text-conditional image gener- ation with clip latents

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.977758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.817211Z digest=sha256:5f40f760fe6d8049121679c615e6b29a2cee45e062d731baa987888ae2107e58

Observation bbd40059-bcac-475d-9ee8-36fcaa52e920 · outbound

This paper cites Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Linguistic bind- ing in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.893358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.822841Z digest=sha256:cf75053e7ed3314b59e5fd5b713527df1e55a72f9563ccead45bc940f02ee961

Observation 3dcf6604-d04d-4d1a-a5cb-732fd86de452 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.826536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.826536Z digest=sha256:7958fd85411daebada1bba14fc631ff64bded0a9ae0e05f368cab18334717704

Observation b67d03b8-8170-4235-af1c-9b8d201f5b6d · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation U- net: Convolutional networks for biomedical image segmen- tation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.829757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.829757Z digest=sha256:be52798eff1fcf2d9d20148fad375085fac053f01c89b64fb7f558991597447b

Observation 833800c9-b2ad-4aa3-8cf2-26678c44e1f4 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Photorealistic text-to-image diffusion models with deep language understanding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.801556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.833381Z digest=sha256:30366744f050fe1665eed9fe566ae6353c5cffebbdfc06aa7e7e46a0ba060fed

Observation 23863b69-e10b-4b33-8ced-ae041f090424 · outbound

This paper cites Denois- ing diffusion implicit models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Denois- ing diffusion implicit models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.788266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.837770Z digest=sha256:97d63aa3311581a04c99d4d1ff91c421b094982017054f9d30114b06ef08c366

Observation d781a978-54d4-47d8-bc54-39682746c2ee · outbound

This paper cites Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.775875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.842029Z digest=sha256:9497add06e1134e939fa37328cf077c530672fe0fbafa30a9921146f977926d9

Observation 0935ab0c-1cbc-422a-a4a6-4218777fd227 · outbound

This paper cites Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.762730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.845366Z digest=sha256:aaea6b80a6cb57e4e20fd5149e75f0716e296a09fabe75db6465db7d14057598

Observation eaeb710e-71b3-4f0b-ba2c-9a265ee8c7ff · outbound

This paper cites Hairclipv2: Unifying hair editing via proxy feature blending.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Hairclipv2: Unifying hair editing via proxy feature blending

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.749619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.848565Z digest=sha256:f129bdc76cef8960b64e7979df3f2f05e1e019d57dd643ca22650900a2ba1107

Observation 386e4eb0-ce3a-4e38-9ec9-68c1cb288040 · outbound

This paper cites Investigating Prompt Engineering in Diffusion Models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Investigating Prompt Engineering in Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:38.935021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:38.935021Z digest=sha256:0428ef868dc790bfa60d5a9f169c0b4a8fb015a06c5ceea3a6c42f03c64f4d48

Observation a3b47955-6427-45c6-b4fa-3806bf67e7cf · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Attngan: Fine- grained text to image generation with attentional generative adversarial networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:39.031605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:39.031605Z digest=sha256:4935e720fd080793c90819701fda0d085e4505369ed5beaabd957744c3735098

Observation 3937ba5c-5a9e-4c13-9691-b1a363dc67b4 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Scaling autoregressive models for content-rich text-to-image generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.706860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.099350Z digest=sha256:467005162a81faee9be62de69bd513509418c22046eb97b38c6e298dc337f960

Observation 4a5d5aa7-326c-437b-b9d7-3813dbe6e9f3 · outbound

This paper cites Cross-modal contrastive learning for text-to- image generation.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Cross-modal contrastive learning for text-to- image generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:23:39.127432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:23:39.127432Z digest=sha256:415e3dba0d6b5a28301cf2e9053f6b29c8ea9001602dc2882a91f013918fff5f

Observation 34a71818-393a-4627-8d68-66ed0062a84a · outbound

This paper cites Enhancing semantic fidelity in text- to-image synthesis: Attention regulation in diffusion mod- els.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Enhancing semantic fidelity in text- to-image synthesis: Attention regulation in diffusion mod- els

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.567076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.172903Z digest=sha256:9962bb33948eff5d3ec152bf9374f5d59b6b8a0ad17cebe3100a969613ee0071

Observation 9d8882ee-abb8-4a05-8a66-9349763eba9c · outbound

This paper cites Object- conditioned energy-based attention map alignment in text-to- image diffusion models.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Object- conditioned energy-based attention map alignment in text-to- image diffusion models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.551247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.177983Z digest=sha256:ebf9b08382298848d27adaf41d934ccea17a75929cbeadda1874c387c5c961df

Observation f5e1e28f-bba3-4172-917f-852acd81b526 · outbound

This paper cites an unresolved cited work.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:23:39.536358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.182240Z digest=sha256:2e9f06905e0ca58c4b269cd1780ccee8844ebfce6c3a5c9e561d02e8672305ff

Observation 6a8b929b-d4ec-465c-8218-775c44242288 · outbound

This paper cites a chicken and a duck and a goose.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation a chicken and a duck and a goose

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:39.523545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.186305Z digest=sha256:e177cdaa9893c74b4a1a5ec2e7b425cc049be8cf8d3fec313f3b492315be36e2

Observation 265dfc8c-1306-4a5d-bcba-b7c41d48760d · outbound

This paper cites an unresolved cited work.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:23:39.472133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.190833Z digest=sha256:04c805df27623d822580ecf29770c59a7ec10bf19b249f0cf63aabf5d67c2c8e

Observation 9e4b1081-bce4-4cd1-9985-1ab3b9884c33 · outbound

This paper cites Bad seed, rejected sampling!.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation Bad seed, rejected sampling!

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T11:23:39.353278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:39.194911Z digest=sha256:ccaa2e3cb26b6884bcea6197819c7025029870befb9a2289adc6e414b8dab253

Observation d0ead241-0e83-4425-a886-40783f5d320b · outbound

This paper cites 7, 8, 12.

Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation 7, 8, 12

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:23:40.041015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T11:23:38.637135Z digest=sha256:93b29e85594bae00db5fc11c7493078deb183a8d6f6c7df70f812e63e9dfe85a

Pith citing papers

Observation 64cd234c-1b99-421e-83c6-3df2b3ec6c14 · inbound

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects cites this paper.

Self-Cross Diffusion Guidance for Text-to-Image Synthesis of Similar Subjects Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:55.020093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:55.020093Z digest=sha256:31bafb9cc200167fc3fd78929cebe17f872beb8c391164ee309c72a829a9c2e2

Observation 856ae90a-91db-4aef-bce8-f6a7f4ba09ff · inbound

JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models cites this paper.

JEDI: The Force of Jensen-Shannon Divergence in Disentangling Diffusion Models Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:36.425140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:21:36.425140Z digest=sha256:8da9495bab63fc9601c62b80725f40455e3846bdc1b1b62611ddd2ec9cd3a179

Observation 4de35d14-f731-4746-ba70-481e00658dbf · inbound

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling cites this paper.

Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:19:24.182520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:19:24.182520Z digest=sha256:1eb98c6b7cbe4d7fd2b030cdada5b50d3407f4e9262505579e226c54c815820f

Observation 9bfe246e-1ad4-4f52-8f8e-d93c200fa060 · inbound

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models cites this paper.

Detail++: Training-Free Detail Enhancer for T2I Diffusion Models Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:53.470401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:53.470401Z digest=sha256:fae0c26205680f569bc28716f8247d87013a3f723aae2861cb03bd75643109fb

Observation 50c1ff24-d2f5-4c3c-a041-996742e60fc0 · inbound

LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering cites this paper.

LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:02:50.403741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-05T22:02:50.152765Z digest=sha256:bdd79a0e4fb33efae47a79034f31c1d8484966ac4f30a001aa82d9d8bb9b2f30

Observation d5ea9f5f-5257-4c32-a219-cd3e9235ba5b · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Enhancing MMDiT-Based Text-to-Image Models for Similar Subject Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:30.592173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:30.592173Z digest=sha256:6a92dcfaa8ba99ae1a5d89e955828ff0005305c7db0a9194d91d31d864b79e3c