Pith. sign in

Paper Citation Record · LEDGER

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

As of 21 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2506.23418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23418 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:48:57.591417Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact6
  • verified fuzzy27
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b45c9dc-5db5-4982-9ef1-e50edfe83359 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.808315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.808315Z digest=sha256:1d3ffddc33b4e8c71b5a139ce067abd65ff022cd6bfccfc7a92f8277074e4f68

Observation 3f71b3e6-8ec6-463b-bbec-afd2add3b53b · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.867916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.867916Z digest=sha256:1d3c060834ab4be4e9a9958b9bb2bfd8832b994ac8a1fea3df6a721007c8175c

Observation 07799072-ad50-41b4-b658-3d6ceaf1db9d · outbound

This paper cites Kandinsky 3.0 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Kandinsky 3.0 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.976891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.976891Z digest=sha256:66182a88e0cbd3c5adca7b8b63ac0117fca9af3276e983d0bd006e46e4162f2e

Observation f2a69bfd-0dca-4dd9-8e21-23820b2719b6 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Finite-time analysis of the multiarmed bandit problem

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.033654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.033654Z digest=sha256:dd9f4a994a4dc86d38bc5783885722629d6894ea636b57bd91597f0d21b3fada

Observation 6a8c774a-95af-48c2-849f-65c0d9e7309f · outbound

This paper cites Spatext: Spatio-textual representation for controllable image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Spatext: Spatio-textual representation for controllable image generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.080754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.080754Z digest=sha256:1aeddfa78a51216e78b48e637cccea8cb420029d5685cac60248346d63aaa561

Observation 4b624120-e17e-4a45-9d74-32299bfb1238 · outbound

This paper cites Cc3d: Layout-conditioned generation of compositional 3d scenes.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Cc3d: Layout-conditioned generation of compositional 3d scenes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.127391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.127391Z digest=sha256:b16f55ac1c2631cf90ec4d2a876880ecc0c1cb3c44e8ea4a9e60fd52736869f2

Observation 2e40a2e8-8777-44e5-873e-7d0bd3629d80 · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.167595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.167595Z digest=sha256:b1d1975e3cd8a539533d63467f852b224cd4f65d790b58689234bcc050f06be8

Observation c3b64dc5-5646-4eef-900f-5df72f601870 · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.252650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.252650Z digest=sha256:88fdc15175ed930783c182fdf6280d38a641b6304416e95f34b324702ef8ccf2

Observation 41a8f068-b396-4a9b-b4e4-133192444c86 · outbound

This paper cites Improving image generation with better captions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving image generation with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.298620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.298620Z digest=sha256:3de3704ead8bf62db8f66e1bd705488500ce9ec5419e5960dbc19a7b73cf98b3

Observation 4fd1b3e2-497d-4ced-ae61-9af9773af035 · outbound

This paper cites Make It Count: Text-to-Image Generation with an Accurate Number of Objects.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make It Count: Text-to-Image Generation with an Accurate Number of Objects

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.348576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.348576Z digest=sha256:2b84a6c5a77bfafeba52849c5841468475850d72b32743945960acdbb67c177d

Observation 26a059a3-eb12-4ace-b2ed-612a9e13c0e9 · outbound

This paper cites Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.878208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.432547Z digest=sha256:166af9e0282605b5cd712938d699a57cb794ac0eba5decbc9814ef3f5c69c343

Observation b3a0d950-88b9-4ea0-9329-aa5f0ab4ffb8 · outbound

This paper cites Getting it Right: Improving Spatial Consistency in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Getting it Right: Improving Spatial Consistency in Text-to-Image Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.518955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.518955Z digest=sha256:f98d3ef776984280b58dfe1c92f172d2f632548c33d78c4dbcc4c6544aead239

Observation 4bfb679c-f849-4704-bf91-e714a875ec45 · outbound

This paper cites Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.719627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.601008Z digest=sha256:1b67301da6dae8f0395afc433d612547600535f3c97729b1cb4658cb01fef49e

Observation e1d998d1-f294-4f63-905f-d4b06ff9f1f6 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.657856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.657856Z digest=sha256:a75d69ea6bfae1fc382dae90d7c9cbe3a9aa868b027b23fe25dacfa5318b05af

Observation aaae8dd0-4517-47b7-bf31-a63a6426e88d · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.736340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.736340Z digest=sha256:3b0b2335ab00190c70b2242e5fab3179fddfcc23e983698be709cd1acdca6dfa

Observation 307c10fb-40cd-45ef-8cf9-05c055f67e7e · outbound

This paper cites Training-free layout control with cross-attention guidance.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free layout control with cross-attention guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.564234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.783311Z digest=sha256:5dfd463cc946407e5b402b40e70fcf24f7b701642feddf8377e69690a37637f7

Observation f60b0b08-79b3-4092-81a7-67085a3a6130 · outbound

This paper cites LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.832817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.832817Z digest=sha256:f428112e9f5a47560b1c93c66660defd9027956277d6fa67291110a7832531b7

Observation 4a510a3f-4052-400d-b187-98e87f161fd8 · outbound

This paper cites Dall·e mini, 7 2021.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dall·e mini, 7 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.406246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.883431Z digest=sha256:f8a931c348a00b01af1ffdcc78d0fd91172857ef73e44d4cd2d363e44f6072eb

Observation 9765d374-7cba-47c7-84b2-6d72e643d53c · outbound

This paper cites CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.971405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.971405Z digest=sha256:becb9fd8f9b9f37312d1718e83bd83a1b81413cdf13d7aaa8dbd31b078e5b7ab

Observation c7458bd6-c8d0-4088-bc99-df1b62ad21c2 · outbound

This paper cites Training-free structured diffusion guidance for compositional text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free structured diffusion guidance for compositional text-to-image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:52.059608Z digest=sha256:9f5e5864de2fd9a3431b6289dd904c1340a6a8579fcb83884b69771ac884e85b

Observation 6e88afe7-aa0c-466e-b77f-134e9213febc · outbound

This paper cites Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.182184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.182184Z digest=sha256:71495371ec57897461d6a0f18fb7ec716442629efd3bc2745419928bcd70b7f2

Observation f2b3f176-fa36-40c8-99f9-e08ba899c89e · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.298002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.298002Z digest=sha256:57d36370dfb47cdbb6ab02a4ca59b5054d6e5ffa131d816f13ef2db3f3b1fa37

Observation 54e5f73d-cf88-4460-bc27-fa90af3275fe · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.391265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.391265Z digest=sha256:f1550f132fa261b63cd3d90ec0a72c0fa03d92704cf5607c7912314da732c187

Observation a97cb007-cd59-49c9-8682-85279d24b011 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.464757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.464757Z digest=sha256:64b2b5f871e637513d12f9656470bb2842bb083c17841514dd6211117ca1187c

Observation 56e3b960-01db-4269-88bc-a7be5e2209ab · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.605567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.605567Z digest=sha256:26436dd5540a1192edd1d4e9b4435047d0f14dcff34b54c139efc1b2bb4da0f3

Observation a5d12599-6650-44dc-a625-5dbe2c18cc58 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.737177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.737177Z digest=sha256:c19a7aa6918a33ab08c535263e04d0c61b50decbcff74f19fac05cdf46552fe9

Observation a7b727f7-47fc-4c9b-83c7-40f10f905237 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.842219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.842219Z digest=sha256:4e33186665b256a85c40cd3e896d3b9e76ddfeefa9c10f3f4d1561d996fdba60

Observation b6111a05-8052-4a32-8d30-9ecd28e2c8e2 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.928167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.928167Z digest=sha256:59e9478fffe97596a8132b2cda3945698ee1a3f5d7a3f6e32ec53de59ac6a594

Observation f00c05be-3b62-4de4-a843-a95554a680d8 · outbound

This paper cites Token merging for training-free semantic binding in text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Token merging for training-free semantic binding in text-to-image synthesis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.036753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.000087Z digest=sha256:21eb8c4d66d74ada5acca53adbc7e1bebce296bb9b4cbd8cd94b6aeea1c95299

Observation 79b60aa4-8021-40ee-a135-487356241c7b · outbound

This paper cites A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.751381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.058729Z digest=sha256:7d8c1bb3a30bcc2879750108c0ec83eb31a6869154faf655791be8b0ad892695

Observation 68e853c2-b619-4f9d-95ab-a9eda965bb51 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.829715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.116358Z digest=sha256:299c219c0593269dc82104b34c9aea41d0225091f33f2e18c6dfa770fcee57d3

Observation 8206daf3-8e7a-48dc-9aa8-fead9a8a5ba7 · outbound

This paper cites An information-theoretic evaluation of generative models in learning multi-modal distributions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models An information-theoretic evaluation of generative models in learning multi-modal distributions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.654824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.228368Z digest=sha256:60f8181fad05f05678f5574869f526cf3bf56d52caae4a99b209786625f7572f

Observation 07d48034-0798-4fe7-9a7d-d13cc77c49dc · outbound

This paper cites Rethinking FID: Towards a Better Evaluation Metric for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Rethinking FID: Towards a Better Evaluation Metric for Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.316391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.316391Z digest=sha256:2a34eb7711093a09e18fdceb9aa34e8cc1ea5a8cdd93e8e4b85b563d9f1ea9dd

Observation f2b20c63-ef67-4c6f-8e17-fdad4f6363f6 · outbound

This paper cites What's ``up'' with vision-language models? investigating their struggle with spatial reasoning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models What's ``up'' with vision-language models? investigating their struggle with spatial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.501071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.352085Z digest=sha256:d2cf04a310545e8723f7163a21c09cc9c591c1b64f9438a5ee16703f7f2fe752

Observation 568f43b0-6aeb-4c70-a4c0-43df7e04bb1f · outbound

This paper cites If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.437835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.437835Z digest=sha256:b832730362dab3b7b6193d4c716c0068845b3dd9379eba47560215ba93a276fd

Observation 71edc826-1f08-423d-bc26-8dfe123c5cfb · outbound

This paper cites BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.598720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.533737Z digest=sha256:f862277b32ba9493420b784831fe04479c8ba0a910064e753860dcd4c313f657

Observation 65b24792-db47-47ac-a5cb-977279bd3622 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dense text-to-image generation with attention modulation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.338882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.586132Z digest=sha256:71c9f37c00699a7ecfd14614410a42c9bf16acca538c0001364582f30b79061e

Observation 60788508-c8c4-4c60-b0e9-270c434757eb · outbound

This paper cites Segment Anything.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Segment Anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.722447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.722447Z digest=sha256:c718ec0d8ebac98b5254b191ba32dd4e1e5cfc384fb1fcb2ba82b7b8d8302a3c

Observation 632e5ff8-76bb-4f4e-89de-f5cc318033e6 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improved Precision and Recall Metric for Assessing Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.852853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.852853Z digest=sha256:68d64327ea13ac0a0ed8e6f72c8b9abaca1a4710cb528f1a2472977a78e77286

Observation 938ff7f5-8e90-47a5-b9f5-08632237ec65 · outbound

This paper cites Laion-coco 600m.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Laion-coco 600m

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.158280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.993200Z digest=sha256:99f55dd8851aa5a169fdf2b691dc9514f5ff519e7b3d05ffa89965661bae7310

Observation cc4b0226-9d4c-407a-8619-11b898df9d86 · outbound

This paper cites GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:48:58.472999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.063558Z digest=sha256:96e900b5fb307304f3f038a6ce8bb5b8b6cf995ed69d644b0ebfa0900728d85d

Observation 83dd898d-37de-4d8f-b1b2-ea746ca26f09 · outbound

This paper cites GLIGEN: Open-Set Grounded Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIGEN: Open-Set Grounded Text-to-Image Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.226577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.226577Z digest=sha256:ae40fab9c5367da6e2a27a8fea68a41405678bd9d9dad17d7e29b21a8ab98864

Observation 0981c3cf-42c0-4f5a-ad88-a0fa44bdb9bb · outbound

This paper cites Divide & bind your attention for improved generative semantic nursing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Divide & bind your attention for improved generative semantic nursing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.947376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.383202Z digest=sha256:95ed5c6e8b0ba31dc0d881ab54d3bb0c69f23f5f3a9cd37f31e95209c4594511

Observation 163ad252-7a8e-46fe-966d-63458b240b4e · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.541171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.541171Z digest=sha256:31bdedc36460896238d5a85ce25a7cc556afa48fb0317699cd95b591be98147a

Observation 1b6393a3-ac2e-4b9f-b839-c7024428db99 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Microsoft COCO: Common Objects in Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.670059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.670059Z digest=sha256:e86e4b0fa0e6a47482c0446d14154bc364231015cae8b59c1b5ecd592a8f2869

Observation babd993d-63a6-4fb9-89a9-c3b2f7b3f4cd · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.778142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.778142Z digest=sha256:8a107c63edb1a8bb533d46685144b0818c6ecb2bd278a26b0b8d788de6a2034f

Observation 0c9d4880-0781-4eaa-a364-e680ac9a6c13 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Compositional Visual Generation with Composable Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.880447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.880447Z digest=sha256:74968f4138f53fb087352cbd7e1001cf06991897604f1a3732ba10d4d95924a1

Observation ced64cf3-11ee-4234-afc1-7abfb9ed26c6 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.783156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.923631Z digest=sha256:81ee6739bbb7838612e0d6a750bafa0a7433793df3a1447b2f9442ab43b91502

Observation 9fd69e75-d069-47bf-9d8d-e0c8a10bb6cf · outbound

This paper cites Correcting diffusion generation through resampling.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Correcting diffusion generation through resampling

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.584642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.042044Z digest=sha256:6b44ec1cb241c944ad7e249e4f074eccdb0c05b8ea6be00c1db5c946a1ca068f

Observation 9090d239-61cb-4e85-98b5-3771c792f67e · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.131271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.131271Z digest=sha256:51bfcee48687118e1f20ea3bcc872d579b3cab7547a2be8bc39556fc70ad46d0

Observation 1a67cb10-e9b4-462b-8869-d8d9ce9ea2b3 · outbound

This paper cites Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.305767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.186979Z digest=sha256:3c4bce23427ca8bc873cf572bd1a9cab957edc89f9393bc905b0c8ab9683addd

Observation b8d16741-9771-454f-b974-e8b3bc0b739c · outbound

This paper cites Scaling open-vocabulary object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Scaling open-vocabulary object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.393772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.224103Z digest=sha256:83d6ce684dfe4d658af7ecca40422ab6d3197fe09d61766e06305a0dd632dd6d

Observation 9970ec12-3997-4701-9dba-82053fae0d7e · outbound

This paper cites Simple Open-Vocabulary Object Detection with Vision Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.316835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.316835Z digest=sha256:a8ee34b699c946cc69367924fd6618c271d5213a0715cacf686530d7ee4cb9c6

Observation a378cbee-6979-48cf-a0c9-9d1ed0c2c8aa · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.441530Z digest=sha256:ec51452c9fb0ec2b5fc378e0b3d44df34b1c4bc908917fcba618d2d7aed5b74c

Observation 35e4c707-9e52-4348-b0d1-5155491146cf · outbound

This paper cites GPT-4 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GPT-4 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.619085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.619085Z digest=sha256:3ef8cbc90342628f32431c0da7d497a6f2a8a9284060532b73c0107acda5c65b

Observation 84c356ae-bf5c-4a8f-b989-9ee6242478a9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models DINOv2: Learning Robust Visual Features without Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.723845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.723845Z digest=sha256:262cff0dc5d24ceceeeef2e6757032f7dc18679130f928e9f4a580e727043a6d

Observation 676cee09-5884-4021-9fad-a5e72b17e219 · outbound

This paper cites Grounded Text-to-Image Synthesis with Attention Refocusing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded Text-to-Image Synthesis with Attention Refocusing

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.795064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.795064Z digest=sha256:a7f6118552859b2983ca78e4da0068f6f98e2e3bcf4d36daf4f1fc83bcad5675

Observation 721246fa-07dd-4586-a22c-97860a7034c4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.869024Z digest=sha256:23d8352692a44148aac3f623c66711f22052cc095694f94d68033c1eab8ffb91

Observation 3a181d03-2143-46e2-898a-5e4a2f6f696a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.192900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.965615Z digest=sha256:1e5e5a09a4e7aadf6100663e0b312e2da930fee79b2a8f8f687a2668ec21ce95

Observation e6aace94-dd9d-4064-93eb-a33d6b4d5ae1 · outbound

This paper cites Zero-shot text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Zero-shot text-to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.928478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.023411Z digest=sha256:e14deb800f5db617fa6211261ebdf03cf897cefd763b2de269db21c69a47677c

Observation 26b51eb6-89dd-43c0-83c6-b8ba3c63a021 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.058839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.058839Z digest=sha256:17932d9f0e061a258668d3e450e65a727acbab1e72e3c7875958e8e8c3d6d8be

Observation f2edd81a-8cec-4252-a975-0b63c4d33ac0 · outbound

This paper cites Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.717211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.114052Z digest=sha256:6ae39fc877ebb9c711cd806f4cb504d99dda82ea4e127ec06324e4ed87ec1c94

Observation 367c2b0b-cd2b-4eb0-9ee8-07b77b7c9b72 · outbound

This paper cites Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.511252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.170409Z digest=sha256:f54a348a1dd281f1e434a2467c9ccb6cf2904acd2544d498e7ddbc4c7a471105

Observation 112b0fc5-524c-4449-a5eb-e9e146ef55fd · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.243827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.243827Z digest=sha256:0c7effa65b7370389a4d4c1e4f3910eca5762a47010f329ee9c2cdd2030db7b5

Observation 9bf64bec-9478-44dd-9edb-b24615ee4229 · outbound

This paper cites Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.335950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.338961Z digest=sha256:603981bbedc8402483273f64055811b6012bea7f53946964f9e1747ff48b0426

Observation dc1a2e9c-df3d-4cad-9d28-43c19116b4c3 · outbound

This paper cites Some aspects of the sequential design of experiments.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Some aspects of the sequential design of experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.381683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.381683Z digest=sha256:7974d1b4c4ca60104e4ee51a8f02c3f18cc6c81b72b2fa2f9113ad162f481f40

Observation dce2eed6-d5f8-4927-b29d-0b19ead8005f · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models High-resolution image synthesis with latent diffusion models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.449019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.449019Z digest=sha256:05a8ccbd717da2adb9ae70795f716fc12ffc3fa70b1befdb37863e058ec4fb80

Observation e97464c5-dbec-4752-bd3f-816cbb486125 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.178187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.503848Z digest=sha256:92509abc8bce5265ae088463f3dd8d5fe9fe0bd58abbb5791d5b0a7ede320891

Observation 1878b31b-1707-4e95-b649-8c5be5da182b · outbound

This paper cites Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.979688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.540544Z digest=sha256:0f35d9ba94816f679e8c62fa793c1145f363aefa0643d2bd62b3166ca5442af8

Observation e6c76451-9de6-4185-a098-92aa33429390 · outbound

This paper cites InstanceDiffusion: Instance-level Control for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models InstanceDiffusion: Instance-level Control for Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.611116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.611116Z digest=sha256:83631649a55d917981ce8cc57d63010233e8095720b6f399e541621058531e1c

Observation 6eea2c92-3000-4d91-a3bb-cca1487973e0 · outbound

This paper cites Wolfe and Robert V.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Wolfe and Robert V

Reference 71

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:48:58.107706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.663174Z digest=sha256:ff9294ea04bb547d771f55ad6e16cdd77232ef63c62ea48129bfc2df5674e6f7

Observation aa2d77d7-1af9-455d-87fa-796a078c9105 · outbound

This paper cites R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.745122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.745122Z digest=sha256:7210d3159218a80c022713c6c49b9f7ebc7504da37f505aa1a5c94a96838889f

Observation 1d63d2f5-54b2-4226-bed7-c60f7208a5cd · outbound

This paper cites BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.931564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.816954Z digest=sha256:35c320998a2eb30a5842278678667fe53833865f1c903101e2de2d02a112b514

Observation 8ba51f86-cc4a-4f76-a2b8-a0ccd79790fb · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.834159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.852398Z digest=sha256:cbc1dbd6dc7874e13a03e713e6261fe6567fd76ce35a1cda1d092dcc8ace3ebe

Observation ff379478-dc49-45a9-9c2d-0c34e5ca44a1 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.923723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.923723Z digest=sha256:def8c5325f3c8a0d47ef985a5195f9d74d4d6e4a6115a97d6cd90a7dd385a89e

Observation 855cada1-bf5a-43e5-8e87-dbdad5b0b6fc · outbound

This paper cites Reco: Region-controlled text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Reco: Region-controlled text-to-image generation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.681822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.998259Z digest=sha256:5547e587ea13d489d0230b287dacd7a158211eac0819f7bb4a5306057f80f129

Observation 3c0c1880-0a35-4d0d-a0ed-89c56d17e0ef · outbound

This paper cites Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.035500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.035500Z digest=sha256:a6725bbc4f953821869bb994f49aae43b4faa7006153c3e557c3937a9e6fd647

Observation 0ecde0f2-5b63-4da1-a25c-e919df18b343 · outbound

This paper cites Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.092979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.092979Z digest=sha256:f94bbf4a884ae6dd19919785c0eb5a1e024cf8b0b21e1dbf845c8bc9eb0472a7

Observation 6c6e4dad-b9be-4ec0-9c4e-a2e4a8147b8d · outbound

This paper cites Multi-grained vision language pre-training: Aligning texts with visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Multi-grained vision language pre-training: Aligning texts with visual concepts

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.553613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.225701Z digest=sha256:9cb8ade57697b702b3bd49987ecf75774a4f631e0fde4adef487ebff23346ad5

Observation dbe8a91e-ea7b-491b-89ed-248bec555a13 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Adding Conditional Control to Text-to-Image Diffusion Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.313620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.313620Z digest=sha256:22c5da61cab3e011784a942a6cce1e84b1f68d782d873e6a1fdc28a9c01f372f

Observation 2075568a-7757-4427-bc8b-f412c48e1914 · outbound

This paper cites Realcompo: Balancing realism and compositionality improves text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Realcompo: Balancing realism and compositionality improves text-to-image diffusion models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.423169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.350636Z digest=sha256:a6dd6020f3eaaa1fc893f7b68d1c6137120216612f0542b4963e7874e20e9eca

Observation 640cee54-8bbc-4849-9c19-455d032b15ab · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.426716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.426716Z digest=sha256:35c151e92cde9e9e732aaf3d2e15f0d46a56bf1c6400a25321f79cb3214306de

Observation ac7f8dd9-14e8-4c32-a600-b96082be45b4 · outbound

This paper cites LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.770163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.484646Z digest=sha256:f630c2b57c15549a182fda3c807a19f171f3ccfcdc5d132a3f21c1109676d950

Observation 7924c320-10cd-4992-9f9a-459d6d3807d2 · outbound

This paper cites a henb \.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models a henb \

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.258838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.523921Z digest=sha256:1853c8fe9e8c6959b8f062f71abbff3205aef2b38a71401f3e9c194e9caefd12

Observation 092761a5-b219-4d41-b024-75209b95bd11 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.591417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.591417Z digest=sha256:f3e554239a84ad2a280d20db407c881accb6d371c6dff97e7fa6abd663da0cd3

Pith citing papers

No inbound Pith citation observations are available.