Pith. sign in

Paper Citation Record · LEDGER

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

As of 9 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2506.23418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23418 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:48:57.591417Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact6
  • verified fuzzy27
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b45c9dc-5db5-4982-9ef1-e50edfe83359 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.808315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.808315Z digest=sha256:b4ef243ba43566a8036f56868c9fd1423d8be0aed3138e206b158cb8333b0539

Observation 3f71b3e6-8ec6-463b-bbec-afd2add3b53b · outbound

This paper cites A-star: Test-time attention segregation and retention for text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A-star: Test-time attention segregation and retention for text-to-image synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.867916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.867916Z digest=sha256:73db298f735ca4b6f48c2e07cfafb4fca6d30fb76e418c261145fe81d9bcbf62

Observation 07799072-ad50-41b4-b658-3d6ceaf1db9d · outbound

This paper cites Kandinsky 3.0 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Kandinsky 3.0 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:50.976891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:50.976891Z digest=sha256:f08bbc8a790d01a0939d43f21834e7790ed3db2bec40e90637d8e129ffeefc59

Observation f2a69bfd-0dca-4dd9-8e21-23820b2719b6 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Finite-time analysis of the multiarmed bandit problem

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.033654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.033654Z digest=sha256:2d6a846a82222feedb7c1ae17039930aa720c4afa65bf441f4e876db30ac3c51

Observation 6a8c774a-95af-48c2-849f-65c0d9e7309f · outbound

This paper cites Spatext: Spatio-textual representation for controllable image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Spatext: Spatio-textual representation for controllable image generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.080754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.080754Z digest=sha256:e0a199c7c57846d84d91bc1bc104b4ff8ef9f415c3065c4625c137b0c17a1b95

Observation 4b624120-e17e-4a45-9d74-32299bfb1238 · outbound

This paper cites Cc3d: Layout-conditioned generation of compositional 3d scenes.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Cc3d: Layout-conditioned generation of compositional 3d scenes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.127391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.127391Z digest=sha256:430e04a8edf62def7cb88003bce40625a02e58f552ea0889485960186eb3a71f

Observation 2e40a2e8-8777-44e5-873e-7d0bd3629d80 · outbound

This paper cites Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hrs-bench: Holistic, reliable and scalable benchmark for text-to-image models, 2023

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.167595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.167595Z digest=sha256:619df32af22d0a8a4a4c7f026885b978668e21c6143d3215df0d00131d107338

Observation c3b64dc5-5646-4eef-900f-5df72f601870 · outbound

This paper cites MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.252650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.252650Z digest=sha256:a316564326d897d38bece0fcf393b51cdfd5575e7123a71b9c66c32b7acfdfd4

Observation 41a8f068-b396-4a9b-b4e4-133192444c86 · outbound

This paper cites Improving image generation with better captions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving image generation with better captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.298620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.298620Z digest=sha256:ad78bcb66b00a6f268f4b8f0b9dd7e8ee54c672fedd2b3dfdcc1d3128218c980

Observation 4fd1b3e2-497d-4ced-ae61-9af9773af035 · outbound

This paper cites Make It Count: Text-to-Image Generation with an Accurate Number of Objects.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make It Count: Text-to-Image Generation with an Accurate Number of Objects

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.348576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.348576Z digest=sha256:f9d0f961990adc16a1d5a252f0f9f4829ccd2247f5df0fd09f939b2908aff291

Observation 26a059a3-eb12-4ace-b2ed-612a9e13c0e9 · outbound

This paper cites Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Conceptual 12M : Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.878208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.432547Z digest=sha256:2ee9559f07c240e4c8cd3776fa0ea147fc4d31678444b0851908c96cd43ee417

Observation b3a0d950-88b9-4ea0-9329-aa5f0ab4ffb8 · outbound

This paper cites Getting it Right: Improving Spatial Consistency in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Getting it Right: Improving Spatial Consistency in Text-to-Image Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.518955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.518955Z digest=sha256:bf63aadcc9431e51f922f62c3371a8ca7073586733e9ee2c410a659220d4f641

Observation 4bfb679c-f849-4704-bf91-e714a875ec45 · outbound

This paper cites Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Attend-and-excite: Attention-based semantic guidance for text-to-image diffusion models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.719627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.601008Z digest=sha256:9e9decc09537e4d2813de53bf64cbf3d0f206b0f5b6649c5bd47b9b8306db30e

Observation e1d998d1-f294-4f63-905f-d4b06ff9f1f6 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.657856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.657856Z digest=sha256:e8d26c09f0c9cb63bbefac000d5c6eac2ffd83815aad636e41633561fa35a574

Observation aaae8dd0-4517-47b7-bf31-a63a6426e88d · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.736340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.736340Z digest=sha256:f71d4e1dafd0c1059f26794b3ad67f37633065aaef1fd597aa097ef585f9179b

Observation 307c10fb-40cd-45ef-8cf9-05c055f67e7e · outbound

This paper cites Training-free layout control with cross-attention guidance.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free layout control with cross-attention guidance

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.564234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.783311Z digest=sha256:dd58b4c333d5538a0823558df721c1c4963f4249b63ecc9d7c9e06642c63a6a5

Observation f60b0b08-79b3-4092-81a7-67085a3a6130 · outbound

This paper cites LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffuse: Adapting Foundational Diffusion Models for Layout-to-Image Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.832817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.832817Z digest=sha256:9c1632fb0fd081a2bce278ce70f834092e6e7cf3a1bdeb0fda6f0b02f0fe41ab

Observation 4a510a3f-4052-400d-b187-98e87f161fd8 · outbound

This paper cites Dall·e mini, 7 2021.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dall·e mini, 7 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.406246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:51.883431Z digest=sha256:c2d463d782e37444fc61a162c36e9cfc0277ded38a3ee5cb7fd23c476b0d5052

Observation 9765d374-7cba-47c7-84b2-6d72e643d53c · outbound

This paper cites CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:51.971405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:51.971405Z digest=sha256:5c020011c2a1d476e9b844aeb64071d6ad6d187e3273e99239395865681eba95

Observation c7458bd6-c8d0-4088-bc99-df1b62ad21c2 · outbound

This paper cites Training-free structured diffusion guidance for compositional text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Training-free structured diffusion guidance for compositional text-to-image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.187547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:52.059608Z digest=sha256:cb539fe2ca16059c05976f36e7dedc55ddf2235eaf73019d1fdf525076e28077

Observation 6e88afe7-aa0c-466e-b77f-134e9213febc · outbound

This paper cites Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Make-A-Scene: Scene-Based Text-to-Image Generation with Human Priors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.182184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.182184Z digest=sha256:e3718f21542577e0b66356dcead71205cf379d2d51c35fea6916fca1f141c96d

Observation f2b3f176-fa36-40c8-99f9-e08ba899c89e · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.298002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.298002Z digest=sha256:5a19f1df515651501e302e5795c2d3645689d11b0c60d835494018109b5a8459

Observation 54e5f73d-cf88-4460-bc27-fa90af3275fe · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.391265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.391265Z digest=sha256:e6fe1da18395dbe39ef3f453126a2cf348ada11d7141b8a9b46c70e818818d05

Observation a97cb007-cd59-49c9-8682-85279d24b011 · outbound

This paper cites Gradient Guidance for Diffusion Models: An Optimization Perspective.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Gradient Guidance for Diffusion Models: An Optimization Perspective

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.464757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.464757Z digest=sha256:b13e633df745515c5ed878fb61109ded5938f1121b94438bb164aa136b013b18

Observation 56e3b960-01db-4269-88bc-a7be5e2209ab · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.605567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.605567Z digest=sha256:4b3652519c101198ecdde724e7394daef919d78380c9fb6ee629e79945f898ab

Observation a5d12599-6650-44dc-a625-5dbe2c18cc58 · outbound

This paper cites GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.737177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.737177Z digest=sha256:c755929636692690ff72846a675e966da6491edebd63993ac4168c5465f685cf

Observation a7b727f7-47fc-4c9b-83c7-40f10f905237 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.842219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.842219Z digest=sha256:874293dcd73dec57a1a503e574849c75029032af5d68f3195444d54124691724

Observation b6111a05-8052-4a32-8d30-9ecd28e2c8e2 · outbound

This paper cites Probability inequalities for sums of bounded random variables.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Probability inequalities for sums of bounded random variables

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.928167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.928167Z digest=sha256:3e8b4dded6cec16c31916062d0b943bbcc0c0d13f0a066adfbf74e24b4b44f5d

Observation f00c05be-3b62-4de4-a843-a95554a680d8 · outbound

This paper cites Token merging for training-free semantic binding in text-to-image synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Token merging for training-free semantic binding in text-to-image synthesis

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:03.036753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.000087Z digest=sha256:7272dd9f24530e53a27d3ef95fd131c491427658f15dd26beed41a72ba395f61

Observation 79b60aa4-8021-40ee-a135-487356241c7b · outbound

This paper cites A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models A Multi-Armed Bandit Approach to Online Selection and Evaluation of Generative Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.751381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.058729Z digest=sha256:ee2d2ee49edf4c943b3b0bbd45af0f2cf1274b104dc1907c49b8206e445ba38a

Observation 68e853c2-b619-4f9d-95ab-a9eda965bb51 · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.829715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.116358Z digest=sha256:895d219a3ec100b00fb7979dff34a80596c841c2609393c2b2bd5196ba847cad

Observation 8206daf3-8e7a-48dc-9aa8-fead9a8a5ba7 · outbound

This paper cites An information-theoretic evaluation of generative models in learning multi-modal distributions.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models An information-theoretic evaluation of generative models in learning multi-modal distributions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.654824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.228368Z digest=sha256:cef420f3f3f88c0d46badca0c76b24327971cb58df427088c62d1fade30ce5da

Observation 07d48034-0798-4fe7-9a7d-d13cc77c49dc · outbound

This paper cites Rethinking FID: Towards a Better Evaluation Metric for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Rethinking FID: Towards a Better Evaluation Metric for Image Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.316391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.316391Z digest=sha256:3a6f8905d04f1bfef00e86fa6d77868f070d30252e8775db90629e6896b16ba7

Observation f2b20c63-ef67-4c6f-8e17-fdad4f6363f6 · outbound

This paper cites What's ``up'' with vision-language models? investigating their struggle with spatial reasoning.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models What's ``up'' with vision-language models? investigating their struggle with spatial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.501071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.352085Z digest=sha256:e952eb53a9095835733ab0f6e3b5534ca95729b39a055af54c07634720156f81

Observation 568f43b0-6aeb-4c70-a4c0-43df7e04bb1f · outbound

This paper cites If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models If at First You Don't Succeed, Try, Try Again: Faithful Diffusion-based Text-to-Image Generation by Selection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.437835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.437835Z digest=sha256:1bc4af525215687aa02041b9801016151684e5fdea637238f5e90ad1867fcaef

Observation 71edc826-1f08-423d-bc26-8dfe123c5cfb · outbound

This paper cites BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BeyondScene: Higher-Resolution Human-Centric Scene Generation With Pretrained Diffusion

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.598720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.533737Z digest=sha256:4ecabaca3525b61a2206382191ab1c897efd2829acaf3bef05a3d6ba4f3bb112

Observation 65b24792-db47-47ac-a5cb-977279bd3622 · outbound

This paper cites Dense text-to-image generation with attention modulation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dense text-to-image generation with attention modulation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.338882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.586132Z digest=sha256:2a6c446ceb1edff420166a3323dfac7a9c57463c3b77175da35f98d4cff886d1

Observation 60788508-c8c4-4c60-b0e9-270c434757eb · outbound

This paper cites Segment Anything.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Segment Anything

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.722447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.722447Z digest=sha256:0b8bbf1bbf7d5c138f5bcdfb758aa62363b82179abd6102fb5d0e45eee38c158

Observation 632e5ff8-76bb-4f4e-89de-f5cc318033e6 · outbound

This paper cites Improved Precision and Recall Metric for Assessing Generative Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improved Precision and Recall Metric for Assessing Generative Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:53.852853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:53.852853Z digest=sha256:22a5b11200f7fdba1240e7b416c572de79c54066fbc903f8769207e0b26a2540

Observation 938ff7f5-8e90-47a5-b9f5-08632237ec65 · outbound

This paper cites Laion-coco 600m.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Laion-coco 600m

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:02.158280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:53.993200Z digest=sha256:77aff18fe7a1f2ff95f6359be2ecd797e63e50f07c18a9f9a20490a4519c4a21

Observation cc4b0226-9d4c-407a-8619-11b898df9d86 · outbound

This paper cites GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:48:58.472999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.063558Z digest=sha256:77141688464cbc9b08eba38e3f5845d9f26f04637bee4adb6d96f88fb4f22b44

Observation 83dd898d-37de-4d8f-b1b2-ea746ca26f09 · outbound

This paper cites GLIGEN: Open-Set Grounded Text-to-Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIGEN: Open-Set Grounded Text-to-Image Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.226577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.226577Z digest=sha256:e625beb6a68dd57e7d05469ed8496a392f63fa97b9974ed11960941737641628

Observation 0981c3cf-42c0-4f5a-ad88-a0fa44bdb9bb · outbound

This paper cites Divide & bind your attention for improved generative semantic nursing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Divide & bind your attention for improved generative semantic nursing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.947376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.383202Z digest=sha256:6e636a2a25e66714aee6fda92d12fb25d41246f550d1ba06d67c2ee731d9d3f0

Observation 163ad252-7a8e-46fe-966d-63458b240b4e · outbound

This paper cites LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.541171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.541171Z digest=sha256:f455fafb6eaaf196c390b848c7bc39d40cc9e978bf486e42a43f77b06cf1486c

Observation 1b6393a3-ac2e-4b9f-b839-c7024428db99 · outbound

This paper cites Microsoft COCO: Common Objects in Context.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Microsoft COCO: Common Objects in Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.670059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.670059Z digest=sha256:ea7d56722d77eedca4b996f81ed43d154927529fdc22670d51596c82fbac7bd3

Observation babd993d-63a6-4fb9-89a9-c3b2f7b3f4cd · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.778142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.778142Z digest=sha256:29c02675a5321ad472d1843ee9db5ec90206524795d283de2a2a9b2a5557f15c

Observation 0c9d4880-0781-4eaa-a364-e680ac9a6c13 · outbound

This paper cites Compositional Visual Generation with Composable Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Compositional Visual Generation with Composable Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.880447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.880447Z digest=sha256:0260b7d66826bb014fc5a7fe8be9968704e76b72fdc89c7151878a5697118de0

Observation ced64cf3-11ee-4234-afc1-7abfb9ed26c6 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.783156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:54.923631Z digest=sha256:f95f75412dffcde6f6748b9a561d694ce89d45457d15358de542cb4dcee0cd5c

Observation 9fd69e75-d069-47bf-9d8d-e0c8a10bb6cf · outbound

This paper cites Correcting diffusion generation through resampling.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Correcting diffusion generation through resampling

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.584642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.042044Z digest=sha256:891894527bdcd03f6d0dca59ce0b35f9c927a6faf194661e115de3817ce55be1

Observation 9090d239-61cb-4e85-98b5-3771c792f67e · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.131271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.131271Z digest=sha256:2330a7a0b6f7f35cea4b84759eb917ebdc4c78d1f021b7f6029029a0b62b991c

Observation 1a67cb10-e9b4-462b-8869-d8d9ce9ea2b3 · outbound

This paper cites Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Diffusion Beats Autoregressive: An Evaluation of Compositional Generation in Text-to-Image Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:58.305767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.186979Z digest=sha256:616ef7f7df75e054fd9b7f57d0347806064b43c9382d78d9637ee89721b15aef

Observation b8d16741-9771-454f-b974-e8b3bc0b739c · outbound

This paper cites Scaling open-vocabulary object detection.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Scaling open-vocabulary object detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.393772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.224103Z digest=sha256:283abce3414e882df804f78e11a76488f6e75b55e6cec4dc3f308cd8944e19c1

Observation 9970ec12-3997-4701-9dba-82053fae0d7e · outbound

This paper cites Simple Open-Vocabulary Object Detection with Vision Transformers.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Simple Open-Vocabulary Object Detection with Vision Transformers

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.316835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.316835Z digest=sha256:faaacbe813d05b4f5b1453b7e8d674638796d6e20aa384c4b5ee39eb384bee2a

Observation a378cbee-6979-48cf-a0c9-9d1ed0c2c8aa · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.441530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.441530Z digest=sha256:0bcf2f3c65ffea82857e9809a81f59d7b36226a2d7b92990ad4aafcc12a2a311

Observation 35e4c707-9e52-4348-b0d1-5155491146cf · outbound

This paper cites GPT-4 Technical Report.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models GPT-4 Technical Report

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.619085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.619085Z digest=sha256:5be7191ddbe7b98748156f39dde288ed202f029b5ecbf99fc47c602656b7ec02

Observation 84c356ae-bf5c-4a8f-b989-9ee6242478a9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models DINOv2: Learning Robust Visual Features without Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.723845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.723845Z digest=sha256:790e36f3cb57183fc47051c217a1ed7f2bcb2e81b51b80d36f60eaa0c7672a5d

Observation 676cee09-5884-4021-9fad-a5e72b17e219 · outbound

This paper cites Grounded Text-to-Image Synthesis with Attention Refocusing.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded Text-to-Image Synthesis with Attention Refocusing

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.795064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.795064Z digest=sha256:f0c39114b1de6023268cb2588bf2b5bf22b8cc786a429c1ea710c8f7bd6b3fa4

Observation 721246fa-07dd-4586-a22c-97860a7034c4 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:55.869024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:55.869024Z digest=sha256:883ea5fdbaf1c9466e67c1e797397df83af4a57907e652b386928b8b6a7a6371

Observation 3a181d03-2143-46e2-898a-5e4a2f6f696a · outbound

This paper cites Learning transferable visual models from natural language supervision.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:01.192900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:55.965615Z digest=sha256:31f5c3def990874c0b06ec50c8f9cf9081cb4f627b95fb2cae3156dbc9f0b11b

Observation e6aace94-dd9d-4064-93eb-a33d6b4d5ae1 · outbound

This paper cites Zero-shot text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Zero-shot text-to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.928478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.023411Z digest=sha256:d3aea9d3af1dcf54c81de9999c751c3d4da5bf8986a153aae18d4e444a5c21f8

Observation 26b51eb6-89dd-43c0-83c6-b8ba3c63a021 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.058839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.058839Z digest=sha256:93e068546b74947c8e70b0c285817866a046c7b67503c4a525b909c5b29fcad1

Observation f2edd81a-8cec-4252-a975-0b63c4d33ac0 · outbound

This paper cites Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Linguistic binding in diffusion models: Enhancing attribute correspondence through attention map alignment

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.717211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.114052Z digest=sha256:c3790f00b27b16320a58b7ed10291107fbacf9cb5d5a7e320d5967052c80e692

Observation 367c2b0b-cd2b-4eb0-9ee8-07b77b7c9b72 · outbound

This paper cites Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Arkhipkin, Igor Pavlov, Ilya Ryabov, Angelina Kuts, Alexander Panchenko, Andrey Kuznetsov, and Denis Dimitrov

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.511252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.170409Z digest=sha256:6418157b6a6292f0b271bd67999cfc0d3bbde933642f603e596bc63390317477

Observation 112b0fc5-524c-4449-a5eb-e9e146ef55fd · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.243827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.243827Z digest=sha256:a2211f748ea410bd656e2ecc5c3a02100bc2a87aad3bf3da322c6420807b3274

Observation 9bf64bec-9478-44dd-9edb-b24615ee4229 · outbound

This paper cites Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Be more diverse than the most diverse: Optimal mixtures of generative models via mixture- UCB bandit algorithms

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.335950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.338961Z digest=sha256:2c8b4df70e148543829ccea8e0f2c0f955d7f68c00f1464233c8aea1b1310672

Observation dc1a2e9c-df3d-4cad-9d28-43c19116b4c3 · outbound

This paper cites Some aspects of the sequential design of experiments.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Some aspects of the sequential design of experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.381683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.381683Z digest=sha256:8e8e3ba912877310b195a682a57909ce2d5c7eb3ed037d19b988217ca380cc14

Observation dce2eed6-d5f8-4927-b29d-0b19ead8005f · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models High-resolution image synthesis with latent diffusion models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.449019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.449019Z digest=sha256:f8945514363848e268a811770680fd3cf58b8a4620493f6f4ebc94c3108f9a62

Observation e97464c5-dbec-4752-bd3f-816cbb486125 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:49:00.178187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.503848Z digest=sha256:70253de4d31893bcb5584ea5e2cae00bbeb16e470088f91a0883e43278f284d1

Observation 1878b31b-1707-4e95-b649-8c5be5da182b · outbound

This paper cites Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Predicated diffusion: Predicate logic-based attention guidance for text-to-image diffusion models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.979688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.540544Z digest=sha256:851f8ddc82b720f30d594ed3f313cbc8f18ce7aa60261f20001fca3a2252a90c

Observation e6c76451-9de6-4185-a098-92aa33429390 · outbound

This paper cites InstanceDiffusion: Instance-level Control for Image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models InstanceDiffusion: Instance-level Control for Image Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.611116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.611116Z digest=sha256:67a20787099fd0242adb3ae52ecf7922faaba61b54f9bc520c4058efd0ed6e39

Observation 6eea2c92-3000-4d91-a3bb-cca1487973e0 · outbound

This paper cites Wolfe and Robert V.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Wolfe and Robert V

Reference 71

Resolution
verified exact
raw_fallback, observed 2026-08-06T21:48:58.107706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.663174Z digest=sha256:df4fe6afad6a56cfbc3f851efa6c0941e94aef5d1b8e0b9693811f86f5b15520

Observation aa2d77d7-1af9-455d-87fa-796a078c9105 · outbound

This paper cites R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models R&B: Region and Boundary Aware Zero-shot Grounded Text-to-image Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.745122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.745122Z digest=sha256:45de61294c9e20661f02aec0ceca9d37f3f2bdf27df705e8c44850559665ebf3

Observation 1d63d2f5-54b2-4226-bed7-c60f7208a5cd · outbound

This paper cites BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.931564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.816954Z digest=sha256:60369f75e49d5348445f296bacf000f0f73a4e31271c54e4a8ad1aaad492ead9

Observation 8ba51f86-cc4a-4f76-a2b8-a0ccd79790fb · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.834159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.852398Z digest=sha256:803019b1e19fda415176575e802e990d74f60c49a2127684d8623501d46c46cc

Observation ff379478-dc49-45a9-9c2d-0c34e5ca44a1 · outbound

This paper cites Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:56.923723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:56.923723Z digest=sha256:e733f833b6887e9237005fb68cf2304e265466587b6f4a787f516519e4015ad8

Observation 855cada1-bf5a-43e5-8e87-dbdad5b0b6fc · outbound

This paper cites Reco: Region-controlled text-to-image generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Reco: Region-controlled text-to-image generation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.681822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:56.998259Z digest=sha256:0cdb7e517777f5e8a304ab7196bedc058eb16050fd0f1443e897d60ebdfdc1da

Observation 3c0c1880-0a35-4d0d-a0ed-89c56d17e0ef · outbound

This paper cites Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.035500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.035500Z digest=sha256:e571647281f742fc692f41d57853d8202db4d16bc25f296054f0c8bce00b0475

Observation 0ecde0f2-5b63-4da1-a25c-e919df18b343 · outbound

This paper cites Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Improving Compositional Attribute Binding in Text-to-Image Generative Models via Enhanced Text Embeddings

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.092979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.092979Z digest=sha256:5ab054060c7d63d178b2a6c5bfad1eac07091cfa65beea2d689361b6897a1be7

Observation 6c6e4dad-b9be-4ec0-9c4e-a2e4a8147b8d · outbound

This paper cites Multi-grained vision language pre-training: Aligning texts with visual concepts.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Multi-grained vision language pre-training: Aligning texts with visual concepts

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.553613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.225701Z digest=sha256:01add62c9be407faacfb91bf4983dacac6618ce27c9cf1474291b6236d4b721a

Observation dbe8a91e-ea7b-491b-89ed-248bec555a13 · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Adding Conditional Control to Text-to-Image Diffusion Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.313620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.313620Z digest=sha256:f300fe00547d2dcd63dd8d5e6d57ef73839c524fbb508c360725e608ecd9add0

Observation 2075568a-7757-4427-bc8b-f412c48e1914 · outbound

This paper cites Realcompo: Balancing realism and compositionality improves text-to-image diffusion models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Realcompo: Balancing realism and compositionality improves text-to-image diffusion models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.423169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.350636Z digest=sha256:0749efbe81c9b717d6d2bd9975f33de1c14f2bb49a1e431e3ae1edad64f3a22c

Observation 640cee54-8bbc-4849-9c19-455d032b15ab · outbound

This paper cites Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Enhancing Semantic Fidelity in Text-to-Image Synthesis: Attention Regulation in Diffusion Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.426716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.426716Z digest=sha256:7d2b5149ea46b0708c05a7e9e429cd510f5caff9f16a794065f7cbbfefd2c8f4

Observation ac7f8dd9-14e8-4c32-a600-b96082be45b4 · outbound

This paper cites LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models LayoutDiffusion: Controllable Diffusion Model for Layout-to-image Generation

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:48:57.770163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.484646Z digest=sha256:781005ed6092f303ec86d881b4c3b1293f445bf7761983373fd5423a38e1c22a

Observation 7924c320-10cd-4992-9f9a-459d6d3807d2 · outbound

This paper cites a henb \.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models a henb \

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:48:59.258838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T21:48:57.523921Z digest=sha256:b86317435cfb488449833bc5111da9ac007fc160b0edbb27ce3f2083fcdd5718

Observation 092761a5-b219-4d41-b024-75209b95bd11 · outbound

This paper cites write newline.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:57.591417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:57.591417Z digest=sha256:89e02d77be33ccf7bc1746026e9b1574462b6a1b264d0d6055cae32ece135365

Pith citing papers

No inbound Pith citation observations are available.