Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Spatial Relationships in Text-to-Image Generation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2212.10015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.10015 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:19:10.904066Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:29.849367Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5692ab96-29d4-4585-ae9c-5f51b1e7b0b4 · inbound

Exploring Spatial Language Grounding Through Referring Expressions cites this paper.

Exploring Spatial Language Grounding Through Referring Expressions Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 166

Resolution
unresolved
no resolver link, observed 2026-08-09T11:19:10.904066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:19:10.904066Z digest=sha256:f67d6598d1deb098fe7467ba673fd9f21f64e56c050479ec54a1df8e324dc3ba

Observation c3e72552-b9b7-4add-a107-f91363c9ae06 · inbound

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? cites this paper.

Can Diffusion Models Learn Hidden Inter-Feature Rules Behind Images? Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T21:50:24.505718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:50:24.505718Z digest=sha256:f2221d618e485a193339fa04e7081e498998188d02a8acfb819404475d01c22a

Observation c9468441-2afa-4be0-b5c9-cf182aa22214 · inbound

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation cites this paper.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.923026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.923026Z digest=sha256:75e93e849937b0b73ca5d0f0b115d1483dbb6f770c87a37913af0eebd8b33d2c

Observation fbaee488-b011-4168-9d0d-9bf5f8ac655b · inbound

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models cites this paper.

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:41.429534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:41.429534Z digest=sha256:2520b37f0d929f86503dd223e58e0e5ff4c5449ef3746a48d217594ad3785fb9

Observation d00ac191-626e-49e9-bd97-7f9e6761377d · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:46.511922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:46.511922Z digest=sha256:15814c803391a054bec91ce22c21d220fd91ac62bb0e9cddaf90e5c037d20732

Observation 54e5f73d-cf88-4460-bc27-fa90af3275fe · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:52.391265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:52.391265Z digest=sha256:e6fe1da18395dbe39ef3f453126a2cf348ada11d7141b8a9b46c70e818818d05

Observation 8dda9b9c-2267-4826-8203-8d08ffa3620f · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:42.353509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:42.353509Z digest=sha256:98674a4b635a838cadf40173ddedb6545987d6c94af7668e46a11b6279bc4893

Observation 9d0e6edc-0bf6-4db3-bc1a-0f0fca961dd1 · inbound

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas cites this paper.

Canvas3D: Empowering Precise Spatial Control for Image Generation with Constraints from a 3D Virtual Canvas Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:23:09.608412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:23:09.608412Z digest=sha256:2ee87065854d7cf8f7b68eb4647e667e07d0151ac03cf7bd8a547c3a72a538e9

Observation a5df3ae9-484b-47a3-9946-033742454da1 · inbound

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation cites this paper.

Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:25.229741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:25.229741Z digest=sha256:9c3bce0e05223c9496c8d6c9306f7e3bfce49e3fc204f1c196b043876c98a53e

Observation 4cba0c9c-9294-49d7-96a7-d81c09a42921 · inbound

CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration cites this paper.

CARINOX: Inference-time Scaling with Category-Aware Reward-based Initial Noise Optimization and Exploration Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:26:33.618048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T15:25:39.116881Z digest=sha256:9e74baef474b0870af5e80d80b6505df9ae9e8019b3639ec63e8a0b243ab1e85

Observation 4fbf225c-09f3-4334-a9f3-ca043d6b7fa5 · inbound

AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs cites this paper.

AMVICC: A Novel Benchmark for Cross-Modal Failure Mode Profiling for VLMs and IGMs Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T09:35:27.378827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T09:35:27.378827Z digest=sha256:1bfeeaf68c5359b440e32663ec99f0db83a198cdbea0b2b9590961da63e7f8d0

Observation da2973fe-c940-41b8-a44c-8499cff71122 · inbound

Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk? cites this paper.

Exploring the AI Obedience: Why is Generating a Pure Color Image Harder than CyberPunk? Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:16:31.387201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:15:11.575505Z digest=sha256:b23ce603f36e3abfc5eda963e47cd41d35defc713951c4be7f2c08846b02c5e3

Observation 63def0fd-170a-45bf-a37f-3b9949a7e061 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.397746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:2e8bb8e73e56a3733cd1adf3e8d8d782be174922342730cbb61662eaf93eddbb

Observation 9a3cbfd8-8b21-4830-ad48-68a08784bdde · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:bd61749283d2c057ccc07c1d844855997515c87c2e76dd52268210135613ef90

Observation c5b2d4fb-eb60-40ee-838d-538e2cb0d628 · inbound

EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation cites this paper.

EPIC: Efficient Predicate-Guided Inference-Time Control for Compositional Text-to-Image Generation Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:12:22.668531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:10:52.567634Z digest=sha256:119dfb99e8b335f26b53b4e70fce719e141096d33eb6a97daa1e91d304df5c2a

Observation 300a6ef3-f683-48d3-a286-9aa5671031eb · inbound

Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection cites this paper.

Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:28:05.540179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T06:25:32.694239Z digest=sha256:a6a9a9a981296caa1be53150beab4032f216dd463712f532180a2aacff1e6939

Observation db924f41-522a-4d10-aa6d-5074d461b1f2 · inbound

A Systematic Study of Behavioral Cloning for Scientific Data Annotation cites this paper.

A Systematic Study of Behavioral Cloning for Scientific Data Annotation Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 268

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T16:23:39.188094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T16:23:08.402194Z digest=sha256:ec4b5b26c0230d9f2fe88a5660b67bb484ce6194870238d2af66ea72183c2533

Observation b3361d10-8a55-4354-b500-ec4d341334e8 · inbound

Compositionality Emerges in a Narrow Depth-Connectivity Regime: Architecture Constraints and Solution Manifolds cites this paper.

Compositionality Emerges in a Narrow Depth-Connectivity Regime: Architecture Constraints and Solution Manifolds Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.854512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:22:56.676469Z digest=sha256:ef67cf3be0607bd21b1fc9740092cb2e17ee2a22b583dfe0e588c59764dd6d2b

Observation 3e74535b-719e-49d1-a2a2-c3d88c0dfca8 · inbound

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion cites this paper.

ELDiff: When Evidential Learning Meets Text-to-Image Diffusion Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:29.850942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:40:14.587143Z digest=sha256:726338836a83814842e8fd4dda54347e3d13c76b4b9a3b0e52aa581a30ad59d9

Observation 36b8fd0f-250a-427f-9c62-25dba55a608d · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:1da77fea7c11c7d125be7e5e7241c089dbeb215446646fba5bc2e7e1cdbd1c4a

Observation d397bc5e-c954-44b6-a872-c5f290657a1a · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.518805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.518805Z digest=sha256:430a7c7032b73f4189823d5f62dbb8b2569af20a8ce7bca00ce924bd44cff342