Pith. sign in

Paper Citation Record · LEDGER

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

As of 10 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2505.18730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18730 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:21.502877Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:37.003251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:15:41.320900Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb155788-5860-4057-8397-0021a4ce904e · outbound

This paper cites Tallyqa: Answering complex counting questions.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Tallyqa: Answering complex counting questions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:13.694400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:13.694400Z digest=sha256:8939d6023d8cc5307dae1a73ce5b269140ae5ff32671f841f31c6cc06d45c40c

Observation b9288481-74b3-4e63-8eb6-e7786d218de3 · outbound

This paper cites Improving image generation with better captions.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Improving image generation with better captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:27.077673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:13.847917Z digest=sha256:3e03e0a9d215d93b5a7d6bb1d60163476297100e558ea95097d19a36598f3b32

Observation b22070d4-e572-4756-8565-068c657e214f · outbound

This paper cites Measuring Progress in Fine-grained Vision-and-Language Understanding.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Measuring Progress in Fine-grained Vision-and-Language Understanding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:29:22.567556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:13.981587Z digest=sha256:7aee3571b39de8245cec6e2071034d50d06c33122e444a102089d44f804aeaf0

Observation 26f778c1-79de-4b56-be96-4f45fd88478d · outbound

This paper cites Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.112716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.112716Z digest=sha256:21e96b84228cea6f5bfdc68582ff17171c5d4be1acaa2347d2b2f20d1291f038

Observation 983e1a34-ec16-4d73-985d-19c899a95292 · outbound

This paper cites Diffusion bridges vector quantized variational autoencoders.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Diffusion bridges vector quantized variational autoencoders

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:29:22.316111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:14.286282Z digest=sha256:0284af49119ce2b030b1f9a76e134b309ad3a61fc10db7e064321a5ad0fee91c

Observation 5ee96bd6-5812-4111-9840-580c0171524a · outbound

This paper cites Cogview: Mastering text-to-image generation via transformers.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Cogview: Mastering text-to-image generation via transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.417588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.417588Z digest=sha256:ab0ef89ffb512a7fe72cacd19577a4b30dab238174b4d0cb3e41c374ddc28091

Observation dedadfc0-9772-4ce7-a207-fa949c6364cb · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.601166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.601166Z digest=sha256:790557052ba1138771739433577c5f9f58734035c33ffb799ebf7bf978a71752

Observation cc46d3a9-9e75-4f0f-83e2-6b198eacd0d1 · outbound

This paper cites Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Commonsense-T2I Challenge: Can Text-to-Image Generation Models Understand Commonsense?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.789184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.789184Z digest=sha256:06110c1021ed1b95bdd96361f8e044e538d77a099e707a3604905ae6fca87311

Observation c9468441-2afa-4be0-b5c9-cf182aa22214 · outbound

This paper cites Benchmarking Spatial Relationships in Text-to-Image Generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Benchmarking Spatial Relationships in Text-to-Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:14.923026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:14.923026Z digest=sha256:75e93e849937b0b73ca5d0f0b115d1483dbb6f770c87a37913af0eebd8b33d2c

Observation 59ccdda0-f430-4c38-9009-11a1a97e28b7 · outbound

This paper cites Generative adversarial networks.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Generative adversarial networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:26.796147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:15.109768Z digest=sha256:5880e8f9c42273e3a3d9956fa6482a437b1c0c87b42f878daf82df3ab3ac0a7e

Observation 147352c5-6109-4beb-b5ee-c404ccb4391a · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:15.301188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:15.301188Z digest=sha256:aa48c4cc393c4fe14ac19defddfacc68deae48d08460429fcce40c171555a3c9

Observation 43fa08c3-38d4-4b2b-9a3c-32c0d08ae7e0 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 30, 2017

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:15.489580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:15.489580Z digest=sha256:c165c40093f0b2e7f0096f2b9295ad7b8adb65c0bc023f63c85088e306039c82

Observation 09441fd1-98fd-44c3-bd7b-73ffe3d9233b · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:15.681993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:15.681993Z digest=sha256:1c82c92e84060dc02df55907c4c5dad05a3ed91a1bbe6397f10215875944c020

Observation 931fe707-d246-42bf-9148-193bf9ab403e · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:15.867570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:15.867570Z digest=sha256:889ddcab85371cb17c7b0ebddfab204d8836a01fc144bc43c59b5c347f85bd7c

Observation 44eaefcc-fe82-4fde-9909-4827c573cea9 · outbound

This paper cites Evaluating numerical reasoning in text-to-image models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Evaluating numerical reasoning in text-to-image models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:26.572628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:15.998258Z digest=sha256:507172511caa836d75b2ba6c49a8467c5d61da7924790b698ff3fa11deab0a20

Observation 1988dbe4-d4fd-455a-bd7e-c2f5b4bcc060 · outbound

This paper cites Diffusionclip: Text-guided diffusion models for robust image manipulation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Diffusionclip: Text-guided diffusion models for robust image manipulation

Reference 16

Resolution
verified exact
raw_fallback, observed 2026-08-07T14:29:22.049661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:16.164128Z digest=sha256:2dd8a85aab32b49500de7d21d541f39293d9304fb91bc3837210520777d4c6f0

Observation f8548f64-4ed7-47d4-8d13-c88d1a28bc68 · outbound

This paper cites Pick-a- pic: An open dataset of user preferences for text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Pick-a- pic: An open dataset of user preferences for text-to-image generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:16.279484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:16.279484Z digest=sha256:8e49768c878e549215eaba11b2de84537102b4e43a3ab15ed90294dd586f8144

Observation 1e5abc39-9021-49d6-8b63-15ce95ed4fee · outbound

This paper cites VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:16.408545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:16.408545Z digest=sha256:c5e47ce67244eb2e3b25937f0e69991a8dd669d79df247f881a13527b1186120

Observation fd29f28e-7ba9-4377-881e-7345fd63229c · outbound

This paper cites Multi-concept customization of text-to-image diffusion.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Multi-concept customization of text-to-image diffusion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:16.572717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:16.572717Z digest=sha256:1c48aee4b44054b2f05083a8d1d2528fab9a9bf2f3e126e6b227b2a4abbe91ee

Observation 49365c10-8081-4ec1-b56c-12845b9c04e1 · outbound

This paper cites Holistic evaluation of text-to-image models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Holistic evaluation of text-to-image models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:26.317345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:16.687718Z digest=sha256:c324a585da3f89ef661c7f72282472346fff7cdded1a56b13c3883a4c8246b10

Observation 1bd714eb-f125-4c00-9927-2646a537b6fc · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:16.838722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:16.838722Z digest=sha256:bad9e58b12e2ac2044f3967dca3f17f7e3aeda9bb6e981ff587ddcc667ec854f

Observation b4ba232f-7829-4ec7-94a6-3846824810e9 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Evaluating and improving compositional text-to-visual generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:26.109268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:16.980088Z digest=sha256:1b4723ff6cadda1cc1102cbce7abb658b24bca9142de4b0b93cbd483094ed4e8

Observation 6dc4ac3c-753d-4ec2-ae99-68fa65054518 · outbound

This paper cites Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Blip-diffusion: Pre-trained subject representation for controllable text-to-image generation and editing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:17.166743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:17.166743Z digest=sha256:560a68ca6e99427a34d768a29c4698e5c3b4e9894368d91c625dc8efe5584530

Observation 798f96f2-cacf-46a2-943a-da52e98e7db8 · outbound

This paper cites Science-t2i: Addressing scientific illusions in image synthesis.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Science-t2i: Addressing scientific illusions in image synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:17.284669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:17.284669Z digest=sha256:d175473c7e7b8b91a8a0c60325f05afadfb8964940e18665fbd75639c33dcb9f

Observation 9f06d282-874c-4a0d-b22c-04d94b4a4482 · outbound

This paper cites Rich human feedback for text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Rich human feedback for text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:25.806499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:17.436476Z digest=sha256:f6eb6173c97a1e418658f94a60b129097a66b138a1f691cce0ac337e2c0df5b4

Observation bfdddb92-6456-4b80-84e6-9c188efbb1ea · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Evaluating text-to-visual generation with image-to-text generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:17.595148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:17.595148Z digest=sha256:f32ed44440c903d42281499c78159783d54e92c725b63339c12bddc22e6ee731

Observation 481a1896-9b42-41f6-a953-80a733cd2052 · outbound

This paper cites Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Llmscore: Unveiling the power of large language models in text-to-image synthesis evaluation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:25.552190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:17.729078Z digest=sha256:5cde2491ebbca5936c81ceacf621a5766549c1b573577833984490238ef4d9dd

Observation 68867e7c-2fb7-40e9-a590-99e9727bb743 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:17.887379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:17.887379Z digest=sha256:de492f31d334da8c679c155ed2e8ce415c06d84a311f6369f1f152ac04b127ae

Observation d461ebed-ec12-46be-b344-a95a16f4f7af · outbound

This paper cites PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:18.079146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:18.079146Z digest=sha256:f458f5e19346e5fc6323a85ceb08e6257dfcb9c26cdc6f6b8df5c6ac00876899

Observation 3964ce9f-9f17-4662-ace8-aa6ea3b9e62b · outbound

This paper cites Midjourney version 6, 2024.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Midjourney version 6, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:25.220668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:18.238625Z digest=sha256:f31db6ef6ce00703e8e36230da7a8acfc4e6e906996801a4ac1d981e8f368228

Observation eae11b18-8e14-493c-8ac3-d3666dbe26ce · outbound

This paper cites Addendum to gpt-4o system card: Native image generation, 2025.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Addendum to gpt-4o system card: Native image generation, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:24.828748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:18.374045Z digest=sha256:9d6e063b5d4685b8892830580e82a299374a49b11839ab5a746d6923ab988abf

Observation b1fa7751-f450-4eb1-9bf9-585305419fb4 · outbound

This paper cites Toward verifiable and reproducible human evaluation for text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Toward verifiable and reproducible human evaluation for text-to-image generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:24.454859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:18.523882Z digest=sha256:83e5d3ad812e61d76d372931c03a1bc52bf71c4adb1ce55145757d8065e44047

Observation 37a47874-8eea-404d-a9a3-198a3b995e48 · outbound

This paper cites Teaching clip to count to ten.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Teaching clip to count to ten

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:24.122426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:18.652281Z digest=sha256:559fbb150b915b758a64793a0ecba88c62a9777e1f85cdf4d896ea47d5d28c61

Observation c9382c27-0787-40e4-a11d-9b9e6779c11c · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:18.821920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:18.821920Z digest=sha256:5f2be90da5de6e23d79d47312f163a9f03b444601a0f52f0615d9b58aebb2213

Observation 7fb2f27d-6d47-4896-9d19-36e92326f8d1 · outbound

This paper cites Zero-shot text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Zero-shot text-to-image generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:18.979999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:18.979999Z digest=sha256:7dcea4d8d6ad5420a0f85e689e058feec547b8cb263603c1f4131f6d4fe931ea

Observation d76aa6c1-2180-4ba8-a0ad-e45426ed8678 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation High-resolution image synthesis with latent diffusion models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:19.152449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:19.152449Z digest=sha256:4d45342a5dc4281059ad4b4f5971b232c9aab35f2b42fa53618a431865a30575

Observation f54582fd-e8b0-48f6-9018-350a703b5364 · outbound

This paper cites Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Dream- booth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:23.828286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:19.316019Z digest=sha256:53a3972ce2983eb19c967ce068b985042b3e71bfccdd3994e6cd9f00c634908e

Observation 95d758fe-e73d-41a6-861f-0b67054236a1 · outbound

This paper cites Photorealistic text-to- image diffusion models with deep language understanding.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Photorealistic text-to- image diffusion models with deep language understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:19.487869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:19.487869Z digest=sha256:73b63a9c3a05a29ceeb4d69376f48e536f4458ad2aa53c9479798df0fee00f19

Observation 97a2557f-3e34-439b-b6c5-11430e8c3237 · outbound

This paper cites Improved techniques for training gans.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Improved techniques for training gans

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:19.673272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:19.673272Z digest=sha256:3e543a29da5fbe83c4b3c99c798ac02577f892f05c7332e52cc203980986f716

Observation fe780c4f-4ce7-4489-93e9-279dcbd7eb5a · outbound

This paper cites Enhancing image generation by fusing auto encoder & transformative generation approach.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Enhancing image generation by fusing auto encoder & transformative generation approach

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:23.446628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:19.806945Z digest=sha256:52664c76d49f957f5f3c48cd775ce6f8973d2f3a60cc9fff06e753c91904dc47

Observation ab8faaa4-7d4b-4eca-b572-142f7c2330a8 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:19.937516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:19.937516Z digest=sha256:6ac94979ef768318c88be65054cd5a9db6bf3aaeea8f5b9bd3275717ae331f4a

Observation e47f5798-5c80-4869-97a2-f217263df9e6 · outbound

This paper cites Conceptnet 5.5: An open multilingual graph of general knowledge.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Conceptnet 5.5: An open multilingual graph of general knowledge

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:23.159746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:20.089494Z digest=sha256:980ae7ae46ebf81b9e007868a4c78c542e09a1914e64891d567dbec02dba768e

Observation d68228e6-44a5-494f-af9d-d2e7fcc388e7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.226654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.226654Z digest=sha256:8c32e3193f41c141384c747b5134c30fc931ebc60a41a732ec0848bcc5efe244

Observation 8e87ac9c-38de-46f8-a40f-655a398cedf0 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.385035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.385035Z digest=sha256:71d6017c82dca3be000dd9257a79d176aa63f35c44e7d39668b0d0b38df61858

Observation b020c638-43fa-41e1-b516-a4ebdaa05ea0 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.529173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.529173Z digest=sha256:9a38d26b231c777314a64d62c63881efbda6462ea514f06003d7dcde81a9fe50

Observation 9113fd68-0001-425a-adfe-b640bd8c608d · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.704757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.704757Z digest=sha256:d0031c9f1f8f24dde337bcdfc2f9259fd76579494aa13900b8ef511498c1b335

Observation fef0d297-ac87-43a2-930d-ca4374b4f151 · outbound

This paper cites ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation ConceptMix: A Compositional Image Generation Benchmark with Controllable Difficulty

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.813586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.813586Z digest=sha256:c8c61a08e33728557fb88209ea8c0811855879df2c4cab60a201102f49315f70

Observation 50fd0ce5-2fc9-4961-bc78-26bd14e21a66 · outbound

This paper cites Imagereward: Learning and evaluating human preferences for text-to-image generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Imagereward: Learning and evaluating human preferences for text-to-image generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:20.939357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:20.939357Z digest=sha256:4b9becb134a9cc4403774174b8ecc1f2f1e6e1aa684d33598d8d42171b824aea

Observation 622168f4-d83d-4245-b848-2eba141ad681 · outbound

This paper cites What you see is what you read? improving text-image alignment evaluation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation What you see is what you read? improving text-image alignment evaluation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:22.890638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:21.069686Z digest=sha256:4386907a165bfb91351db11eda882db193fbfc4ec0036076e2c55783b931c521

Observation ac01512f-127b-4625-9706-5d0f8abe1555 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:21.180996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:21.180996Z digest=sha256:60fcf456ae6530fe8f1bf61155d4366eaf6d08391787ea6e00f5c5c99fc2bcc3

Observation 6c99276e-8e7d-462f-b1c3-637c0ae3a42c · outbound

This paper cites Sigmoid loss for language image pre-training.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation Sigmoid loss for language image pre-training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:21.304097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:21.304097Z digest=sha256:85650186c5f4e30ee00ede57e36aee20283ad31911ba8d8f5be967184bc2803d

Observation 8dfe7e35-1325-4ff5-8fce-280fbc210314 · outbound

This paper cites GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation GPT-4V(ision) as a Generalist Evaluator for Vision-Language Tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:21.420118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:21.420118Z digest=sha256:63b7c458b6b6f7379b2ada5b78aa018a936f03c169383530c6a1bd1659f02642

Observation 53f5a135-43f9-47b5-b72e-db4660a2314f · outbound

This paper cites A Contrastive Compositional Benchmark for Text-to-Image Synthesis: A Study with Unified Text-to-Image Fidelity Metrics.

Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation A Contrastive Compositional Benchmark for Text-to-Image Synthesis: A Study with Unified Text-to-Image Fidelity Metrics

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:21.502877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:21.502877Z digest=sha256:a34337ffeda550b76724d3f618e6c49df210ae7a99e2789b6f8f16a1e4c59bce

Pith citing papers

Observation f4cced5a-a747-4f11-ab60-922e101f0195 · inbound

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation cites this paper.

ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Align Beyond Prompts: Evaluating World Knowledge Alignment in Text-to-Image Generation

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:15:41.386031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:15:37.003251Z digest=sha256:fb21a7956702529297c40d5b5cd61e1804b30e98962dc7bf84988774ff1a3e7b