Pith. sign in

Paper Citation Record · LEDGER

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 10 inbound Pith citation observations for arXiv:2412.18150.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18150 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:02:21.265519Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:26.460713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T09:04:45.978885Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy31
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 307de585-4f03-4f32-96f1-07fbf6d6be76 · outbound

This paper cites GPT-4 Technical Report.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.851688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.851688Z digest=sha256:df7cec04cfa6a132c5de84c516ab52edf7b98f90a415ba2b09d619146fba5a10

Observation ad95ae33-07e9-4e5f-90e2-3a866e39fc66 · outbound

This paper cites Kandinsky 3.0 Technical Report.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Kandinsky 3.0 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.857101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.857101Z digest=sha256:05efdb6057e05422fb60cb2763393022e77f935d45fbe99890ce0b017c8570a5

Observation fbda3c3d-9c86-458f-90b2-542372b9899f · outbound

This paper cites an unresolved cited work.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:02:22.745634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.862210Z digest=sha256:4b059c3bf49f50c1c391a9e741881b8db290d44758abb61ad26451ab4e051ac6

Observation 76909cde-a0cb-4a6f-ba78-e441b7868eff · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.871423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.871423Z digest=sha256:c7bd91ea196e3a7c3b067db375850720297f32683a9f9f615803033fb6692f9d

Observation 7b580605-deab-43ad-8b41-8a96588eb7a5 · outbound

This paper cites PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.879214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.879214Z digest=sha256:9565ceb5f92cb4b93ef8c82b1060e3b5fc8b34144fe0d40b822b35a99519e82e

Observation cb1c4a1b-9fee-49e9-9f1c-cffd3c41a872 · outbound

This paper cites Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Pixart-α: Fast training of diffusion trans- former for photorealistic text-to-image synthesis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.718503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.886875Z digest=sha256:eb52f6757e6cdcf552b79039d503eddb780a908b9876cf16e8959d66e3382960

Observation feea6f55-9308-4406-b65d-a3da77087c31 · outbound

This paper cites Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Dall-eval: Probing the reasoning skills and social biases of text-to- image generation models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.694638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.894200Z digest=sha256:fe124d9f6caa5178aa2a7385cd1d6d3b2fead41c314a325da25e174659e5521b

Observation c4652b9c-b24b-40c6-b5a5-2ff49340a79d · outbound

This paper cites If-i-xl-v1.0.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation If-i-xl-v1.0

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.675019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.904653Z digest=sha256:fbfa1320876bb5679efac5ef27b6bf7867d6c77323cdc0428573ce74c64640a9

Observation 8602313f-eb29-4ee1-9494-e29fc6325577 · outbound

This paper cites Dreamina.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Dreamina

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.657134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.912278Z digest=sha256:8ea590dc01ee3322b6f2aa450c06e9eacb368cf90afeed0883ed1051ffcb5557

Observation 8c2fe1ef-eeb8-4120-a6e2-afbf475a170a · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.637201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.918953Z digest=sha256:bc7431985a583caaf6feffe6f627649a180d1696ef718e49f65bfed4a67746e9

Observation 3da1b9f5-41a1-4ca2-9ce8-4e57d21880be · outbound

This paper cites Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.925998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.925998Z digest=sha256:12130ed440ea64086a29fbbf83f5ed2098ab5fd62b995328adfa94211cbc5db9

Observation 3e78fa1f-f29e-411d-8a9e-907b01b74f5f · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.937773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.937773Z digest=sha256:effe3c518c06fa80cf6388912689cfa38d5463195a27e95d79a74908439d40db

Observation 455001f5-1923-493b-adae-2ab67bd2dda2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.607738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.948276Z digest=sha256:032beb07c5ba18ea90975d005ccba8a23ba2b1b6bd8acb41d26c8942f2faece7

Observation 55b022e9-6f76-4d76-a03b-436436feaf0e · outbound

This paper cites Denoising diffu- sion probabilistic models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Denoising diffu- sion probabilistic models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.588710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.959363Z digest=sha256:49580dee5000c5dfaeba1c100cb02e4b14695b5b49cbba20a1a3fe89678ce71b

Observation 2523502c-afa7-4a2d-91c2-7d540c0db447 · outbound

This paper cites Midjourney.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Midjourney

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.564578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.965988Z digest=sha256:e2a99db87bb8b8d3e7f0cd369a05d4c6b8e280a4f1b8d0ae009669592a58f13e

Observation 7037f01e-36fa-4e62-b6b6-918a8050cb37 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.538887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.972653Z digest=sha256:b16858cba044a99d1f7af83a01d25cbfb50f4a2aa5e7980e4897d4576893e9d0

Observation a98bdefe-d7df-4fb8-9920-c50cde40387a · outbound

This paper cites T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation T2i-compbench: A comprehensive benchmark for open-world compositional text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.497299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.978469Z digest=sha256:622228c48d46c4ad8fcc3b1e78e64b7cfb515a4cea9e0b253600030b06781e47

Observation 2600e025-9e03-4bb4-bc86-c62ce442f7d9 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.463105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.986792Z digest=sha256:ff738d7d3147f2ed84e1287062cb9022e7c0db3985b439f82eeaaadde560f09f

Observation 23a18070-6271-4b03-87ec-cd4747603d55 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.437220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:20.992411Z digest=sha256:ccfab2292b6b344510ebc1968655f7e373c29d9f6a25b3ac4a584b1bd2a2ebb3

Observation cb6b0374-f776-4503-8711-0803822a4cb4 · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:20.998785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:20.998785Z digest=sha256:3f547715cd0810eb34e2918b7917172ada2b00940b616c4e8397b7df3bc80f2e

Observation be980eac-6c0c-4333-92a6-fdd6b999f4bc · outbound

This paper cites Evaluating and improving composi- tional text-to-visual generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Evaluating and improving composi- tional text-to-visual generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.401777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.005276Z digest=sha256:36cb0d2023b97d5aa81d60acb6dfb0104a7e096bb732930adc4c25f53490e242

Observation f89b545d-0bd5-4bf8-ac60-28d90d10c2d7 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.012013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.012013Z digest=sha256:90a19e4a91d8fb9d8aaf2bef136c9d9a174f2b9d46552f90b981084cfc78de08

Observation 4d610c48-2b54-43ab-8d34-229cab943032 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.021586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.021586Z digest=sha256:7fe739097a90f9507ca01c314f7586f1f522136bab4bfc9c11c44e8d52d979b5

Observation 6decca13-abc0-427c-a26b-d644163b295b · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.029310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.029310Z digest=sha256:3b89330cca663f294001fff7b0b2be3289844a59290b6ee6a95fb65043964069

Observation 5cb528c2-36a5-4668-866e-1f1b969f270d · outbound

This paper cites Rich human feedback for text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Rich human feedback for text-to-image generation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.356693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.036558Z digest=sha256:01e84a92645e102865e2afbc2e4163485a4cd8f1cec1d4d5b92655cb39ef913c

Observation 47dd1d66-3a5d-4564-83eb-8f86b908a486 · outbound

This paper cites SDXL-Lightning: Progressive Adversarial Diffusion Distillation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.042923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.042923Z digest=sha256:3b80e2e379a9de3bafd5555c06bc7ab5ff734d17c29f3f4f4c7d24e69eb127e7

Observation 93f56d51-6c78-40f0-a08f-3f7866d4b12b · outbound

This paper cites Microsoft coco: Common objects in context.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Microsoft coco: Common objects in context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.051255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.051255Z digest=sha256:dd829b860f7d6c28be2bde5aaa5989d32ce1505f21d07f940362a5eed3a34158

Observation fe456f20-c1d2-4ce5-9daa-aacb5dbf0b22 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.311721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.059585Z digest=sha256:26713a768fe304db4295ab0224c25c85adc62d4f052006b155a3c827eb7af2d7

Observation bbf407be-a7e0-4ae2-a39d-3e819f06ae8e · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.071349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.071349Z digest=sha256:ddb2025ba7e4ec82d441e55906b8ef7903716730affdc040404c1e315e515f73

Observation cc432f8e-9bcb-4ad5-8dc5-06bbb44ffefd · outbound

This paper cites Scalable diffusion models with transformers.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scalable diffusion models with transformers

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.280148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.079970Z digest=sha256:898d5c8609f1072faf3319586cdda9a6a6bc6b83220a94fb810b5bd0d1e7b92f

Observation 594a4e38-b197-4553-95b2-511c39c83f54 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.249478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.089847Z digest=sha256:44a111b095bf8d4c526637a3873f5352bb5cc2344affe2f5f8dbfb91076be320

Observation 8eedc475-0ccb-406e-8f84-98bbc525c32a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.097690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.097690Z digest=sha256:677645885b90c334560d7908b53eacc4bfe1270f6b9c36c3116159418ea4cb33

Observation 1e5849c8-95e4-48d9-bdec-b36e2f2386d4 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation High-resolution image syn- thesis with latent diffusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.221393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.103787Z digest=sha256:cb1ee5686fa2ca8ba0bdb7d41941b79f860949a6af897155ef3c337ebfb84afa

Observation 7f5035af-9229-42c9-b7b7-3e740d5f785c · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Photorealistic text-to-image diffusion models with deep language understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.188309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.111622Z digest=sha256:275f9fffd31c102f9866b5545fc61e1a7e09e2841bad46c5f7e86dd323778c25

Observation d49c2c9b-c03a-409a-910a-1e5532028457 · outbound

This paper cites Improved techniques for training gans.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Improved techniques for training gans

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.157587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.117927Z digest=sha256:37b30855e4f45b5c05074ea338ea2243d167373b45ed8e7e04031efe9b483e39

Observation dcd8729b-d21b-4149-8598-4bab0acc0422 · outbound

This paper cites Adversarial diffusion distillation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Adversarial diffusion distillation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.125556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.126537Z digest=sha256:84cd903aaf675801e307afc5900cfd7bea14ee7b84f00d6d04c29d41d3e72c99

Observation 2e0a4791-8562-48ad-8c09-f65c5e542251 · outbound

This paper cites EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.134390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.134390Z digest=sha256:9940fa742faf168e28d52b5ed6548369d19d3b3a4c2e0c44cef48db3ccd78b2e

Observation ec67be34-20d5-4d73-8b90-7687ac6dd39c · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.146494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.146494Z digest=sha256:a9a2d40ad46c1b4ca5bbe9095cc3cb014bfb5508a1fba46a591a234b31750a8f

Observation 5611f23b-ea4e-4db1-b371-7ab0b513843e · outbound

This paper cites Shaping datasets: Optimal data selection for spe- cific target distributions across dimensions.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Shaping datasets: Optimal data selection for spe- cific target distributions across dimensions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.090609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.155621Z digest=sha256:ccaea02b8364ee26171a786b1bf6ea5e024e94d96bbc75a2f27b00849325d005

Observation d903efde-b1a8-4d39-83ca-20655a7eb158 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.167608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.167608Z digest=sha256:229b11983f14c3b85a164ebfeb5a74113ad640b42c3b61c64b55daecd351ad12

Observation 824b497c-b825-43cc-8ecd-15fad71e2ded · outbound

This paper cites Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Diffu- siondb: A large-scale prompt gallery dataset for text-to- image generative models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.056488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.177037Z digest=sha256:44d1d707de52b21729d5aee958ef1d9cb2aa2d1e32a0d40c40a7c666b40f24f0

Observation 8b36e445-d6cc-49f1-9bc1-a2fdc6827740 · outbound

This paper cites Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Revisiting Text-to-Image Evaluation with Gecko: On Metrics, Prompts, and Human Ratings

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.183802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.183802Z digest=sha256:0603bb088bfda0caba5d9d5887a8810aefd173a42bc081b3cb77bc3855a36929

Observation e73ef5b9-5d03-40f0-9ce5-f3d0c7663fa5 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.191344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.191344Z digest=sha256:54625982ba86a4fe3ba101e24da0ab899963c1b150355ad791dffb72d2ca359a

Observation fce54b42-52b2-4ef2-a7f1-a8f8d5e8fc25 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:22.003846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.197982Z digest=sha256:d728a145bc2df749be453176e1ffadd588a20d2b228f942262d5331af8f0ce78

Observation 0668094b-562a-4122-8821-04ba1d075eac · outbound

This paper cites What you see is what you read? improving text- image alignment evaluation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation What you see is what you read? improving text- image alignment evaluation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.970675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.206424Z digest=sha256:d4894c7aafc511167855831ff3eca44e9f92bdab1295f503e17130066d64ee7f

Observation b767f916-d5ab-4753-b078-ea09397f85ca · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T05:02:21.213390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:02:21.213390Z digest=sha256:9751ff5bfe967f5c08e3939b202083a3f3ddfaa8fddf65676b09395b1616ee94

Observation 15f99ce7-d24c-4a56-b0cb-d1f90b984ea8 · outbound

This paper cites Scaling autoregressive models for content-rich text-to-image generation.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Scaling autoregressive models for content-rich text-to-image generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.947283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.230479Z digest=sha256:c8c5caf5f10411be9028dd43cd1a77063ed8e223147725515a61161fcf869964

Observation 080476d3-f4b0-4183-bb0f-b3ee56730743 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation The unreasonable effectiveness of deep features as a perceptual metric

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.921018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.237859Z digest=sha256:baf74f7c9208059f5edc35bd6d780794eba0c161b76e32cb6143aba61226034a

Observation d590b938-a58a-41b9-bf4c-abc96f605fdb · outbound

This paper cites In data collection details, we de- tail the classification and sampling of real user prompts in Sections 8.1 and 8.2, ensuring diversity and balance among the real prompts.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation In data collection details, we de- tail the classification and sampling of real user prompts in Sections 8.1 and 8.2, ensuring diversity and balance among the real prompts

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.894425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.243077Z digest=sha256:b43975a6d492b431c5fa49179546d6748aa400e92e1e9ca443ce761b12afd17e

Observation 5eb7cd80-4821-41c8-bbe2-19b2ca55d1fa · outbound

This paper cites 1 cat and some dogs.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation 1 cat and some dogs

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.865554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.251563Z digest=sha256:9ce8433ac94aaf0f1c03d7f03c21753fa8f5879a43a6336b51387fd0ef2848de

Observation 7a7ddfa5-7fda-4b6e-b168-15ffacdffd25 · outbound

This paper cites Yes.” and “No.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Yes.” and “No

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.841468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.258732Z digest=sha256:641524d4f0a35673930f9c27282d0b39cb68d6ccbdc7e47148d45027da567857

Observation f9334116-0c08-4333-94c8-1cfcc72905b6 · outbound

This paper cites Evaluation of image-text alignment across different T2I models.

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation Evaluation of image-text alignment across different T2I models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:02:21.821794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-11T05:02:21.265519Z digest=sha256:7612414105e7f51e4ab2c8519570daae62f1c357afc08f82dfbb67606e6a2b60

Pith citing papers

Observation 332d43b9-40a9-4b4d-9326-6a9fb54f5748 · inbound

Unified Reward Model for Multimodal Understanding and Generation cites this paper.

Unified Reward Model for Multimodal Understanding and Generation EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T00:44:30.874520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T00:44:30.558048Z digest=sha256:b78a151223339f31f6cdf44cef0a49c3dd86abc4ad9965976d4df0f3941fd16f

Observation 875f6439-a1f1-4702-ac45-987eb212db60 · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T08:27:36.292569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:57b2aeaf3c89191ab98663a2f26613d29430be99f7fe9193447ff2bc40bbf365

Observation 24604afa-ef4f-44c1-9f38-99a20c6f70e7 · inbound

Seedream 3.0 Technical Report cites this paper.

Seedream 3.0 Technical Report EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:55:38.744353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T07:55:38.690569Z digest=sha256:b2594d8d391065a7fdad8dbbd1c0958efc51d95df483ca3e13a5439999acb808

Observation 73ddecb0-e4a0-4d62-b9b0-dec03dc27092 · inbound

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching cites this paper.

Instruction-augmented Multimodal Alignment for Image-Text and Element Matching EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:26.460713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:26.460713Z digest=sha256:a597284074d94eb6b96c816b1f1d9caa39515b5899c30fd12bfae6c0c92d8056

Observation 0717342c-df3f-486f-b6ef-76aa03d125a6 · inbound

LLM Code Customization with Visual Results: A Benchmark on TikZ cites this paper.

LLM Code Customization with Visual Results: A Benchmark on TikZ EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T23:37:38.074067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:37:38.074067Z digest=sha256:7501bd771516bb862549d52dc77bad5e27bda86dd48c176c510dc6fda3c8aec8

Observation 6c5980de-0401-4a87-bc22-9002791225ac · inbound

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation cites this paper.

LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:46.138692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:46.138692Z digest=sha256:f0b26c7a40856ee15d1c4cce7fe336f29134c6cea28dc47f3a070b8ff4caefbc

Observation 20e1c83d-9c5e-4fbe-8e3d-fccc0850c1db · inbound

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models cites this paper.

MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:33.225548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:33.225548Z digest=sha256:27e737d21c1aa3776523a201cdfd70ae0db6d4ddf22317fec43bd85a1c51407a

Observation 0a48ac7b-7546-4399-aed5-fbcb6dd31e9c · inbound

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images cites this paper.

ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:27:52.889724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:27:52.889724Z digest=sha256:f0bba395f3682967115387ffb2a5aea7ec80b3a030bfdc0b7288100fe823dc03

Observation f050f52b-e189-4d85-9826-0aa32acfc810 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.692481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.692481Z digest=sha256:96ea2ec303008e397a487c2dadfb86730be73da0c515d14a5a46713fc4ae8d05

Observation de3d8133-aed9-47b5-abd9-3b631ea87206 · inbound

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model cites this paper.

BEiTScore: Reference-free Image Captioning Evaluation with an Efficient Cross-Encoder Model EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:04:45.982530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T09:01:24.453821Z digest=sha256:1a9bd2a31a5f77621ce1ca9314c62fb8099b6c771169b4df16ca681e5d86410d