Pith. sign in

Paper Citation Record · LEDGER

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models

As of 11 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 1 inbound Pith citation observation for arXiv:2501.13920.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.13920 v1

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:31:35.809732Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:48:52.817656Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:48:57.874105Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact3
  • verified fuzzy30
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 862253b4-62c8-4cbd-9fd6-5301e0fe6adc · outbound

This paper cites Syn- thesizing cta image data for type-b aortic dissection using stable diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Syn- thesizing cta image data for type-b aortic dissection using stable diffusion models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.401810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.401810Z digest=sha256:b8e18b582e8f7d86ce94efbc4f7959c56368489c5364b0d4f069e662bd338762

Observation d2d7f9c2-31dd-48ae-881a-478678888e4e · outbound

This paper cites Unified Pre-training for Program Understanding and Generation.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Unified Pre-training for Program Understanding and Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.407394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.407394Z digest=sha256:382978d0966b65e87e46da2005d735aff0e35da279df177639833b1811c7cd00

Observation a1d3c295-4764-486c-b8e1-2106b10a0e77 · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Blended diffusion for text-driven editing of natural images

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.412820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.412820Z digest=sha256:b96771f119929d5378330b466216a44fea2af2c64a137e009650dc73586af42a

Observation 8dd1f75a-a8e9-42ca-b993-8976e69e658e · outbound

This paper cites Automatic Table completion using Knowledge Base.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Automatic Table completion using Knowledge Base

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:31:36.452343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.417658Z digest=sha256:ade1109518f74d6d18b47dd9ffc71a2751cc1e8ee44ab505e20c71e8e80f3591

Observation 5ad07537-9853-4ed4-97fc-daf005fdf57a · outbound

This paper cites Label-efficient semantic segmentation with diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Label-efficient semantic segmentation with diffusion models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.423057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.423057Z digest=sha256:824c9dd0d670411017357856e7df17263d8162672f95ca4246222f5e1fd1d798

Observation 4e21189d-6680-483d-a129-9d67b1ee0c34 · outbound

This paper cites Improving image generation with better captions.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Improving image generation with better captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.433441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.433441Z digest=sha256:2addaee7a318ea4669fb7223ed8870ba32ad3b8544efb07cca05cc8a81000cf9

Observation e242587f-7cc7-4e46-919e-a70d793ef20b · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.437991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.437991Z digest=sha256:05bace823a9e349803f6f3b58185066787604dcd78169ce4fea8148532576f9f

Observation 06f4cdec-faf2-4218-b027-7a9520b427d0 · outbound

This paper cites Video generation models as world simulators.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Video generation models as world simulators

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.442731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.442731Z digest=sha256:002d218c520e9da77489e7687fa33479918e38589cef40cfb63715fabaa5a4db

Observation d0af0974-de91-4bed-84f7-8409db5b665e · outbound

This paper cites Language models are few-shot learners.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Language models are few-shot learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.447538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.447538Z digest=sha256:7b72c14051a7ec341200f8e565655e33a923233267aca3788ef8697730b1fd2d

Observation 80e43f30-cf00-4f2b-a0af-d85dffaf16e9 · outbound

This paper cites Diffusion Self-Distillation for Zero-Shot Customized Image Generation.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Diffusion Self-Distillation for Zero-Shot Customized Image Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.451938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.451938Z digest=sha256:0b0f2ee2cfad5ce06beff9692d7999f690001d01b04cfa41d8e109cfa1e6f4b8

Observation 105ae17e-4228-49cd-98cc-acb5810de5b6 · outbound

This paper cites Training-free Regional Prompting for Diffusion Transformers.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Training-free Regional Prompting for Diffusion Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.456972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.456972Z digest=sha256:a8458eecb71ddff4eda9762430f310b020c6c4664f03adf881c7383f50549939

Observation f2ac595c-65e3-45d2-a375-7c7ecaebb6fd · outbound

This paper cites Topiq: A top-down approach from semantics to distortions for image quality assessment.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Topiq: A top-down approach from semantics to distortions for image quality assessment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.461745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.461745Z digest=sha256:68f088f5c9f1bd73b3a4ad3e1270981a2f2fd5dc77ecbfe52af96956c9290838

Observation 7964e635-75c6-41b9-bc9c-1f5fad5aa14c · outbound

This paper cites Evaluating Large Language Models Trained on Code.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Evaluating Large Language Models Trained on Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.465712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.465712Z digest=sha256:0dd84b25746b0c4911136cf8762a8822db63373f23d0e2f4f43c50332bf6ade8

Observation 9534fc3a-78a6-4e30-b4ad-e373b1731340 · outbound

This paper cites Graphics- dreamer: Image to 3d generation with physical consistency.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Graphics- dreamer: Image to 3d generation with physical consistency

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.470156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.470156Z digest=sha256:763e4dcf66bed49fa74002cfde85a878fd6b5fa14971d4ff59c747d5ac5f02c3

Observation 4e88b836-a7ee-422d-97f7-135aff87a4b8 · outbound

This paper cites Zhang, Gaofeng Meng, Xinyu Xiao, and Jian Sun.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Zhang, Gaofeng Meng, Xinyu Xiao, and Jian Sun

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.474133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.474133Z digest=sha256:12406bdad13c4214b9590063410ef642ddbacde5ce40bfd14631adcf9689e2ed

Observation f5771222-731c-49fd-925d-4a1773631b48 · outbound

This paper cites CoLay: Controllable Layout Generation through Multi-conditional Latent Diffusion.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models CoLay: Controllable Layout Generation through Multi-conditional Latent Diffusion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.478214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.478214Z digest=sha256:9f7b4ddc9bf0066d6c64e4ddf72daa5fbb4739895e9f0ba423736583f94513e7

Observation 8140b057-ee5f-40aa-b8e0-02c39edfa78d · outbound

This paper cites Dibia and Ça ˘gatay Demiralp.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Dibia and Ça ˘gatay Demiralp

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.109749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.482708Z digest=sha256:d08a5cfb5782c377e246a519eecaa2db98d5da5fcb731df09a0f4fd0fa8720ad

Observation 5150b6c2-4075-4fa7-9b3f-8409af47a5d0 · outbound

This paper cites CodeBERT: A Pre-Trained Model for Programming and Natural Languages.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.486946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.486946Z digest=sha256:ad00a2af3c4e902845cc8aeefab9163c121847f26aee9f3315a49f4c25ba3006

Observation 7fab635d-fefe-42f2-9af2-8cd4db2a551a · outbound

This paper cites Flux, 2024.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Flux, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.093738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.491457Z digest=sha256:4b2d0c305cb8f76f2d494117a25620815940acb8f24aa22be32fd71df0e5f6ae

Observation 05781f8e-0607-4959-8b40-86ae62c311a3 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.495651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.495651Z digest=sha256:75ea284eb1f8fa737852e90a3c7bcd9b5ed2e84c1c1aaa86233a73a6bd8c2532

Observation 86cbbaa3-b068-402a-ae93-7afc7168ddb0 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.500088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.500088Z digest=sha256:b1eb135fb4e723fc70bf2dc4cfb08f9c1e999ebc288c3f74273c8a6bc3e44aa1

Observation 3e0f86c3-9a47-48b3-9a75-7ff8e71c7a35 · outbound

This paper cites Seer: Language Instructed Video Prediction with Latent Diffusion Models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Seer: Language Instructed Video Prediction with Latent Diffusion Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.504476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.504476Z digest=sha256:f58a977df46a077697797ece015f5ee6df22bee87e2e6972059a727432b63901

Observation 8ae5599d-a1e4-4458-b728-f4cfa8134bdc · outbound

This paper cites Sciverse: Multimodal scientific benchmark for large models, 2024.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Sciverse: Multimodal scientific benchmark for large models, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.077001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.508873Z digest=sha256:2d665acb73b164baef9e949fa2ce349d583a6aa0d9286a5a3a8c76cfd87abeb5

Observation d8d4fc67-4aef-42b6-96ed-76d06ed00ded · outbound

This paper cites Joint-mae: 2d-3d joint masked autoencoders for 3d point cloud pre-training.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Joint-mae: 2d-3d joint masked autoencoders for 3d point cloud pre-training

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.058655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.513060Z digest=sha256:19e221a24d76dbf070143117758ab6fec1b924da9a4df675bac0976bb7cf5b8c

Observation 147724f4-f3a0-4a5f-928d-2feeb035c551 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.517200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.517200Z digest=sha256:7de3aa1c8d44dbe2dd3774e75223037a635fca4bae22c098a1d2915d33730094

Observation cc9c52d3-c551-460e-b68f-35e491761680 · outbound

This paper cites SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.521349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.521349Z digest=sha256:677fda4d604998dfe59b6add18367c0cc8d8efa5b91d087cba325e8bddd20233

Observation d6bdb15f-500f-41e2-a7af-ba64b05088d1 · outbound

This paper cites ImageBind-LLM: Multi-modality Instruction Tuning.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models ImageBind-LLM: Multi-modality Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.526133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.526133Z digest=sha256:cfa68ffcd899af072ae00b433fb5b64db0b65f567825de7e075ed1fb88f92fef

Observation 0c1c9088-f367-4185-a4df-9eb0f23c626f · outbound

This paper cites Capa: Carve-n-paint synthesis for efficient 4k textured mesh generation.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Capa: Carve-n-paint synthesis for efficient 4k textured mesh generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.041094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.531024Z digest=sha256:543735f957f7e5e54bd5c0d48f80685c4c2c3173ec9221bad0b6d655425b337e

Observation 2491b58c-dd4d-4536-9df0-82487c0cb798 · outbound

This paper cites Video diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Video diffusion models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.535443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.535443Z digest=sha256:f21274ba418220f53e3fd38f6cbc44dea3efa7610618d146fae2dfaff3bf4e36

Observation 44796fe6-984d-42d7-ac1f-f0d1858ba8b4 · outbound

This paper cites In-Context LoRA for Diffusion Transformers.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models In-Context LoRA for Diffusion Transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.539884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.539884Z digest=sha256:1c7bf90c7bd483648d3422df617ce313ac03e6e621c6f0ccccb71232441852dd

Observation 7ec050bf-8549-454b-8210-ebef4108b854 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.544537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.544537Z digest=sha256:6eda5282b6ab0f32a968d540660d74151ce65430f9e7f7c77e82488c897c58c6

Observation 247d54e7-73b0-4a5a-b922-40f814791996 · outbound

This paper cites Tech: Text-guided reconstruction of lifelike clothed humans.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Tech: Text-guided reconstruction of lifelike clothed humans

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:37.002565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.548914Z digest=sha256:32518b6b009ad6e17a38442d878808c8a5a814ce224cc526fe321d1afb765436

Observation d4086814-3efe-411b-98d0-f9ea7b35fcb1 · outbound

This paper cites Rehg, and Varun Jampani.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Rehg, and Varun Jampani

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.985935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.553559Z digest=sha256:d473394c1e361a391f949b8fa9bc6ffbfff59ee50ae3d4ba3075fe97c0a7e683

Observation 5aa19789-73e0-43a6-b0b9-99bc5d412ad5 · outbound

This paper cites Unifying layout generation with a decoupled diffusion model.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Unifying layout generation with a decoupled diffusion model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.969397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.558057Z digest=sha256:70e6e7c824ae7812962fd4fb84501a390a5a67a44faf84d0f549941f7d2b3cb3

Observation 36e2c03b-547d-4c1f-b183-40f47dd13f4b · outbound

This paper cites Ideogram2.0, 2024.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Ideogram2.0, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.950921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.562579Z digest=sha256:9ff263dbca91c655e5d90de76d53476447191b5de0a39c66398960f793211ad8

Observation 63155cf2-d014-4471-960a-082242593a0e · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.567023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.567023Z digest=sha256:1d3ddb40eaebca6e79eb95839acee3480ac9bcc065fc78f99078314980f98459

Observation 489dd71e-a705-4ba6-ac3b-c1c860b62ce9 · outbound

This paper cites Imagic: Text-based real image editing with diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Imagic: Text-based real image editing with diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.931978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.571600Z digest=sha256:72c61f88dceb3479fcecfddde289c49efb9fc7952e6c6ea765393e42f22fa0d9

Observation 7f880129-2c4a-49fc-809f-d17b4ca9d61f · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Repurposing diffusion-based image generators for monocular depth estimation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.915483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.575904Z digest=sha256:9fa52c507afcc6406dc5cf8e505d134c9b27ba63322191b9d2a3c1b9ba2ce49a

Observation d13fcc26-ef6a-43bc-b5e7-1b979fb97596 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Musiq: Multi-scale image quality transformer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.898901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.580435Z digest=sha256:8a288ceb62904f475fd54d93d59f6d54e6f1ed247f5ceaca62918075adb8248a

Observation 811f0100-f0b2-431b-ad9b-c0f0d4ead4d1 · outbound

This paper cites David: Modeling dynamic affordance of 3d objects using pre-trained video diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models David: Modeling dynamic affordance of 3d objects using pre-trained video diffusion models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.881426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.584914Z digest=sha256:e82219a670dae7c3116fb9688693ddb643ebe74d6a2cc4284b386d10de011a8f

Observation 244b0d09-3906-49f1-ac21-4741d6ebffbd · outbound

This paper cites Diffwave: A versatile diffusion model for audio synthesis.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Diffwave: A versatile diffusion model for audio synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.589188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.589188Z digest=sha256:9194d01f2a1ac178074e64cd3d0473f6c3d8982f977af3328356f39b6d7eb79d

Observation 205081ec-5308-4aa1-afcc-78da1f34a302 · outbound

This paper cites Exploiting diffusion prior for general- izable dense prediction.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Exploiting diffusion prior for general- izable dense prediction

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.851931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.593421Z digest=sha256:5b5742766222e9658ac1a42874fef5da59b8eb2e46a5f6b55f004b5d20d9694f

Observation 1e7fa420-0d1b-4873-b353-6313d8aa430e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.597651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.597651Z digest=sha256:310025c3eb04941a030dc71a7c31e419474e85f54082312d62a7271abe5166aa

Observation 700fca19-cb51-4cbe-809b-15818d86dd94 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.602049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.602049Z digest=sha256:9eae46ad887505fe96d9d4980d627624b4fce2faa231ed9dcf9fa5948ea2ba56

Observation 96e906d5-9858-4aab-90bd-73308d73cd47 · outbound

This paper cites Simavatar: Simulation-ready avatars with layered hair and clothing.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Simavatar: Simulation-ready avatars with layered hair and clothing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.834494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.606391Z digest=sha256:5b6486e4a399129ab8c8f046991d4bdc2bcda8d94883d0af71d0b426c1ff4a1d

Observation f17e0d10-a7ab-4a2b-a132-fbce646c57a0 · outbound

This paper cites Anid: How far are we? evaluating the discrepancies between ai-synthesized images and natural images through multimodal guidance.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Anid: How far are we? evaluating the discrepancies between ai-synthesized images and natural images through multimodal guidance

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.815302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.610295Z digest=sha256:bbcc8a044117c2898b84c7a95279749ce9f937f06ce13ec5230eb7cc92bd5113

Observation 402fa52d-148c-47fe-acaf-54f89dc06237 · outbound

This paper cites Rupprecht.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Rupprecht

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.795707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.614286Z digest=sha256:e15016ae2fada70380dd751e866dcf73c23e3cc480cf12c9642c6f513bf5f891

Observation 38b49547-1091-4bd6-b74f-452fa692f9eb · outbound

This paper cites A Survey of Deep Learning for Mathematical Reasoning.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models A Survey of Deep Learning for Mathematical Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.618310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.618310Z digest=sha256:a2c543c9604ce331f9c1a2deef800c07139bef429badcb28ad160a8d038ffe9e

Observation c0f63fa5-dac7-4971-9900-94b9f83e513c · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.623430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.623430Z digest=sha256:27168c6e88f493d61b724fece8e1ec20f8c19e90eb7ffac6e0d0a8028c4bb08a

Observation e3834a71-6331-41e5-8498-6e9496cb6f5c · outbound

This paper cites Leveraging confident image regions for source-free domain-adaptive object detection.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Leveraging confident image regions for source-free domain-adaptive object detection

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.775995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.628107Z digest=sha256:d1af1294358a12e95f81659f1e0f0796ebf59083235f9ec3fa0e3c8676ade4d2

Observation c534d2b8-ca3b-426e-a7c1-0566c96bcc72 · outbound

This paper cites PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.632349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.632349Z digest=sha256:0c25a04cdc809a9597dc84e48cd4667aa7bf8119c472b9dd35fbe1559982260e

Observation b6312deb-a5b3-427d-9030-602d20374be6 · outbound

This paper cites One-d-piece: Image tokenizer meets quality-controllable compression.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models One-d-piece: Image tokenizer meets quality-controllable compression

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.758142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.636970Z digest=sha256:1ae9b729142c4b8bd427b304bcd82c95ee3abdcdfc5d2c4c921ff6dd899223b8

Observation 2b70e9c7-6e42-4182-b8f3-58ce41681ff9 · outbound

This paper cites Do generative video models learn physical principles from watching videos? 2025.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Do generative video models learn physical principles from watching videos? 2025

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.741662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.641397Z digest=sha256:b3bc95a4897a0dcb3b5410dd5e190cb80260cd08dd4530c5f8ef333e9829ecec

Observation 6593149f-f4c1-4dd0-b36d-91e886732784 · outbound

This paper cites Boosting text-to-image generation via multilingual prompting in large multimodal models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Boosting text-to-image generation via multilingual prompting in large multimodal models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.725137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.645810Z digest=sha256:6cd394f33783b86398cd929aed5ffd0b476f574e3e529053855e1ad01a130bc2

Observation c470553e-010b-402d-b389-34db4fd07ec1 · outbound

This paper cites GPT-4o System Card.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models GPT-4o System Card

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.650268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.650268Z digest=sha256:3786acfa64baf21e332d7e9b0e9658d29d6a54fe2a642e4cc294db66627dbcd8

Observation 914e05fd-07ed-45a4-8e56-8661da6a052a · outbound

This paper cites MathBERT: A Pre-Trained Model for Mathematical Formula Understanding.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models MathBERT: A Pre-Trained Model for Mathematical Formula Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.654643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.654643Z digest=sha256:69e56667fbcd40f749aafb2c4f00666cbc52e4980c530a9f2e1aae9d0a3cd4c4

Observation 38d77a10-99b7-4912-9f10-476d052decf8 · outbound

This paper cites Mutualforce: Mutual-aware enhancement for 4d radar-lidar 3d object detection.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Mutualforce: Mutual-aware enhancement for 4d radar-lidar 3d object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.709105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.659513Z digest=sha256:25281afeb23c6c2b46b110cb6b3044caa10955fc1929afac66d3a1e8dc8e203f

Observation 690bd2c5-eacc-43f6-bd1e-5014b3fd2460 · outbound

This paper cites Improving language understanding by generative pre-training.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Improving language understanding by generative pre-training

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.664152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.664152Z digest=sha256:4c5a297d2cf386967e79d079a81018cf7daab026d40ff55b8fd1c930685343e0

Observation 40d37746-71c2-4d30-aba8-17ce8211d244 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Learning Transferable Visual Models From Natural Language Supervision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.668603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.668603Z digest=sha256:9d3c1ec5739fd0ee25cf6aafcece54e11d06fbdf0e3f53a9f7ee3e16ac11f1ac

Observation 26c27ef8-0c9e-4bb6-80f4-eeb9918ac54c · outbound

This paper cites Language models are unsupervised multitask learners.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Language models are unsupervised multitask learners

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.673281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.673281Z digest=sha256:39428c9d44523b537c58ef6188fe6a92e3a60b2b8c8fae0e2f37f22888921927

Observation 921a0ba6-b59e-41a2-93db-ccacbb4f7ae7 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.677631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.677631Z digest=sha256:0b927bbabe504a8b194e3a5a688459d7cd02e0a8c94f67d3fdd088e3368ced25

Observation 99d78afd-5cf8-4cb9-aa39-2ed2c5f8a74d · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.682278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.682278Z digest=sha256:2c3be6a989e43d820da66ffb40b5847938b08f33c9fc68e8f17cbca29d973863

Observation a840cd9d-5bc9-4df5-9460-126a38fa7bd9 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models High- resolution image synthesis with latent diffusion models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.660747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.686705Z digest=sha256:419c03f2656b12b768d384c736a6b4a46660d1ef8004e4fa3a4749769fa28e51

Observation ed916e73-e527-48eb-ac84-db23ef176aab · outbound

This paper cites Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Generating Analytic Specifications for Data Visualization from Natural Language Queries using Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.691354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.691354Z digest=sha256:d14ca3d3db0543de6f0840feb6730889d6bb2c79e9daddc1c67a3a45edfcf213

Observation 47a61c8e-ddd2-4a96-9875-ff48b88a75a4 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Photorealistic text-to-image diffusion models with deep language understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.696353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.696353Z digest=sha256:642d7998e0ae7fadc23f41284aeeeb68bbf9698992361a905d0f097ecd3bf431

Observation c1bf9d84-aced-4f62-be4b-18bb08317731 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.700579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.700579Z digest=sha256:91d21a32cf412724285e161412b115d3e88b0af5c5c4c92f3bd18911f16bbedb

Observation f54c9011-7ba0-4318-9659-e3d16a50ebf4 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.705131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.705131Z digest=sha256:8edac89f645f328980625823fe8e86163f0d3de6acb13b480372609a3c4fcefb

Observation e327c015-69cd-44d6-95a8-04179afb4994 · outbound

This paper cites Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.709560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.709560Z digest=sha256:6310d96a4b5c0d29bd75b5b3477093248a8ea67afd69ff513fba46efa49c9315

Observation 3c1ace7a-dbec-4dcc-b897-386a7e70bb96 · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.713925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.713925Z digest=sha256:5838b66dfd1ff4d65384b4c2c328c8f6ee67c1016210f56d1b13986729ec57a1

Observation dbefce17-3ceb-40f0-b84b-30c72b83f12e · outbound

This paper cites Human motion diffusion model.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Human motion diffusion model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.718409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.718409Z digest=sha256:49ec10a32c75ecd42855e6bd4f9371fe4553bc4b969ddd1e59ebe2be08954d6f

Observation f07f4f2f-a5c2-4f98-b73a-a130ab840b4e · outbound

This paper cites CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.722593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.722593Z digest=sha256:80d1075211ec8c202bbae9932573523662bebef34b1b7640fde911bca89906f3

Observation 3c91327d-04a7-4606-b7da-2643fa0b2a23 · outbound

This paper cites Translating Math Formula Images to LaTeX Sequences Using Deep Neural Networks with Sequence-level Training.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Translating Math Formula Images to LaTeX Sequences Using Deep Neural Networks with Sequence-level Training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.726865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.726865Z digest=sha256:d4c5636f3e067a8833a3911284dd7163ef631ee2b50b39a42194583aa96e88bb

Observation 9c7320db-a18f-40dc-90ce-193852b85683 · outbound

This paper cites On AI-Inspired UI-Design.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models On AI-Inspired UI-Design

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:31:35.959457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.732113Z digest=sha256:bfd330df348db547cd0509511851b8b68b1df1378be83a7b8826ae4770773679

Observation e6b2ca0a-5bb8-4ce8-92b1-3c3fb81e8cd0 · outbound

This paper cites Boosting gui prototyping with diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Boosting gui prototyping with diffusion models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.622654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.736798Z digest=sha256:51685f798e155e2200236c2d94cdf930f9d4ba0b36d016c6e22961a9c85419d6

Observation 37bdb7c2-420f-470e-ae00-56a8729e9854 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.741713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.741713Z digest=sha256:70ffb916ede955b6e19e90eaf9145f67dc6bce9b8d382c1df6a4c77399444c36

Observation 4d3f3e3c-69e1-4e78-8e12-fac062797ca7 · outbound

This paper cites Diffir: Efficient diffusion model for image restoration.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Diffir: Efficient diffusion model for image restoration

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.746673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.746673Z digest=sha256:5990490d8dc08a3deab783a953650c731c2be5cd298ca669f820f03eed10f0ca

Observation 89843210-b96e-4734-9fc1-afe37ee1c9b6 · outbound

This paper cites MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models MuLan: Adapting Multilingual Diffusion Models for Hundreds of Languages with Negligible Cost

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-10T15:31:35.919905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.751516Z digest=sha256:d7a5552cdf01a960ce41a756ccb993217eccb1f31d23e9f8edfdde7e47523a24

Observation 751dd917-2e38-4677-a8e6-880905b685b4 · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to-image diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Open-vocabulary panoptic segmentation with text-to-image diffusion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.596727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.756473Z digest=sha256:fb8e9d938b677cf30b965fe14c4156583e65e3afebd933d0279261f5d4541476

Observation af2a47b7-d719-4a24-96f4-e7a87ae3edb8 · outbound

This paper cites Synthesizing Tabular Data using Generative Adversarial Networks.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Synthesizing Tabular Data using Generative Adversarial Networks

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.761133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.761133Z digest=sha256:2587dca53c4992bbaf8232bf44c43479216bdaf35a1e839b44bf7f1229265b53

Observation dee41503-f2a4-4c2e-8572-922e89f6f9c5 · outbound

This paper cites Xu and et al.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Xu and et al

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.581076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.765923Z digest=sha256:75b0bf9c656192931105296e490c97eef7625db35a01df2b036b82d9b7b01b5f

Observation b6b14f4c-f832-497e-9eec-abb12bf15d91 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.770683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.770683Z digest=sha256:3f823bbfdb5db9228ac6a208745a75d449bfc21d95708ed049b9b648dff25435

Observation 0fe6c0ad-7f5a-4a62-bb82-a0ced303a218 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.775497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.775497Z digest=sha256:f27226280a916d2f994fe6fda062b91531ae068e77c85810ff9d58be73ebc90e

Observation 23677610-3f0e-4ecb-b59e-53e60d0dfee9 · outbound

This paper cites Follow-your- multipose: Tuning-free multi-character text-to-video generation via pose guidance.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Follow-your- multipose: Tuning-free multi-character text-to-video generation via pose guidance

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.565682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.780169Z digest=sha256:26fe2c9f0306dae3200877d12bc892dda9c6102065a35d638e1644996d149090

Observation 9ac4499c-3b46-4c57-80c5-cde2616de3dd · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Adding conditional control to text-to-image diffusion models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.784204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.784204Z digest=sha256:efc2d1d3ce179dee22e349b5083d13b3d4be6716e46269fe067c16515e08ba51

Observation 4fe56703-2d65-4690-b28f-40ea1291acd0 · outbound

This paper cites Motiondiffuse: Text-driven human motion generation with diffusion model.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Motiondiffuse: Text-driven human motion generation with diffusion model

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.539602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.788248Z digest=sha256:5def36d1f1d2701eb9b316c4cd08673e3ec71cc74116154c71d12becf6235fe9

Observation 3cb1bebd-d2bd-4f68-a3a8-3382c9a35359 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Llama-adapter: Efficient fine-tuning of large language models with zero-initialized attention

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.792632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.792632Z digest=sha256:12967d65a49c23555dd396b143107220f979835a0380ee1a4c03fa852818a2e3

Observation 86197403-3d8a-4e1a-a67d-9118edc42e66 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? ECCV 2024, 2024

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.796763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.796763Z digest=sha256:068793f609929372e58bbca1f1f8e5e314511e605b1593e5d2ddfe6178b4fad3

Observation 96e24ed7-194e-405b-b6ff-256442c2c3e9 · outbound

This paper cites Personalize segment anything model with one shot.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Personalize segment anything model with one shot

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.502943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.800849Z digest=sha256:d1da49232cc69a5adc974fdfba98df4405c2a512c52fbc2aaa3aafd897298aca

Observation 8bed7407-1a0f-4229-a0b6-dad16cd87fed · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-10T15:31:35.805509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:31:35.805509Z digest=sha256:6fb3679bad50cb93d56f80334a2257e75e339cd65be8096a31eb4001a5c19dfc

Observation 2f3ee3ad-1107-41aa-b0bf-b922f252f2a9 · outbound

This paper cites Blind image quality assessment using a deep bilinear convolutional neural network.

IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models Blind image quality assessment using a deep bilinear convolutional neural network

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T15:31:36.487759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T15:31:35.809732Z digest=sha256:0f54bfdd3722eab856cc8d5fed9f71801cd928d3c89fa9a8be1954dc168701c6

Pith citing papers

Observation 43c04451-4193-4182-b1fc-579930d9d841 · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation IMAGINE-E: Image Generation Intelligence Evaluation of State-of-the-art Text-to-Image Models

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:48:57.976466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-07T12:48:52.817656Z digest=sha256:7701ddb5f788d6caec33a2f16f22c8773c1bbf6529ddae54a9e695ce1c86714d