Pith. sign in

Paper Citation Record · LEDGER

Evaluating Text-to-Visual Generation with Image-to-Text Generation

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.01291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01291 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 50 of 50 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:00:06.795847Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T05:54:33.534825Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1ef24b2-4ccd-43b1-aa21-889efb51e838 · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:34:37.681410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:04be06da8cab4ad95946d6508166c2bc0cef8af102b1a1df310d0479c71f49d8

Observation 96c745dd-1da5-44d2-9fc6-372ca1e8588a · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.431602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:5165ce56730838c0d87cd73a6e5c7e670bc4e97ce3eae4f6811a15c3f8cbf26d

Observation 6368c9ef-7521-4c5e-88f0-bc0195210df1 · inbound

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation cites this paper.

Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:39:59.994659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T14:39:59.870039Z digest=sha256:3bfb91c7465add86ddb30442dcb9d9c746096e1493826ef656abd229b145ddfe

Observation 9059e584-5189-493e-bcfd-23d985e94404 · inbound

SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing cites this paper.

SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:06.795847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:06.795847Z digest=sha256:7b634b1b47173caad730b71cff76de1afbfaf78e831b0bc2cd05480ffb445180

Observation dc21e26a-3f0f-4fd5-ac97-0fb053ab1d2c · inbound

EvalGIM: A Library for Evaluating Generative Image Models cites this paper.

EvalGIM: A Library for Evaluating Generative Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T15:50:16.782526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:50:16.782526Z digest=sha256:7eaae5b8fa2181aa67343f4b4cf60287fdfce175243d3bfc81fe14052da59375

Observation 34d0c6d6-07e8-4428-bc38-eb571007e9e0 · inbound

T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation cites this paper.

T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T17:05:50.701277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:05:50.701277Z digest=sha256:0e747b67dfc1f7e40ad5b7000c903fab6b223bc57439fa325bb309f2f0206fb0

Observation 18a26062-eb2c-4113-b79a-684088bde622 · inbound

REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations cites this paper.

REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T04:23:01.155996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:23:01.155996Z digest=sha256:1d54d0b6c039f8650fb644493259bf31fdf5468e829301fa1ca814bda01854d6

Observation 6e0ec614-d48d-4cff-9a29-373da8aad3e3 · inbound

AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion cites this paper.

AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-08T19:37:19.811973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:37:19.811973Z digest=sha256:56b35d50a117764b1ea26f5d57f229bd3958e99b918bb06e9d8b99986896c408

Observation d03f741a-1100-4821-bbbf-9beb1d706efe · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.675197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:d57c8019fb3aa995e9e11de1ca5225398c18b68d18f8e47b39f39a66335abc30

Observation 98c45371-cbd6-4c4e-9547-7022b77eace7 · inbound

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model cites this paper.

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T08:27:36.315598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T08:27:36.242416Z digest=sha256:f1e4e60d085408a23ad5a9f8e8f5a29ddae97da9228291eddb5608cefe73d442

Observation b9b180b0-a051-498e-a748-04f6ced31318 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.897426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:ca3e5c138b18b3ad598decfd4c5fe0cae6590640132394ac6170e63d20847c99

Observation 0082d370-eec2-4389-a745-f24588fd8b08 · inbound

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs cites this paper.

TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:46.589608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:46.589608Z digest=sha256:9316688af8c68797152efb8a36f1e0271c51474ac12b4227115363d8be5e9366

Observation f79b9c52-9e3f-46e8-9e9a-96587df0c713 · inbound

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction cites this paper.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.662605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.662605Z digest=sha256:40b045fb06a4d1c5a1d72dc5978133e5df69283596255aff672a1f86b3b8489a

Observation 703f16e1-cf9f-40a6-b78c-b678b5ea79e7 · inbound

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models cites this paper.

Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:14:42.136106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:14:42.136106Z digest=sha256:9c5ffe5f38ca9bab56451db6e27b860e6f6a0938f47b854b274c97feb2bd4a00

Observation babd993d-63a6-4fb9-89a9-c3b2f7b3f4cd · inbound

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models cites this paper.

Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:54.778142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:48:54.778142Z digest=sha256:29c02675a5321ad472d1843ee9db5ec90206524795d283de2a2a9b2a5557f15c

Observation df232a2a-1dff-4718-8067-9a7be0277a95 · inbound

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models cites this paper.

Towards Evaluating Robustness of Prompt Adherence in Text to Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:43.571364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:43.571364Z digest=sha256:124dd54ce3d53777cf3da567cc9e5f18f9b46822451e7f2810661998b27b04ef

Observation 67375651-a6ec-40c5-abf9-2ca8678f0965 · inbound

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback cites this paper.

StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:52:46.101151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:52:46.101151Z digest=sha256:f3dbb5a7e6d80d885154e9c878df1579241c1a94cbd243e9b110edb024e9368b

Observation 15f05b24-c517-4cf2-a712-3970452c8e49 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.931759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.931759Z digest=sha256:82e36683d5e9724163ef912d30032f1413aba628df045f576ff19fdf42aa7791

Observation 3f1284d8-18ed-4fd3-9214-978f4bab7dea · inbound

Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation cites this paper.

Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:15:31.732222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-18T00:13:15.606804Z digest=sha256:83155a6793197694ca169d988b1e1bf44c6ab584ba2c6d636b4207c216da133c

Observation 1bb93abf-c03d-43d4-affd-f5561060af96 · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:50:15.301742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T20:45:15.034196Z digest=sha256:6ad9a8145d167660816fbee73486659dc6d87f056934179f529bb5e0035f5a7f

Observation e7ee3650-3c75-4ebc-9673-43051eeabe99 · inbound

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle cites this paper.

FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:20:28.805804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T07:17:26.381232Z digest=sha256:7106a751cb99c505516c6489e4395196062120619fddf7e4736f64006e6e6772

Observation 82693605-52dd-49eb-9436-2e8bfd40d697 · inbound

GeoLoom: High-quality Geometric Diagram Generation from Textual Input cites this paper.

GeoLoom: High-quality Geometric Diagram Generation from Textual Input Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:49:11.784837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:49:11.784837Z digest=sha256:eb890ae3db4b0694be8a461c02c60d40d34d8d7560071b72d5e12aa05975b196

Observation 92cbd170-d91e-4abd-9425-d1a11748ed30 · inbound

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics cites this paper.

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T11:52:45.814202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:52:45.814202Z digest=sha256:d514755a1149d7ab3bca550921ad6ff4dd1a768e47ba27cced7c070879852b86

Observation 549c1047-dbb4-4048-99ee-a67f443d2dee · inbound

Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution cites this paper.

Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:32:36.336186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T08:32:16.532299Z digest=sha256:1ed26536abd9bffaa8458bb6757166e5f53803f88db06add4a4df2b9191494e4

Observation 1b0ba704-665b-4cb9-9591-83fd4ffb39b0 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:13:24.810235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:1cd124b7ce894af04f3c7fc8ef83d23937a9ba18c4b1e8ec915f15472965ce24

Observation ce9ea81a-ed56-4c08-84ba-c45c4ab4dbec · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:03:20.088448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T20:59:52.448832Z digest=sha256:0d82d07fbbba0f08c3dea64c1e1dcfef2baff7bf7ec173b376b3632dc335baf5

Observation 64b1904c-15dc-4891-a05f-7ac9c262e9f1 · inbound

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics cites this paper.

Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-21T10:40:00.846699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-21T10:35:39.269869Z digest=sha256:ba10486702acd556e5b7d1b4c7519ce8b62e506f818433b66c57cf67da735b7e

Observation 1f043b2d-0ca3-427e-adbe-25b5684622cd · inbound

Generative Simulation for Policy Learning in Physical Human-Robot Interaction cites this paper.

Generative Simulation for Policy Learning in Physical Human-Robot Interaction Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:35:57.556101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T17:06:58.099881Z digest=sha256:ec340af09b7c2d64ad4f4ed977f38093953888ba97649c3918292e6ef697162f

Observation 7ec12550-e300-40a6-a992-993cfbec5fa2 · inbound

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies cites this paper.

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.978677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:46:16.091095Z digest=sha256:818ef6d03d115ebb65feecbf0959d9905b15a7a45459c0bfd5047b931fb51bab

Observation a5f563db-b279-4e67-a07d-ecd8a3ea374a · inbound

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies cites this paper.

RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:29:49.685564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T06:28:08.468872Z digest=sha256:64dc22989b492b48e5d16105416dde3a5d85223bc9bfcb42c6b56d1303e84c4f

Observation 3d6ad382-a72e-4b9e-9ff7-eeb5aa843de3 · inbound

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference cites this paper.

Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.797365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:38:00.522094Z digest=sha256:07ef216cfe189a9ca029b56eff0f772935e6685b9421875376d7ea1f9f7c1821

Observation eec96584-3a22-41e4-972c-84cb1601bb34 · inbound

HumanScore: Benchmarking Human Motions in Generated Videos cites this paper.

HumanScore: Benchmarking Human Motions in Generated Videos Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:41:04.395939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T01:17:16.513512Z digest=sha256:7fb9bb4d4bb0eb7f5ac887732b82727555cf36b8b2958700e77b63a819a24c04

Observation 610a4531-40be-463b-9e7c-4064aefe9996 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.455477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:7b14ce3ef09c0f70f9fc8709f0325359538be81d321ea7ffbaee164cb0f0eae2

Observation b398580e-9eb8-4496-950a-496a77b3ae8a · inbound

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning cites this paper.

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:59.750940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T01:51:42.304051Z digest=sha256:ed7114e77d2c0bd26cc4976979b100fd346df1cf382f5062df65ecfc68f1d142

Observation e5c59264-63a6-4e64-847a-012716b9adaf · inbound

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation cites this paper.

MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T06:51:11.484171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T06:46:19.367708Z digest=sha256:cc8f1f4ae458aabcd43960dea6a7160131729ff3716fc884de5e75b460c09850

Observation deae0353-38df-47b4-be9b-7786fa88ebb8 · inbound

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration cites this paper.

CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T05:45:23.595429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-25T05:44:58.627525Z digest=sha256:b3407174cef1252c27d7a2b9d0dc0e71fde13b76023712a212739f872c64a315

Observation 0703f8fc-ca3f-4759-8fba-3b45f87205d4 · inbound

OctoT2I: A Self-Evolving Agentic Text-to-Image Router cites this paper.

OctoT2I: A Self-Evolving Agentic Text-to-Image Router Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:22.511477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T14:22:57.822123Z digest=sha256:5df8ce30ef4af8862e381b172730ee5d5c316b7a1524061a670f4c33cb7d7b60

Observation 78283488-ab90-42ba-84cb-6e676fe9e416 · inbound

Drifting Preference Optimization for One-Step Generative Models cites this paper.

Drifting Preference Optimization for One-Step Generative Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.554631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T15:20:00.341773Z digest=sha256:cb82d822ca30a3c78c4f047e8029ff9712a1d9f4d0efc3267dae179e8cd0d49f

Observation 402ab300-422d-4360-b0b4-3566229a7903 · inbound

Aggregating LLM-Based Weak Verifiers for Spatial Layout Generation cites this paper.

Aggregating LLM-Based Weak Verifiers for Spatial Layout Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:36:55.559355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-28T03:11:18.416545Z digest=sha256:533dbf152ebcc0a7550c0ddcd713872d5fa0e4c51b14b9904b0641353f305850

Observation b28c1deb-60bd-412f-aded-8a884883f08e · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:27.975927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:30:57.001021Z digest=sha256:fed50fe6036206ab66ff579d7bf8dc34c8dc178c3b39866eb6461b287b99b7d9

Observation 58db80f9-b654-423c-876f-d2abe913eee2 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-15T10:53:37.186361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:53:37.186361Z digest=sha256:6b197484ad05d8351d6b47ad7b82b06498d2d4d55be01b54d3f88e6f9690bebb

Observation a9ba6b3d-2670-41f1-a467-67172dc8d38e · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:25:57.825432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T02:03:45.564122Z digest=sha256:2ce2e77e782d9bb6684e4c4d3171a8e72b24b7991b044efbb754b4d555952cd3

Observation a49818c1-a1da-4c19-9e60-7f3ea8b127af · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:40.661088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T06:25:58.872140Z digest=sha256:dd0e3c2446889557ff840c2c88c0ea6660e85337695b0fafe60ad91addda7f30

Observation 84bc719c-7b65-4e39-a14b-8bf0d68da4b5 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:57:22.814090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T20:52:28.444524Z digest=sha256:8e0437fa38f01ebd68a0eb5d0ebacd78e157f4fce9fb9c6d4a48c3d10cdcc5b2

Observation e713bdb3-097c-42a3-9755-f167132cdfa1 · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T22:49:00.907109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-03T22:44:16.272541Z digest=sha256:6baddbf507febe9cddf1b9f021159e94eeb818b36e7778ca3f50d7429737db36

Observation 44293852-cd35-4a06-bd9e-3516eb7cdeaa · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T17:14:19.770867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:14:19.770867Z digest=sha256:38e57ae2614f72d24ca129a3e3085511f2b261444dcbd3021f29257100394911

Observation 885af771-e829-4bee-b993-0be55637f63f · inbound

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments cites this paper.

MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T10:00:11.935029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:00:11.935029Z digest=sha256:b70cf46658a798e81ce9410479c865990b5620da6968cc476bb4f163c15d4782

Observation 36b15485-8046-4abd-9af9-931f173d5ae2 · inbound

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders cites this paper.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.536126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:84a30d9ef83148b24c6414fb12b4679ccf4acb6706a63e0a8a4f16ddaefe0be9

Observation 72d22ed5-b880-49c3-82d2-6868f18c65ad · inbound

PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation cites this paper.

PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T10:49:44.320793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:49:44.320793Z digest=sha256:46b9fd3894d987e639e1264e03ebf8a83cc484185cc3bcf2b5bf9bb211357b62

Observation 938fa610-5582-41c0-bb9f-4b6442b8553a · inbound

Importance-Aware OBS Pruning for Diffusion Models cites this paper.

Importance-Aware OBS Pruning for Diffusion Models Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-01T10:59:55.585977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:59:55.585977Z digest=sha256:e4f16c4cfa580709f0d1f2c2a74405278959c2781bc126a9ca19a53147f05f32