Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 50 inbound Pith citation observations for arXiv:2404.01291.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:00:06.795847Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T05:54:33.534825Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d1ef24b2-4ccd-43b1-aa21-889efb51e838 · inbound
VideoPhy: Evaluating Physical Commonsense for Video Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96c745dd-1da5-44d2-9fc6-372ca1e8588a · inbound
VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6368c9ef-7521-4c5e-88f0-bc0195210df1 · inbound
Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9059e584-5189-493e-bcfd-23d985e94404 · inbound
SimAvatar: Simulation-Ready Avatars with Layered Hair and Clothing Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc21e26a-3f0f-4fd5-ac97-0fb053ab1d2c · inbound
EvalGIM: A Library for Evaluating Generative Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d0c6d6-07e8-4428-bc38-eb571007e9e0 · inbound
T2ISafety: Benchmark for Assessing Fairness, Toxicity, and Privacy in Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a26062-eb2c-4113-b79a-684088bde622 · inbound
REALEDIT: Reddit Edits As a Large-scale Empirical Dataset for Image Transformations Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e0ec614-d48d-4cff-9a29-373da8aad3e3 · inbound
AutoSketch: VLM-assisted Style-Aware Vector Sketch Completion Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03f741a-1100-4821-bbbf-9beb1d706efe · inbound
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 98c45371-cbd6-4c4e-9547-7022b77eace7 · inbound
Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9b180b0-a051-498e-a748-04f6ced31318 · inbound
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0082d370-eec2-4389-a745-f24588fd8b08 · inbound
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f79b9c52-9e3f-46e8-9e9a-96587df0c713 · inbound
RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703f16e1-cf9f-40a6-b78c-b678b5ea79e7 · inbound
Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation babd993d-63a6-4fb9-89a9-c3b2f7b3f4cd · inbound
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df232a2a-1dff-4718-8067-9a7be0277a95 · inbound
Towards Evaluating Robustness of Prompt Adherence in Text to Image Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67375651-a6ec-40c5-abf9-2ca8678f0965 · inbound
StyleTailor: Towards Personalized Fashion Styling via Hierarchical Negative Feedback Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15f05b24-c517-4cf2-a712-3970452c8e49 · inbound
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1284d8-18ed-4fd3-9214-978f4bab7dea · inbound
Determinism of Randomness: Prompt-Residual Seed Shaping for Diffusion Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1bb93abf-c03d-43d4-affd-f5561060af96 · inbound
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e7ee3650-3c75-4ebc-9673-43051eeabe99 · inbound
FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82693605-52dd-49eb-9436-2e8bfd40d697 · inbound
GeoLoom: High-quality Geometric Diagram Generation from Textual Input Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92cbd170-d91e-4abd-9425-d1a11748ed30 · inbound
Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 549c1047-dbb4-4048-99ee-a67f443d2dee · inbound
Tiled Prompts: Overcoming Prompt Misguidance in Image and Video Super-Resolution Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b0ba704-665b-4cb9-9591-83fd4ffb39b0 · inbound
Multimodal Language Models Cannot Spot Spatial Inconsistencies Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce9ea81a-ed56-4c08-84ba-c45c4ab4dbec · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 64b1904c-15dc-4891-a05f-7ac9c262e9f1 · inbound
Evaluating AI-Generated Images of Cultural Artifacts with Community-Informed Rubrics Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1f043b2d-0ca3-427e-adbe-25b5684622cd · inbound
Generative Simulation for Policy Learning in Physical Human-Robot Interaction Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ec12550-e300-40a6-a992-993cfbec5fa2 · inbound
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5f563db-b279-4e67-a07d-ecd8a3ea374a · inbound
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3d6ad382-a72e-4b9e-9ff7-eeb5aa843de3 · inbound
Revisiting Compositionality in Dual-Encoder Vision-Language Models: The Role of Inference Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eec96584-3a22-41e4-972c-84cb1601bb34 · inbound
HumanScore: Benchmarking Human Motions in Generated Videos Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 610a4531-40be-463b-9e7c-4064aefe9996 · inbound
Building a Precise Video Language with Human-AI Oversight Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b398580e-9eb8-4496-950a-496a77b3ae8a · inbound
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e5c59264-63a6-4e64-847a-012716b9adaf · inbound
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation deae0353-38df-47b4-be9b-7786fa88ebb8 · inbound
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0703f8fc-ca3f-4759-8fba-3b45f87205d4 · inbound
OctoT2I: A Self-Evolving Agentic Text-to-Image Router Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 78283488-ab90-42ba-84cb-6e676fe9e416 · inbound
Drifting Preference Optimization for One-Step Generative Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 402ab300-422d-4360-b0b4-3566229a7903 · inbound
Aggregating LLM-Based Weak Verifiers for Spatial Layout Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b28c1deb-60bd-412f-aded-8a884883f08e · inbound
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 58db80f9-b654-423c-876f-d2abe913eee2 · inbound
Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9ba6b3d-2670-41f1-a467-67172dc8d38e · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a49818c1-a1da-4c19-9e60-7f3ea8b127af · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84bc719c-7b65-4e39-a14b-8bf0d68da4b5 · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e713bdb3-097c-42a3-9755-f167132cdfa1 · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44293852-cd35-4a06-bd9e-3516eb7cdeaa · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 885af771-e829-4bee-b993-0be55637f63f · inbound
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b15485-8046-4abd-9af9-931f173d5ae2 · inbound
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 72d22ed5-b880-49c3-82d2-6868f18c65ad · inbound
PoseAlign: Sculpting Pose-Consistent Meshes via Text-Guided Deformation Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938fa610-5582-41c0-bb9f-4b6442b8553a · inbound
Importance-Aware OBS Pruning for Diffusion Models Evaluating Text-to-Visual Generation with Image-to-Text Generation
Reference 113
Source-reported events for the cited work
Unavailable: canonical work link unavailable.