Pith. sign in

Paper Citation Record · LEDGER

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

As of 8 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 34 inbound Pith citation observations for arXiv:2508.09987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09987 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:44:18.490290Z

measured 121 of 121 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:47:10.728416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.706550Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8617fb45-85f5-4cac-b6d2-2f2376b89421 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.478127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.478127Z digest=sha256:c8c6d419592ffdd821ac5a4e388522fc41d12250c4e10c11cb7d3eee5c6cd2df

Observation 971a41a5-30bb-4562-b464-0237a213c23b · outbound

This paper cites Sd3-medium.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sd3-medium

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.536355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.536355Z digest=sha256:a2342600de66cf4b8b19733c9c6d3f87a2ab81bd9fdcce69c2deb828ed4b9cfd

Observation 97555354-04a7-49f3-9c76-1e9d20533c84 · outbound

This paper cites Qwen Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.622883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.622883Z digest=sha256:dfcbd5b23d3955222d5f60dcbf519b04dbc0308501c4c5b8698a7aee9f909f2c

Observation b99b6dda-b6a7-4f10-985a-3a24a173b7f9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.688300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.688300Z digest=sha256:0b97a751a227b58446edf794c86996bbf835a58dbf9b6b51d5ea6abdcce03a0b

Observation f773e167-5a88-44cd-be4e-e7ebc581fa0d · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Instructpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.745556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.745556Z digest=sha256:d795db5d4c2d7d9e163ccd66b85765df9d13501df53072c226ebe58386dfe7d5

Observation 34d256f9-ab71-4f3d-a38d-cc6df7905906 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.827910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.827910Z digest=sha256:b0cb7a5d0e07d853a71d9984ecf42883ddb47c5c3c46e2fe9ddf3594360a493c

Observation fb06e333-9da3-4435-ba91-0b2c8208054e · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.890407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.890407Z digest=sha256:6565e67d7f14b9b7336586074af6eb7f92f00b7fc53d39dacfc2cd9de84e748f

Observation 0a9f6bb9-d2f5-455e-8ebc-1323b3f2d242 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.965681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.965681Z digest=sha256:e815061d20abf30e37cce2b8e3f95d80555a0fcb54e1c19dd1cb332a594e1a9c

Observation 97a4601e-976b-4aee-a89a-3b446283c17c · outbound

This paper cites ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.031053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.031053Z digest=sha256:f7e2f1011d7fbe63a0bbae454cddf2a4dcaf2ef3301b51d51f93a1b9549192b1

Observation cc1c951d-86d4-43dc-88e6-8323404fe7e1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.140230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.140230Z digest=sha256:6c166f86b99725eb2479ba3721447b9dc15558d988f0fc7f2956bcaff1630ce5

Observation a9b7a0c5-af64-46f3-9b24-bc9e6904c3cb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sharegpt4video: Improving video understanding and generation with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.209115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.209115Z digest=sha256:176eaad29b8670a374a9c27b24d5a3a0ea2951785199572c4d3d18d048018e3a

Observation 0413e382-52db-46f9-8719-9efc0395d285 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.307662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.307662Z digest=sha256:6324b5c10c2bdc10aedd46b29e59f1ba6d2e89fa96e0b54e0d4b532aaf5c6a7d

Observation 669b36a1-b70e-4040-8632-c5031034c7f0 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.389528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.389528Z digest=sha256:0b108ef790198c22ac77c62948dcadf6cf0c70465ddfc23e11ef86c6888af2f5

Observation a496aa47-6186-4223-bf92-5645788ca429 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.474169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.474169Z digest=sha256:c32c5764a08dd9a1e77b981c5e4deb19824c9e0ec8d46087b891effce3186497

Observation e08b543a-3a34-488d-a6e8-45e8dd3a0fbf · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.546425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.546425Z digest=sha256:890053096c3289e228ce4896ae13a4a8da6baf575a638996ed89bad510c7ec94

Observation 2495bef3-5bb7-47dd-8736-453fa1f105bd · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Autoregressive Video Generation without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.606456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.606456Z digest=sha256:d88abfb7dcf6218f3bf8145957de3494bfafe45720af15bd7a08cefd426df88b

Observation 21cde564-0d55-4a8e-a3ae-3cdf3dcd7b23 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.699910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.699910Z digest=sha256:0399c2465015dbbc12b43391ffb289c97fd283196c2d76147e172ab26229af1b

Observation 46ce3066-ded4-48b6-a4fc-a590c1e153e8 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.819994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.819994Z digest=sha256:4a25a8113dd18346098fefe076c2291daabe3abf31d264afa14467a1787b903e

Observation 013bb6d9-fef6-4856-a0fd-0524bf938b8f · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.883278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.883278Z digest=sha256:4bdd2ad7e8c87cff95356e5c197094c8915580fa9c0d55baf8706bd395417c7f

Observation 0c5d3d43-a58b-42ee-8789-60210e68f368 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.523614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:11.977252Z digest=sha256:3f79eb898c785edeea129576ea81cac8a4c5255c961088439d8e8ca1da6eb8fe

Observation 6f1fba65-6838-45c5-9b27-3fb226a9fba5 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.088623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.088623Z digest=sha256:9d57e62fcd4d4de30f30aa71e4eb3d9fcc021372784655ebf73075f333479e02

Observation c8cae32e-ceaa-4ea0-aeb3-6c90f0bf8e42 · outbound

This paper cites Gemini 2.0 flash.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gemini 2.0 flash

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.361409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.163832Z digest=sha256:61988a5167a6ddedbd33e47bcb184c657b89c03d923b9aacabf3f61d11258952

Observation 9254bbaa-5f05-4d51-bb8e-8dda23a2fa65 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.258440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.258440Z digest=sha256:3b2b509bca1deadef4071ffabf8466005649b72f35632e4ea58489e6888a8ba3

Observation 3cb084c7-d4b8-4bcc-b561-a4bbe46415cd · outbound

This paper cites PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.369370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.369370Z digest=sha256:9966b2577f455e099a95cddebae0f53a4393135064c95258c3578d981aff1f9a

Observation 1c41b22a-8423-4dd8-9ff9-b14b3a0b11d8 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.432203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.432203Z digest=sha256:a912913996f43e84b2252d9bea8a2e0a6ee28cacf58423828418c742dfb81b5f

Observation 5ac129a1-469f-4d0e-8853-84325d4ba9fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.504055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.504055Z digest=sha256:c3c4f69751e0f8a62efba0f9edef0ce831d478785527fde8cdb404a06ea31a61

Observation 97dabc6a-db8a-40ef-beeb-1ff6075eb843 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.564771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.564771Z digest=sha256:7ffbd95905bdc58a0b9aaea6c62e23e5621562aa59068eb9d614bb77037004f2

Observation 7a515664-b7a0-4767-a102-1e81076cb6f6 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.131160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.627331Z digest=sha256:05f77d393263d155a35eea8f1be1fe381863fa42bb41b6ee8e0eabcf51c7ae70

Observation 536242bd-79f5-4edf-b5cc-89f572b67e5e · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.680706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.680706Z digest=sha256:dbd96eb95bed0c91c048f0267fd20c928c3967b38dc76a8835a3d1d0f0baf9c2

Observation c7d45430-4670-46bf-abdf-01758b3f0373 · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.755158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.755158Z digest=sha256:7f9d38d444669451db1bdc28e1bf300fe65a134f1074d12b01209e65ab009899

Observation b065b8fc-0f30-42c5-adfa-00e05b1bc5f6 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.841584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.841584Z digest=sha256:fc092db350d7f64f4b608b900bb39033a1a33338c8f28f40991f731f7b8b9ba3

Observation ac2ac4a8-37f7-4a10-9308-141a84c5d824 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.948621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.948621Z digest=sha256:90ee6f9c4b6974bf660fef36d3124ad5d79455e41ed9bfddd1d85536edb9f3c5

Observation 537b8e27-1d48-45fe-9349-229f5710d9d2 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.035019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.035019Z digest=sha256:09dd800c5a96f1c2304779e7a37a13270fc5dd2a918d6cf84f30f17880e5f23e

Observation eeee2d81-7577-4268-aec4-0d0e8cb0bc80 · outbound

This paper cites Auto-Encoding Variational Bayes.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Auto-Encoding Variational Bayes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.131965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.131965Z digest=sha256:91cf2fcbbb1666703ce184d0a3e8924ea33d7fe67b8b66f1a95a034c0a39631c

Observation 9f0e7380-5565-4469-8545-0d46e55d4818 · outbound

This paper cites Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.239307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.239307Z digest=sha256:a911f5351a12f6db778094e50ba1715ee7bc07201f749f5f1ec8b9e9702c21b8

Observation d76b024a-e6ab-4322-a1b3-20081f7708fd · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.299556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.299556Z digest=sha256:41b6f997d89531fe8a58a74072cff2f3d019b6fc4067912dedb6a952bf70e5db

Observation d33fa69b-7a28-4909-9114-fbebdc2da3be · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.367107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.367107Z digest=sha256:08f46d18b125a8ac6a32f9fce7a0f5e4881089c44f2347e31f612db53b371f32

Observation e8fecf9f-ad82-47a4-839e-4c0e0c6197b2 · outbound

This paper cites CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.455328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.455328Z digest=sha256:158ec981cb32d2e635e8ac9de92583a4f83d4083d8819b2813d9a29ec0d92f4a

Observation e7854012-5878-4145-a2f3-f0b127f40fdd · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.554546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.554546Z digest=sha256:dff9d707856cde7edb359a5fce70c06088f0e1b2679a740dc45ddc4ad70c7df1

Observation 1edcc029-c7a2-4424-9eaa-778cfe197606 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.670112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.670112Z digest=sha256:43fc0ba33a5c1f917bbd57419d4890c8f6625d8e6b35b5db68762de0f12a8e52

Observation 96bb787e-a117-4b9c-b4d2-0490c681795f · outbound

This paper cites Microsoft coco: Common objects in context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Microsoft coco: Common objects in context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.852362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.852362Z digest=sha256:c54fa021b9a820079d80918b40488999d9bf2ea74ff5a8a4f44ab4f64727ed33

Observation 15f05b24-c517-4cf2-a712-3970452c8e49 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.931759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.931759Z digest=sha256:526b26a82a09bf40a4a4420dd75b93eb98cdd1a1946d469c0b563fa47ff9a365

Observation ef264d48-ff5c-4941-9c5c-0fde57f60077 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.745660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.038559Z digest=sha256:65d0f46df01ff2ced572a23bb4ee50a6831e374c2663f4c0394976d9fc1f37ca

Observation b3b2deda-1d67-44f3-ae94-ec315b1903d6 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.104263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.104263Z digest=sha256:f9d6e6585d4ac4e7b6bfb9b1674d7cd970ff18c3e1c49835dcdf9a4144cf8b09

Observation 29b277f6-3478-4afb-b12d-ca87597a64cc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.255400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.255400Z digest=sha256:d7823ff5ccb46c084468d0596b79c53bf803a57071721e33b9a1a75130f9f177

Observation b74d1110-41e6-40c1-9fd0-3fde59ad31c3 · outbound

This paper cites The Llama 3 Herd of Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.390201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.390201Z digest=sha256:0819caeb1e69e94547932536ace83e89242b4718a5b1190cabb7b88c319d8ccb

Observation e491591a-c5fb-4ea4-9fb1-c8ef4e85a808 · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.465050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.465050Z digest=sha256:ed7e4b55120c50054815b924c9931f87d22ebd2f53ffe9b4704cb2988b0bd60b

Observation 8b0ca092-b47e-4b6a-b559-4a2cc6f71a60 · outbound

This paper cites GPT-4 Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.544936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.544936Z digest=sha256:4ace57c391ff848ddf92eb0ed266cd58a198d6bc7d2d46995465b2c2eec394c5

Observation 6b7be0b3-184d-4b69-8eb5-b3ca8f3a3970 · outbound

This paper cites GPT-4V(ision) system card, 2023 c.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4V(ision) system card, 2023 c

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.415811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.654402Z digest=sha256:3ca97236980c541d22ba81ce9f9b5ce894b6bbff19fa4e4809bfe4238ce428b3

Observation 700e44bb-11b0-4977-9a4b-b32311c363a4 · outbound

This paper cites Dall·e 3.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Dall·e 3

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.813143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.813143Z digest=sha256:daadba4b8ae75ca57662ddfec35a19d45e4a1018daad1f18c724d09cd646abe8

Observation 786b6bb7-1fe6-4836-9248-a1b61e0b8f3a · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:22.066520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.924248Z digest=sha256:b2fce5455be1dddca8747cf230023db14c824db4a03cdabdc15e38d2c2426972

Observation ce6d28a4-dd5a-4a0d-b3b8-3b3878bf56ec · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:21.653553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.018814Z digest=sha256:b49b1c390658348f3e43b3e3ce54c7a77a60ccbb56b86b6e8e3e8e73846985f5

Observation bc58b3d3-28a5-44a2-9759-bd83c6352742 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfer between Modalities with MetaQueries

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.111083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.111083Z digest=sha256:6f0c90f54a03dfc6e1ae9ac0b00912c0564e8e5be8f1cd09b663dba153116a44

Observation 1268b0aa-837c-497b-9fbd-7c962c5667d5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.208196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.208196Z digest=sha256:a43dce44f13117d2a71dbeae1f480b73e325457cca89938ba10f948163254d8d

Observation d626d2a7-e5f8-4924-b316-49e5c9a29f2b · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.328873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.342084Z digest=sha256:7c4e7f67a3111a17cf0d5e494162da3dd35ec94e06a93bf025bccd3db60f3a30

Observation d4ce61b1-0977-44b7-943a-a93b7694132a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.466929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.466929Z digest=sha256:a0ec858e467e4885b667091a9fa1cee4f4a356daae18be93849a42137766bf00

Observation 1565b050-13d7-44d7-89b9-1d9e7d97a88b · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gpqa: A graduate-level google-proof q&a benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.560365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.560365Z digest=sha256:2fe0bde72b3b92ebf65ce1c25214ba200ae3a213274febcab1a2ba595f439785

Observation e043e2f4-65c6-4329-a8f4-45e308aabdc9 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.096417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.634551Z digest=sha256:b64e0fc06da347954f8396b81d1a8bdc6ce1c5241def3793e2ed3e697ab7e899

Observation 7e1b5794-57e4-43d6-ae34-5f736ffe59bc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.943984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.688147Z digest=sha256:ddc836f52ffafd9a85e2525018b96fa102c620047935248033c062dc6609a7c2

Observation 86ca084c-fab6-4139-a63d-6d3ac42082a7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Journeydb: A benchmark for generative image understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.762368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.762368Z digest=sha256:4aefa789416164c44092e8730944c593fbf3a9391d699ebe9428fe82a1ad3f41

Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · outbound

This paper cites Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.870519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.870519Z digest=sha256:0360373a98af270eda950b324925b6e0adedfeba4e77c9a2ecd3fc6e35e57b3f

Observation eec3108d-0d5d-42b9-8fe0-56002d0fe716 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.057133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.057133Z digest=sha256:b51fcfc07d36b3089d7211b7cac6dd783e9a6b58a165895f587a2367c75e6227

Observation 2adcf704-cc20-496a-9476-b17d332f7b92 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emu3: Next-Token Prediction is All You Need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.154776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.154776Z digest=sha256:56d88e0dd3529064e41662d0f7aea5c4581f03a98e078976a82a446230c1d7b2

Observation ccef5db5-15eb-4134-9cfd-857a8cfa91a8 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.254186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.254186Z digest=sha256:8c22c79b8eecdb123887f2b29a4d359b5cc12fa5df8072088ff433a01b245aaa

Observation 31c843ec-e643-4c38-9c8d-3749a0c7960a · outbound

This paper cites TIIF-Bench: How Does Your T2I Model Follow Your Instructions?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation TIIF-Bench: How Does Your T2I Model Follow Your Instructions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.332457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.332457Z digest=sha256:33db794a14fdf9eb1fdf50ca76bce0046c0ac702208ed40efa95a56e3dc3b8ff

Observation b202c619-3cc5-410c-a751-e1392b7e1a01 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.378281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.378281Z digest=sha256:99d26dbc58d953f0644674c23579669b075d375472989e5f43000322b6fb2e0c

Observation 276f6e35-9182-4dfc-a0df-9f322ea10bf2 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.475472Z digest=sha256:b30828d1575cd3db52d205939eb4229d2cfb51a8d7cb51b805d5e819a160e2a0

Observation 6770393b-c88b-4663-8b23-dc95edadc091 · outbound

This paper cites Less-to-More Generalization: Unlocking More Controllability by In-Context Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.671224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.671224Z digest=sha256:8d4753bfd7bd97a7805a41970d57ded096c0e9f34432c297b9d8155dd761aca3

Observation eeea55b8-05d8-4c65-b7ac-5545717ab100 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.792997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.792997Z digest=sha256:b12b92681b8b557d104055f8eaf2bfc6707bcb415947528b50beabf6f1ed8d66

Observation 22868f17-5beb-4301-b17d-3d16e94d28f9 · outbound

This paper cites Omnigen: Unified image generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Omnigen: Unified image generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.728214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:16.871964Z digest=sha256:d7d98bb80c0f9c24353da14d3b55c8f98de0d66febc6bf8291ea1336c048c224

Observation d3fbf88e-5d6a-48df-a184-85ccf5454910 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.966007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.966007Z digest=sha256:c7634f215326beb91066a746069b69bd9b41eb41f9088b85e6ecbfee7d04f645

Observation 33200aa2-a6a0-47a6-a36b-2cecc935da35 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.042652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.042652Z digest=sha256:2b6dc689ab218702b900f1e7207a18f7341abc176dafa45cf8cc0845df5d6671

Observation 6a496748-550e-4b94-9a5a-8e115ba8c9a7 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.158713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.158713Z digest=sha256:9e4634449463c67210c1c0914ee84dcda26ea7c2a79658ccad98881c4b1c4a22

Observation c0c1295e-055b-4fd9-a112-8a432002e628 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.270406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.270406Z digest=sha256:5903254532ebb3060c1ba6d36f2abd4c83139ca4fc3fcaf1b9aa4dbcd2ca24fb

Observation 8ab861fb-602b-4843-91ac-80f88e1e7445 · outbound

This paper cites The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.513604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.333380Z digest=sha256:a8a9608fd8e14b6e2311c1aeb42525f1cf77ed72072a2a10c0b60f0509dd43fc

Observation 152a51c7-8452-4b4f-a891-a2d42955a03b · outbound

This paper cites Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.319888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.442397Z digest=sha256:390abbac0329bff3a8347067412bd4021dcb650abd8a5b1095ff8b995af9504e

Observation 48755bd2-c57a-48c5-aec5-8dd61a2ccc70 · outbound

This paper cites LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.585671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.585671Z digest=sha256:fd044dc7c18107752093d159142eb626540f280ecc93d70323230c938d3288fe

Observation 8b0dddf8-44ef-465b-a070-e08ecac09a8b · outbound

This paper cites Sigmoid loss for language image pre-training.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sigmoid loss for language image pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.123840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.696061Z digest=sha256:f59023ea4f54b8e8da4a1d5eae6b5c0b56dd2c8c699b25676bdc54cdebbaf576

Observation 71c09a01-7924-40c1-98d8-da3eda218ec8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.843196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.843196Z digest=sha256:0fed1aa48e6ac3a066da1c23857e773d951d30029719d46fe7fce9c1a1e5d2fd

Observation 368ac5c2-1d7a-43de-aef5-1e26b2a2556e · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Adding conditional control to text-to-image diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.903316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.937041Z digest=sha256:09328119c4979ea2286c08528d3b359f651b672b9fedd73e19bc5c54dc933f04

Observation 1fbe3f38-a77e-41fb-a23d-3a9c4e3cf118 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.701308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.003309Z digest=sha256:15b2e6aeaea063029a7dd28eadb46cdc295d9ebdf7453964638adfd0f62bc489

Observation 1c842d2f-a386-4325-bfd3-8d7b4aa80f9d · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.084915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.084915Z digest=sha256:00b97d2aa2a1a97d8587a4994eda95060a85c012a42feb30a808b152ddf3bdc0

Observation 57992e02-9dd8-4f1c-b753-d67c0251ccf2 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.171068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.171068Z digest=sha256:a0cfcdc37c99fecdb4830723fc0286819d27fffb9aac8e8e92a06f84b195a8bc

Observation 5fe73f18-f381-49f7-b573-9eb9369a696d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.224864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.224864Z digest=sha256:b7cf76b33284928bab77fae198522476b815781d8c4230503e49e6f99bd87180

Observation 6210cb88-0136-4d27-ac1b-3b0f2b9bb64d · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lumina-next: Making lumina-t2x stronger and faster with next-dit

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.536676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.313183Z digest=sha256:3977438faf19d0f20f990890760da107ef473fb77db6fea50d5f5277deb07c6d

Observation 598ab3a7-4c0c-4cc4-a8a6-16ea8eceeea2 · outbound

This paper cites EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.395972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.395972Z digest=sha256:637935d6b51096c689cdcbda1b86e13d8b6a0edd74fcccd2b6b2616c58fe16a7

Observation 0ff93221-48a4-4c12-9888-aff39271040f · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.490290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.490290Z digest=sha256:7d1cb197788cd656c7ecf500e81bbe975343ec376b28f3847c9debfb01566755

Pith citing papers

Observation 930cac71-df1c-4ade-85a8-a73f533250b2 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.728416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.728416Z digest=sha256:d6e1fc2a889c96f4c2ba41828111df2e5657c6fc2d8a30b41e03981f74322d2b

Observation 9077d993-71cf-4d75-a590-62ba585e9563 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.408189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:efd0c98248e2b9706c76cb33bbb1bc72dad435055ce1574599e0a9f29d17cdff

Observation ee11affc-1452-4348-b2b3-13d6e7e85949 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.467647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:03bcc782ea1d455642419581a061743c0b69eb212f6ec8ff8274ee8da80f1596

Observation 3727f093-580b-42bd-b965-57a7b0d45c8e · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:04:10.562099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:00:22.875471Z digest=sha256:69fed7d9f94e8ccd8b395f271dc8de3a42872f3358632860b4fd009e4aa1cee4

Observation 4b0d842f-3bb5-43bd-b628-7734f017fe59 · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:46.874348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:46.874348Z digest=sha256:f9382d1f204a57822aa06e4ceccec1509aadaf6c634d28cd602c846cb93ee8d7

Observation 431cceef-dc42-4ef3-a2e3-06fdd4459c29 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:58.290824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:58.290824Z digest=sha256:8eca86727ef2bb3ef566fd81f31e2154806f64c5fac0309fedf6af38617088c0

Observation 87d8668a-d32e-40c4-b5ad-1c8834f1d5ef · inbound

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing cites this paper.

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:21:55.099717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:21:41.483427Z digest=sha256:64becedfecde187db22d56fe03e822f8cd0984fbaaaa52534d63dfc24cd4ee91

Observation 578ef19e-c334-4815-b475-25e8afa0b1a0 · inbound

IncreFA: Breaking the Static Wall of Generative Model Attribution cites this paper.

IncreFA: Breaking the Static Wall of Generative Model Attribution Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.504198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:35:46.381465Z digest=sha256:b79d6d4fb4ba9a140649311a6fb6d455f5cb1e10bb48c4d6e749b7085a3209aa

Observation 3d432ef3-d8fb-47df-b223-a5d3b74ead1e · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.636695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:f6f8fd5a69d715bc2104fbc79278f9a9bbd69ef1d5bf059446c4e1da4b58f0c2

Observation 3eec311c-890e-447f-aa40-f6e2275e829d · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.357764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:8cce5a6d61edba46bee1ba9b67dcb9d50d3d6b18314cb7c7ffa9f99541824795

Observation 04c45296-5970-461d-b729-da28578ca476 · inbound

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection cites this paper.

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:35.272848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T19:00:45.628715Z digest=sha256:96a31297964cc4546e697f1154f34bde5425fc5fcc92d504037057b34f75212a

Observation a69c4dbc-3e41-46e2-8883-0988202182e8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:44.257603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:fb5a9c5d8714e4f268f5f657d20920a5d29905c1fcd573559d66c65474baf047

Observation e6f3e836-e3b6-44dd-b520-5a428ec346ac · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:29.604208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:7c3b8ea641b745205b342be4e6cc34d614d2cab2853f29464d88da82d87d107b

Observation 5a634008-15f1-4d99-9bdb-a99ff3c88d2d · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.888423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:7fe134bfa2c30fdb0af0ede5ef9c925fcd549cf123517b25bf2e0d28fbc1e468

Observation 9cb5b6af-48b2-49d6-8066-3d516d37c56f · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.959544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:d9c7ee205860bedb893b68bdd328046fbf6f16ce05712a75176f331a9b03a738

Observation 8b878cdd-2cd2-48db-8273-c4b711f3440d · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.420134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:cd20912c981d1183af7163c5893e9a789caeda2c296b8192279ba4372462be4d

Observation 9f65ce3f-0c10-407a-ab3a-167a1e28387d · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.717435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:14805ec39248dcb75f903c7b0182472867974dadb89feb67d93b6e23b9b1a9b7

Observation d80e7656-d23a-42e3-8b1d-56b62a9a65ae · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.544636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:497c7afe597348a60ac4f374c30d906028944156c1848e2d9e55526cd8471040

Observation 7b84827b-75db-4d6c-856e-0c62ab14ec03 · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.705633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:10565032080db2b85090c1862868ad7bdd0ae86b47a83f068c0399da49342c5d

Observation 22edd390-9aec-43c3-90e6-85f988829d6e · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.346678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:b7edf4d3a119331f035d44ffc43fc87b824f43fe6894bc16767a4c35c1ce8c7e

Observation 820dacfe-2939-4c1e-835f-8d8cdf1b3416 · inbound

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset cites this paper.

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.730880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T04:21:05.811662Z digest=sha256:8b259fa0546854bc5b40ea661d651116eb4fa2267789c09f6a8ff8c5b3f78cd5

Observation c0059742-8e6d-4142-a723-dbbc63bd864c · inbound

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection cites this paper.

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.105435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:59:44.436264Z digest=sha256:348792f88b6daae3162084672b28afb39ee8848008a8db9b935c9b24c19ebb10

Observation d90f1426-8151-49e6-9048-e1ba002e6484 · inbound

GenClaw: Code-Driven Agentic Image Generation cites this paper.

GenClaw: Code-Driven Agentic Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.991999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T07:36:54.292448Z digest=sha256:9da2818a69496b997ea2f8f70c167306727ebd7430459b7e02c908650e185e0b

Observation c4e0c9e5-a423-4714-88cd-6dd7da9438c5 · inbound

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? cites this paper.

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.150642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T10:48:18.630777Z digest=sha256:56a60ce90e9e2c7e8824c06e2749eb088e45c484cfd9558a4adf714a339cbb15

Observation abf57ffa-361c-4e17-9e30-b478241da339 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.135045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:ab6ed5f257600b7c462cb64678d3f156eb03cd3b9d78b09412cfe6aade1c2ce7

Observation 560ea876-8511-4e1e-867e-76970259df4b · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.708346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:a2c9cae08a4069ac13ffbb0e4492f1d123ea4a49efa0d9747a1b68f690633dc8

Observation f3feccb5-939d-4652-a186-3fb74814f15f · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.702108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:19:21.433049Z digest=sha256:404642887e7171957dc2b6ec5bf7b58fd7ea63cbd61549cdc7caf3936744f5f9

Observation f665122f-5ab2-4958-a065-9e5e97710d82 · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:03:52.050125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T04:58:40.065005Z digest=sha256:8bd0f1d0785d49bdfad0b89dd7d790e720f8bb43d7f27d0693e116407db61f57

Observation 4bc666a7-0f71-40ba-82a3-fcba477715e7 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.742715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:a60ee7da078b85f743b5bfc2b493a52c1129f438ecd0e996eee842d433113ad0

Observation d10be50f-ff23-4c43-9fd5-9c5543cfb045 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.288631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:bbce38d58ee59186f921db9a7b694ed82b8639de3c72df7c672011d85040f890

Observation 6d962e50-6fa4-47a8-9724-bb4ce6867b2d · inbound

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling cites this paper.

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:47:42.650721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:47:42.650721Z digest=sha256:1ee4f07df838474bc8561bc6dfb030f39b324251423898d04297f9ec5536fc0d

Observation a22c9aec-68da-4707-8639-adf7a373f25f · inbound

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis cites this paper.

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:26.424119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:26.424119Z digest=sha256:cc8bac347d921fe9b3b85970143ce2e635a1299092d15a584f0cce0061dfd624

Observation e6e8cc14-ba48-4b0f-8e93-6b5832611c76 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.772449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.772449Z digest=sha256:49d6007296fc0b60f734bf72e76e78ee352ae1159e56e2b7c4c551c5bf75e8b0

Observation e2f9fab1-341b-4ec4-b735-2dbbf7c15ceb · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.496196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.496196Z digest=sha256:93793ca1b1fbda00b0e58790910bec4181ecf75ff1acf45e7b066882966c993c