Pith. sign in

Paper Citation Record · LEDGER

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

As of 15 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 35 inbound Pith citation observations for arXiv:2508.09987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.09987 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:44:18.490290Z

measured 122 of 122 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:41:34.355920Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.706550Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8617fb45-85f5-4cac-b6d2-2f2376b89421 · outbound

This paper cites Lawrence Zitnick, Devi Parikh, and Dhruv Batra.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lawrence Zitnick, Devi Parikh, and Dhruv Batra

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.478127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.478127Z digest=sha256:5df3e85dd4ad9ff539edfbbc5486896b239d54ad84a353349a50d235163e605e

Observation 971a41a5-30bb-4562-b464-0237a213c23b · outbound

This paper cites Sd3-medium.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sd3-medium

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.536355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.536355Z digest=sha256:72200dde73c81f9b6f65a7b037114592595f0ff9193c4dd17662ce0ea5965436

Observation 97555354-04a7-49f3-9c76-1e9d20533c84 · outbound

This paper cites Qwen Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.622883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.622883Z digest=sha256:4ef19898d0d5f5decac7c114d24e95616064ccee095e53daf4136cd7d4a95cb7

Observation b99b6dda-b6a7-4f10-985a-3a24a173b7f9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.688300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.688300Z digest=sha256:5fdb26235ff90e7e558ff997e6c073d70f4c019a0138dc5739208d3f97073156

Observation f773e167-5a88-44cd-be4e-e7ebc581fa0d · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Instructpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.745556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.745556Z digest=sha256:3d58a3171440555f812443b3ef98d38f4b09a086607ae0e6a335406ecbefcbb9

Observation 34d256f9-ab71-4f3d-a38d-cc6df7905906 · outbound

This paper cites Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Allava: Harnessing gpt4v-synthesized data for a lite vision-language model, 2024 a

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.827910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.827910Z digest=sha256:8c5a7dcdc7d82e74ae2e65ed1dd41702a2f7f516b123ea91a7e02a3de0d6e511

Observation fb06e333-9da3-4435-ba91-0b2c8208054e · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.890407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.890407Z digest=sha256:46cc45a94a16f0efcf6a5bffe21edb2ddd0d88cfc051a72b3db5ce74c91651b4

Observation 0a9f6bb9-d2f5-455e-8ebc-1323b3f2d242 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:10.965681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:10.965681Z digest=sha256:cd94d861b59591f2e2fd0bf4cd53d8c2fef7cd0266262fe181d661e04327954e

Observation 97a4601e-976b-4aee-a89a-3b446283c17c · outbound

This paper cites ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.031053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.031053Z digest=sha256:1fcd25148b346c7168aa7752c6f84eb0c266ff8b6ebe57726a1065583198dcc6

Observation cc1c951d-86d4-43dc-88e6-8323404fe7e1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.140230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.140230Z digest=sha256:22e973af6e6b3000ef2b48838cc61254adbfeca0405539ba81eb66e8e1898a9a

Observation a9b7a0c5-af64-46f3-9b24-bc9e6904c3cb · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sharegpt4video: Improving video understanding and generation with better captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.209115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.209115Z digest=sha256:0dd5ca67a6b385fe9f4f622659acc21a103352e81bcde9654509444ed128fc70

Observation 0413e382-52db-46f9-8719-9efc0395d285 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.307662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.307662Z digest=sha256:76e7b7e482ab64a98727e811ab4e13ac467f10ecdfde7818cc220ae5b9715ce1

Observation 669b36a1-b70e-4040-8632-c5031034c7f0 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.389528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.389528Z digest=sha256:309f0587bf313ab0dc5561fb2429d7e60e23aa5ecc9d1ded135ac4eff4b25eb9

Observation a496aa47-6186-4223-bf92-5645788ca429 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Training Verifiers to Solve Math Word Problems

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.474169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.474169Z digest=sha256:ad63debcfd8820b476f59549cce3b58906477b19f5bb9ba1653cc7f73102bef2

Observation e08b543a-3a34-488d-a6e8-45e8dd3a0fbf · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emerging Properties in Unified Multimodal Pretraining

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.546425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.546425Z digest=sha256:0972107d37663a289c89b9a99c59369746c346a5ed734a2d65c1ead69c472410

Observation 2495bef3-5bb7-47dd-8736-453fa1f105bd · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Autoregressive Video Generation without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.606456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.606456Z digest=sha256:0ad4d4e041bfa2bb74efb4c62c620c7060083cd8004a37cde8bfe9a0f5e9fdae

Observation 21cde564-0d55-4a8e-a3ae-3cdf3dcd7b23 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.699910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.699910Z digest=sha256:554b5771d04ca40b293a07fb3b21cc902b0c900e0a3299118b5d5e1fc16cd13b

Observation 46ce3066-ded4-48b6-a4fc-a590c1e153e8 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.819994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.819994Z digest=sha256:f1e8bff271a5c5b089f9948b1b2f9cec0948d2681276bbe475bb01bfc9b96527

Observation 013bb6d9-fef6-4856-a0fd-0524bf938b8f · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:11.883278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:11.883278Z digest=sha256:8738f69ae3393fe4e68a04748f8b0e4edce0bc18d93ba044697f2d17ea1f887c

Observation 0c5d3d43-a58b-42ee-8789-60210e68f368 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.523614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:11.977252Z digest=sha256:82cf776ef6bf4b368f0c6c31a8f12b397ed12d82edfddd2549c3014a001bc344

Observation 6f1fba65-6838-45c5-9b27-3fb226a9fba5 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.088623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.088623Z digest=sha256:5545c9f3753910d9164b390aa70ec41de4bf6a44b9c8a6d991c55241746fa5ee

Observation c8cae32e-ceaa-4ea0-aeb3-6c90f0bf8e42 · outbound

This paper cites Gemini 2.0 flash.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gemini 2.0 flash

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.361409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.163832Z digest=sha256:481682a043e04664a6c73fdbd96f440abeb88accc2440c484f2eb92970f25519

Observation 9254bbaa-5f05-4d51-bb8e-8dda23a2fa65 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.258440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.258440Z digest=sha256:6546c270e5bad7fcc5c5019d6229f7a15e9547c0404cfff82ffc140540eb514c

Observation 3cb084c7-d4b8-4bcc-b561-a4bbe46415cd · outbound

This paper cites PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.369370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.369370Z digest=sha256:50bc57838488c8462f20e529f7bf0588421c048fd3c7c7a9719e49dccde44761

Observation 1c41b22a-8423-4dd8-9ff9-b14b3a0b11d8 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.432203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.432203Z digest=sha256:af49f450957968e975495f0603d395f276f97454d48730613a5a669e05f4fc3c

Observation 5ac129a1-469f-4d0e-8853-84325d4ba9fd · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Measuring Mathematical Problem Solving With the MATH Dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.504055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.504055Z digest=sha256:319c76196eaaeabaec2cc7c92a4418be20be4a13740a8a67ae8bad14c920c4ed

Observation 97dabc6a-db8a-40ef-beeb-1ff6075eb843 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.564771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.564771Z digest=sha256:320655a27405bb12ea351a4395c1d6d08af6e0d611bdf003f189e610330cb648

Observation 7a515664-b7a0-4767-a102-1e81076cb6f6 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:23.131160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:12.627331Z digest=sha256:5548470a090674fc9e90c23c2ad0622cfcbc3ea2b6cf0bfd3a630fece9115c72

Observation 536242bd-79f5-4edf-b5cc-89f572b67e5e · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.680706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.680706Z digest=sha256:70491aafdc42e1ee439b855eb18471dc40ef5d4bccaaaca7f4f824e5aaba8cae

Observation c7d45430-4670-46bf-abdf-01758b3f0373 · outbound

This paper cites CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.755158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.755158Z digest=sha256:259aecacdda323c41bdebbe7e26b224a490e662355a269215c71dfffb7ccd0f7

Observation b065b8fc-0f30-42c5-adfa-00e05b1bc5f6 · outbound

This paper cites MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.841584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.841584Z digest=sha256:7a5ccccf26a4b8f95c558a027ec0c9396e0a5fc40fd0431f8238cf96e020eaeb

Observation ac2ac4a8-37f7-4a10-9308-141a84c5d824 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:12.948621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:12.948621Z digest=sha256:6900e40a13e9bdd703dba448d77252be436cd94d30f4621e2760157651f2abb9

Observation 537b8e27-1d48-45fe-9349-229f5710d9d2 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.035019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.035019Z digest=sha256:b2c1f10508c590c64e4d8cb10884ff6f5de29d5cf0b2bad67e7a791d9790c4cf

Observation eeee2d81-7577-4268-aec4-0d0e8cb0bc80 · outbound

This paper cites Auto-Encoding Variational Bayes.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Auto-Encoding Variational Bayes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.131965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.131965Z digest=sha256:515b7f5a380ec26b7701ba3789da375158f76419f9f9ea1455abf421f6dec8f4

Observation 9f0e7380-5565-4469-8545-0d46e55d4818 · outbound

This paper cites Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Viescore: Towards explainable metrics for conditional image synthesis evaluation, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.239307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.239307Z digest=sha256:100b9e588fc25bf1c72f6258b30cf52e4263ab6ecbcfd9582ed9870479a93942

Observation d76b024a-e6ab-4322-a1b3-20081f7708fd · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.299556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.299556Z digest=sha256:8115708e5bc58a476395a5299d3c7cf207eb9427bef5057bdf9dfcf43d2a33f5

Observation d33fa69b-7a28-4909-9114-fbebdc2da3be · outbound

This paper cites GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.367107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.367107Z digest=sha256:5c237f299ae7568ace39939d5a95c016011c57173bcddafec7b7e2c8d11673a6

Observation e8fecf9f-ad82-47a4-839e-4c0e0c6197b2 · outbound

This paper cites CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation CrossViewDiff: A Cross-View Diffusion Model for Satellite-to-Street View Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.455328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.455328Z digest=sha256:3f240a2a04c9f62b7f51920458c9ba8dc6d1354dd0cd8103357405b9200d2fc8

Observation e7854012-5878-4145-a2f3-f0b127f40fdd · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.554546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.554546Z digest=sha256:eb4439d09a6f4b87b3207f3772cf3b4301908441e4f7f0c75c34c17afe2b33f0

Observation 1edcc029-c7a2-4424-9eaa-778cfe197606 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.670112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.670112Z digest=sha256:0c2f70d899a76f6f9bc9e88dca0378b92f331f9637dd27e8c1181c105b835e05

Observation 96bb787e-a117-4b9c-b4d2-0490c681795f · outbound

This paper cites Microsoft coco: Common objects in context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Microsoft coco: Common objects in context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.852362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.852362Z digest=sha256:6bcf6960acdaf309f98390a44d80d746e601a928bf04b5943a0ff1c52ea453d5

Observation 15f05b24-c517-4cf2-a712-3970452c8e49 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:13.931759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:13.931759Z digest=sha256:4400f6d9701d9fe99bc4326ad7db9b7ca1c6aa49f563e127d236fafecc208367

Observation ef264d48-ff5c-4941-9c5c-0fde57f60077 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.745660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.038559Z digest=sha256:49f1cecdbe5df6f2a31e502a47c60b52432901818d62c162901cb21f501080ac

Observation b3b2deda-1d67-44f3-ae94-ec315b1903d6 · outbound

This paper cites Visual instruction tuning.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.104263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.104263Z digest=sha256:dcd2efefaf5bd2aafd491112a3d711d35fe89cc1b577b363c7f73737e6b9a53f

Observation 29b277f6-3478-4afb-b12d-ca87597a64cc · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.255400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.255400Z digest=sha256:2d1159b3f8a0f6cc90be19a15ec33f33e8e5b04eee3f25d3150de0cbfef165eb

Observation b74d1110-41e6-40c1-9fd0-3fde59ad31c3 · outbound

This paper cites The Llama 3 Herd of Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.390201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.390201Z digest=sha256:617e9cabaefbe3c81fbcfdcd7e868ffeb7ac94d81c80f41610dc95eb01fd134d

Observation e491591a-c5fb-4ea4-9fb1-c8ef4e85a808 · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.465050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.465050Z digest=sha256:8e4acfe770ddcf3dc28290358af1d7ba5044b8bdf46c558a3558d1a8262f13c2

Observation 8b0ca092-b47e-4b6a-b559-4a2cc6f71a60 · outbound

This paper cites GPT-4 Technical Report.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.544936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.544936Z digest=sha256:f9a66a8c08c5469bf0e0c832eac18267ebfd47c4f7fc0a1b5c42e5128078f93f

Observation 6b7be0b3-184d-4b69-8eb5-b3ca8f3a3970 · outbound

This paper cites GPT-4V(ision) system card, 2023 c.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-4V(ision) system card, 2023 c

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:22.415811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.654402Z digest=sha256:ec890465eb2e8dd4c17a4563652993f66d5542b3c66a3bfae8b3e51877db9cf6

Observation 700e44bb-11b0-4977-9a4b-b32311c363a4 · outbound

This paper cites Dall·e 3.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Dall·e 3

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:14.813143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:14.813143Z digest=sha256:a51d8ddd622e151158d8aa0a12f099ebd746619f980725a071d73940c43f030e

Observation 786b6bb7-1fe6-4836-9248-a1b61e0b8f3a · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:22.066520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:14.924248Z digest=sha256:06f98cd5be716ecb3b2fec2e6326ec9ceb7c512bb6779830a2e005b47382787d

Observation ce6d28a4-dd5a-4a0d-b3b8-3b3878bf56ec · outbound

This paper cites an unresolved cited work.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:44:21.653553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.018814Z digest=sha256:a5d70c36f1177c87d71604274a6190d04027d710bc79d7097c3852765397eb87

Observation bc58b3d3-28a5-44a2-9759-bd83c6352742 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfer between Modalities with MetaQueries

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.111083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.111083Z digest=sha256:2f62509da525a7c4e4d24476d81377f0277bb16036a3d1f99492f1f33aaaaa49

Observation 1268b0aa-837c-497b-9fbd-7c962c5667d5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.208196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.208196Z digest=sha256:298fa87a5e420f69f1f2dba1d180d08aba2ec53e7a4365f7e82e8b886f049f3a

Observation d626d2a7-e5f8-4924-b316-49e5c9a29f2b · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.328873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.342084Z digest=sha256:0edc88b93f683c93dfbecd6c9f2af735212aaa59971394256bc9328cedaf714d

Observation d4ce61b1-0977-44b7-943a-a93b7694132a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.466929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.466929Z digest=sha256:d30ea0fd6619131fa30e4d1117030adeda5034177dd42d782b669aed90f9dd4b

Observation 1565b050-13d7-44d7-89b9-1d9e7d97a88b · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Gpqa: A graduate-level google-proof q&a benchmark

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.560365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.560365Z digest=sha256:8a9bf0150a7f3bdf9bb4129ac4e08955903d755438eb680fffcbc3177a865746

Observation e043e2f4-65c6-4329-a8f4-45e308aabdc9 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:21.096417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.634551Z digest=sha256:039f34836c138716beb20ed5807ce6de57a4c59fcbe97200de67786ce06865ba

Observation 7e1b5794-57e4-43d6-ae34-5f736ffe59bc · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation High-resolution image synthesis with latent diffusion models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.943984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:15.688147Z digest=sha256:2e26e0b39de9aa02766e5b7b6627962d8f10a667dd70d62e71fc901ca778a1f3

Observation 86ca084c-fab6-4139-a63d-6d3ac42082a7 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Journeydb: A benchmark for generative image understanding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.762368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.762368Z digest=sha256:3faf69c4acbaa2a582e91a8068087837822255d92bfde5ab9d3347ae5eded85f

Observation 0751e862-9eb7-42f8-90f8-edd4ff0e57a6 · outbound

This paper cites Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:15.870519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:15.870519Z digest=sha256:38035954745bf425096716aad66baf30aec39baf8bc6a15abc8db0ef76140298

Observation eec3108d-0d5d-42b9-8fe0-56002d0fe716 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.057133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.057133Z digest=sha256:e046a1e3631b5161fe44aaddd628803f41ad11fd520f39872fdf07e9a459a529

Observation 2adcf704-cc20-496a-9476-b17d332f7b92 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Emu3: Next-Token Prediction is All You Need

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.154776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.154776Z digest=sha256:3e681e5d4db15cee94a06639657ad83b8bd1978ab5c915c980ac1c44aaea79a2

Observation ccef5db5-15eb-4134-9cfd-857a8cfa91a8 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.254186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.254186Z digest=sha256:dd65395a89c9bc10f461c414c75277ef4a4bb60a6b7a937f8888a7a5ba17a253

Observation 31c843ec-e643-4c38-9c8d-3749a0c7960a · outbound

This paper cites TIIF-Bench: How Does Your T2I Model Follow Your Instructions?.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation TIIF-Bench: How Does Your T2I Model Follow Your Instructions?

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.332457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.332457Z digest=sha256:692ca4e8557ec01eb07f37fcaac4802cb0264ff33dded133bdeabe4ba5f63233

Observation b202c619-3cc5-410c-a751-e1392b7e1a01 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.378281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.378281Z digest=sha256:d9674f420f8caf506e9d9db80b22e620f29b1cad07e73fa71fe3027c403030d8

Observation 276f6e35-9182-4dfc-a0df-9f322ea10bf2 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.475472Z digest=sha256:38bea404031620698cc9ab34fc2f932830b72b5da2d9c979f7737b34ccc689b6

Observation 6770393b-c88b-4663-8b23-dc95edadc091 · outbound

This paper cites Less-to-More Generalization: Unlocking More Controllability by In-Context Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.671224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.671224Z digest=sha256:0bb5f360ef452cf0a0e82c974a4dd076235cf4d06a9c4849f8917a9cbec3157b

Observation eeea55b8-05d8-4c65-b7ac-5545717ab100 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.792997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.792997Z digest=sha256:5d80f4f749d46d6d4e3a35f0175333fd6b1831d0c77f68dd0af3aef677556f61

Observation 22868f17-5beb-4301-b17d-3d16e94d28f9 · outbound

This paper cites Omnigen: Unified image generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Omnigen: Unified image generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.728214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:16.871964Z digest=sha256:7d9369deab765813408764de2ff0757de323dd97607bd513f38f46beb14470d7

Observation d3fbf88e-5d6a-48df-a184-85ccf5454910 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.966007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.966007Z digest=sha256:51d61d1965b26c087c42d346e8f2491030a9245d757fceb90488252ec122caea

Observation 33200aa2-a6a0-47a6-a36b-2cecc935da35 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.042652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.042652Z digest=sha256:352f66d4c83a91c768840100ca2f7c4b923d616e69d8e97998f45c21b969c583

Observation 6a496748-550e-4b94-9a5a-8e115ba8c9a7 · outbound

This paper cites PointLLM: Empowering Large Language Models to Understand Point Clouds.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation PointLLM: Empowering Large Language Models to Understand Point Clouds

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.158713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.158713Z digest=sha256:483b30459c5dffdda44e8c236ebc2e427d53562d18a50a2af22f65bf47b09eb1

Observation c0c1295e-055b-4fd9-a112-8a432002e628 · outbound

This paper cites GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.270406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.270406Z digest=sha256:d6f89c273f15d4e563a5060986095c3203d9f0b8a5b083ad38df3ae2be4c036e

Observation 8ab861fb-602b-4843-91ac-80f88e1e7445 · outbound

This paper cites The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation The fabrication of reality and fantasy: Scene generation with llm-assisted prompt interpretation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.513604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.333380Z digest=sha256:6c9a5114dff45870b26d230ac8ef5fe6643c5d5d5f95b04d0a2687940627c101

Observation 152a51c7-8452-4b4f-a891-a2d42955a03b · outbound

This paper cites Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Skydiffusion: Street-to-satellite image synthesis with diffusion models and bev paradigm

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.319888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.442397Z digest=sha256:688100dc5dcc2c55e390dad7a6db61ce78b60c3be6a0e7d0ce02b8fc38023dd2

Observation 48755bd2-c57a-48c5-aec5-8dd61a2ccc70 · outbound

This paper cites LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation LOKI: A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.585671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.585671Z digest=sha256:36c9b4e9cb27d5d8891cb722469cc73ff7ea7d0634eb4f9f94a69c03d385a40a

Observation 8b0dddf8-44ef-465b-a070-e08ecac09a8b · outbound

This paper cites Sigmoid loss for language image pre-training.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Sigmoid loss for language image pre-training

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:20.123840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.696061Z digest=sha256:77a48602cc647dfad2567023bdb9aaf43a59a6e4eff72b75475c1cb1d8506e5b

Observation 71c09a01-7924-40c1-98d8-da3eda218ec8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:17.843196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:17.843196Z digest=sha256:1050576b3c863106163012fa305eee8b3b28879ba2005ea05ac638af82f9f6a6

Observation 368ac5c2-1d7a-43de-aef5-1e26b2a2556e · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Adding conditional control to text-to-image diffusion models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.903316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:17.937041Z digest=sha256:bf8dacec6a7017c500857d51794df7a2207771fcd081f97f5f5082f9d090da47

Observation 1fbe3f38-a77e-41fb-a23d-3a9c4e3cf118 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In European Conference on Computer Vision, pages 169--186

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.701308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.003309Z digest=sha256:80aa0b270d2da9635a05b3eba7fff2c74a75b6dd5f066e10d11b34c762e55aba

Observation 1c842d2f-a386-4325-bfd3-8d7b4aa80f9d · outbound

This paper cites MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.084915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.084915Z digest=sha256:c1fec1be92b8bd3394f3c059437cebdacfe589f6bf61f16ec370c90c7ca6b029

Observation 57992e02-9dd8-4f1c-b753-d67c0251ccf2 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.171068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.171068Z digest=sha256:087cd480a11bc562ccc281b0c4fb0c252b90c0a8bd08773bf77a067249ed6390

Observation 5fe73f18-f381-49f7-b573-9eb9369a696d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.224864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.224864Z digest=sha256:183bfc44366191deff1ab839e3308f551ce2e62392b66c798fb5260ef6d0ad2b

Observation 6210cb88-0136-4d27-ac1b-3b0f2b9bb64d · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation Lumina-next: Making lumina-t2x stronger and faster with next-dit

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:44:19.536676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-05T20:44:18.313183Z digest=sha256:c2390a2af3deac4e8157e359c2c8ea40837d118fa6ee64c70be592363c0b078a

Observation 598ab3a7-4c0c-4cc4-a8a6-16ea8eceeea2 · outbound

This paper cites EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.395972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.395972Z digest=sha256:a5d0b1c9c2f0c5a91b76bcedecb16d0285f47ea31484a95ff4cdcf19585f1f0d

Observation 0ff93221-48a4-4c12-9888-aff39271040f · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:18.490290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:18.490290Z digest=sha256:4f585d9d862f07af6e60da605c541f276e84fd9af899ba39749d1a0d6f131dd0

Pith citing papers

Observation 930cac71-df1c-4ade-85a8-a73f533250b2 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:10.728416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:10.728416Z digest=sha256:ced4f47fc3001867ed65820563c35346eb360a3e8f467e8b8ce8b8fd47f3d407

Observation 9077d993-71cf-4d75-a590-62ba585e9563 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.408189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:664e7a3b9fe632bcb32b824de3826def0f2f36f291445ae9cc69b6f3cb8f1fd7

Observation ee11affc-1452-4348-b2b3-13d6e7e85949 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.467647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:8c2c7a1b90a908c3769f18902ab96f26ca0ad3b71c5d36fcf981f07ab804ce07

Observation 3727f093-580b-42bd-b965-57a7b0d45c8e · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:04:10.562099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T13:00:22.875471Z digest=sha256:5cd4ccfc15a1b27ab154886c3600490af91606903c1bbb62b98582fc14819343

Observation 4b0d842f-3bb5-43bd-b628-7734f017fe59 · inbound

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation cites this paper.

Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers for Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T03:50:46.874348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:50:46.874348Z digest=sha256:89d48e6aa672bfafe6faf8046f715ea7baebfecef19e54c8b7b2b28c6af05aa2

Observation 431cceef-dc42-4ef3-a2e3-06fdd4459c29 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:58.290824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:58.290824Z digest=sha256:6f8dedf330f017e7d6099524154d026295e10833efea1994c137d1a2dcc71581

Observation 87d8668a-d32e-40c4-b5ad-1c8834f1d5ef · inbound

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing cites this paper.

LIVE: Leveraging Image Manipulation Priors for Instruction-based Video Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:21:55.099717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T07:21:41.483427Z digest=sha256:e6c07be5b8eec2fa3201843adb16ef8a3c454d2ac673f19798a9252825b9be20

Observation 578ef19e-c334-4815-b475-25e8afa0b1a0 · inbound

IncreFA: Breaking the Static Wall of Generative Model Attribution cites this paper.

IncreFA: Breaking the Static Wall of Generative Model Attribution Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.504198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:35:46.381465Z digest=sha256:c60877eed3f8c3bbefcefa1e0ab76de18b94df0024f91c0f29da4abe53fc26b7

Observation 3d432ef3-d8fb-47df-b223-a5d3b74ead1e · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.636695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:e5f5f5a8cf0b2aaaeb49c14bf64bc10dc50497c0c9c21b77545a4357b7778b69

Observation 3eec311c-890e-447f-aa40-f6e2275e829d · inbound

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model cites this paper.

LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:49:48.357764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:49:38.156237Z digest=sha256:5b94bed201411b402b6588d092786c780b7e76afcd7bab3e43d2b243d744353e

Observation 04c45296-5970-461d-b729-da28578ca476 · inbound

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection cites this paper.

Automated In-the-Wild Data Collection for Continual AI Generated Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:35.272848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T19:00:45.628715Z digest=sha256:a33e1de7cdf62731620a583cc3e85bbd509fea9e30183b6c72b4d26255f19ad2

Observation a69c4dbc-3e41-46e2-8883-0988202182e8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:44.257603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:504975ea407cde7f99321b86c772793847d289188ffb2f2f1fcae715f3b723de

Observation e6f3e836-e3b6-44dd-b520-5a428ec346ac · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:32:29.604208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:3f29a0ffdb36727cdcffd25d54ef9571bcafaf0d143b64b8989ea97db10f2256

Observation 5a634008-15f1-4d99-9bdb-a99ff3c88d2d · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.888423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:22cfe044d2a95e45491c88b412fdb23f3ab9df2bb72fd234d507236faaa80bfe

Observation 9cb5b6af-48b2-49d6-8066-3d516d37c56f · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:23.959544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:0160e08e3f7e6e6458a6b7b28b8f196240df34d7c9246bc8530743006358d361

Observation 8b878cdd-2cd2-48db-8273-c4b711f3440d · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:03:03.420134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:96518fc99e91073cdce7173da77fb32a8a4b5eb30be68f78419944e512fa4722

Observation 9f65ce3f-0c10-407a-ab3a-167a1e28387d · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.717435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:76724e70c84150507b7c56492ceb252c54d79f74d97f4be98f8172e08aa4bc8f

Observation d80e7656-d23a-42e3-8b1d-56b62a9a65ae · inbound

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation cites this paper.

InsightTok: Improving Text and Face Fidelity in Discrete Tokenization for Autoregressive Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:13:30.544636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T02:10:18.004724Z digest=sha256:cd6fb8beb61cb34735308cbe37c3d5bf165d9846a43ce8a5e5207aa06dd0efe6

Observation 7b84827b-75db-4d6c-856e-0c62ab14ec03 · inbound

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning cites this paper.

Unlocking Complex Visual Generation via Closed-Loop Verified Reasoning Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:37:39.705633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:35:43.166697Z digest=sha256:ded5785753f3415d29367f2ded85d9eda7c663d15fd8ca0ebd168ecf40d7aed6

Observation 22edd390-9aec-43c3-90e6-85f988829d6e · inbound

TextSculptor: Training and Benchmarking Scene Text Editing cites this paper.

TextSculptor: Training and Benchmarking Scene Text Editing Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:19:39.346678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T05:16:43.756525Z digest=sha256:39cebee15efac63ac41b862f81a2f0e422af5dee043edfdd2b77df338ef1cc9f

Observation 820dacfe-2939-4c1e-835f-8d8cdf1b3416 · inbound

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset cites this paper.

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.730880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T04:21:05.811662Z digest=sha256:c4dd6eede597ba291dfd3f970db4aedb297c566e502c727c33204754f16dd521

Observation c0059742-8e6d-4142-a723-dbbc63bd864c · inbound

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection cites this paper.

FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:14.105435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T07:59:44.436264Z digest=sha256:6b01bb87febb7291f6b1997b25c7759417551a719357f39cdf5dcaedc71ebc7a

Observation d90f1426-8151-49e6-9048-e1ba002e6484 · inbound

GenClaw: Code-Driven Agentic Image Generation cites this paper.

GenClaw: Code-Driven Agentic Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.991999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T07:36:54.292448Z digest=sha256:9735a1bbfc36232dec8b22ef8ced2d7e85de62f9e959103d9a718deb316f7fbe

Observation c4e0c9e5-a423-4714-88cd-6dd7da9438c5 · inbound

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? cites this paper.

MemoGen: Can Past Experience Improve Future Text-to-Image Generation? Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.150642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-28T10:48:18.630777Z digest=sha256:8442451f487a4dabf43704514e845729518440a12ab1badd08945dafbd422e1e

Observation abf57ffa-361c-4e17-9e30-b478241da339 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:58.135045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:d361f0fa7c9495c49ddf8fdbb45a71e769d7a6a7111827048180113b1a6507f6

Observation 560ea876-8511-4e1e-867e-76970259df4b · inbound

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation cites this paper.

DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:10:09.708346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-25T19:00:23.260939Z digest=sha256:210af6c5d3b63686c523f6e7f6d1767f6c5742d8a21c75b49c0fc02bf5bc2582

Observation f3feccb5-939d-4652-a186-3fb74814f15f · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.702108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T05:19:21.433049Z digest=sha256:39882ba2f5ddd9f6cc56fc74215479a6ef045bc9e400b1dd6ce62b5cd65a801f

Observation f665122f-5ab2-4958-a065-9e5e97710d82 · inbound

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation cites this paper.

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T19:03:52.050125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T04:58:40.065005Z digest=sha256:fc8b6f2baba6243566b4ec0ec9f6219f6013392640097d0579bf29ac740979c3

Observation 4bc666a7-0f71-40ba-82a3-fcba477715e7 · inbound

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation cites this paper.

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:54:20.742715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T06:48:22.594002Z digest=sha256:5beb30c1a3a3edcd838e9c7f7399431f3de6346b3fcc82a2391c52e69fd3df15

Observation d10be50f-ff23-4c43-9fd5-9c5543cfb045 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.288631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:c991880a98b3e6a132c389676e844c7ef4986e6fa701b9e27100c1ad77e564a0

Observation 6d962e50-6fa4-47a8-9724-bb4ce6867b2d · inbound

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling cites this paper.

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:47:42.650721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:47:42.650721Z digest=sha256:4d492c75f8f7633a164672955bbccd34982b5965b3f2ac7e6bb484932beb5309

Observation a22c9aec-68da-4707-8639-adf7a373f25f · inbound

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis cites this paper.

ExpertVerse: A General-Purpose Benchmark for Expert-Level Reasoning in Knowledge-Intensive Visual Synthesis Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T12:46:26.424119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:46:26.424119Z digest=sha256:597b16aa3eb57e8049eb1ce0adebf33984db76c6519826525a4e650bbfc0127c

Observation e6e8cc14-ba48-4b0f-8e93-6b5832611c76 · inbound

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation cites this paper.

Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:14:09.772449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:14:09.772449Z digest=sha256:1b133cb3bf6d0785494dd25cc9705393819f9dbca869ae2bbab9f27d4342a4d3

Observation e2f9fab1-341b-4ec4-b735-2dbbf7c15ceb · inbound

Amortized Moment Matching for Visual Generation cites this paper.

Amortized Moment Matching for Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 151

Resolution
unresolved
no resolver link, observed 2026-07-30T18:58:28.496196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:58:28.496196Z digest=sha256:22f0a19c4cc83258128c51957e50082c712ffca676a98bc4238dc2f4c36803c4

Observation 1e3c2a70-445c-49fa-a886-e9efa78515e5 · inbound

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework cites this paper.

Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:41:34.355920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:41:34.355920Z digest=sha256:89853e63b3d8f6751542fa3ae0b8abbc42713d2a7de5f84aab7fc4c597df68f4