Pith. sign in

Paper Citation Record · LEDGER

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

As of 15 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2411.16749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16749 v2

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:03:29.034095Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:02:55.814289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T07:39:50.478310Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact2
  • verified fuzzy35
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c692dc30-8d04-46fc-bdb9-69b11c94c281 · outbound

This paper cites Qwen Technical Report.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:26.456495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:26.456495Z digest=sha256:01867f572dd3d4a5de73f1a4f72721921c911d824c90da9e4249e6c3c8466f2a

Observation 05e9e2ff-15f9-4df0-9a12-3885a126b03b · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:26.521824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:26.521824Z digest=sha256:8187ff8234aa534037b8c4d859366e9767c0d3fe6f542e18273d2e418cadc785

Observation 12fd3cd8-e7a3-4c30-ac7b-2e13521fdd51 · outbound

This paper cites Zero-shot composed image retrieval with textual inversion, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Zero-shot composed image retrieval with textual inversion, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:34.462169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.604774Z digest=sha256:8982d88c35845dc793dc34f5e631f71ab9558cdfcd5d8235f7a21ee596242fb1

Observation 1fceac45-1586-4f64-9669-e0b3e690454e · outbound

This paper cites Murphy, William T.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Murphy, William T

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:26.623147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:26.623147Z digest=sha256:b40648371f4d1c074531b79badb26034540f163b60bfeb2edc04ef50c1247df4

Observation f9676e7d-26e8-41a9-ace2-ad83eeb116f1 · outbound

This paper cites Training-Free Layout Control with Cross-Attention Guidance.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Training-Free Layout Control with Cross-Attention Guidance

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:26.633424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:26.633424Z digest=sha256:7743126081d831b3985e65ce7276a7064d7669a93582279d43942ff4454e3e13

Observation 30d5f4ab-e8c4-4845-80e6-4428c1afc9c1 · outbound

This paper cites an unresolved cited work.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:03:34.246347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.696735Z digest=sha256:97da880dfc4786d0a3478b4443649a0ee7a473638a22aaae4d62466eb54b545f

Observation 86fb60b6-fb77-4d46-ad0e-0327c9ca0e31 · outbound

This paper cites Auto cherry-picker: Learning from high-quality generative data driven by lan- guage, 2024.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Auto cherry-picker: Learning from high-quality generative data driven by lan- guage, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:34.207549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.821036Z digest=sha256:fc40f7d5940443bc56b7c6e91dc4914e7f13bfc6c0b069fc8401ad4348661a46

Observation 57cbed7a-a7c8-4af1-95c7-dad21678ad53 · outbound

This paper cites Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Style injec- tion in diffusion: A training-free approach for adapting large- scale diffusion models for style transfer

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:34.163509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.897043Z digest=sha256:8e74291e63a362e3a3f3a00558a18c83cfc47034549e87740e323c9cb657ac9c

Observation 7fbaf30c-0d5d-4a4f-b529-a8115df8bdb5 · outbound

This paper cites Roboflow 100: A rich, multi-domain object detection benchmark, 2022.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Roboflow 100: A rich, multi-domain object detection benchmark, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:34.108061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.920948Z digest=sha256:fc947e46b4a9266e973ab59fd456e4c8d42b6816725a234978ef36d04651d84a

Observation 725a48ba-ac3f-4282-aa4a-a4760ac979ef · outbound

This paper cites Diffusion models beat gans on image synthesis, 2021.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Diffusion models beat gans on image synthesis, 2021

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:26.932515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:26.932515Z digest=sha256:4d2178939bf86c3bf9ccedea912fb06fb1e054c8b817bfa5d494df32fb5234fe

Observation 2b2d9cbc-24d2-422d-8402-23e9f07a5016 · outbound

This paper cites Hierarchical text-conditional image generation with clip latents, 2022.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Hierarchical text-conditional image generation with clip latents, 2022

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:34.027916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.946007Z digest=sha256:7f481ed379b94481e2b3c0a7b5d437e6f4b628836c6fc8fa1a49b9ebb0e8ff2c

Observation 4331ec7b-9037-4d7d-9f39-83d9892c0d0b · outbound

This paper cites Everingham, L.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Everingham, L

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.890273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:26.997032Z digest=sha256:0cad802da0dfc217705b819904652718fabecb3b0dd2fff1e0967c54f26241b6

Observation 3bfac62c-0fa0-45e1-b224-9716fe6e27ea · outbound

This paper cites Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Mme: A compre- hensive evaluation benchmark for multimodal large language models, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.088187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.088187Z digest=sha256:2d1742bb52174f18ab93ccbfcae240ed3d69d5659c87c7a6ac9b68d5891e8327

Observation 4bb681df-7e46-47aa-895e-db80bd997c94 · outbound

This paper cites Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Cross-Domain Few-Shot Object Detection via Enhanced Open-Set Object Detector

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:03:30.502490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.118808Z digest=sha256:8f2b1bd2579ec28c0a53cf403088a7fc51acd601b23ac21d7aa574817e933477

Observation b384a23f-7d64-4873-825a-a005b7c34e3d · outbound

This paper cites Apollo home page.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Apollo home page

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.746541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.133870Z digest=sha256:5a958cbe5829b94dc1478b4ebc5716402374790c44fae24af9eda2c865bd7c00

Observation fc22c40a-94d4-4411-ab47-e9b6483d592e · outbound

This paper cites Classifier-Free Diffusion Guidance.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.143840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.143840Z digest=sha256:bdc783395bc2d2b3d5815c246baf6aac4742ede26b4fda4ade91b02e0d7ceb04

Observation 5d6d2f82-53ec-40de-b0ed-6d409125ef26 · outbound

This paper cites Denoising dif- fusion probabilistic models.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Denoising dif- fusion probabilistic models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.218020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.218020Z digest=sha256:a606b15aef8b1a0069df819136f6cf522b6badb703ae4f9ec372ffdd6b6a6fc8

Observation b98fe489-8dcc-4ac2-9b0f-810d4a896028 · outbound

This paper cites Egtr: Extracting graph from trans- former for scene graph generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Egtr: Extracting graph from trans- former for scene graph generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.622574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.314114Z digest=sha256:bba74337dc531ff80bcfe5b30b8f7af0ee95284e5e6ad2c6b4a33d852ca1f1d6

Observation 4fa718c8-11e7-49a5-9a2f-cc964bfb94a5 · outbound

This paper cites Cross-domain weakly-supervised object de- tection through progressive domain adaptation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Cross-domain weakly-supervised object de- tection through progressive domain adaptation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.556492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.331224Z digest=sha256:4bbf8bcf64459db479697fa7a6f201fd286da43d9b3fc7bfd3451e8cdaaf84c0

Observation 9fea8582-4a9d-4cba-be0c-4c3f05fe742c · outbound

This paper cites YOLOv5 by Ultralytics.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks YOLOv5 by Ultralytics

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.502944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.355230Z digest=sha256:9d46d6cf387bd14e2aed24ae48743e85bd5bded0cfe4bc1a4a9d1264f28dd023

Observation 0bd6db3c-721e-42d1-82a3-881ec9ed0645 · outbound

This paper cites FOCUS: Familiar objects in common and uncommon settings.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks FOCUS: Familiar objects in common and uncommon settings

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.443212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.371263Z digest=sha256:2e31ec893e92f60e14922a08166d1d02531f1b50583afd39cdaee238e59af484

Observation d1e5a5f1-0f2c-4d5a-9c33-08189376a4d2 · outbound

This paper cites Segment Anything.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Segment Anything

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.408609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.408609Z digest=sha256:40a46da5a22b4ec0ab2ce882293997eef501549db5623ce1c4d2ecca302aaa78

Observation d990bf1c-0b0f-4314-8820-2da0ebba2b21 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Pick-a-pic: An open dataset of user preferences for text-to-image generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.471889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.471889Z digest=sha256:d90bbac18e7f8700b60fa8d3590b086656f6ec5c9c777e4f671dc03f60a4f784

Observation 6604a584-c5e4-4b53-aa79-057edbbcdd07 · outbound

This paper cites ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks ELEVATER: A Benchmark and Toolkit for Evaluating Language-Augmented Visual Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-12T14:03:30.305941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.535050Z digest=sha256:bcc3fb7e5cff7ed2bd004675ad0be6325cb0f2e50a92ec0f5baecdaa18bee7ae

Observation 94809d2d-bf34-453e-acf5-884c507b947c · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.552055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.552055Z digest=sha256:2ab974ef153848833d8ff0a76360d1cbac225b94eb311bcb60111c6fc7d61f70

Observation 672024d6-6fdd-40d0-b38d-b47ee26ff9a8 · outbound

This paper cites Grounded language-image pre-training.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Grounded language-image pre-training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.259659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.567422Z digest=sha256:b69502579b517e4376bc39f7ea9b0c3efeb9a2cd126766f1329af8ccce8332b7

Observation 8e810f95-7c1a-4a0d-9de2-aa8f2c5990a1 · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Gligen: Open-set grounded text-to-image generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.580116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.580116Z digest=sha256:c1e995ad4f971d6cd1b0f58d0f18f63819b5980e64aceadf0f35d4c66e018033

Observation b6400b1a-6787-4076-92df-1fd2960d76d1 · outbound

This paper cites Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.632029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.632029Z digest=sha256:e85200578a2e9f929d79f62ee6e206cc928a0a7b0ac0782408ffcdc7240c0fc8

Observation a757d1eb-0aec-482d-88ec-ae79e205b068 · outbound

This paper cites Caphuman: Capture your moments in parallel uni- verses.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Caphuman: Capture your moments in parallel uni- verses

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.698541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.698541Z digest=sha256:b84d59f28bbecc21949177ff1020b4f8206a0cc0cb7f1fd9f2e0122e0665406e

Observation d680e15a-5900-426e-a21e-67303e07ebca · outbound

This paper cites Lawrence Zitnick, and Piotr Doll ´ar.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Lawrence Zitnick, and Piotr Doll ´ar

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.093601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.720318Z digest=sha256:2fcec8fd6a8e315abb629b0e030bc7d32aa60b8b4d7f08de30c52bea6d4985ad

Observation 868fa835-a716-4f6f-9411-8aee9c97ffde · outbound

This paper cites DIAGen: Semantically Diverse Image Augmentation with Generative Models for Few-Shot Learning.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks DIAGen: Semantically Diverse Image Augmentation with Generative Models for Few-Shot Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.739980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.739980Z digest=sha256:9670cf58d15d9f75f423da49b184981a5f6b04f0101028521eb2c44b8e6d141c

Observation 02d73e1a-f0b5-413e-afc0-d045992f6ac3 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Improved baselines with visual instruction tuning, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:33.033757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.755471Z digest=sha256:fb3eae6fac77618e5c3e4f62e4bdd6058d8ee233969cc394083041f03672c71f

Observation 68a6655e-c7ba-48c1-9842-8fd608740192 · outbound

This paper cites Visual instruction tuning, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Visual instruction tuning, 2023

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.823858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.823858Z digest=sha256:ab476142cfc1e25e31e4a3e0359b423e9bcb0796e6ab97869b0009f6a719838f

Observation 8906a387-71f8-44dd-805b-c5b6a8bf93a2 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.883139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.883139Z digest=sha256:d2fb0691dfe785172ef3083a461db5a51f1a042ab7f2fb0f6c9dc2abe6858c8c

Observation 6cccad7c-37c5-4f6f-89c5-8a0898859cbd · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.893766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.893766Z digest=sha256:4fd16b272bd37ef0b319e3b06af9046f9b5eb949ccb902326bbee40dc68dcca9

Observation c0510e0d-9415-4383-9f60-4f2291ab2036 · outbound

This paper cites Image retrieval on real-life images with pre-trained vision-and-language models.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Image retrieval on real-life images with pre-trained vision-and-language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.897542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:27.902034Z digest=sha256:8cc0c58e216c7d2640af3d1abcadb9f133bfd4d649fe9d32386e1fd66cdeaacd

Observation 92301f46-ef64-48b7-b9d5-80f8c10b6a9f · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.913396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.913396Z digest=sha256:64f4e09ba8fa236a476e2a9d81e77c8f44dc818a4a7555bbc1da57e046ea7a86

Observation f5d3620f-a2bc-4caa-9616-117ccdcbcb9b · outbound

This paper cites Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Glide: Towards photorealistic image generation and editing with text-guided diffusion models, 2022

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:27.993872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:27.993872Z digest=sha256:fef2088a9e6cf4d86532ffb7b9a916fa059547f75ee406718329f445fbabbc23

Observation 9081faca-b818-44b1-80a7-f9976f73494b · outbound

This paper cites Gpt-4 technical report, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Gpt-4 technical report, 2023

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.819405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.072563Z digest=sha256:ca95ef642c42343b57b7832d5e2cce932324b83ec023a78f596d029e1337adaf

Observation b9b9f30c-02a5-4ecf-8ea5-926de6c119f0 · outbound

This paper cites Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.753579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.098326Z digest=sha256:0411259722140e0e1eec68e3e252ba01630f29c14d2858a214b5df560358fb0f

Observation 77bdaf5e-5f3f-411d-a735-ea8d85940848 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.583296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.105091Z digest=sha256:c24bf1abecd4dd6db3c9fb36e4d04fa9068660f7d5a021bdfa2d3b0f242396da

Observation 87495f66-569f-4ae2-9092-b39838d75d2b · outbound

This paper cites Generative ad- versarial text to image synthesis, 2016.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Generative ad- versarial text to image synthesis, 2016

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.547836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.125663Z digest=sha256:657a9f340ef9f938fcf63bba3c95b7041ab845d0d483c467f96fa917a0987ec3

Observation 1d5a05aa-53b5-45ad-a7fb-bb717b5f2590 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks High-resolution image syn- thesis with latent diffusion models, 2021

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.225160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.225160Z digest=sha256:d108068035bd1e2f1a488a2a5bc24affe0cba1356bb0f793d70d937accd4434b

Observation b9275d70-af62-44e8-a70a-15ef44e87a99 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.276830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.276830Z digest=sha256:cfafdfa2fdad1893355ef037fef984badca4f250baf3c441d1bc3e638fa7f54b

Observation 66b0a1ed-7dc2-46f1-b642-0a0b7da3ce48 · outbound

This paper cites Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David Fleet, and Mohammad Norouzi.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Denton, Seyed Kamyar Seyed Ghasemipour, Burcu Karagol Ayan, Seyedeh Sara Mahdavi, Raphael Gontijo Lopes, Tim Salimans, Jonathan Ho, David Fleet, and Mohammad Norouzi

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.311524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.311524Z digest=sha256:38dd1a72241d6718d9447dbdb3b1e88d2ada5b0b4605188383be1cfbb20d41de

Observation 05aeeb2c-1ae7-460e-b985-211d8ff29cea · outbound

This paper cites Denoising Diffusion Implicit Models.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Denoising Diffusion Implicit Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.325435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.325435Z digest=sha256:6a370c1ae3d7ef17161c15cc04767d32587d93c68bd24362c8118b1739b70822

Observation e2ac4473-ffae-4975-bfb9-b9eb4824860e · outbound

This paper cites Gen2det: Generate to detect, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Gen2det: Generate to detect, 2023

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.384266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.335230Z digest=sha256:059fe618414bd9be1e0801a45e31f201bb1658dcfdaae08a5a3a17320e9f8f25

Observation 44feffb4-403d-4823-a13d-286d07063256 · outbound

This paper cites Effective data augmentation with diffusion models, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Effective data augmentation with diffusion models, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.267410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.344898Z digest=sha256:1600f4fa6562bccfde5178ee357dc56ab48a9e102cbd7938e421e5671967d049

Observation 630417ef-43f3-4def-bb5e-3f5a037b4189 · outbound

This paper cites Attention is all you need.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.356046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.356046Z digest=sha256:504f4dca4ebddb6934b8d1e6c286c329c1e0db75f5ff237bcbab816b61ac03f1

Observation 0c230471-ae9e-4077-bead-c08e9c77ba6b · outbound

This paper cites YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.170435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.366056Z digest=sha256:98614db8ffbf4fdf233cb7843a06c3fbfdba1a1b5bc2860cb4eb4a1691ce1d14

Observation a2c6d3a8-2416-485a-a91f-d8128744806a · outbound

This paper cites Learning from synthetic data for crowd counting in the wild, 2019.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Learning from synthetic data for crowd counting in the wild, 2019

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:32.106032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.414014Z digest=sha256:88b8f3d2b2cf8ef092c1b0ae8ebb36547b49f23e5fc9f9dd9e0c89bf45c587fa

Observation d2519556-a068-41a8-b835-fa463d309491 · outbound

This paper cites Instancediffusion: Instance-level control for image generation, 2024.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Instancediffusion: Instance-level control for image generation, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.493864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.493864Z digest=sha256:ca2a48b550c0555418cfeab9af034421b63d5a316836d39988740e681bca5c6a

Observation 514d7a75-729a-4785-8607-2ed0d3bb1d8b · outbound

This paper cites Fine-grained prototypes distillation for few-shot object de- tection.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Fine-grained prototypes distillation for few-shot object de- tection

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.893574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.533632Z digest=sha256:e74325ae8c224ba3c450a391fb9333e8a1e283c1027ebb679eda22e0c3971820

Observation 6cdd4fc8-b4e4-4bd1-b7d3-b37a6a950138 · outbound

This paper cites ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.547482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.547482Z digest=sha256:783da23f86a338dc1d782d6bf9550c1ecb64a8f6817ecc6ca574624196014652

Observation 097caa9f-a2f3-4d66-80f5-74b614ede8d3 · outbound

This paper cites Cashman, and Jamie Shotton.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Cashman, and Jamie Shotton

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.688006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.557863Z digest=sha256:b81a5058213eab77a220ddf11a6d45c9b1085cc2ad470e13d734b7d1fbe31a0a

Observation 1fcc220b-2b9f-48d7-b0a9-9b329eaf17c0 · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.567989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.567989Z digest=sha256:fa135ed7a89276a1a7c933a160c89c4619bde081f1deb12d170e032f4b12af8d

Observation 5a1387d4-dfc1-454c-b7a9-08eb179b69c8 · outbound

This paper cites BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks BoxDiff: Text-to-Image Synthesis with Training-Free Box-Constrained Diffusion

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.583234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.583234Z digest=sha256:400962faa58f01f1d3efa315899c5b7d9d6541c2ed2d8294ee60679cfc7c8a96

Observation a2c6a7ab-555b-45ce-ac4b-dc398d456c8e · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation, 2023.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Imagere- ward: Learning and evaluating human preferences for text- to-image generation, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.579976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.597956Z digest=sha256:412720323a1d79730a56c2cd7d5a7b2442b905a19019986ff65686268b5a82ca

Observation f36b568b-2786-4b0d-ab84-207ac6ce8456 · outbound

This paper cites Attngan: Fine- grained text to image generation with attentional generative adversarial networks, 2017.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Attngan: Fine- grained text to image generation with attentional generative adversarial networks, 2017

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.525743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.614952Z digest=sha256:ddb659163ee0600b887f9ca4fd7421f12d5278141a4e2843299ad4bb961e7da3

Observation 6a75bd0a-58ae-4df3-b0bc-088e5495643f · outbound

This paper cites Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image re- trieval.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Ldre: Llm-based divergent reasoning and ensemble for zero-shot composed image re- trieval

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.491714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.652368Z digest=sha256:572015e1d66493332a2f7a3ae94fd2010e4dea03c83b89cfe8d0a7009442fdb2

Observation cefe5213-571d-4075-85e9-fa4563dec278 · outbound

This paper cites GLIPv2: Unifying Localization and Vision-Language Understanding.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks GLIPv2: Unifying Localization and Vision-Language Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.738271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.738271Z digest=sha256:1024dfaa3108e295d117e935c0097c4acf2b4926d0c53bd7d90fda24f92aa239

Observation 12682082-1eb5-40a3-bdd7-f5849f50a3a0 · outbound

This paper cites Dynrefer: Delving into region-level multi-modality tasks via dynamic resolution,.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Dynrefer: Delving into region-level multi-modality tasks via dynamic resolution,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.327190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.819762Z digest=sha256:0269810c95c1ffe4ae69f2de7c14ce40806e09163b2544720fd59aed180b77d0

Observation cf6ff3b9-0cc2-4568-b37e-683e08f0dc1c · outbound

This paper cites Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Layoutdiffusion: Controllable diffu- sion model for layout-to-image generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.221819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.842953Z digest=sha256:e60c00e131cdc2495e38cab135ca816ced0a3511fdc665692d46037c26a465f7

Observation 276af15a-9816-481e-b84d-d75d68c2fa70 · outbound

This paper cites Pyramid Diffusion Models For Low-light Image Enhancement.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Pyramid Diffusion Models For Low-light Image Enhancement

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.852307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.852307Z digest=sha256:7742c72bc9291464ce8a483353feae86f0a803804baffa1352df3e58e0f0736b

Observation f14d6669-2162-4006-80e6-e51d6e6be41b · outbound

This paper cites Migc: Multi-instance generation controller for text-to-image synthesis.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Migc: Multi-instance generation controller for text-to-image synthesis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.185926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.864617Z digest=sha256:582197a91f9b558f3d5602be8f92903fb9944eac04e774e7454c25a0785d3903

Observation f09b7ed7-b6f1-431f-b1df-c58e619096fb · outbound

This paper cites MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks MIGC++: Advanced Multi-Instance Generation Controller for Image Synthesis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.875664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.875664Z digest=sha256:d763b6609deb40eadf471fa03d1b0e2ea6051a148ee1ed81a8c125a519c14bc1

Observation 7caf1579-74f0-4b62-9afd-110d02338df8 · outbound

This paper cites 3dis: Depth-driven decoupled instance synthesis for text-to-image generation.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks 3dis: Depth-driven decoupled instance synthesis for text-to-image generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T14:03:28.932870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:03:28.932870Z digest=sha256:0bcb52eb37cf92566469b50658bddc852c8b2ea93739fe58c68ec31646ac13b6

Observation 220c78ab-c58c-4ee3-824c-44ce1abc0362 · outbound

This paper cites Odgen: Domain-specific object detection data generation with diffusion models, 2024.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Odgen: Domain-specific object detection data generation with diffusion models, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.111447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.970830Z digest=sha256:7deb76fb000df5ab3cdd7f53b7f119763a302aedd0e95b0ab16265fec57a13fe

Observation 8770c905-370a-4bcc-b38b-b0fb47b48218 · outbound

This paper cites Imagine a basic scene, including whether the scene is near or far, and give a scene label, such as on the grass or in the hospital.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Imagine a basic scene, including whether the scene is near or far, and give a scene label, such as on the grass or in the hospital

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:31.051485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:28.990852Z digest=sha256:771e27ee473fd96755d13c65ff7b0dd12649e218dd6663085379636c931b0f64

Observation 9d62d14f-82bd-46ae-9124-c91781b2037c · outbound

This paper cites users may provide reference images and modify the content of the reference images.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks users may provide reference images and modify the content of the reference images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:30.904796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:29.010667Z digest=sha256:efb1a13cc1e34ae8bf70e03f5061cd1bd511729a7906aa2735dbd20813b274ff

Observation 95af2ccc-a133-4f01-aef9-57dc3d639933 · outbound

This paper cites an unresolved cited work.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:03:30.813039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:29.025215Z digest=sha256:947e6d289552e0a13311d6a2e51852228e9433f824c2e60035491f013090bf59

Observation 8a5bc7cf-b773-40af-9680-26cedd6cc491 · outbound

This paper cites Summarize output in format: … Make sure that each layout contains only one instance.

AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks Summarize output in format: … Make sure that each layout contains only one instance

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:03:30.744950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T14:03:29.034095Z digest=sha256:70d2790535fdc0016e4a3767afd11fc45177b26e03cb17b0737e736decbadf6d

Pith citing papers

Observation 33eaaa53-e5a5-4bf7-b5de-37e8634de503 · inbound

Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy cites this paper.

Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:02:55.814289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:02:55.814289Z digest=sha256:a730730a51f20e23136a304f2cf9a7e95965a03d52860ec4c63d4dca73c18c19

Observation 52e3cd3d-fc9b-450e-9da5-55576307077c · inbound

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts cites this paper.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.480308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:aff5c57386e0b3e6c7938846fb693e314e93d5d3c5a4541266742019ff4580f8

Observation 01d80f55-4586-43b5-b625-b49aebed3ee0 · inbound

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details cites this paper.

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details AnySynth: Harnessing the Power of Image Synthetic Data Generation for Generalized Vision-Language Tasks

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:20:54.236052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:11:43.172296Z digest=sha256:ee7b7a5a68cc61623ccc3abca530289502378b839c5235c3722675ba7b888cbb