Pith. sign in

Paper Citation Record · LEDGER

Dual Diffusion for Unified Image Generation and Understanding

As of 14 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 11 inbound Pith citation observations for arXiv:2501.00289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00289 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:02:07.964380Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:28:32.362554Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T13:02:18.322222Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved66
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 12614b94-8db3-4552-9103-3fcda37994f4 · outbound

This paper cites GPT-4 Technical Report.

Dual Diffusion for Unified Image Generation and Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.535461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.535461Z digest=sha256:57277f8a437afb648fdeb800508a23d77aad6fd6ba2b944af191832825254d56

Observation 4f830245-6bd7-45d8-b696-4ff51a1c28c3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Dual Diffusion for Unified Image Generation and Understanding Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.541244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.541244Z digest=sha256:03fc759e3601b216a5945abf00e25ee37fd4ed72b50ad90c236969b9f981dd83

Observation ad0f78d8-c3f1-4973-973d-25b316d7e2ff · outbound

This paper cites Reverse-time diffusion equation mod- els.

Dual Diffusion for Unified Image Generation and Understanding Reverse-time diffusion equation mod- els

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.546269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.546269Z digest=sha256:eeea9cf87126d3166a5fc5a683220a2946436a6353cdfbaa7084bf52595347ad

Observation ca9ba559-4f16-4a77-b559-01469055e3e7 · outbound

This paper cites Structured denoising dif- fusion models in discrete state-spaces.

Dual Diffusion for Unified Image Generation and Understanding Structured denoising dif- fusion models in discrete state-spaces

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.551047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.551047Z digest=sha256:2b8775f9693678d83bd74a09378182f7d2d6ee8dfecb6e1b5f4426ef3b9f1471

Observation 56943249-3212-46e4-9730-4315a4b6d84b · outbound

This paper cites Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023.

Dual Diffusion for Unified Image Generation and Understanding Openflamingo: An open-source frame- work for training large autoregressive vision-language mod- els, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.556093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.556093Z digest=sha256:119304cb659bd50fa01cfe2ef34bc94ab136518f5803a683f1db0c2048381d51

Observation c1965964-ac7f-4f26-9c84-5c1b953a397b · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Dual Diffusion for Unified Image Generation and Understanding Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.561312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.561312Z digest=sha256:3daf5ee8b562f9a2bde3756f69fc8878699bbfce0a4760a8f40077ce8b3abea0

Observation bd605e9f-67f2-4df8-bbc1-2ab7cac49b69 · outbound

This paper cites One transformer fits all distributions in multi-modal diffu- sion at scale.

Dual Diffusion for Unified Image Generation and Understanding One transformer fits all distributions in multi-modal diffu- sion at scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.566136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.566136Z digest=sha256:4e0ccdd0e0acb90e2f470cff6c67af4fad23936df34701eb9d5bb9a01345c12d

Observation 09d5f859-714c-4b32-a2ea-a9a90571c2ac · outbound

This paper cites Vizwiz: nearly real-time answers to visual questions.

Dual Diffusion for Unified Image Generation and Understanding Vizwiz: nearly real-time answers to visual questions

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.571001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.571001Z digest=sha256:2e30da23c95128e7e5d13395965e33c64e94f03bae2b2dc61217a055b8f39592

Observation cd54cb74-3fc5-49fb-ad6a-79fb2f58ec0a · outbound

This paper cites Language Models are Few-Shot Learners.

Dual Diffusion for Unified Image Generation and Understanding Language Models are Few-Shot Learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.576429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.576429Z digest=sha256:509bfb94581c2eb86b9b29ec3c0704fcf07cc25ecc21dda29213ca270efd992f

Observation 7aee87cd-613b-4e54-9fcc-4bd4d9b978bd · outbound

This paper cites Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,.

Dual Diffusion for Unified Image Generation and Understanding Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.581347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.581347Z digest=sha256:3c7d2488b41d7850230cc52c6953ae444f755373f4db384997cd01cff52d6c46

Observation 59817b7a-a16f-4ccd-9710-f1310c4e2075 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Dual Diffusion for Unified Image Generation and Understanding ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.586294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.586294Z digest=sha256:796023aa2170ed47e6d3123a6a81941e91870ed420d4a7644ff5466c6b7f6cc6

Observation 7203fae7-e095-4923-9c80-a057ab80e1cd · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

Dual Diffusion for Unified Image Generation and Understanding Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.591748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.591748Z digest=sha256:f2d63fd34983b9f7fbfdabb64e02bd42ff04e261ee6f64a17d7045a47611082d

Observation 2faa5857-6ba0-4b48-ad13-9d3c01f7b4fc · outbound

This paper cites Ex- panding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025.

Dual Diffusion for Unified Image Generation and Understanding Ex- panding performance boundaries of open-source multimodal models with model, data, and test-time scaling, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.596687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.596687Z digest=sha256:f50846fdecfdb647af46d06d7cd2aa4a8c5cc98ba12434cee541faf9bdd17329

Observation de9daf0c-0f99-488d-b378-75d101d30b85 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Dual Diffusion for Unified Image Generation and Understanding Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.601154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.601154Z digest=sha256:213db27201e180613d046128245abab0aaad44ee06846df9a69786927ccaee88

Observation 090fbbe5-ba13-48da-a809-22bfe8708d5c · outbound

This paper cites Diffusion models beat gans on image synthesis.

Dual Diffusion for Unified Image Generation and Understanding Diffusion models beat gans on image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.606056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.606056Z digest=sha256:eaa901093e87499445d27fc044e7d7beb6de168d0c6479bf401781e6b3cdfa60

Observation 26b5011a-6190-45d1-95a2-7f294344c698 · outbound

This paper cites Continuous diffusion for categorical data.

Dual Diffusion for Unified Image Generation and Understanding Continuous diffusion for categorical data

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.610619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.610619Z digest=sha256:c7761717143b04c75e886b7820d1e5aa6473642c6e9a2b979750272eba180e4c

Observation d0df6326-2066-45a5-ab1e-fcf6cf96f285 · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

Dual Diffusion for Unified Image Generation and Understanding DreamLLM: Synergistic multimodal com- prehension and creation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.615721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.615721Z digest=sha256:0d1ce00e0baa6a854acaee56079dd185ebcc99fc41cce0ce1096cde63b19c4c8

Observation cd51f741-1ab5-49f0-946c-b8753c1bee38 · outbound

This paper cites The Llama 3 Herd of Models.

Dual Diffusion for Unified Image Generation and Understanding The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.620267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.620267Z digest=sha256:bb562eb53f0a826e478ca6540a72525b58c51715fe3082695c1766fde724fe5e

Observation ebce373a-e02e-4d81-b674-7c46aac7f1d7 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Dual Diffusion for Unified Image Generation and Understanding Taming transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.625688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.625688Z digest=sha256:df8530624cc6960be1f4cddba78574c09122fec694478838f8f72d528386d5cc

Observation 930f4408-39f7-43c5-82a4-c79502c14177 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Dual Diffusion for Unified Image Generation and Understanding Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.140942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.630405Z digest=sha256:7378474ec860eeed6a06c3c1eed08af0578a7932bf391083498da4dbf5b2f6af

Observation a4d15cd5-6240-4594-ae61-5553ae042622 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Dual Diffusion for Unified Image Generation and Understanding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.635023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.635023Z digest=sha256:9c30032eb91ee1380132d4faf42af140aede3fd92df366a38a1721ba13b1d876

Observation 78465120-2039-486d-8f58-0bce3a27cb57 · outbound

This paper cites Datacomp: In search of the next generation of multimodal datasets, 2023.

Dual Diffusion for Unified Image Generation and Understanding Datacomp: In search of the next generation of multimodal datasets, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.124939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.639730Z digest=sha256:06c8488d4e0827f885cf57f3fc585f300d93ea7cfadd661e4cb1ef25dbc4edf0

Observation 70ca07cd-d460-41aa-aafc-a56c2210b7c7 · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

Dual Diffusion for Unified Image Generation and Understanding Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.643640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.643640Z digest=sha256:88b547e1985a2af57acdaccff4f61b09291402abd68d87cde2f242ff38893068

Observation c3f792e5-a7e0-47b7-a4e7-37dc30453433 · outbound

This paper cites Discrete Flow Matching.

Dual Diffusion for Unified Image Generation and Understanding Discrete Flow Matching

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.647704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.647704Z digest=sha256:18dba3033e3b98a0868fcbb88809d896070da35ef815243c2b11c2119b68775f

Observation 1cf65f1b-7790-4aa7-8ee8-aa3128b26611 · outbound

This paper cites Planting a SEED of Vision in Large Language Model.

Dual Diffusion for Unified Image Generation and Understanding Planting a SEED of Vision in Large Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.652832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.652832Z digest=sha256:3b9590cead6cb988d4ac048689cb034062582e895b01b53e043421ecfe765d86

Observation 19b6c957-0c07-4a83-a829-506788e29008 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment, 2023.

Dual Diffusion for Unified Image Generation and Understanding Geneval: An object-focused framework for evaluating text- to-image alignment, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.656919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.656919Z digest=sha256:84c76b0ccc6dc3cef2dbcce936b3f95db0176d8e5282d4c0680f44d0ab43c47e

Observation 4de7c55c-9195-4211-8e07-c22878fc2d6e · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Dual Diffusion for Unified Image Generation and Understanding Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.096918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.660973Z digest=sha256:ea5989acd72d00dfbd991d2cfd2f7a365b14591a00df94e880b80cabb26885a5

Observation 4cb9e3ad-bcfc-4183-be5f-ffe778fe1486 · outbound

This paper cites Likelihood- based diffusion language models.

Dual Diffusion for Unified Image Generation and Understanding Likelihood- based diffusion language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.079956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.664995Z digest=sha256:bf4b26cb289e3312cb09198837b7cdec103b12b514d2e4db8668eeca7597c3c8

Observation 9e65b7ce-101b-4ed8-8b00-b0d12639ba9f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Dual Diffusion for Unified Image Generation and Understanding Classifier-Free Diffusion Guidance

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.669512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.669512Z digest=sha256:4986962302dd08be0cd2fe41b95ae6f03967433700278ab40b47be9339ce57b9

Observation 849ff9ac-7ad1-4e6c-9deb-4d0c1c8f1937 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Dual Diffusion for Unified Image Generation and Understanding Denoising diffu- sion probabilistic models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.062955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.674312Z digest=sha256:e31a52375c4ff6f7b8183d58bc451d525ca7fe799a3cac14befa1feafb6d276a

Observation 3c038f1c-7be4-44df-8f1d-10a7c2728172 · outbound

This paper cites T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion.

Dual Diffusion for Unified Image Generation and Understanding T2i-compbench: A comprehensive bench- mark for open-world compositional text-to-image genera- tion

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.046610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.678758Z digest=sha256:a2798a3ba2dcb1881d2f6690e617b0fa6b2b7c5a8c5312b65797ec26e70b9583

Observation 339486c2-7de3-4ebc-89db-faf2eb3fcdf9 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Dual Diffusion for Unified Image Generation and Understanding Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.683123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.683123Z digest=sha256:6c8048497037a09f977b3d44d970abfd7802975108f6ddead0ff62fd5e0da00e

Observation 17cfd9b0-443e-4cd4-bbd1-3db1ba19e0df · outbound

This paper cites The Platonic Representation Hypothesis.

Dual Diffusion for Unified Image Generation and Understanding The Platonic Representation Hypothesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.687815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.687815Z digest=sha256:8a68ece12e20a6c9efb039e293966bd9fa1d65f922327ee9cc8c375e30506826

Observation aff5a9fe-a7dc-4ae5-94ea-5f27130a97b6 · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

Dual Diffusion for Unified Image Generation and Understanding Elucidating the Design Space of Diffusion-Based Generative Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.692404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.692404Z digest=sha256:f4fd080313e090eee6ceba1c293e216c440a34d77b505966793b3378705d900d

Observation 532711f8-ed2a-4204-bd54-9568dc147d8f · outbound

This paper cites Variational diffusion models.

Dual Diffusion for Unified Image Generation and Understanding Variational diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:09.017992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.697148Z digest=sha256:76408042bc1d0f731fba05950463cbbec27de602d01d110a7f640d60a40cb558

Observation 3922031b-bf94-46f9-8381-46ca209261f4 · outbound

This paper cites Auto-Encoding Variational Bayes.

Dual Diffusion for Unified Image Generation and Understanding Auto-Encoding Variational Bayes

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.701668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.701668Z digest=sha256:816674aa167e3dfb94ee271dccd5362c490c4b0271e63b50f6a8e1430bc3a98d

Observation 6dfc06ce-04d2-4071-852d-491caa67c15e · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Dual Diffusion for Unified Image Generation and Understanding The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.706461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.706461Z digest=sha256:2c37ba12de00054a99d43e1ebb5c23b0817ba7afc9aff3e6efb3c37023001344

Observation 8bea8a73-09bd-4243-86d2-3e5ee57e83d3 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Dual Diffusion for Unified Image Generation and Understanding Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.711246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.711246Z digest=sha256:a84eff1e61aa7b99091ff3bb1c1000fb7fc250b6e405ee1b964511f47d2b7627

Observation 15968370-18d9-4465-bb75-559a4d0ddd94 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

Dual Diffusion for Unified Image Generation and Understanding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.715985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.715985Z digest=sha256:97f6e4f1cea4ed4d8df086198f88fca166fc342b013f7261c004e15db6cf6d60

Observation 7d208d2e-32fd-4b93-b2a4-6b727ecee7b4 · outbound

This paper cites Likelihood train- ing of cascaded diffusion models via hierarchical volume- preserving maps.

Dual Diffusion for Unified Image Generation and Understanding Likelihood train- ing of cascaded diffusion models via hierarchical volume- preserving maps

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.983628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.720826Z digest=sha256:058443dee7711609865d0ebf5151501e57b145f8308983b28d76fe38a9e7536c

Observation 20515266-7ab0-4758-b556-20ccb6555b9b · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Dual Diffusion for Unified Image Generation and Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.725598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.725598Z digest=sha256:54dc084eaed31de211881e7ce5f695f386e74058fa685785a71e5475cf46cd67

Observation 97be30ac-5b86-452a-bebb-329aa00c7287 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

Dual Diffusion for Unified Image Generation and Understanding Autoregressive Image Generation without Vector Quantization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.730197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.730197Z digest=sha256:3da8f7772d6f2583671ba85ac7b7ebc751dcf0cd26ba0bb642f30f7008ffe668

Observation f64cbffb-6e9c-4e02-b47d-a5808e2e5de0 · outbound

This paper cites Diffusion-lm improves control- lable text generation.

Dual Diffusion for Unified Image Generation and Understanding Diffusion-lm improves control- lable text generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.959255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.734889Z digest=sha256:7017a4c3af72ad0b58dca30598e7370193e7cbf0a584583284f57b965decea32

Observation 104eaeb3-cee6-4a01-ba8e-7d73a36759c7 · outbound

This paper cites Diffusion-lm improves control- lable text generation.

Dual Diffusion for Unified Image Generation and Understanding Diffusion-lm improves control- lable text generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.739365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.739365Z digest=sha256:06d24741005074244491ecf4f23226a0626aa3412d5ba4b957dfd8a72af5ec09

Observation c19cb48d-4b04-4567-aa31-da0a91370c34 · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Dual Diffusion for Unified Image Generation and Understanding What If We Recaption Billions of Web Images with LLaMA-3?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.743838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.743838Z digest=sha256:a3cc0ba237bb24ad668052368a1a2a94e9206648b8b0a0cca8aea253ba5a6e02

Observation 4c961acc-18f1-44db-8fd7-ee2acfd6ec59 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Dual Diffusion for Unified Image Generation and Understanding Evaluating object hallucination in large vision-language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.933495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.748528Z digest=sha256:050fa26f8d7a3db2fd64b68f13b08332214a7d1dbdb7ddeab51b49744b166f4b

Observation cf28e946-5e85-473b-a009-df6110367d9d · outbound

This paper cites Microsoft coco: Common objects in context.

Dual Diffusion for Unified Image Generation and Understanding Microsoft coco: Common objects in context

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.752756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.752756Z digest=sha256:a24ccbfee7905f5cd55ddf1ee694e636adc8439e69d2f9aff3f203f8f7ffe781

Observation 97384505-e0a4-4802-aa41-59ab308a0144 · outbound

This paper cites Flow Matching for Generative Modeling.

Dual Diffusion for Unified Image Generation and Understanding Flow Matching for Generative Modeling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.757149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.757149Z digest=sha256:807e897f2c08f0f270907ce24a595a5a8a48ced1ef6100b4e2dd0fa322781f14

Observation cd9f8e1a-d532-4776-a7bd-1c0c12cf4239 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Dual Diffusion for Unified Image Generation and Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.762283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.762283Z digest=sha256:087a1c44b24b09b23a53dcb9a2d18d23a60b91359664fc3f798f65ee6e9d0566

Observation 11c76f7b-f8aa-4da4-9d17-5fa2ca516084 · outbound

This paper cites Visual instruction tuning.

Dual Diffusion for Unified Image Generation and Understanding Visual instruction tuning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.896134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.766506Z digest=sha256:e87889de6be792d754ad7512321d3f88bfa2ab858ed701492d5544533c2c820c

Observation 526a1bdb-27f3-4409-a0c3-67a857194a3b · outbound

This paper cites World model on million-length video and language with blockwise ringattention, 2024.

Dual Diffusion for Unified Image Generation and Understanding World model on million-length video and language with blockwise ringattention, 2024

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.880360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.770647Z digest=sha256:bb1ea43a370a6eed825fadd73c6ab90dfc0e4bc0eee70997c3c1917765af3a56

Observation bae71810-f060-4f91-ab9b-8b4ce1c5aa4b · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

Dual Diffusion for Unified Image Generation and Understanding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.774842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.774842Z digest=sha256:3541e0204d8a68dbb52d0d620da2bb1bcbe2b016bf32d7671bad93fbb8d58172

Observation c8aac922-88d2-4efc-a815-2823e36b8326 · outbound

This paper cites Flow straight and fast: Learning to generate and transfer data with rectified flow.

Dual Diffusion for Unified Image Generation and Understanding Flow straight and fast: Learning to generate and transfer data with rectified flow

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.779204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.779204Z digest=sha256:4ab5fc2f64095fb4431299ca3a3a9050355b4ac14a406cb62802eedb5b7ab636

Observation c71bd467-dcdd-46b6-94df-917dc4bb56ac · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Dual Diffusion for Unified Image Generation and Understanding Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.783211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.783211Z digest=sha256:a8c626506bee17e5af020c42cd7e9dc4eec6bcdf728e1e3117b036dae9597fdd

Observation 5b4b32d0-1100-457d-8898-d59f75a540e3 · outbound

This paper cites Latent diffusion for language generation.

Dual Diffusion for Unified Image Generation and Understanding Latent diffusion for language generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.855152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.787945Z digest=sha256:4c8f4b9a77f70df2384a39ecfbd924d0fa8abc359d6c2adfcda09d1e5bc494f4

Observation 2ebabc2a-7cc5-46b4-8fc2-fa2d594ed6cf · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Dual Diffusion for Unified Image Generation and Understanding Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.792083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.792083Z digest=sha256:fd5e6ce109552595fb91c5150c0b12274baf5d29d0e6f935fbddbe06ef27b886

Observation e08a9967-99b0-421c-86ce-53545f525c9c · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Dual Diffusion for Unified Image Generation and Understanding Concrete score matching: Generalized score matching for discrete data

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.829639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.796738Z digest=sha256:45bc19fedb8d1322c8f865bd173dd0bbbf3203aae1968d7928ccab5eb50e89f6

Observation 2d0eafc0-37c1-4a2a-8b4a-d4fd2131db73 · outbound

This paper cites Improved denoising diffusion probabilistic models.

Dual Diffusion for Unified Image Generation and Understanding Improved denoising diffusion probabilistic models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.801299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.801299Z digest=sha256:6896d43e5bfac9029ea771ffc2a9b12fbf9d552ab4c3e5ca2e820bedb151450f

Observation 3428c539-73a8-4fe9-a21c-a568fc103fb3 · outbound

This paper cites Scalable diffusion models with transformers.

Dual Diffusion for Unified Image Generation and Understanding Scalable diffusion models with transformers

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.805758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.805758Z digest=sha256:0abf5b5e4a33d58babe86f2ecdbcbb8d2d19deab50190c35d880c203fd8a69e9

Observation a3344b70-868b-40d7-8a1b-7eba5905ce38 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Dual Diffusion for Unified Image Generation and Understanding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.812547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.812547Z digest=sha256:aea3f37bea2fa67083fd60bb98c0a60b1097dde9052e04b80308859f9aaaa0eb

Observation ecad3bcf-0eff-44c1-9b2d-f9f8166f9037 · outbound

This paper cites Improving language understanding by generative pre-training.

Dual Diffusion for Unified Image Generation and Understanding Improving language understanding by generative pre-training

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.793632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.818164Z digest=sha256:6611dd122c0bdc794a804d4364060c3f2f797cb8e154bb45dcc2f25d3ed6613e

Observation 8fc8d875-4d0e-4d33-852a-cd441fbb3441 · outbound

This paper cites Language models are unsu- pervised multitask learners.

Dual Diffusion for Unified Image Generation and Understanding Language models are unsu- pervised multitask learners

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.822762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.822762Z digest=sha256:40a8ba405db57bcd6a059651b378e70b9572bf92f747a36fce1d039ed7079816

Observation 4782ba2c-048e-4375-a8ee-a82646ce6121 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Dual Diffusion for Unified Image Generation and Understanding Learning transferable visual models from natural language supervi- sion

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.767699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.827412Z digest=sha256:55971a7158561eb13ca165cf84270b45a6839d285acca15e307cd7729c89aa42

Observation 6318971a-872a-4f1e-aac3-58850c1bc3d5 · outbound

This paper cites an unresolved cited work.

Dual Diffusion for Unified Image Generation and Understanding Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.833069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.833069Z digest=sha256:eb879c405dfd1818963656d7f5d10a05a378364930a8abef0d85bbf6db99356f

Observation e59f7b03-a240-4b35-bbca-59ea593e955a · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Dual Diffusion for Unified Image Generation and Understanding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.837870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.837870Z digest=sha256:37d64118a74d23aea3f1f22e131324b7ff4aa32565d1dfdacbd92baba8bc2216

Observation f2535679-94ca-49e0-9e45-5118b9863964 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Dual Diffusion for Unified Image Generation and Understanding High-resolution image synthesis with latent diffusion models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.842619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.842619Z digest=sha256:54ce9ee25a6820a8c0fdc29db9da46395b46816865eb844e753e2d5f915d2d14

Observation b6ca8d73-2488-44ef-8329-5b8461d41994 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Dual Diffusion for Unified Image Generation and Understanding Photorealistic text-to-image diffusion models with deep language understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.730758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.847157Z digest=sha256:706ab4c6d7aef00f680c37ce9490d4efe580bf53b04298b1323fb4312329054d

Observation f69cd402-92ef-4d14-9658-0d9c7d1c8b49 · outbound

This paper cites Simple and Effective Masked Diffusion Language Models.

Dual Diffusion for Unified Image Generation and Understanding Simple and Effective Masked Diffusion Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.851664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.851664Z digest=sha256:f1aba2f884dd5030f092b2f9603585dacef0f49334451d20148570794df65f08

Observation 9b150719-b58e-4344-a8e8-2ea6f480e9f6 · outbound

This paper cites an unresolved cited work.

Dual Diffusion for Unified Image Generation and Understanding Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:02:08.715296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.856391Z digest=sha256:aff8dab728c7c4d2edd30dab05df09597365f4d2b3d989740e4e3ba0b31d0caa

Observation 9d1ca3b6-2f14-4707-a057-10c6f6454802 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Dual Diffusion for Unified Image Generation and Understanding Deep unsupervised learning using nonequilibrium thermodynamics

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.699612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.860406Z digest=sha256:87edc3658abcbe51b58fdebc4648ac7fdbba8e03fcc522a0b6f2088ec45982c1

Observation f9163c04-2049-4aac-8b53-22507e1425d3 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Dual Diffusion for Unified Image Generation and Understanding Score-Based Generative Modeling through Stochastic Differential Equations

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.864348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.864348Z digest=sha256:9dc9d5a583292cac0fe6a4d01a918ff2236ec17d5047eecd534112149b30cb06

Observation 6652dbdd-6f6f-4f01-a86b-98a2ff2e03c0 · outbound

This paper cites Maximum likelihood training of score-based diffusion mod- els.

Dual Diffusion for Unified Image Generation and Understanding Maximum likelihood training of score-based diffusion mod- els

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.682194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.868491Z digest=sha256:161730314bb7e43ca4e24eb9cae65fb00ec2e989255ef29737d6f7b77690d0c8

Observation ccf16b99-5962-467b-8070-713973b7da2f · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Dual Diffusion for Unified Image Generation and Understanding Emu: Generative Pretraining in Multimodality

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.872745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.872745Z digest=sha256:c8243cd361008a78210ea2decf43236610a8031c6c91a3a449afeb1fdf9e943e

Observation 4fc79b0a-0f60-4fd0-8c0d-667845e8388b · outbound

This paper cites Any-to-any generation via composable diffu- sion.

Dual Diffusion for Unified Image Generation and Understanding Any-to-any generation via composable diffu- sion

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.667185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.876852Z digest=sha256:fb06ebfb2d9066484b983276afcedcfe6a5483933060d1548c988514dc8cd685

Observation 1a3a1287-3d8e-446a-af6a-7de4476f5680 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Dual Diffusion for Unified Image Generation and Understanding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.880690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.880690Z digest=sha256:bc7bfda80077d7cb0b7ae375cab4a3e186002b0ffe2a33aeb01a13ba93c85852

Observation 5554944d-8fef-4dbc-a107-a310e4eb2d63 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Dual Diffusion for Unified Image Generation and Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.885317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.885317Z digest=sha256:3bc0fb4ec32cf29c89246e08e2285952b9522cb37a4fb13e4592fe4f099fde23

Observation 7a85e97e-667e-4a98-aeff-3fc37818342f · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Dual Diffusion for Unified Image Generation and Understanding Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.890346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.890346Z digest=sha256:7f819ba4dba2aa732353855b6a3a63876919551f2fc4a46dc3b0e126d8d822ac

Observation 0ce767ee-4f28-479e-9a5b-7d095aa1768b · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Dual Diffusion for Unified Image Generation and Understanding Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.895344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.895344Z digest=sha256:f7c6646e9cfe7d769d176272e829eb1db10a69a633afbab2f2cc47e0b140d7ce

Observation 49080754-2798-46a1-b8bc-35d5a6a1faca · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Dual Diffusion for Unified Image Generation and Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.900245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.900245Z digest=sha256:77a9d5e412309b92f34f4d485fba18544f6d0260aa055c4728a46b8cfff74e48

Observation 1d1253d4-a196-485f-987d-2ec1352ebd18 · outbound

This paper cites A connection between score matching and denoising autoencoders.

Dual Diffusion for Unified Image Generation and Understanding A connection between score matching and denoising autoencoders

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.652669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.905064Z digest=sha256:c83592bfa55d7942290ca49749d99202467729f59c849d8a0e92d832a8f0ebd1

Observation 6633b1ed-d3a6-4578-add5-08de184f0d66 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Dual Diffusion for Unified Image Generation and Understanding Emu3: Next-Token Prediction is All You Need

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.909908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.909908Z digest=sha256:b3b121cf0746c29a56ea55480f6a872d80f6c178924a06c18c53744f94f6feea

Observation 8af2dd09-1b19-40ad-90c6-8cb3addcb339 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Dual Diffusion for Unified Image Generation and Understanding VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.914690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.914690Z digest=sha256:db4474ffb0bc20ab1bbf4aa701faed58c2f573b3cd8e16b4d946a405a4541833

Observation f4bcfe59-c79a-4ba7-afc4-d7bead0b7692 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Dual Diffusion for Unified Image Generation and Understanding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.919728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.919728Z digest=sha256:6952f4fc3f8266e81126c4fc0b8f5e8ad44203f5027ec1bd8811601a243747fa

Observation 7edc67d6-4819-48cd-a2de-de9b643e65fe · outbound

This paper cites Versatile diffusion: Text, images and variations all in one diffusion model.

Dual Diffusion for Unified Image Generation and Understanding Versatile diffusion: Text, images and variations all in one diffusion model

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.636588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.924609Z digest=sha256:b0adc7643b44493ab488951df8bc6b4382ede4a7a015ad38b48d426c7704c0fe

Observation 135d5202-5355-4b3f-b18e-10559816eb6e · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Dual Diffusion for Unified Image Generation and Understanding Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.929381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.929381Z digest=sha256:c99c4a00d9b156f82ed456581291f272667ff10ffeed2a8b85c378e9312dfb16

Observation 0cba4a55-bae1-4657-9598-9a2761189983 · outbound

This paper cites Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023.

Dual Diffusion for Unified Image Generation and Understanding Scaling autoregressive multi-modal mod- els: Pretraining and instruction tuning, 2023

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.621701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.934634Z digest=sha256:b0e4067d8aa4889843d17e3eb0fb6daa5eaa8c35172336d6de3c30c1c39b0c38

Observation 646f88c1-ec84-4785-a9a0-ca8a56d6aacd · outbound

This paper cites Sigmoid loss for language image pre-training,.

Dual Diffusion for Unified Image Generation and Understanding Sigmoid loss for language image pre-training,

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.939321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.939321Z digest=sha256:f19669989850fc1d90bd23322b3591d47fe85e4e9f93cf39a22f3d2a49c901a8

Observation 9a96a108-076f-4901-a72c-4f033c1eab6a · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Dual Diffusion for Unified Image Generation and Understanding Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.943972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.943972Z digest=sha256:54ab582d047d7e7d9f9deba1a2fe6f7ffc1ac630a1518e82923a222bef2317b7

Observation ed18e898-53b6-4a40-80b4-02a19d08ff69 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Dual Diffusion for Unified Image Generation and Understanding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-10T23:02:07.949502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:02:07.949502Z digest=sha256:28f1f977751824919650287098be85f6c1c1d06fef87cadeb8018d4f3eccc2b2

Observation 893d3773-da98-4aff-ad5c-b2ad65017bd2 · outbound

This paper cites Dual pretrainContinued pretrainInstruct.

Dual Diffusion for Unified Image Generation and Understanding Dual pretrainContinued pretrainInstruct

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.597705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.954670Z digest=sha256:dcdec1c2b362d548f1497216c0f4f816873121db68b18fa98e4df55a81692ad6

Observation 981ccc90-16f0-4585-9ab8-e18da6861801 · outbound

This paper cites (B) FID ↓ SD-XL [60] Diff.

Dual Diffusion for Unified Image Generation and Understanding (B) FID ↓ SD-XL [60] Diff

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T23:02:08.582647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.959469Z digest=sha256:2a23aeec0bbd3a59df5db3c6546587b2e07a96a31df0c77d70c06246eb9e6b76

Observation 84498b3d-5063-4915-ba59-cbe31d579ee6 · outbound

This paper cites On T2I CompBench, we find that after dual diffusion fine tuning the model performs worse in texture.

Dual Diffusion for Unified Image Generation and Understanding On T2I CompBench, we find that after dual diffusion fine tuning the model performs worse in texture

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:02:08.566612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T23:02:07.964380Z digest=sha256:e1db43e6ba3da139ae9175dbc86e4d089c3b6b0a31f56045dfbc1cc1e621ef50

Pith citing papers

Observation eb9db4b8-1c2d-402f-870d-5a2b59dde72d · inbound

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling cites this paper.

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling Dual Diffusion for Unified Image Generation and Understanding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:14:53.136264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T08:14:52.890145Z digest=sha256:a0f3eef2934d663f11f423ba47edafcdc8761f4b6752acc71f4f8884a1b43f66

Observation 49bb7643-8ddb-44c6-a10d-9aacbc2e1f4c · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Dual Diffusion for Unified Image Generation and Understanding

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T16:24:27.656162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:588e0cc170678fe70c8b8bcf85237bd9a25e3fd9a4cb837859f8d48bc117c9aa

Observation 0563063a-900d-40b3-9637-a4bb099a7ec9 · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Dual Diffusion for Unified Image Generation and Understanding

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:24:04.880861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:0cdb33dfe99f739527d8e4456117c9d04fea7247aad004c4ecaae4d03e652cc4

Observation b40dcffc-9bae-4483-bd3b-485d13ea5e07 · inbound

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning cites this paper.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Dual Diffusion for Unified Image Generation and Understanding

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.166172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3a265645f43b35edcd1b55dc7ee5fec3c90ba76becc6f307d3626bf92e3d1f72

Observation ed0f780d-9212-47a9-96ab-f8b587539ce6 · inbound

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks cites this paper.

OmniGenBench: A Benchmark for Omnipotent Multimodal Generation across 50+ Tasks Dual Diffusion for Unified Image Generation and Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:28:32.362554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:28:32.362554Z digest=sha256:ae222fcdad9d1abf9eadbd3ad10806c98acc9e9c6fa10712a7330cb055ce6a3e

Observation ba11db8f-ae42-4811-9e68-11470bdeffb9 · inbound

Jodi: Unification of Visual Generation and Understanding via Joint Modeling cites this paper.

Jodi: Unification of Visual Generation and Understanding via Joint Modeling Dual Diffusion for Unified Image Generation and Understanding

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:22.394934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:25:22.394934Z digest=sha256:a167277bdb1fb566e7338ce52d34f82aa06d3d649b94efaf6340120e648a83c6

Observation 45345e19-f584-4399-ae16-ff1d85fe254c · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Dual Diffusion for Unified Image Generation and Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:04:56.901976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:04:56.901976Z digest=sha256:8796bdd8495d644ef9108a2413c647814f38633a14b4bd91f0c1c64539d0fca7

Observation 00e001e7-be36-48f5-adc0-f3ca80abc84f · inbound

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model cites this paper.

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model Dual Diffusion for Unified Image Generation and Understanding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.324887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T12:59:31.454155Z digest=sha256:0203c64041b4612ac5ea30b00aea1b813329c2a4fac65bb86723614e554b00ba

Observation e2932b53-4de1-4e27-9e66-765eb64805cf · inbound

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL cites this paper.

ReasonGen-R1: CoT for Autoregressive Image generation models through SFT and RL Dual Diffusion for Unified Image Generation and Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:19:22.862697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:19:22.862697Z digest=sha256:e81088052cf80e5ea2748111940990a76288a579b2d7c917a93011e57eb637fb

Observation b74913d7-a687-49da-bcfd-170f68407a83 · inbound

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards cites this paper.

SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards Dual Diffusion for Unified Image Generation and Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:27.042171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:27.042171Z digest=sha256:52f72081d210b193c2fa1822fcb49317017d01d1e055e6bd769c33fdbf289562

Observation 0ac33dff-0e29-4430-b42d-934d3e4af888 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Dual Diffusion for Unified Image Generation and Understanding

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:51:15.997772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:8690cdca0e3693e295a803d5b0dfbb7217f0bec69bc5695c7f8665f58e2ff26e