Pith. sign in

Paper Citation Record · LEDGER

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation

As of 13 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.21611.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.21611 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-22T09:25:19.066598Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact23
  • verified fuzzy15
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9ee4736-29c4-4eb7-a044-59007715992e · outbound

This paper cites Ming-Omni: A Unified Multimodal Model for Perception and Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ming-Omni: A Unified Multimodal Model for Perception and Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.784211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5a7f07c0e6502fd9fc2bfea603991422eccbe5f3435092e31c9c5afea0687cec

Observation b5095ec4-53ae-4cf5-945c-c4decadf84b6 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Nougat: Neural Optical Understanding for Academic Documents

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.807752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:3725eba0e3efd56b82c995cf90494b9dabc7b469cff4983fc8a047368fdcea97

Observation 8ce526d6-a5da-497b-a961-57b2ea8a45b8 · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.779221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:331dee14a56fee1cf3299483dd3f38078f92b899778b42d6b9cf79a135430d56

Observation 28a45d0d-6e1c-47ce-88c4-d9bff06cbdb3 · outbound

This paper cites Textdiffuser: Diffusion models as text painters.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Textdiffuser: Diffusion models as text painters

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.105030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:68838ca601e76cbe9c0802dd6d6792b632e4396e2c8763f7fb9477b60fbd2350

Observation b7b9a8ed-a566-4ab9-9b41-df3fb1a561a9 · outbound

This paper cites Anydoor: Zero- shot object-level image customization.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Anydoor: Zero- shot object-level image customization

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.101857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:9d1a8a3bcc28f384e53ca44828666e37f9dc92d44d046cfa69b98df353b7263d

Observation 5d7c2be3-5219-4df0-a6e0-dd70cddb09a1 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.789280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:144699ec5d5fe4520f6da231b3b187142b36ef63e6ce7c5a0a57364c3a4d9225

Observation 8524f1f3-0e75-474c-b4f6-398f929e8a88 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Diffusion Models Beat GANs on Image Synthesis

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.793708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:694b6638b56753eca7540756ca44351d26fcc2e0e3895881a1e2663e37130b20

Observation 83f599ec-a41a-4cfb-ac98-f26ecb5157bf · outbound

This paper cites Demystifying Flux Architecture.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Demystifying Flux Architecture

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.803649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:0cf768e67767ec8e222b01faf36ab0761c63436462384b40816898c698895a82

Observation 9488ac53-35bc-469f-a3b8-3201efc5f28c · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.798428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:78eb37baead8115f7a6b633a99557a04edc7a243fff8ee8e69207ef50401a0f3

Observation e4866763-e7c6-4584-aac1-374522c0023a · outbound

This paper cites Denoising Diffusion Probabilistic Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Denoising Diffusion Probabilistic Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.773573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:7d53f5c055ce428ae80a3bee2ac27851c30279e653bb717fd890b05a1d022a98

Observation d7e8133b-e81c-4b8d-b985-1469f3022c78 · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.108241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:42f595ca0c29425c45482ac2b30c5dabf8ccbe4629749fa84f674158734a4a35

Observation bd08cfa4-afd8-4379-8129-51fe1fea28b8 · outbound

This paper cites A style-based generator architecture for generative adversar- ial networks.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A style-based generator architecture for generative adversar- ial networks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.078006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:a1865a6491167589172ae0e123d9c3031a5b2aa801f30bfa168e5c733470ea3c

Observation 2e377043-cd8f-49d2-9bab-2a21f7d1a85a · outbound

This paper cites Musiq: Multi-scale image quality transformer.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Musiq: Multi-scale image quality transformer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.091733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:01696a1b7722498dea3f8b4a6c0e40dfc8882e50348e6be5269cc29e9f877568

Observation 9b53c889-ee76-4457-9a18-170b68a99b45 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.727279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:faf86291a0d5a9f320b4a76756201abf5f48f491784c8e21d3b014a4af2d710d

Observation 653956ce-3acc-4992-8711-f570c1e5b97a · outbound

This paper cites Gligen: Open-set grounded text-to-image generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Gligen: Open-set grounded text-to-image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.074980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:8f322650284ea0921a062327cff226de61fd7955d83e7214330c260eddefa7af

Observation 99ca2bf9-9662-4d98-8565-4c15972f8df9 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.722625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:0b7e018fc642cc85960c11ccb76417a717e5643d183531d00d74f1c512ec2b27

Observation 37d2f4f1-d39b-437b-a7c6-82ea543d0408 · outbound

This paper cites Repaint: Inpainting using denoising diffusion probabilistic models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Repaint: Inpainting using denoising diffusion probabilistic models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.081507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:6b19743f74b9194af7ded8a472eb83ee3dc61484456f3fbefb3d3304579b9847

Observation c62f74dd-006a-4a95-800c-4c0ba9772ca1 · outbound

This paper cites GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation GlyphDraw: Seamlessly Rendering Text with Intricate Spatial Structures in Text-to-Image Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.712176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:c85d5760fac45f3369655956777810655ff86cb8270a0bf237d9d033a78c584e

Observation 9aa96945-010b-4485-a1d9-8efec66cfa4e · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Semantic image synthesis with spatially-adaptive normalization

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.064550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:f72113dc6871d2ad1c2f51c0a2f67d52b1773f63d3e0163bd748ae3406abfd01

Observation 2f98f065-2389-4ab9-9e81-2186815c78df · outbound

This paper cites Scalable diffusion models with transformers.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scalable diffusion models with transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.067867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:a5d6c9ae88806b0a6f7488683fdbb11400ed21497e65f35e93509c651da24895

Observation 1c763e88-d7a7-4421-975b-499e995bb117 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Learning Transferable Visual Models From Natural Language Supervision

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.751151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:8dafe0fe5e7de3210036e19a267c41e790ce0526e155d4c27b53c83a78e2d2a7

Observation d520d741-dfc8-46af-be69-95eb11f1ca9c · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.759730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:b858e75dd7c4174250d1c8a0673729024b621990dfeb2033d5cf088445ad21b1

Observation 42a2e913-e131-413e-a98e-a300fd8a95da · outbound

This paper cites Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Scaling Up Models and Data with $\texttt{t5x}$ and $\texttt{seqio}$

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.769472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:d4e145863b2206ece17b3f8f57937bb05cab9c772661e89e94f32d2a95e8dedd

Observation 354ddd2e-350f-488f-88f1-7f3289730960 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation High-Resolution Image Synthesis with Latent Diffusion Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.755259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:2dd6c65c22711b4d3e3e3621ed3616782b3324d9abee49e9a5b78f759d996931

Observation 9c25e2c3-6a69-4367-af4b-7e84cb1db5a1 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.071652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:7cd6b150d8256088eda17dab990e52cb13cee9640c1082ed68d910972c6eb00e

Observation dc7f86de-d30c-4621-9a55-0870c79384f0 · outbound

This paper cites Ominicontrol: Min- imal and universal control for diffusion transformer.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Ominicontrol: Min- imal and universal control for diffusion transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.088603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:7986777df9ba193dc288303030fc2c6ad0db734e5e7855db8ad3fc4bf2ed3e54

Observation d3e474c7-4cb7-4026-9265-8ff4885844b6 · outbound

This paper cites Omni-video: Democra- tizing unified video understanding and generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omni-video: Democra- tizing unified video understanding and generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.764404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:1dc6851201241c4ba33d06ab3a47c35915937a15436bed8b26adc269acba21e1

Observation c5fc9805-f99b-445c-814b-aba84d2adc10 · outbound

This paper cites AnyText: Multilingual Visual Text Generation And Editing.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation AnyText: Multilingual Visual Text Generation And Editing

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.717505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:37144b6447a2aed9e8fa82dac732c67abbfe4219811bb833305aebdbeea1f464

Observation 53fcbf81-7705-44f2-bdd4-9d5f88fa8717 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.695715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:1cdc68889682593be3bae6fe5f6d72e9121f8afbc0ca3f3c4e93f61f63817826

Observation d1d634a9-f8ad-49a2-a030-998f632dc134 · outbound

This paper cites DeepSeek-OCR: Contexts Optical Compression.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation DeepSeek-OCR: Contexts Optical Compression

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.706394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:b66e27a79655fe93b1b43de8705dc7cd005d35828d6b399e0f9f90d31bbaedc6

Observation 8c761652-e10c-45df-983f-fb3b13f4f356 · outbound

This paper cites Deepseek-ocr 2: Visual causal flow.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Deepseek-ocr 2: Visual causal flow

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:26:20.701364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:376829d8551f1c694108a59c5eb4ebb91cdd7714cb7241b8a3df7f2f4f8a6639

Observation b8705968-57a0-4b7e-b311-35719e7ce089 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.737616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:4a111f6817db9a10c661ac84fdd05da1b825cd00e827a3723f7f978b91957355

Observation 8f3d3cf9-c942-44af-ad24-af88ed4a70f8 · outbound

This paper cites Omnigen: Unified image generation.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Omnigen: Unified image generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.085223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:bff54bce078223c928cf431f85f46169803dee5d15530109511a0fb6e257136c

Observation ab611996-7029-4139-8fd1-d9aa9b7c8fdb · outbound

This paper cites Paint by example: Exemplar-based image editing with diffusion models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Paint by example: Exemplar-based image editing with diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.095108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:dd232eefe5c115bf7b65a2ebc7f94dc16b344348683493b6ddef661174a28de2

Observation fbfdd351-9b9c-4eae-a68f-3166f4fc26da · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.741995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:eb823ca3cdeb5c2527a8608c913655eb4b52af83f85e548f6231c640fe210d2c

Observation 915eae3d-299d-4543-ab68-ea78f422d9ed · outbound

This paper cites Adding Conditional Control to Text-to-Image Diffusion Models.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Adding Conditional Control to Text-to-Image Diffusion Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-22T09:26:20.746997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:dcef6f3db2a2d1cbf1c34e09a3cc3ee928bd056565ad58420ed7ebb175b0f19f

Observation 0677353e-6415-4072-a3fe-cf3a01ae01dd · outbound

This paper cites Unpaired image-to-image translation using cycle-consistent adversarial networks.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unpaired image-to-image translation using cycle-consistent adversarial networks

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.098748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:2c31bf86b182d19a2634849d2745d11f920b555a377b845cd83bc190c4dfe10e

Observation 307974bf-dfd1-48a8-81d6-67dfa5595704 · outbound

This paper cites A task is worth one word: Learning with task prompts for high-quality versatile image inpainting.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation A task is worth one word: Learning with task prompts for high-quality versatile image inpainting

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-22T09:26:21.061206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:2cb071e0bd777ce82fc3ff6fdd59ebbd4530c9d45f46e377091912bad9306b77

Observation 7bbe8cac-2642-4c61-a2d0-8a9a4f521ca5 · outbound

This paper cites an unresolved cited work.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-22T09:26:21.054117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:43e9a442466be8cddaaa749c93c3fbc98379a7321200033fb5476b52b0fe161b

Observation 6ccd17de-a318-4717-b189-794b0442c827 · outbound

This paper cites an unresolved cited work.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-22T09:26:21.057550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:44462e17d787d5203c8d5b3e934d4038f856f796950b2a9e4f91398eb02df7a9

Observation 9726b270-0018-4994-b3af-627d55015f88 · outbound

This paper cites with CLIP semantic instruction.

UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation with CLIP semantic instruction

Reference 42

Resolution
malformed identifier
arxiv_id, observed 2026-05-22T09:26:20.732509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:25:19.066598Z digest=sha256:5de581112bb9f798302167883fc8a02bbe151c261cb75e55b1dc94e5c7e1e822

Pith citing papers

No inbound Pith citation observations are available.