Pith. sign in

Paper Citation Record · LEDGER

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.14555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14555 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:26:57.014104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.608389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e345701-4149-431f-a655-25e323bdb32e · inbound

A Survey on Pre-Trained Diffusion Model Distillations cites this paper.

A Survey on Pre-Trained Diffusion Model Distillations A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:26:57.014104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:26:57.014104Z digest=sha256:c2a05953f0c2445c9fdaa6082e45f66a61048a1220df3428d2d2f0978d943e76

Observation 878b40ac-f11b-43b7-a642-64d481dac74d · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:18.230361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:18.230361Z digest=sha256:a711e5136ea85bfc89234e9e9dca2784562ce23d7fb32de49f55a4ee764c06a9

Observation 6c2fbbb3-ed15-49b7-9928-5a3138001b26 · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:34.712223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:34.712223Z digest=sha256:da57499d316ed6481e016f4374a2143f523447030166fca68593b2aaed853ce9

Observation 40b743f7-bc70-444f-8b57-053c8fc6fc83 · inbound

AnyI2V: Animating Any Conditional Image with Motion Control cites this paper.

AnyI2V: Animating Any Conditional Image with Motion Control A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.267209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.267209Z digest=sha256:8bcd60a6b333f9a5b38f47bcc1ccb6aff844db4fd64e8d2d9e7a442742fd5ece

Observation ba30208d-2536-458a-92f1-067bd37636db · inbound

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors cites this paper.

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:25.575695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:25.575695Z digest=sha256:f06b5932f275c7258170189d90ee3244fc7db53cacaca4c31dec16b25083e980

Observation 8165f4f7-8e44-44bb-922b-a229563d7772 · inbound

CharaConsist: Fine-Grained Consistent Character Generation cites this paper.

CharaConsist: Fine-Grained Consistent Character Generation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:56.151197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:10:56.151197Z digest=sha256:c72ad263cb0b0e8be091656f7dc7918bff3f65b43b8b2f491f16bfd56f5dab6f

Observation e6af47d5-4af3-440e-8548-3eb7ebda4aa4 · inbound

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding cites this paper.

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.614068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:52:26.262129Z digest=sha256:e26c2f73c5a827fb03ba84839a5963ba0e3b12622d36e2c09905e1f0fc021e0c

Observation 59122a93-fd2f-47a9-9bbf-ac778dbaa19e · inbound

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects cites this paper.

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:29.034804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:10:35.320940Z digest=sha256:ce313903b8032a2873fec69eb0f680ad57f828095d1af99ae6e2c8f4898d0012

Observation 892abf8a-cee7-46cd-9aea-c74da73112c2 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.239593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:c2294c315241901c170b8eac39d5857dd6460e60c290aeb546d1e1e12db46347

Observation 4396eee7-f712-44f0-95e4-134bbbd06b38 · inbound

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking cites this paper.

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.005615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T23:06:29.200541Z digest=sha256:746a2e54225e11b07d3ca628cae25d490373818d61f6c573755dd6752f555e6e

Observation 0b2d1af0-f569-4eb2-8d1a-0abb21da0e73 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.609856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:c076ffdd4c5a0ea6a72562104ff6da21465c74ab24a115f51b3336143f96cfe0

Observation 434f46b2-f6e4-45c2-bbba-52c4b15aa7ec · inbound

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers cites this paper.

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:45:18.568969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:45:18.568969Z digest=sha256:f1e5888338828d8de3f68bfe21a60bddbd1209db81e9e91b22d76efb1f84d338