Pith. sign in

Paper Citation Record · LEDGER

A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2406.14555.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.14555 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:26:57.014104Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:48:32.608389Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e345701-4149-431f-a655-25e323bdb32e · inbound

A Survey on Pre-Trained Diffusion Model Distillations cites this paper.

A Survey on Pre-Trained Diffusion Model Distillations A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:26:57.014104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T05:26:57.014104Z digest=sha256:277b1470591ccb568ac2d89e715466d995da1bc36a7aff5c519f097572ba6947

Observation 878b40ac-f11b-43b7-a642-64d481dac74d · inbound

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models cites this paper.

KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:18.230361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:18.230361Z digest=sha256:0d60107ccb7bc477d2fba6238431da0e3e933272149c5849ce1d5ad809adffdd

Observation 6c2fbbb3-ed15-49b7-9928-5a3138001b26 · inbound

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs cites this paper.

MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:34.712223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:34.712223Z digest=sha256:ad28f66bce5579a45b7dee81e9ae2bfa23c5c0f4a70f7d611c302d02d75fde4a

Observation 40b743f7-bc70-444f-8b57-053c8fc6fc83 · inbound

AnyI2V: Animating Any Conditional Image with Motion Control cites this paper.

AnyI2V: Animating Any Conditional Image with Motion Control A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.267209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.267209Z digest=sha256:948ccfbc76b64d2af252abd6334d5b8e2ba8a111e309b041274f1c5ec9e54d02

Observation ba30208d-2536-458a-92f1-067bd37636db · inbound

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors cites this paper.

Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:34:25.575695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:34:25.575695Z digest=sha256:d10a312a838efc00051bf34ff6c6df8295461c748639202895f4b906157fef33

Observation 8165f4f7-8e44-44bb-922b-a229563d7772 · inbound

CharaConsist: Fine-Grained Consistent Character Generation cites this paper.

CharaConsist: Fine-Grained Consistent Character Generation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:10:56.151197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:10:56.151197Z digest=sha256:933cb0c11a387e19311abe996dcd7617217edca99975c7371f1fe489c6d1cb69

Observation e6af47d5-4af3-440e-8548-3eb7ebda4aa4 · inbound

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding cites this paper.

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:15.614068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T16:52:26.262129Z digest=sha256:591b051e379452342d32ea6f22891a655c6f4976e05811aedf7273d6ff900727

Observation 59122a93-fd2f-47a9-9bbf-ac778dbaa19e · inbound

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects cites this paper.

3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:29.034804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:10:35.320940Z digest=sha256:6351fb7828e78cbb414f567bffa7cd7aa10e17277ba0f0398f1b46f27fc41931

Observation 892abf8a-cee7-46cd-9aea-c74da73112c2 · inbound

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models cites this paper.

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.239593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:01:01.427273Z digest=sha256:f4128e04f9b40c5769ffab21ff026e5c3c5ba6e3bcb5abf5519f2e44bfc2655f

Observation 4396eee7-f712-44f0-95e4-134bbbd06b38 · inbound

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking cites this paper.

Rethinking Scribble-Guided Image Editing: Generalization, Instruction Adherence, and Multi-Tasking A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:14:02.005615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:06:29.200541Z digest=sha256:d5354bb439825c584205d659ee6b50e26bc4c15768a69d5ed0bd89c12f513f97

Observation 0b2d1af0-f569-4eb2-8d1a-0abb21da0e73 · inbound

Seek to Segment: Active Perception for Panoramic Referring Segmentation cites this paper.

Seek to Segment: Active Perception for Panoramic Referring Segmentation A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:48:32.609856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T14:39:22.617747Z digest=sha256:8aab2b3efa79976c32fb7e896c0258f7d4594758c8c9e2a4072e31c3df82221f

Observation 434f46b2-f6e4-45c2-bbba-52c4b15aa7ec · inbound

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers cites this paper.

Appearance Pointers -- Multimodal Region Control of Diffusion Transformers A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T12:45:18.568969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:45:18.568969Z digest=sha256:fe4f380572c8c3674167aba16f8128b1a697a15b1de32c250a0da9b1c43da710