Pith. sign in

Paper Citation Record · LEDGER

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2505.19149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19149 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:23:11.635467Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:56:46.514177Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:36:57.287687Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e6166ce-f527-4459-b734-4e4f87ae8b6e · outbound

This paper cites Blended diffusion for text-driven editing of natural images.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Blended diffusion for text-driven editing of natural images

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.290015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.290015Z digest=sha256:61b7943263ae0d3042361c55a1274b34a67457581ed46d4cc347014a750254a7

Observation f7b4572f-e89d-450a-bb12-a64c87fed13e · outbound

This paper cites HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.407350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.407350Z digest=sha256:6fd829719279cf7403c5d3a051f138f73c030627ff5ef917cd698d47b1c082d4

Observation b6e93817-f659-499d-9034-4f6d70423030 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Instructpix2pix: Learning to follow image editing instructions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.484280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.484280Z digest=sha256:d3cc4536a86b2ca8a8413494d6e191e853c0fc8cd4e6fd4998e2273750f508fa

Observation d1b14510-6aab-402a-ac90-5245746e5a62 · outbound

This paper cites Training-free layout control with cross-attention guidance.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Training-free layout control with cross-attention guidance

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.552410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.552410Z digest=sha256:f780c6bb6a50b082d58f93b39df120e6041450a4304d68646f7eeedaf7df17d7

Observation 52d020b0-924a-4840-82e2-70cde09142ac · outbound

This paper cites Zero-shot Image Editing with Reference Imitation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Zero-shot Image Editing with Reference Imitation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:23:12.316270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:05.622338Z digest=sha256:7803ae7973098399b7ec6600768380fb14d5066c7e17295ba8afa81a6994bcf0

Observation ad20f985-a1de-444b-8ac2-69340bf30b0c · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Anydoor: Zero-shot object-level image customization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.729789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.729789Z digest=sha256:0ac95b47df2aaf5d334e77a6d952849e8c78f11f58c98430370a27b23de5205a

Observation d692a309-d3a0-4e9a-873e-1454d442229b · outbound

This paper cites Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Region-Aware Text-to-Image Generation via Hard Binding and Soft Refinement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:05.812424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:05.812424Z digest=sha256:64409c456c30de283e9cec8c975caa7ea302073cd1ee7eacb80b1aaebdf4437a

Observation bb31ca85-84df-4703-b8fa-12cece736a53 · outbound

This paper cites Swiftbrush v2: Make your one-step diffusion model better than its teacher.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Swiftbrush v2: Make your one-step diffusion model better than its teacher

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.375314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:05.893851Z digest=sha256:855470033b7738bb46157139ece773c35d6c4e3040c2d72e7b08b329e931389e

Observation beb1240c-aab3-4b4b-9585-24117ff3a691 · outbound

This paper cites Turboedit: Text- based image editing using few-step diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Turboedit: Text- based image editing using few-step diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:05.995294Z digest=sha256:1357361fbe4d7358a5d145836a0a7f4c16500f9abcba1f9e390d95e83077ab7f

Observation cc29225c-ccd1-442b-a46c-3a0f1384b12e · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Textcrafter: Accurately rendering multiple texts in complex visual scenes.arXiv preprint arXiv:2503.23461, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.129326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.129326Z digest=sha256:33ee546fb9a8b2c1b75d5a4a1b65330658e4ae0acbb1f5a56e5d0adb8479e120

Observation 3be98bb8-76f9-43c3-ac61-cbe3ad2df277 · outbound

This paper cites Complex multistep image-editing dataset, 2025.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Complex multistep image-editing dataset, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.353470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:06.203916Z digest=sha256:a9efb439563cb506e8ba32de36b353e5dc63f46d18e80496b45c98779dc1df93

Observation bc80ed0c-6205-467f-a165-7c0712f4bede · outbound

This paper cites Guiding Instruction-based Image Editing via Multimodal Large Language Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Guiding Instruction-based Image Editing via Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.307126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.307126Z digest=sha256:572248a1e86e97426803ed71a3f12e80d5c851fde40f2205d330bdad3dc5e07d

Observation 6179cd9c-4a9f-4af8-9af2-4803e939fce2 · outbound

This paper cites Renoise: Real image inversion through iterative noising.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Renoise: Real image inversion through iterative noising

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.408275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.408275Z digest=sha256:d036036a7daa5ebd2d2d66f470acaa484035ce9fe2b34d289d06b531d31cfb11

Observation 4fb036e1-7338-4489-b762-6b96481cbebf · outbound

This paper cites Generative adversarial nets.Advances in neural information processing systems, 27, 2014.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generative adversarial nets.Advances in neural information processing systems, 27, 2014

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.481834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.481834Z digest=sha256:dc43ef9fee23bdc3020b3d39cbba169ab7fdc1059ae6ff5dc87cf6b122ac9a7b

Observation 18a4508d-02b4-4113-8c94-b06b31da0179 · outbound

This paper cites Multi-Reward as Condition for Instruction-based Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Multi-Reward as Condition for Instruction-based Image Editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.580357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.580357Z digest=sha256:6385d4c68451bfa9fbe650181cfd8e24c2b4bd7379535c622da7d0e307ea5d1d

Observation e89e5aa7-2461-4333-884d-72f237573a52 · outbound

This paper cites FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FreeEdit: Mask-free Reference-based Image Editing with Multi-modal Instruction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.656115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.656115Z digest=sha256:647c5e412099f1b3d282c441277db481f203af9ce52ca9e35bc0b1423748ae56

Observation e34b6a98-679d-4106-b753-83c1d60c63ed · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.733485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.733485Z digest=sha256:ad85023dff3a3671ee8c01e8150ab818dde3b1ff71553bd2335d5fdc583881b7

Observation bd5d451a-9dc0-4e6b-894f-3e0d9697428b · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.825599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.825599Z digest=sha256:46977945c86fce1083fe0e2d7584eebbb32271cf07c43ca451bf65cf945105e8

Observation 0b967774-43c1-4d18-b6c2-5b40dc7ce02b · outbound

This paper cites Smartedit: Exploring complex instruction- based image editing with multimodal large language models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Smartedit: Exploring complex instruction- based image editing with multimodal large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:06.886247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:06.886247Z digest=sha256:e450aaae54f84d045f06b1fcf54210593d93d2ab6adf36a5ca39622ab7ed49be

Observation fc4172b0-2a0d-4d7a-9a6a-f7500c1b3deb · outbound

This paper cites Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Brushnet: A plug-and-play image inpainting model with decomposed dual-branch diffusion

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.323963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:06.997606Z digest=sha256:c73f8b8ea5aebb2ebd7f67f2b636b30de0f86b970afcf8167895e0064fb01f59

Observation 35adf55b-9118-4577-a2c7-70df5da76cb4 · outbound

This paper cites Image Inpainting Models are Effective Tools for Instruction-guided Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image Inpainting Models are Effective Tools for Instruction-guided Image Editing

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:23:11.973222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:07.084405Z digest=sha256:4c044e4f99526dd8a961159119b182047e64eb752beb1aa306a5eaf4a8edd103

Observation 42379728-bd94-4049-a1bd-ee305d5d2994 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection A style-based generator architecture for generative adversarial networks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.148075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.148075Z digest=sha256:fc4db30e25f9ae0b4408eff01d8ae8c038e7859e04811b05e43ff5bd3095f5ba

Observation b5cfbf33-76e4-4636-8adc-368730060563 · outbound

This paper cites Analyzing and improving the image quality of stylegan.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Analyzing and improving the image quality of stylegan

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.213831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.213831Z digest=sha256:c9d89cc209e93be8918e1d68739d6777e5d3764f2d065233c21f2fab1215cfb8

Observation 1bff09a2-e945-41c2-b4f7-5b5baacb48fd · outbound

This paper cites Auto-encoding variational bayes, 2013.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Auto-encoding variational bayes, 2013

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.318655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.318655Z digest=sha256:7ac6e74f5a7c39c3307b1614829a8d7e5714897612267cebfa10da9c0a6633ec

Observation 7cbfeb69-1b04-4dba-b626-80087ca8934b · outbound

This paper cites Generating images with multimodal language models.Advances in Neural Information Processing Systems, 36:21487–21506, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Generating images with multimodal language models.Advances in Neural Information Processing Systems, 36:21487–21506, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.418627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.418627Z digest=sha256:02bbe564b030cee1662f0fe5f32b3ec1d15fe4727d75dc4a0bfc3c9a514c47f6

Observation d5ff60fe-417b-4454-8ebd-7287f5e77646 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection LLaVA-OneVision: Easy Visual Task Transfer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.470068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.470068Z digest=sha256:f9aa4f8de49281bb4a78c06b1b2f4cc77c26f68882534487b65d69506e2b9dce

Observation 6c8b85b2-4ba2-4ead-9039-fa2ec3d732cb · outbound

This paper cites Q-Insight: Understanding Image Quality via Visual Reinforcement Learning.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.552002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.552002Z digest=sha256:658b1ac8f2d84a84890fea923e8642a234e5b312903ab3066fbc87dcd06517ed

Observation 4eb9912e-f721-4025-97c0-40031ef509ed · outbound

This paper cites Resvr: Joint rescaling and viewport rendering of omnidirectional images.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Resvr: Joint rescaling and viewport rendering of omnidirectional images

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.292457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:07.636743Z digest=sha256:a4dabc8a3dee00f0b9114cd71042f98af85593ec3f597d6ce661fd9ae8cf7e41

Observation a4e84820-72e2-4453-b3ea-d26e37bf8cac · outbound

This paper cites OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.710615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.710615Z digest=sha256:bcbb1164a1631e30105df59cf987568a9daf37d53256f22535043e43aa5ddb2a

Observation 1604f9f3-00db-4a10-b30c-a40fd6fab71e · outbound

This paper cites BrushEdit: All-In-One Image Inpainting and Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection BrushEdit: All-In-One Image Inpainting and Editing

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.789904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.789904Z digest=sha256:733fde98399ff300c3cb090b419f7a60bd3b0580cfd98a21ead3ccd5dbef6b41

Observation aff9180f-e37b-4128-92cc-dfa09b296a79 · outbound

This paper cites Adversarial supervision makes layout-to-image diffusion models thrive.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adversarial supervision makes layout-to-image diffusion models thrive

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.281552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:07.849807Z digest=sha256:c0edb8c671b22526902dc0a5a20b3ca8898e34cf180ab3bf1a07c83e85b04cb6

Observation 918830ec-ccf9-4edd-aa77-01b5b8bc6c3d · outbound

This paper cites Drag your noise: Interactive point-based editing via diffusion semantic propagation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your noise: Interactive point-based editing via diffusion semantic propagation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.269834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:07.911521Z digest=sha256:cac4e6f320d3f561564b69d0e3968d7746c8181168c5b63a31193d527f8003a4

Observation 0216285b-ef45-4c89-9d0a-cc1b30b8f39b · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:07.989663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:07.989663Z digest=sha256:71a782088324064c0ad05b28be07c21dc57233b0143fcfdd2fb617ea6321fd56

Observation f93afe6a-2e55-44bd-b5be-7f4a4feb3d03 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Step1X-Edit: A Practical Framework for General Image Editing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.052387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.052387Z digest=sha256:603bb69c08c185e953a1b8a1df1129162bc2be4c1e974744c334daea8170b258

Observation b4db43ad-cf9d-489f-b737-f0535a7cf1c4 · outbound

This paper cites Magicquill: An intelligent interactive image editing system.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicquill: An intelligent interactive image editing system

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.252940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:08.111111Z digest=sha256:bfabaa5d5eb6d19abb01b6c6d51993e7b9b9e0b4c5191db8523b2f6cf426d127

Observation 26de2773-2e35-41dc-80b3-1fe414a377ba · outbound

This paper cites Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Diffeditor: Boosting accuracy and flexibility on diffusion-based image editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.156958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.156958Z digest=sha256:f45b0c6d88fc5b60bc09d71df5de6f955ded0ab57051edb5c6d493895b33ed4b

Observation a1e50af1-e713-44a2-90d2-cf6413222ae1 · outbound

This paper cites Dragondiffusion: Enabling drag-style manipulation on diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragondiffusion: Enabling drag-style manipulation on diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.237351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:08.216468Z digest=sha256:d2173bfa85b07d3034a7494a6912c9abab4dc6cfe09b1cf660fdf2056215a31d

Observation bc87defe-6221-4c52-85a4-5b9f57067a60 · outbound

This paper cites T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection T2i-adapter: Learning adapters to dig out more controllable ability for text-to-image diffusion models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.307630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.307630Z digest=sha256:2bf4e1edbdef9fa44de0b6682a26e0abf380743e483baf1e958287d357f6079b

Observation 709b76fe-869d-46c9-ad80-238c731b80df · outbound

This paper cites Handiffuser: Text-to-image generation with realistic hand appearances.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Handiffuser: Text-to-image generation with realistic hand appearances

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.221725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:08.367757Z digest=sha256:a4e822a0c34edae8e6543a2d4b51545884e81d2818d64eaae604cd236e717890

Observation 8935cc25-1e62-46ae-8909-beac8916a29e · outbound

This paper cites Transfer between Modalities with MetaQueries.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Transfer between Modalities with MetaQueries

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.422122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.422122Z digest=sha256:2089abcf798d7acec6088dc6a33d2ad5618ff3e87ab6544da85efa56d5d1e9ac

Observation 261e470d-a3b1-47ce-b1f6-9b32998d414a · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.494940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.494940Z digest=sha256:dfcc85cf8d17da6976c976d300c644a9b3fb9d69cebaab9c7cbaa767bbbb2a68

Observation 248cb43e-66ce-4cf6-8f70-34f87b5d6722 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.555573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.555573Z digest=sha256:168c5969ba75de6157fa013a3ecc8354f72f6cf2b1d15d7abe63c863a2ea0d45

Observation 1fe2ce5e-4ff4-40db-bbb9-4a029952f774 · outbound

This paper cites Learning transferable visual models from natural language supervision.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Learning transferable visual models from natural language supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.617044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.617044Z digest=sha256:53d2591b19bde9d83eccd00cbe1b09ddc9488856bd4131aee60b54be60278a14

Observation 1266e1bc-ecbe-42e0-a468-dc2f0145448f · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection High- resolution image synthesis with latent diffusion models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.712174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.712174Z digest=sha256:ab5c4b052055548b62ac7b7603ff709294cbb139ec1cecd0a82c6ba7eebf3a57

Observation c1320835-73ca-48a9-a69c-5a5bfcbcd9c2 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Photorealistic text-to-image diffusion models with deep language understanding.Advances in neural information processing systems, 35:36479–36494, 2022

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:08.840800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:08.840800Z digest=sha256:667b7a4e699c698e2fdb00b9f5ea6e1ec5085e357f2b3dd92a3c402a272b0b88

Observation 9e154348-4e85-4f2d-ae28-df9b0607a3ce · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu edit: Precise image editing via recognition and generation tasks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.180735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:08.922371Z digest=sha256:42e81dd95dd483709c1dacbc4d3cab852417bd75fcc377e4eaf744ca2e1f5bc5

Observation 3dbd24ee-55f1-4e00-b648-d9568b50ec81 · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interactive point-based image editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Dragdiffusion: Harnessing diffusion models for interactive point-based image editing

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.001818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.001818Z digest=sha256:59c5eb93ce64f248d207cc7ebed406e243455a2919fe7c8569fe6767e43ce5e2

Observation 74a839d2-7d8f-40ea-b90e-8fc5727eed50 · outbound

This paper cites Insert Anything: Image Insertion via In-Context Editing in DiT.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Insert Anything: Image Insertion via In-Context Editing in DiT

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.117476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.117476Z digest=sha256:39f86066ec6fa9b87f4fb5b8e684d245a609bfd9c8a371e1ce2f904ea230a18a

Observation 38c37376-6b06-426a-9519-dd98316ca5ec · outbound

This paper cites Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Invertible Consistency Distillation for Text-Guided Image Editing in Around 7 Steps

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.205192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.205192Z digest=sha256:4b8ca52df8b5a74594a5fd4525805b920473966a710f6bcddbb3c89643e47822

Observation 1a6114cb-3d3d-428e-8ad6-c3e28434033a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.297602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.297602Z digest=sha256:901ec75f253b9bc91f73dc773bca66fcd47138fb222f3cbf19c00f94a5f64d96

Observation e986d506-1f5c-4c35-b9c7-24a90e475e81 · outbound

This paper cites Gpt-4o system card, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Gpt-4o system card, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.385434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.385434Z digest=sha256:1d5b0f5fce8960080733efc21b58f0218e9210b9ffae16855f75de16b65c6afa

Observation f90c1bc6-8a38-4195-af3e-38350038b4a6 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.580072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.580072Z digest=sha256:b8aa39b9ce0a32beded8927db49b6f2a4923ed380c94a3228fb00b563cf519e7

Observation 40030a32-4388-47ac-be64-d4212c2b8a14 · outbound

This paper cites Plug-and-play diffusion features for text-driven image-to-image translation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Plug-and-play diffusion features for text-driven image-to-image translation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.686407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.686407Z digest=sha256:95fd9e3c1c9736a98f6b1538e74c30b00d111d3eab868a46d1e011c427e5b730

Observation 94e5fa03-dff7-4e81-bb65-58b861e3ee4b · outbound

This paper cites FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection FlexEdit: Marrying Free-Shape Masks to VLLM for Flexible Image Editing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.768033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.768033Z digest=sha256:6cfa3e74fddae29f77343d9e7df376d688f4dde37d0174b9509962489ca1509f

Observation 850fd77c-bdfc-4d80-bbda-a51ff526850d · outbound

This paper cites 360dvd: Controllable panorama video generation with 360-degree video diffusion model.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection 360dvd: Controllable panorama video generation with 360-degree video diffusion model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:09.890616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:09.890616Z digest=sha256:1c16cafc81acfefb3e2fcce1c366cb419c31a7768df40eecc0fc7b828c020b30

Observation 3206f2f9-55c5-4c1a-ba13-1b381ed42fe6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Emu3: Next-Token Prediction is All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.024028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.024028Z digest=sha256:41471b058087e82545b5bfb39b03612a662881601f95385c7105dfc3cf2b40db

Observation d922f159-7107-48d9-afba-24819e1eeba1 · outbound

This paper cites In- stancediffusion: Instance-level control for image generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection In- stancediffusion: Instance-level control for image generation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.144422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:10.133279Z digest=sha256:1e39b25b2d8fa6f4fd0d186bdb23367ba79166b6e9521ff7f3797f8ad47a7392

Observation 6e286aec-fcff-49d6-9e79-82842c48ebc6 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image generation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Genartist: Multimodal llm as an agent for unified image generation and editing.Advances in Neural Information Processing Systems, 37:128374–128395, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:13.051692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:10.221516Z digest=sha256:2484950840c44b94e2bab91d59bcb8ce40a78f8994b1a1f957a5941086b5ce4c

Observation 1516427e-e951-4437-bea2-2cad37912ca3 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600– 612, 2004.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600– 612, 2004

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.370380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.370380Z digest=sha256:c2dbbd605abd4dc65eaa730e00efc3c64120c5237ff29e8bdb141a5aecb622c0

Observation 57fa69e6-d237-49bf-b1a2-53afc3db4e95 · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.504956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.504956Z digest=sha256:c15a81901e54b9cee008694376d73d22ab8bd91da090917985ff1e7e82b9fff2

Observation 9a5c87af-ddff-4ea1-9bcd-82103e720673 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.608892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.608892Z digest=sha256:ec674c20fda266f1bbb3e89fd7eb44f8d4596210569a986529f2287a74d69e7c

Observation d328811f-6996-4525-912b-34f4991f13f7 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.781055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.781055Z digest=sha256:0b83c9686d16a43c73e132896922c43d860fc73dc108adf58179bf763356912b

Observation 07d7c50d-9c33-495c-96bb-081da015947a · outbound

This paper cites AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:10.975139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:10.975139Z digest=sha256:33ca86913b12fda82909f8009d4ca577110194a3311b49459701da9342712242

Observation d1a5f95d-8415-4ef0-9669-7b883dcc3d10 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449, 2023

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.145234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.145234Z digest=sha256:f230d05ccea7e3d5c298bdeeb20711777f372ab2054bb1af2d361d999f2da59d

Observation 3abffb1d-1c68-472d-97e3-a1aa62985ec2 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Adding conditional control to text-to-image diffusion models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:12.809196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:11.296196Z digest=sha256:1395c9ba199644f5162a8933a0570b8f6a9a77f707817dd962434213000148e7

Observation f58d4b08-9b8a-4631-973f-5bfe97a10850 · outbound

This paper cites The unrea- sonable effectiveness of deep features as a perceptual metric.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection The unrea- sonable effectiveness of deep features as a perceptual metric

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.402858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.402858Z digest=sha256:b6aa909722741dd31491a32b0ea0839a339063dc939c8719ae9831aafa150be8

Observation c353a799-e1be-4c50-8eb6-d00667549de2 · outbound

This paper cites GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:11.514069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:11.514069Z digest=sha256:b22917ce9ec212d3364f5beddbee41648e13c9fc18120191037781b35fb8a4a9

Observation c66d4d9f-6be9-41e9-8238-2925bfbf133b · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024.

MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Processing Systems, 37:3058–3093, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:23:12.600148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T14:23:11.635467Z digest=sha256:c4024ef44d266ba804e6b90bbf984bd9965cdbae56ab5376a885d38545b2fdfa

Pith citing papers

Observation 0025a0dc-b722-4309-ab78-5eb73859901f · inbound

TextWand: A Unified Framework for Scene Text Editing cites this paper.

TextWand: A Unified Framework for Scene Text Editing MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:36:57.289066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T01:56:46.514177Z digest=sha256:29506e6591099eef78cc362f5221ec9a0ff6e5d4bedc592eb45f98da43a4dbf7