REVIEW 10 cited by
EditVal: Benchmarking Diffusion Based Text-Guided Image Editing Methods
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
A plethora of text-guided image editing methods have recently been developed by leveraging the impressive capabilities of large-scale diffusion-based generative models such as Imagen and Stable Diffusion. A standardized evaluation protocol, however, does not exist to compare methods across different types of fine-grained edits. To address this gap, we introduce EditVal, a standardized benchmark for quantitatively evaluating text-guided image editing methods. EditVal consists of a curated dataset of images, a set of editable attributes for each image drawn from 13 possible edit types, and an automated evaluation pipeline that uses pre-trained vision-language models to assess the fidelity of generated images for each edit type. We use EditVal to benchmark 8 cutting-edge diffusion-based editing methods including SINE, Imagic and Instruct-Pix2Pix. We complement this with a large-scale human study where we show that EditVall's automated evaluation pipeline is strongly correlated with human-preferences for the edit types we considered. From both the human study and automated evaluation, we find that: (i) Instruct-Pix2Pix, Null-Text and SINE are the top-performing methods averaged across different edit types, however {\it only} Instruct-Pix2Pix and Null-Text are able to preserve original image properties; (ii) Most of the editing methods fail at edits involving spatial operations (e.g., changing the position of an object). (iii) There is no `winner' method which ranks the best individually across a range of different edit types. We hope that our benchmark can pave the way to developing more reliable text-guided image editing tools in the future. We will publicly release EditVal, and all associated code and human-study templates to support these research directions in https://deep-ml-research.github.io/editval/.
Forward citations
Cited by 10 Pith papers
-
Tunable Polariton Canalization in Natural van der Waals Oxide
Untreated alpha-V2O5 exhibits frequency-tunable in-plane polariton canalization with unidirectional Poynting-vector flow, mapped by infrared nano-imaging and a permittivity phase diagram.
-
LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.
-
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
On real Reddit photo-editing requests, human judges prefer human edits over AI edits 66% of the time, and AI editors can satisfactorily handle about 33% of requests.
-
Making Image Editing Easier via Adaptive Task Reformulation with Agentic Executions
An MLLM agent that profiles, routes, and reformulates image-editing queries into better-conditioned multi-step operations consistently improves existing editors on hard cases.
-
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
VDE Bench is a new human-annotated dataset and OCR-based evaluation framework for measuring image editing model performance on bilingual dense visual documents.
-
Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.
-
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
A new human-labeled benchmark shows leading vision-language models are unreliable at judging image edits, and the authors' methods improve artifact detection and difference captioning.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
Immunizing Images from Text to Image Editing via Adversarial Cross-Attention
An imperceptible adversarial noise, computed with a LLaVA caption as a stand-in for the unknown edit prompt, disrupts cross-attention in Stable Diffusion-based editors and makes text-guided edits fail.
-
Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications
A survey of instruction-based image editing plus a new 21-task benchmark, CDD-IIE, on which ten open models are scored by human experts.
Discussion (0). Sign in to comment.