Pith. sign in

REVIEW 5 cited by

$\texttt{Complex-Edit}$: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.13143 v1 pith:3IONYCKK submitted 2025-04-17 cs.CV cs.AI

classification cs.CVcs.AI
keywords editingmodelsbenchmarkcomplexityinstructionsinstructionperformancesynthetic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

We introduce $\texttt{Complex-Edit}$, a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity. To develop this benchmark, we harness GPT-4o to automatically collect a diverse set of editing instructions at scale. Our approach follows a well-structured ``Chain-of-Edit'' pipeline: we first generate individual atomic editing tasks independently and then integrate them to form cohesive, complex instructions. Additionally, we introduce a suite of metrics to assess various aspects of editing performance, along with a VLM-based auto-evaluation pipeline that supports large-scale assessments. Our benchmark yields several notable insights: 1) Open-source models significantly underperform relative to proprietary, closed-source models, with the performance gap widening as instruction complexity increases; 2) Increased instructional complexity primarily impairs the models' ability to retain key elements from the input images and to preserve the overall aesthetic quality; 3) Decomposing a complex instruction into a sequence of atomic steps, executed in a step-by-step manner, substantially degrades performance across multiple metrics; 4) A straightforward Best-of-N selection strategy improves results for both direct editing and the step-by-step sequential approach; and 5) We observe a ``curse of synthetic data'': when synthetic data is involved in model training, the edited images from such models tend to appear increasingly synthetic as the complexity of the editing instructions rises -- a phenomenon that intriguingly also manifests in the latest GPT-4o outputs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    Dress-ED is the first large-scale benchmark unifying virtual try-on, try-off, and text-guided garment editing with 146k verified samples plus a multimodal diffusion baseline.

  2. GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

    cs.CV 2025-07 conditional novelty 6.0 of 10

    The paper introduces GPT-IMAGE-EDIT-1.5M, a 1.5-million-triplet image-editing dataset refined by GPT-4o, and shows that fine-tuning FluxKontext on it achieves state-of-the-art open-source scores on GEdit-EN and Complex-Edit.

  3. ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Introduces a benchmark for chain-dependent image editing instructions plus a region-aware consistency metric, and shows a chain-of-thought prompt improves a Gemini-based editor.

  4. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  5. Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications

    cs.CV 2026-07 conditional novelty 4.0 of 10

    A survey of instruction-based image editing plus a new 21-task benchmark, CDD-IIE, on which ten open models are scored by human experts.

Pith tools