Pith. sign in

REVIEW 8 cited by

HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09990 v1 pith:RBCTCVX5 submitted 2024-04-15 cs.CV cs.AI

classification cs.CVcs.AI
keywords editingimagehigh-qualityhq-editmodelsalignmentdatadataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This study introduces HQ-Edit, a high-quality instruction-based image editing dataset with around 200,000 edits. Unlike prior approaches relying on attribute guidance or human feedback on building datasets, we devise a scalable data collection pipeline leveraging advanced foundation models, namely GPT-4V and DALL-E 3. To ensure its high quality, diverse examples are first collected online, expanded, and then used to create high-quality diptychs featuring input and output images with detailed text prompts, followed by precise alignment ensured through post-processing. In addition, we propose two evaluation metrics, Alignment and Coherence, to quantitatively assess the quality of image edit pairs using GPT-4V. HQ-Edits high-resolution images, rich in detail and accompanied by comprehensive editing prompts, substantially enhance the capabilities of existing image editing models. For example, an HQ-Edit finetuned InstructPix2Pix can attain state-of-the-art image editing performance, even surpassing those models fine-tuned with human-annotated data. The project page is https://thefllood.github.io/HQEdit_web.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.

  2. Making Implicit Preservation Intent Explicit in Conversational Image Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Conversational image editors fail to restore temporarily occluded content; ReSpec fixes this by explicitly selecting historical visual references and rewriting instructions to guide restoration.

  3. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  4. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.

  5. Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.

  6. JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

    cs.CV 2026-07 conditional novelty 5.0 of 10

    JarvisHub open-sources a three-layer canvas-state, protocol-bridge, and agent-runtime harness so multimodal creative agents can inspect and update a shared editable project graph over long workflows.

  7. SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A unified multimodal model can improve its own text-to-image generation by using its understanding module as a rewarder in a global-plus-local reward-weighted training loop.

  8. Towards Efficient Exemplar Based Image Editing with Multimodal VLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ReEdit transfers exemplar-based edits to new images by conditioning Stable Diffusion on a LLaVA-written caption plus a CLIP edit-direction vector, with no per-example optimization.

Pith tools