Pith. sign in

REVIEW 15 cited by

InstructPix2Pix: Learning to Follow Image Editing Instructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.09800 v2 pith:NTOX36GL submitted 2022-11-17 cs.CV cs.AIcs.CLcs.GRcs.LG

classification cs.CVcs.AIcs.CLcs.GRcs.LG
keywords modelinstructionseditingimageimagesdatadiffusionedits
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image. To obtain training data for this problem, we combine the knowledge of two large pretrained models -- a language model (GPT-3) and a text-to-image model (Stable Diffusion) -- to generate a large dataset of image editing examples. Our conditional diffusion model, InstructPix2Pix, is trained on our generated data, and generalizes to real images and user-written instructions at inference time. Since it performs edits in the forward pass and does not require per example fine-tuning or inversion, our model edits images quickly, in a matter of seconds. We show compelling editing results for a diverse collection of input images and written instructions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples

    cs.CV 2025-05 conditional novelty 7.0 of 10

    Mask-guided self-attention fusion creates well-aligned target images that stay visually close to poorly-aligned base images, with full denoising trajectories, and DPO on these pairs improves alignment.

  2. ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval

    cs.CV 2025-05 conditional novelty 7.0 of 10

    ConText-CIR uses a concept-consistency loss plus an LLM-based data pipeline to set new state-of-the-art results on CIRR and CIRCO.

  3. OSVE: One Step Video Editing with One Step Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...

  4. Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration

    cs.DC 2026-07 conditional novelty 6.0 of 10

    Trace-guided fine-grained memory control and offline joint planning raise diffusion serving SLO attainment by up to 3.7× while cutting configuration search from hours to minutes.

  5. Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing

    cs.CV 2026-03 conditional novelty 6.0 of 10

    RL3DEdit fine-tunes FLUX-Kontext with GRPO using VGGT confidence and pose rewards to produce multi-view consistent 3D scene edits in a single pass.

  6. Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Cultural Counterfactuals — same person placed in different cultural contexts — shows that LVLMs vary salary, rent, and character judgments with the depicted religion, nationality, and income level.

  7. Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.

  8. Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An iterative text-image plan generation framework with pseudo-PDDL visual feedback improves multimodal consistency and visual coherence across Mistral-7B, Gemini-1.5, and GPT-4o.

  9. EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new human-labeled benchmark shows leading vision-language models are unreliable at judging image edits, and the authors' methods improve artifact detection and difference captioning.

  10. Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.

  11. Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...

  12. VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics

    cs.LG 2025-06 conditional novelty 5.0 of 10

    VectorEdits is a 271k-pair dataset and benchmark for text-guided vector image editing, and current LLMs fail to outperform a no-edit baseline.

  13. Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations

    cs.MM 2025-06 conditional novelty 5.0 of 10

    MultiFakeVerse provides 845,286 person-centric images edited through VLM-generated instructions; state-of-the-art deepfake detectors and human observers misclassify a large fraction of them.

  14. Pinterest Canvas: Large-Scale Image Generation at Pinterest

    cs.CV 2026-03 conditional novelty 4.0 of 10

    A FLUX-style base diffusion model plus task-specific fine-tunes and product-preserving pipelines yields double-digit Pinterest ads engagement lifts and higher no-defect rates than GPT-Image, FLUX Kontext, and Nano Banana.

  15. DiffIER: Optimizing Diffusion Models with Iterative Error Reduction

    cs.CV 2025-08 reject novelty 4.0 of 10

    DiffIER claims that iteratively minimizing the distance between a diffusion model's predicted noise and a random Gaussian sample at each inference step improves generation quality.

Pith tools