REVIEW 12 cited by
InstructPix2Pix: Learning to Follow Image Editing Instructions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image. To obtain training data for this problem, we combine the knowledge of two large pretrained models -- a language model (GPT-3) and a text-to-image model (Stable Diffusion) -- to generate a large dataset of image editing examples. Our conditional diffusion model, InstructPix2Pix, is trained on our generated data, and generalizes to real images and user-written instructions at inference time. Since it performs edits in the forward pass and does not require per example fine-tuning or inversion, our model edits images quickly, in a matter of seconds. We show compelling editing results for a diverse collection of input images and written instructions.
Forward citations
Cited by 12 Pith papers
-
OSVE: One Step Video Editing with One Step Diffusion Models
OSVE performs text-guided video editing with one-step diffusion models by training a single-pass inversion encoder and unifying frame latents for cross-frame attention, achieving quality comparable to multi-step metho...
-
Xema: Efficient Diffusion Serving through Fine-Grained Memory Management and Auto-Configuration
Trace-guided fine-grained memory control and offline joint planning raise diffusion serving SLO attainment by up to 3.7× while cutting configuration search from hours to minutes.
-
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
RL3DEdit fine-tunes FLUX-Kontext with GRPO using VGGT confidence and pose rewards to produce multi-view consistent 3D scene edits in a single pass.
-
Cultural Counterfactuals: Evaluating Cultural Biases in Large Vision-Language Models with Counterfactual Examples
Cultural Counterfactuals — same person placed in different cultural contexts — shows that LVLMs vary salary, rent, and character judgments with the depicted religion, nationality, and income level.
-
Balancing Preservation and Modification: A Region and Semantic Aware Metric for Instruction-Based Image Editing
A region and semantic aware metric for instruction-based image editing, built from LLM parsing plus detection, segmentation, and CLIP directional similarity, reports the highest human alignment among compared metrics.
-
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
An iterative text-image plan generation framework with pseudo-PDDL visual feedback improves multimodal consistency and visual coherence across Mistral-7B, Gemini-1.5, and GPT-4o.
-
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
A new human-labeled benchmark shows leading vision-language models are unreliable at judging image edits, and the authors' methods improve artifact detection and difference captioning.
-
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
A compact 4B image generation/editing system with a fast one-step VAE, native-resolution packing, RL alignment, and 4-step distillation reports competitive benchmarks against 6B–80B open models.
-
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
ATLAS adds a Think–Plan–Paint loop with shared positional tokens to unified MLLMs, plus RL-based layout alignment, achieving large reported gains over prior layout-based unified models on compositional image generatio...
-
VectorEdits: A Dataset and Benchmark for Instruction-Based Editing of Vector Graphics
VectorEdits is a 271k-pair dataset and benchmark for text-guided vector image editing, and current LLMs fail to outperform a no-edit baseline.
-
Pinterest Canvas: Large-Scale Image Generation at Pinterest
A FLUX-style base diffusion model plus task-specific fine-tunes and product-preserving pipelines yields double-digit Pinterest ads engagement lifts and higher no-defect rates than GPT-Image, FLUX Kontext, and Nano Banana.
-
DiffIER: Optimizing Diffusion Models with Iterative Error Reduction
DiffIER claims that iteratively minimizing the distance between a diffusion model's predicted noise and a random Gaussian sample at each inference step improves generation quality.
Discussion (0). Sign in to comment.