REVIEW 17 cited by
SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this technical report, we introduce SEED-Data-Edit: a unique hybrid dataset for instruction-guided image editing, which aims to facilitate image manipulation using open-form language. SEED-Data-Edit is composed of three distinct types of data: (1) High-quality editing data produced by an automated pipeline, ensuring a substantial volume of diverse image editing pairs. (2) Real-world scenario data collected from the internet, which captures the intricacies of user intentions for promoting the practical application of image editing in the real world. (3) High-precision multi-turn editing data annotated by humans, which involves multiple rounds of edits for simulating iterative editing processes. The combination of these diverse data sources makes SEED-Data-Edit a comprehensive and versatile dataset for training language-guided image editing model. We fine-tune a pretrained Multimodal Large Language Model (MLLM) that unifies comprehension and generation with SEED-Data-Edit. The instruction tuned model demonstrates promising results, indicating the potential and effectiveness of SEED-Data-Edit in advancing the field of instructional image editing. The datasets are released in https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit.
Forward citations
Cited by 17 Pith papers
-
Understanding Generative AI Capabilities in Everyday Image Editing Tasks
On real Reddit photo-editing requests, human judges prefer human edits over AI edits 66% of the time, and AI editors can satisfactorily handle about 33% of requests.
-
LoRA of Change: Learning to Generate LoRA for the Editing Instruction from A Single Before-After Image Pair
A hypernetwork generates a per-instruction LoRA from a before-after image pair, and a reverse training loss allows learning from paired data alone.
-
Under One Sun: Multi-Object Generative Perception of Materials and Illumination
Factorizing video editing into semantic-token anchoring and motion-restoration pre-training produces strong zero-shot and SOTA open-source instruction-guided video edits without heavy external structural priors.
-
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation
A diffusion-transformer framework with VLM-grounded masked attention and VAE dropout improves identity and prompt fidelity for multi-subject image generation.
-
ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.
-
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.
-
ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
Introduces a benchmark for chain-dependent image editing instructions plus a region-aware consistency metric, and shows a chain-of-thought prompt improves a Gemini-based editor.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
Explanatory Instructions: Towards Unified Vision Tasks Understanding and Zero-shot Generalization
Explanatory instructions, detailed text descriptions of image-to-image transformations, are introduced with a 12M-pair dataset and show qualitative evidence of zero-shot generalization on unseen vision tasks.
-
HumanEdit: A High-Quality Human-Rewarded Dataset for Instruction-based Image Editing
HumanEdit provides 5,751 human-annotated, high-resolution image editing pairs with masks and a six-type instruction taxonomy, plus baseline benchmark results.
-
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
A large new benchmark and an offline judge model for open-ended interleaved image-text generation, with IntJudge matching human agreement better than GPT-4o.
-
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
A large automatically collected image editing dataset with 25 editing types and a task-aware diffusion model trained on it achieve new state-of-the-art results on two standard image editing benchmarks.
-
GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment Design
GenTune improves AI image refinement by tracing image regions back to prompt labels and allowing element-level, semantic-guided edits.
-
Ovis-U1 Technical Report
A 3B unified multimodal model with a diffusion decoder and bidirectional refiner achieves competitive understanding, generation, and editing benchmark scores.
-
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.
-
EditAR: Unified Conditional Generation with Autoregressive Models
EditAR shows a single next-token autoregressive model can handle image editing and translation tasks, with competitive FID on translation benchmarks.
-
Hands-off Image Editing: Language-guided Editing without any Task-specific Labeling, Masking or even Training
An instruction-guided image editor that needs no training, labels, or masks: an LLM writes before/after captions and their embedding difference guides Stable Diffusion.
Discussion (0). Continue with ORCID to comment.