Pith. sign in

REVIEW 12 cited by

ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02487 v3 pith:VO4XBESE submitted 2025-01-05 cs.CV

classification cs.CV
keywords modelstasksimageeditingfinetuningfluxmodelstage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We report ACE++, an instruction-based diffusion framework that tackles various image generation and editing tasks. Inspired by the input format for the inpainting task proposed by FLUX.1-Fill-dev, we improve the Long-context Condition Unit (LCU) introduced in ACE and extend this input paradigm to any editing and generation tasks. To take full advantage of image generative priors, we develop a two-stage training scheme to minimize the efforts of finetuning powerful text-to-image diffusion models like FLUX.1-dev. In the first stage, we pre-train the model using task data with the 0-ref tasks from the text-to-image model. There are many models in the community based on the post-training of text-to-image foundational models that meet this training paradigm of the first stage. For example, FLUX.1-Fill-dev deals primarily with painting tasks and can be used as an initialization to accelerate the training process. In the second stage, we finetune the above model to support the general instructions using all tasks defined in ACE. To promote the widespread application of ACE++ in different scenarios, we provide a comprehensive set of models that cover both full finetuning and lightweight finetuning, while considering general applicability and applicability in vertical scenarios. The qualitative analysis showcases the superiority of ACE++ in terms of generating image quality and prompt following ability. Code and models will be available on the project page: https://ali-vilab. github.io/ACE_plus_page/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs

    cs.CV 2025-07 conditional novelty 7.0 of 10

    A large human-annotated benchmark of AI-edited images (EBench-18K) plus a fine-tuned LMM metric (LMM4Edit) that predicts human preference scores across three dimensions and answers editing-specific questions.

  2. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.

  3. Appearance Pointers -- Multimodal Region Control of Diffusion Transformers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Appearance pointers are compact tokens that let a diffusion transformer apply text, image, or combined prompts to specific image regions in a single pass.

  4. iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation

    cs.CV 2025-11 conditional novelty 6.0 of 10

    iMontage repurposes a pretrained video diffusion model to generate coherent yet highly dynamic image sets from arbitrary numbers of input images.

  5. Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SaaS, a self-adaptive attention-scaling method, improves instruction-following fidelity of unified image generation models without training by boosting the cross-attention activation of each sub-instruction in regions...

  6. Dreamland: Controllable World Creation with Simulator and Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A three-stage hybrid pipeline uses an intermediate layered world representation to refine simulator-rendered driving scenes into realistic, controllable images and videos.

  7. Identity-Preserving Text-to-Video Generation via Training-Free Prompt, Image, and Guidance Enhancement

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A training-free prompt, image, and guidance enhancement framework improves face consistency and video quality for identity-preserving text-to-video generation, winning the ACM Multimedia 2025 IPVG challenge.

  8. NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

    cs.CV 2025-08 conditional novelty 5.0 of 10

    NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.

  9. iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

    cs.GR 2025-06 conditional novelty 5.0 of 10

    A two-stage inpainting-based video diffusion transformer that reuses pretrained attention to reenact hand-object interactions with novel objects, reporting SOTA performance on Re-HOLD and a new in-the-wild dataset.

  10. PairEdit: Learning Semantic Variations for Exemplar-based Image Editing

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PairEdit trains two LoRA adapters on a pretrained diffusion model to capture the semantic direction between paired source-target images, enabling text-free, controllable image editing from as few as one pair.

  11. Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.

  12. ComfyMind: Toward General-Purpose Generation via Tree-Based Planning and Reactive Feedback

    cs.AI 2025-05 conditional novelty 5.0 of 10

    A ComfyUI-based multi-agent system with semantic workflow modules and tree-based local-feedback planning reports near-perfect pass rates on ComfyBench and competitive scores on GenEval and Reason-Edit.

Pith tools