Pith. sign in

REVIEW 46 cited by

OminiControl: Minimal and Universal Control for Diffusion Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.15098 v6 pith:NADNMMJF submitted 2024-11-22 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords imagecontrolominicontroltaskstransformerapproacharchitecturalconditioning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present OminiControl, a novel approach that rethinks how image conditions are integrated into Diffusion Transformer (DiT) architectures. Current image conditioning methods either introduce substantial parameter overhead or handle only specific control tasks effectively, limiting their practical versatility. OminiControl addresses these limitations through three key innovations: (1) a minimal architectural design that leverages the DiT's own VAE encoder and transformer blocks, requiring just 0.1% additional parameters; (2) a unified sequence processing strategy that combines condition tokens with image tokens for flexible token interactions; and (3) a dynamic position encoding mechanism that adapts to both spatially-aligned and non-aligned control tasks. Our extensive experiments show that this streamlined approach not only matches but surpasses the performance of specialized methods across multiple conditioning tasks. To overcome data limitations in subject-driven generation, we also introduce Subjects200K, a large-scale dataset of identity-consistent image pairs synthesized using DiT models themselves. This work demonstrates that effective image control can be achieved without architectural complexity, opening new possibilities for efficient and versatile image generation systems.

Discussion (0). Sign in to comment.

Forward citations

Cited by 46 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    DSH-Bench supplies a hierarchical 58-category subject set, difficulty/scenario labels, and a human-aligned SICS metric that exposes systematic failures of 19 subject-driven T2I models.

  2. Story2Board: A Training-Free Approach for Expressive Storyboard Generation

    cs.CV 2025-08 conditional novelty 7.0 of 10

    Story2Board uses reciprocal attention value mixing and latent panel anchoring to generate consistent yet visually diverse storyboards from text without any training.

  3. MultiRef: Controllable Image Generation with Multiple Visual References

    cs.CV 2025-08 conditional novelty 7.0 of 10

    MultiRef-bench shows that current image generators that accept multiple visual references still fail to combine them reliably, with the best tested model OmniGen reaching only 66.6% synthetic and 79.0% real-world alig...

  4. Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention

    cs.CV 2025-06 conditional novelty 7.0 of 10

    IT-Blender blends a real image and a text prompt by injecting clean reference features into a trained attention module, improving disentangled concept blending in SD and FLUX.

  5. InnoText: A Unified Model for Visual Text Generation and Editing

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A unified DiT model with font-size-aware modulation and region-weighted loss outperforms existing visual text generation and editing systems on bilingual benchmarks.

  6. VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A single diffusion model with in-context conditioning and fractional RoPE positions completes videos from arbitrary spatio-temporal image patches.

  7. Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Kontinuous Kontext adds continuous edit-strength control to instruction-based image editing by projecting a scalar strength and text embedding into the modulation space of a Flux Kontext diffusion editor.

  8. UniVideo: Unified Understanding, Generation, and Editing for Videos

    cs.CV 2025-10 conditional novelty 6.0 of 10

    UniVideo combines a frozen MLLM and a video DiT to unify video understanding, generation, in-context editing, visual prompting, and zero-shot free-form video edits under one instruction interface.

  9. FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus

    cs.CV 2025-09 conditional novelty 6.0 of 10

    FocusDPO adds dynamic spatial weighting to preference-based fine-tuning, improving subject fidelity and reducing attribute leakage in multi-subject personalized image generation.

  10. Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Face-MoGLE improves controllable face generation by feeding decoupled binary masks through global and local experts with time- and space-dependent gating in a diffusion transformer.

  11. USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.

  12. Delay-constrained re-entry governs large-scale brain seizures and other network pathologies

    q-bio.NC 2025-08 unverdicted novelty 6.0 of 10

    An epilepsy modeling preprint claims delay-constrained re-entry of traveling excitation drives seizures and predicts 184 recorded seizures, but the submitted full text is an unrelated computer vision paper, so the cla...

  13. Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...

  14. DreamPainter: Image Background Inpainting for E-commerce Scenarios

    cs.CV 2025-08 conditional novelty 6.0 of 10

    DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...

  15. Training-free Geometric Image Editing on Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FreeFine splits geometric image editing into object transformation, source-region inpainting, and target refinement, using temporal attention, local noise, and text guidance in a training-free way.

  16. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  17. FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.

  18. FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A method for multi-subject image personalization that fuses independently trained LoRA modules at inference time on visual autoregressive models.

  19. DreamLight: Towards Harmonious and Consistent Image Relighting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.

  20. PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PolyVivid combines VLLM-based grounding, 3D-RoPE positional encoding, and attention-inherited identity injection to generate customized videos with multiple consistent subjects and text-specified interactions.

  21. UNIC: Unified In-Context Video Editing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.

  22. Image Editing As Programs with Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.

  23. Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A reference-based video editing pipeline that guides cross-image attention with diffusion correspondence, then trains a per-video restoration model to clean up the zero-shot output.

  24. AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    AlginGen improves zero-shot personalized image generation by training a learnable token and a selective attention mask that align textual and visual priors, achieving the best balance of concept preservation and promp...

  25. HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

    cs.CV 2025-05 conditional novelty 6.0 of 10

    HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...

  26. Jodi: Unification of Visual Generation and Understanding via Joint Modeling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single diffusion transformer with role-switch training performs joint generation, controllable generation, and multi-label perception across image and seven label domains.

  27. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

    cs.CV 2025-05 conditional novelty 6.0 of 10

    DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.

  28. OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data

    cs.CV 2025-05 conditional novelty 6.0 of 10

    OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.

  29. KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.

  30. TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Scaling text-condition hidden states by 1.5 in a small, attribute-specific set of MMDiT blocks improves text-image alignment, editing, and speed on SD3.5, FLUX, and Qwen Image with no training.

  31. Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance

    cs.CV 2025-11 conditional novelty 5.0 of 10

    A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.

  32. DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.

  33. NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

    cs.CV 2025-08 conditional novelty 5.0 of 10

    NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.

  34. DivControl: Knowledge Diversion for Controllable Image Generation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DivControl factorizes ControlNet weights via SVD into shared 'learngenes' and condition-specific 'tailors', routed by a text-conditioned gate, enabling unified control and efficient adaptation to new conditions.

  35. WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A text-to-image pipeline that lets users restyle individual characters or regions of artistic typography and iteratively refine them with region-specific prompts.

  36. DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DreamPoster fine-tunes Seedream3.0 with a deconstruction-recaptioning dataset pipeline and a three-stage curriculum to turn image-plus-text inputs into finished posters, reporting substantially higher usability than G...

  37. Ovis-U1 Technical Report

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A 3B unified multimodal model with a diffusion decoder and bidirectional refiner achieves competitive understanding, generation, and editing benchmark scores.

  38. XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    XVerse learns token-specific offsets that modify the text-stream modulation of a diffusion transformer, enabling multi-subject identity and attribute control in image generation.

  39. WordCon: Word-level Typography Control in Scene Text Rendering

    cs.CV 2025-06 conditional novelty 5.0 of 10

    WordCon uses grounding-model masks and two extra losses to fine-tune Flux so that typography can be controlled word by word.

  40. FramePrompt: In-context Controllable Animation with Zero Structural Changes

    cs.GR 2025-06 conditional novelty 5.0 of 10

    FramePrompt turns character animation into a video-continuation task by concatenating reference image, skeleton frames, and target frames into one sequence, then training the pretrained Wan-I2V model to generate only ...

  41. PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework

    cs.CV 2025-06 conditional novelty 5.0 of 10

    PosterCraft improves text-to-poster generation by cascading four stages of training (text rendering, region-weighted fine-tuning, preference optimization, and vision-language feedback), outperforming open-source basel...

  42. FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers

    cs.CV 2025-06 conditional novelty 5.0 of 10

    FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.

  43. ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.

  44. Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.

  45. UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes

    cs.CV 2025-05 conditional novelty 5.0 of 10

    UniTEX generates textures for 3D shapes by predicting continuous volumetric texture functions, bypassing UV maps.

  46. EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation

    cs.CV 2025-09 reject novelty 3.0 of 10

    EditIDv2 fine-tunes only PerceiverAttention cross-attention weights on 3K images to inject editability into Flux-based ID customization, reporting selective gains on the self-proposed IBench benchmark.

Pith tools