REVIEW 46 cited by
OminiControl: Minimal and Universal Control for Diffusion Transformer
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present OminiControl, a novel approach that rethinks how image conditions are integrated into Diffusion Transformer (DiT) architectures. Current image conditioning methods either introduce substantial parameter overhead or handle only specific control tasks effectively, limiting their practical versatility. OminiControl addresses these limitations through three key innovations: (1) a minimal architectural design that leverages the DiT's own VAE encoder and transformer blocks, requiring just 0.1% additional parameters; (2) a unified sequence processing strategy that combines condition tokens with image tokens for flexible token interactions; and (3) a dynamic position encoding mechanism that adapts to both spatially-aligned and non-aligned control tasks. Our extensive experiments show that this streamlined approach not only matches but surpasses the performance of specialized methods across multiple conditioning tasks. To overcome data limitations in subject-driven generation, we also introduce Subjects200K, a large-scale dataset of identity-consistent image pairs synthesized using DiT models themselves. This work demonstrates that effective image control can be achieved without architectural complexity, opening new possibilities for efficient and versatile image generation systems.
Forward citations
Cited by 46 Pith papers
-
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
DSH-Bench supplies a hierarchical 58-category subject set, difficulty/scenario labels, and a human-aligned SICS metric that exposes systematic failures of 19 subject-driven T2I models.
-
Story2Board: A Training-Free Approach for Expressive Storyboard Generation
Story2Board uses reciprocal attention value mixing and latent panel anchoring to generate consistent yet visually diverse storyboards from text without any training.
-
MultiRef: Controllable Image Generation with Multiple Visual References
MultiRef-bench shows that current image generators that accept multiple visual references still fail to combine them reliably, with the best tested model OmniGen reaching only 66.6% synthetic and 79.0% real-world alig...
-
Imagine for Me: Creative Conceptual Blending of Real Images and Text via Blended Attention
IT-Blender blends a real image and a text prompt by injecting clean reference features into a trained attention module, improving disentangled concept blending in SD and FLUX.
-
InnoText: A Unified Model for Visual Text Generation and Editing
A unified DiT model with font-size-aware modulation and region-weighted loss outperforms existing visual text generation and editing systems on bilingual benchmarks.
-
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
A single diffusion model with in-context conditioning and fractional RoPE positions completes videos from arbitrary spatio-temporal image patches.
-
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
Kontinuous Kontext adds continuous edit-strength control to instruction-based image editing by projecting a scalar strength and text embedding into the modulation space of a Flux Kontext diffusion editor.
-
UniVideo: Unified Understanding, Generation, and Editing for Videos
UniVideo combines a frozen MLLM and a video DiT to unify video understanding, generation, in-context editing, visual prompting, and zero-shot free-form video edits under one instruction interface.
-
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus
FocusDPO adds dynamic spatial weighting to preference-based fine-tuning, improving subject fidelity and reducing attribute leakage in multi-subject personalized image generation.
-
Mixture of Global and Local Experts with Diffusion Transformer for Controllable Face Generation
Face-MoGLE improves controllable face generation by feeding decoupled binary masks through global and local experts with time- and space-dependent gating in a diffusion transformer.
-
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
USO trains one DiT model for subject-driven, style-driven, and joint generation by disentangling content and style from triplet data and adding a style-reward objective, claiming SOTA on USO-Bench.
-
Delay-constrained re-entry governs large-scale brain seizures and other network pathologies
An epilepsy modeling preprint claims delay-constrained re-entry of traveling excitation drives seizures and predicts 184 recorded seizures, but the submitted full text is an unrelated computer vision paper, so the cla...
-
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...
-
DreamPainter: Image Background Inpainting for E-commerce Scenarios
DreamPainter introduces a two-stage diffusion framework trained on a new synthetic e-commerce dataset, DreamEcom-400K, that outperforms open-source inpainting baselines on background generation with text and reference...
-
Training-free Geometric Image Editing on Diffusion Models
FreeFine splits geometric image editing into object transformation, source-region inpainting, and target refinement, using temporal attention, local noise, and text guidance in a training-free way.
-
Trade-offs in Image Generation: How Do Different Dimensions Interact?
A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.
-
FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
FreeCus is a training-free method that combines pivotal attention sharing, reversed noise shifting, and MLLM captions to personalize Flux.1 text-to-image generation from a single reference image.
-
FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization
A method for multi-subject image personalization that fuses independently trained LoRA modules at inference time on visual autoregressive models.
-
DreamLight: Towards Harmonious and Consistent Image Relighting
A unified image- and text-based relighting model with direction-biased attention and a wavelet foreground fixer outperforms existing methods on a synthetic relighting benchmark.
-
PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
PolyVivid combines VLLM-based grounding, 3D-RoPE positional encoding, and attention-inherited identity injection to generate customized videos with multiple consistent subjects and text-specified interactions.
-
UNIC: Unified In-Context Video Editing
One diffusion transformer handles ID insert, swap, delete, stylization, propagation, and re-camera control in a single model using in-context token concatenation with task-aware positional encoding and bias.
-
Image Editing As Programs with Diffusion Models
IEAP decomposes complex editing instructions into atomic operations executed sequentially on a diffusion transformer, and reports state-of-the-art results on MagicBrush and AnyEdit.
-
Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing
A reference-based video editing pipeline that guides cross-image attention with diffusion correspondence, then trains a per-video restoration model to clean up the zero-shot output.
-
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
AlginGen improves zero-shot personalized image generation by training a learnable token and a selective attention mask that align textual and visual priors, achieving the best balance of concept preservation and promp...
-
HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
HunyuanVideo-Avatar is an audio-driven video generator that enables emotion-controllable and multi-character animation by injecting character images, routing audio via face masks, and transferring emotion from referen...
-
Jodi: Unification of Visual Generation and Understanding via Joint Modeling
A single diffusion transformer with role-switch training performs joint generation, controllable generation, and multi-label perception across image and seven label domains.
-
DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
DiffDecompose recovers foreground and background layers from alpha-composited images using in-context diffusion with position encoding cloning, trained and evaluated on a new six-task synthetic dataset.
-
OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data
OmniConsistency is a style-agnostic consistency module for Flux that preserves structure and details during stylization with arbitrary LoRAs, reaching GPT-4o-level content consistency.
-
KRIS-Bench: Benchmarking Next-Level Intelligent Image Editing Models
A new benchmark, KRIS-Bench, evaluates image editing models on knowledge-grounded reasoning across factual, conceptual, and procedural tasks, and finds large performance gaps in current models.
-
TexTailor: Inference-Time Textual Guidance Tailoring for Multimodal Diffusion Transformers
Scaling text-condition hidden states by 1.5 in a small, attribute-specific set of MMDiT blocks improves text-image alignment, editing, and speed on SD3.5, FLUX, and Qwen Image with no training.
-
Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
A 2B multimodal model trained with HCM-GRPO, a GRPO variant with partial-credit rewards and hard-case oversampling, outperforms larger models on the authors' private physical-plausibility test set.
-
DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
DreamSwapV performs mask-guided, subject-agnostic subject swapping in videos using multiple conditions, an adaptive mask strategy, and a new benchmark.
-
NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer
NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.
-
DivControl: Knowledge Diversion for Controllable Image Generation
DivControl factorizes ControlNet weights via SVD into shared 'learngenes' and condition-specific 'tailors', routed by a text-conditioned gate, enabling unified control and efficient adaptation to new conditions.
-
WordCraft: Interactive Artistic Typography with Attention Awareness and Noise Blending
A text-to-image pipeline that lets users restyle individual characters or regions of artistic typography and iteratively refine them with region-specific prompts.
-
DreamPoster: A Unified Framework for Image-Conditioned Generative Poster Design
DreamPoster fine-tunes Seedream3.0 with a deconstruction-recaptioning dataset pipeline and a three-stage curriculum to turn image-plus-text inputs into finished posters, reporting substantially higher usability than G...
-
Ovis-U1 Technical Report
A 3B unified multimodal model with a diffusion decoder and bidirectional refiner achieves competitive understanding, generation, and editing benchmark scores.
-
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation
XVerse learns token-specific offsets that modify the text-stream modulation of a diffusion transformer, enabling multi-subject identity and attribute control in image generation.
-
WordCon: Word-level Typography Control in Scene Text Rendering
WordCon uses grounding-model masks and two extra losses to fine-tune Flux so that typography can be controlled word by word.
-
FramePrompt: In-context Controllable Animation with Zero Structural Changes
FramePrompt turns character animation into a video-continuation task by concatenating reference image, skeleton frames, and target frames into one sequence, then training the pretrained Wan-I2V model to generate only ...
-
PosterCraft: Rethinking High-Quality Aesthetic Poster Generation in a Unified Framework
PosterCraft improves text-to-poster generation by cascading four stages of training (text rendering, region-weighted fine-tuning, preference optimization, and vision-language feedback), outperforming open-source basel...
-
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
FullDiT2 accelerates FullDiT-style in-context conditioning for video by dynamic token selection and selective context caching, cutting per-step time by 2-3x with minimal quality loss.
-
ByteMorph: Benchmarking Instruction-Guided Image Editing with Non-Rigid Motions
A released 6.4 million pair dataset and 613 sample benchmark for instruction-guided image editing of non-rigid motions, plus a Flux.1-dev based baseline that outperforms open-source methods on the new benchmark.
-
Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis
A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.
-
UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes
UniTEX generates textures for 3D shapes by predicting continuous volumetric texture functions, bypassing UV maps.
-
EditIDv2: Editable ID Customization with Data-Lubricated ID Feature Integration for Text-to-Image Generation
EditIDv2 fine-tunes only PerceiverAttention cross-attention weights on 3K images to inject editability into Flux-based ID customization, reporting selective gains on the self-proposed IBench benchmark.
Discussion (0). Sign in to comment.