Pith. sign in

REVIEW 10 cited by

Uni-ControlNet: All-in-One Control to Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16322 v3 pith:Y5LSIEFL submitted 2023-05-25 cs.CV cs.GR

classification cs.CVcs.GR
keywords uni-controlnetcontrolsmodelsdiffusiononlytexttext-to-imageadapters
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Text-to-Image diffusion models have made tremendous progress over the past two years, enabling the generation of highly realistic images based on open-domain text descriptions. However, despite their success, text descriptions often struggle to adequately convey detailed controls, even when composed of long and complex texts. Moreover, recent studies have also shown that these models face challenges in understanding such complex texts and generating the corresponding images. Therefore, there is a growing need to enable more control modes beyond text description. In this paper, we introduce Uni-ControlNet, a unified framework that allows for the simultaneous utilization of different local controls (e.g., edge maps, depth map, segmentation masks) and global controls (e.g., CLIP image embeddings) in a flexible and composable manner within one single model. Unlike existing methods, Uni-ControlNet only requires the fine-tuning of two additional adapters upon frozen pre-trained text-to-image diffusion models, eliminating the huge cost of training from scratch. Moreover, thanks to some dedicated adapter designs, Uni-ControlNet only necessitates a constant number (i.e., 2) of adapters, regardless of the number of local or global controls used. This not only reduces the fine-tuning costs and model size, making it more suitable for real-world deployment, but also facilitate composability of different conditions. Through both quantitative and qualitative comparisons, Uni-ControlNet demonstrates its superiority over existing methods in terms of controllability, generation quality and composability. Code is available at \url{https://github.com/ShihaoZhaoZSH/Uni-ControlNet}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Fine-tuning a diffusion transformer with asymmetric LoRA plus a new 24,000-sequence dataset enables multi-domain, step-by-step procedural generation and image-to-process reconstruction.

  2. Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A training-free method that matches sparse-point inter-frame feature correlations to transfer reference motion to generated videos with improved temporal consistency.

  3. Mojito: Motion Trajectory and Intensity Control for Video Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Mojito enables both trajectory and intensity control in text-to-video generation by combining training-free cross-attention guidance with optical-flow-conditioned motion intensity embeddings.

  4. NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

    cs.CV 2025-08 conditional novelty 5.0 of 10

    NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.

  5. Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.

  6. Test-time Conditional Text-to-Image Synthesis Using Diffusion Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    TINTIN conditions Stable Diffusion outputs at test time on color palettes and edge maps by backpropagating losses between decoded images and the condition through the denoising steps.

  7. Enhancing Low-Cost Video Editing with Lightweight Adaptors and Temporal-Aware Inversion

    cs.CV 2025-01 reject novelty 4.0 of 10

    GE-Adapter combines a temporal smoothness loss, bilateral-filtered DDIM inversion, and shared plus frame-specific prompt tokens to improve text-to-video editing, though the reported evidence is inconsistent.

  8. From Text to Pose to Image: Improving Diffusion Model Control and Quality

    cs.CV 2024-11 conditional novelty 4.0 of 10

    A text-to-pose transformer and a face-and-hand-aware pose adapter form a text-to-pose-to-image pipeline for diffusion models, beating the prior adapter baseline on 70 to 78 percent of test cases.

  9. Efficient Temporal Consistency in Diffusion-Based Video Editing with Adaptor Modules: A Theoretical Framework

    cs.CV 2025-04 reject novelty 3.0 of 10

    The paper attempts, but fails, to prove convergence and stability guarantees for adapter-based temporal consistency in diffusion video editing.

  10. Artificial Intelligence for Geometry-Based Feature Extraction, Analysis and Synthesis in Artistic Images: A Survey

    cs.AI 2024-12 conditional novelty 2.0 of 10

    A survey reviewing how geometric features (bounding boxes, keypoints, poses, 3D representations) are used in AI for extracting, analyzing, and synthesizing artistic images, concluding that geometry improves performanc...

Pith tools