REVIEW 18 cited by
Controlling Vision-Language Models for Multi-Task Image Restoration
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degradation-aware vision-language model (DA-CLIP) to better transfer pretrained vision-language models to low-level vision tasks as a multi-task framework for image restoration. More specifically, DA-CLIP trains an additional controller that adapts the fixed CLIP image encoder to predict high-quality feature embeddings. By integrating the embedding into an image restoration network via cross-attention, we are able to pilot the model to learn a high-fidelity image reconstruction. The controller itself will also output a degradation feature that matches the real corruptions of the input, yielding a natural classifier for different degradation types. In addition, we construct a mixed degradation dataset with synthetic captions for DA-CLIP training. Our approach advances state-of-the-art performance on both \emph{degradation-specific} and \emph{unified} image restoration tasks, showing a promising direction of prompting image restoration with large-scale pretrained vision-language models. Our code is available at https://github.com/Algolzw/daclip-uir.
Forward citations
Cited by 18 Pith papers
-
SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning
A spiking neural network with subtractive and additive attention performs all-in-one image restoration in one time step, matching older ANN baselines with much lower estimated energy.
-
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
A vision-language agent trained with SFT plus RL, exploration-driven trajectory perturbation, and adaptive multi-metric rewards learns direct tool selection for composite image restoration, beating training-free agent...
-
TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal
A two-stage prompt-tuning method with low-rank and contrastive prompt enhancement claims all-in-one adverse weather removal at 2.75M parameters.
-
Robust Adverse Weather Removal via Spectral-based Spatial Grouping
SSGformer, an all-in-one transformer combining Sobel and SVD spectral prompts with mask-based group-wise attention, reports state-of-the-art averages on All-weather and WeatherStream.
-
ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
ModalFormer combines nine auxiliary modality features with a cross-modal attention mechanism and reports state-of-the-art low-light image enhancement on LOL-v1, LOL-v2, and SDSD benchmarks.
-
Grounding Degradations in Natural Language for All-In-One Video Restoration
RONIN distills per-frame language descriptions of video degradations into lightweight input-conditioned prompts, achieving all-in-one video restoration without any text encoder or MLLM at inference and outperforming p...
-
4KAgent: Agentic Any Image to 4K Super-Resolution
An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.
-
PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors
PBR-SR super-resolves PBR texture maps (albedo, roughness, metallic, normal) in a zero-shot way by optimizing textures so differentiable renderings match super-resolved multi-view renderings from a pretrained image SR model.
-
CLIP-driven rain perception: Adaptive deraining with pattern-aware network routing and mask-guided cross-attention
A deraining network that routes each rainy image to a sub-network based on CLIP text-image similarity, adds mask-guided cross-attention and a dynamic loss, and reports marginal benchmark improvements.
-
AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation
Training-free AccelAes accelerates DiTs with aesthetic focus masks and step caches, reporting 2.11× speedup and +11.9% ImageReward on Lumina-Next.
-
Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration
Pref-Restore combines AR semantic tokens, a diffusion generator, and DiffusionNFT-style RL to make blind face restoration more consistent, but its deterministic-identity claim is weakened by self-referential rewards a...
-
FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution
FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.
-
UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restoration
UniLDiff combines degradation-aware attention fusion with a detail-aware expert decoder to achieve state-of-the-art perceptual quality on unified image restoration benchmarks.
-
HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement
HVI-CIDNet+ replaces the HSV color plane with polarized hue-saturation coordinates and a learned dark-intensity collapse, then trains a dual-branch transformer-CNN network with CLIP-derived priors for low-light enhancement.
-
Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention
LAformer applies rank-enhanced linear attention to image restoration, reporting state-of-the-art performance across 21 benchmarks with linear-complexity global modeling.
-
Diffusion Once and Done: Degradation-Aware LoRA for Efficient All-in-One Image Restoration
Proposes DOD, a one-step Stable Diffusion model for all-in-one image restoration, but the submitted manuscript text is an unrelated software engineering review, leaving the claim unverifiable.
-
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
Multimodal LLM accuracy can improve on visually degraded images, and a lightweight test-time tuning module that modulates input quality yields small accuracy gains on some benchmarks.
-
From Controlled Scenarios to Real-World: Cross-Domain Degradation Pattern Matching for All-in-One Image Restoration
UDAIR combines codebook quantization, cross-sample contrastive learning, and CORAL-based test-time adaptation to reduce the domain gap in all-in-one image restoration.
Discussion (0). Continue with ORCID to comment.