Pith. sign in

REVIEW 18 cited by

Controlling Vision-Language Models for Multi-Task Image Restoration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.01018 v2 pith:LPZK7BAY submitted 2023-10-02 cs.CV

classification cs.CV
keywords imagerestorationvision-languagemodelsda-clipdegradationtasksclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates dramatically due to corrupted inputs. In this paper, we present a degradation-aware vision-language model (DA-CLIP) to better transfer pretrained vision-language models to low-level vision tasks as a multi-task framework for image restoration. More specifically, DA-CLIP trains an additional controller that adapts the fixed CLIP image encoder to predict high-quality feature embeddings. By integrating the embedding into an image restoration network via cross-attention, we are able to pilot the model to learn a high-fidelity image reconstruction. The controller itself will also output a degradation feature that matches the real corruptions of the input, yielding a natural classifier for different degradation types. In addition, we construct a mixed degradation dataset with synthetic captions for DA-CLIP training. Our approach advances state-of-the-art performance on both \emph{degradation-specific} and \emph{unified} image restoration tasks, showing a promising direction of prompting image restoration with large-scale pretrained vision-language models. Our code is available at https://github.com/Algolzw/daclip-uir.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SpikeRestormer: Towards Energy-Efficient All-in-One Image Restoration via Unified Event Reasoning

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A spiking neural network with subtractive and additive attention performs all-in-one image restoration in one time step, matching older ANN baselines with much lower estimated energy.

  2. TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A vision-language agent trained with SFT plus RL, exploration-driven trajectory perturbation, and adaptive multi-metric rewards learns direct tool selection for composite image restoration, beating training-free agent...

  3. TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A two-stage prompt-tuning method with low-rank and contrastive prompt enhancement claims all-in-one adverse weather removal at 2.75M parameters.

  4. Robust Adverse Weather Removal via Spectral-based Spatial Grouping

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SSGformer, an all-in-one transformer combining Sobel and SVD spectral prompts with mask-based group-wise attention, reports state-of-the-art averages on All-weather and WeatherStream.

  5. ModalFormer: Multimodal Transformer for Low-Light Image Enhancement

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ModalFormer combines nine auxiliary modality features with a cross-modal attention mechanism and reports state-of-the-art low-light image enhancement on LOL-v1, LOL-v2, and SDSD benchmarks.

  6. Grounding Degradations in Natural Language for All-In-One Video Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RONIN distills per-frame language descriptions of video degradations into lightweight input-conditioned prompts, achieving all-in-one video restoration without any text encoder or MLLM at inference and outperforming p...

  7. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  8. PBR-SR: Mesh PBR Texture Super Resolution from 2D Image Priors

    cs.CV 2025-06 conditional novelty 6.0 of 10

    PBR-SR super-resolves PBR texture maps (albedo, roughness, metallic, normal) in a zero-shot way by optimizing textures so differentiable renderings match super-resolved multi-view renderings from a pretrained image SR model.

  9. CLIP-driven rain perception: Adaptive deraining with pattern-aware network routing and mask-guided cross-attention

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A deraining network that routes each rainy image to a sub-network based on CLIP text-image similarity, adds mask-guided cross-attention and a dynamic loss, and reports marginal benchmark improvements.

  10. AccelAes: Accelerating Diffusion Transformers for Training-Free Aesthetic-Enhanced Image Generation

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    Training-free AccelAes accelerates DiTs with aesthetic focus masks and step caches, reporting 2.11× speedup and +11.9% ImageReward on Lumina-Next.

  11. Bridging Information Asymmetry: A Hierarchical Framework for Deterministic Blind Face Restoration

    cs.CV 2026-01 conditional novelty 5.0 of 10

    Pref-Restore combines AR semantic tokens, a diffusion generator, and DiffusionNFT-style RL to make blind face restoration more consistent, but its deterministic-identity claim is weakened by self-referential rewards a...

  12. FS-Diff: Semantic guidance and clarity-aware simultaneous multimodal image fusion and super-resolution

    cs.CV 2025-09 conditional novelty 5.0 of 10

    FS-Diff is a diffusion model that jointly fuses and super-resolves low-resolution multimodal image pairs using clarity-aware CLIP semantics and a bidirectional Mamba feature extractor.

  13. UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restoration

    cs.CV 2025-07 conditional novelty 5.0 of 10

    UniLDiff combines degradation-aware attention fusion with a detail-aware expert decoder to achieve state-of-the-art perceptual quality on unified image restoration benchmarks.

  14. HVI-CIDNet+: Beyond Extreme Darkness for Low-Light Image Enhancement

    cs.CV 2025-07 conditional novelty 5.0 of 10

    HVI-CIDNet+ replaces the HSV color plane with polarized hue-saturation coordinates and a learned dark-intensity collapse, then trains a dual-branch transformer-CNN network with CLIP-derived priors for low-light enhancement.

  15. Fast and Accurate Image Restoration and Generation with Rank Enhanced Linear Attention

    cs.CV 2025-05 conditional novelty 5.0 of 10

    LAformer applies rank-enhanced linear attention to image restoration, reporting state-of-the-art performance across 21 benchmarks with linear-complexity global modeling.

  16. Diffusion Once and Done: Degradation-Aware LoRA for Efficient All-in-One Image Restoration

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    Proposes DOD, a one-step Stable Diffusion model for all-in-one image restoration, but the submitted manuscript text is an unrelated software engineering review, leaving the claim unverifiable.

  17. Demystifying the Visual Quality Paradox in Multimodal Large Language Models

    cs.CV 2025-06 reject novelty 4.0 of 10

    Multimodal LLM accuracy can improve on visually degraded images, and a lightweight test-time tuning module that modulates input quality yields small accuracy gains on some benchmarks.

  18. From Controlled Scenarios to Real-World: Cross-Domain Degradation Pattern Matching for All-in-One Image Restoration

    cs.CV 2025-05 conditional novelty 4.0 of 10

    UDAIR combines codebook quantization, cross-sample contrastive learning, and CORAL-based test-time adaptation to reduce the domain gap in all-in-one image restoration.

Pith tools