Pith. sign in

REVIEW 8 cited by

Quality-Aware Image-Text Alignment for Opinion-Unaware Image Quality Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.11176 v3 pith:AJB7HBAJ submitted 2024-03-17 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagecliphumanquality-awarealignmentimagesopinion-unawarequaliclip
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

No-Reference Image Quality Assessment (NR-IQA) focuses on designing methods to measure image quality in alignment with human perception when a high-quality reference image is unavailable. Most state-of-the-art NR-IQA approaches are opinion-aware, i.e. they require human annotations for training. This dependency limits their scalability and broad applicability. To overcome this limitation, we propose QualiCLIP (Quality-aware CLIP), a CLIP-based self-supervised opinion-unaware approach that does not require human opinions. In particular, we introduce a quality-aware image-text alignment strategy to make CLIP generate quality-aware image representations. Starting from pristine images, we synthetically degrade them with increasing levels of intensity. Then, we train CLIP to rank these degraded images based on their similarity to quality-related antonym text prompts. At the same time, we force CLIP to generate consistent representations for images with similar content and the same level of degradation. Our experiments show that the proposed method improves over existing opinion-unaware approaches across multiple datasets with diverse distortion types. Moreover, despite not requiring human annotations, QualiCLIP achieves excellent performance against supervised opinion-aware methods in cross-dataset experiments, thus demonstrating remarkable generalization capabilities. The code and the model are publicly available at https://github.com/miccunifi/QualiCLIP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Impulse-to-Peak-Output Norm Optimal State-Feedback Control of Linear PDEs

    math.OC 2026-04 unverdicted novelty 7.0 of 10

    Using PIE representations and Lyapunov LMIs with strong duality, the authors give provable I2P-norm bounds and constructive optimal state-feedback for linear PDEs.

  2. Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution

    cs.CV 2026-03 conditional novelty 6.0 of 10

    Asymmetric adversarial distillation plus frequency distribution matching lets DiT models perform one-step Real-ISR without the grid-like periodic artifacts that plague prior one-step DiT distillations.

  3. From Global to Granular: Revealing IQA Model Performance via Correlation Surface

    cs.CV 2026-01 conditional novelty 6.0 of 10

    GMC maps an IQA model's agreement with human scores across the quality-level and quality-difference landscape, exposing local strengths that global PLCC/SRCC hide.

  4. The Devil is in the Darkness: Diffusion-Based Nighttime Dehazing Anchored in Brightness Perception

    cs.CV 2025-06 conditional novelty 6.0 of 10

    DiffND combines a depth- and sky-guided data synthesis pipeline with a diffusion model gated by a brightness perception network to achieve nighttime dehazing with day-level brightness.

  5. Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and Resolution

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A new task and dataset for predicting the scale at which perceived image quality peaks, with a weak-label method that improves several no-reference quality models.

  6. HiRQA: Hierarchical Ranking and Quality Alignment for Opinion-Unaware Image Quality Assessment

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    HiRQA is a self-supervised NR-IQA framework trained on synthetic distortions, using a higher-order ranking loss, embedding distance loss, and text-guided contrastive alignment, claimed to generalize to authentic distortions.

  7. Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A position paper contending that multimedia quality assessment should move beyond scalar Mean Opinion Score toward context-aware, explainable, and multimodal modeling.

  8. A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss

    cs.CV 2025-09 conditional novelty 3.0 of 10

    An ensemble of MobileNetV3-Small and ShuffleNetV2 with a correlation-aware loss and test-time augmentation reaches SRCC 0.9829 and PLCC 0.9894 on the VQualA FIQA validation set.

Pith tools