Pith. sign in

REVIEW 10 cited by

Q-Insight: Understanding Image Quality via Visual Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.22679 v2 pith:GHEZNCL4 submitted 2025-03-28 cs.CV

classification cs.CV
keywords imagequalitydegradationq-insighttasksperceptionreasoningunderstanding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image quality assessment (IQA) focuses on the perceptual visual quality of images, playing a crucial role in downstream tasks such as image reconstruction, compression, and generation. The rapid advancement of multi-modal large language models (MLLMs) has significantly broadened the scope of IQA, moving toward comprehensive image quality understanding that incorporates content analysis, degradation perception, and comparison reasoning beyond mere numerical scoring. Previous MLLM-based methods typically either generate numerical scores lacking interpretability or heavily rely on supervised fine-tuning (SFT) using large-scale annotated datasets to provide descriptive assessments, limiting their flexibility and applicability. In this paper, we propose Q-Insight, a reinforcement learning-based model built upon group relative policy optimization (GRPO), which demonstrates strong visual reasoning capability for image quality understanding while requiring only a limited amount of rating scores and degradation labels. By jointly optimizing score regression and degradation perception tasks with carefully designed reward functions, our approach effectively exploits their mutual benefits for enhanced performance. Extensive experiments demonstrate that Q-Insight substantially outperforms existing state-of-the-art methods in both score regression and degradation perception tasks, while exhibiting impressive zero-shot generalization to comparison reasoning tasks. Code will be available at https://github.com/lwq20020127/Q-Insight.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning

    cs.CV 2026-01 conditional novelty 7.0 of 10

    Zoom-IQA lets a vision-language model iteratively crop and zoom into image regions before giving a quality score, improving reasoning and restoration guidance over single-pass IQA models.

  2. More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models

    cs.CL 2025-05 conditional novelty 7.0 of 10

    Reasoning models trade visual grounding for language-based inference, and this paper measures that trade-off with a new metric and benchmark.

  3. Peak-End-Net: A Peak-End Rule Inspired Framework for Generalizable Video Aesthetic Assessment

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Peak-End-Net uses image-aesthetic priors and peak-end-rule frame weighting to set a new state of the art on VADB and zero-shot DIVIDE-3K video aesthetic assessment.

  4. ME-IQA: Memory-Enhanced Image Quality Assessment via Re-Ranking

    cs.CV 2026-03 conditional novelty 6.0 of 10

    A test-time re-ranking framework that uses a memory of similar images and pairwise VLM comparisons to densify and improve reasoning-based image quality scores.

  5. EmoFeedback$^2$: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback

    cs.CV 2025-11 reject novelty 6.0 of 10

    A closed-loop emotional image generation system uses a fine-tuned vision-language model both as a reinforcement-learning reward and as an iterative prompt refiner, claiming improved valence-arousal fidelity.

  6. VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VQ-Insight uses progressive reinforcement learning with temporal shuffle and task rewards to teach a vision-language model to score and compare AI-generated videos, with gains on multiple video quality benchmarks.

  7. Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Q-Ponder is a two-stage pipeline (distill-then-reinforce) that makes a 7B multimodal model both more accurate at image quality scoring and better at explaining its judgments.

  8. VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A challenge report finding that ensemble-tuned LMMs reach 75.7% accuracy on a 4,000-question visual quality comparison benchmark, only ~0.2 points above a Qwen2.5-VL-72B baseline.

  9. MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection

    cs.CV 2025-05 reject novelty 4.0 of 10

    MIND-Edit combines instruction rewriting with MLLM-derived visual embeddings to guide diffusion-based image editing, but the reported numbers only partly support the claim of state-of-the-art performance.

  10. Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models

    cs.CL 2025-05 conditional novelty 2.0 of 10

    A survey-style position paper claims that reinforcement fine-tuning powers reasoning in multimodal LLMs, summarizing over a hundred recent works and proposing five future research directions.

Pith tools