Pith. sign in

REVIEW 3 cited by

Q-Instruct: Improving Low-level Visual Abilities for Multi-modality Foundation Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.06783 v1 pith:OD4FY62L submitted 2023-11-12 cs.CV cs.MM

classification cs.CVcs.MM
keywords low-levelmodelsvisualfoundationhumanabilitiesappearancefeedbacks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-modality foundation models, as represented by GPT-4V, have brought a new paradigm for low-level visual perception and understanding tasks, that can respond to a broad range of natural human instructions in a model. While existing foundation models have shown exciting potentials on low-level visual tasks, their related abilities are still preliminary and need to be improved. In order to enhance these models, we conduct a large-scale subjective experiment collecting a vast number of real human feedbacks on low-level vision. Each feedback follows a pathway that starts with a detailed description on the low-level visual appearance (*e.g. clarity, color, brightness* of an image, and ends with an overall conclusion, with an average length of 45 words. The constructed **Q-Pathway** dataset includes 58K detailed human feedbacks on 18,973 images with diverse low-level appearance. Moreover, to enable foundation models to robustly respond to diverse types of questions, we design a GPT-participated conversion to process these feedbacks into diverse-format 200K instruction-response pairs. Experimental results indicate that the **Q-Instruct** consistently elevates low-level perception and understanding abilities across several foundational models. We anticipate that our datasets can pave the way for a future that general intelligence can perceive, understand low-level visual appearance and evaluate visual quality like a human. Our dataset, model zoo, and demo is published at: https://q-future.github.io/Q-Instruct.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grounding Degradations in Natural Language for All-In-One Video Restoration

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RONIN distills per-frame language descriptions of video degradations into lightweight input-conditioned prompts, achieving all-in-one video restoration without any text encoder or MLLM at inference and outperforming p...

  2. STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset

    cs.CV 2025-06 conditional novelty 6.0 of 10

    STORM is a new multi-domain ordinal-regression benchmark with coarse-to-fine Chain-of-Thought prompts that improves MLLM zero-shot visual rating, though the 'universal' claim is bounded by its five curated domains.

  3. Parameter-Efficient Adaptation of mPLUG-Owl2 via Pixel-Level Visual Prompts for NR-IQA

    cs.CV 2025-09 conditional novelty 5.0 of 10

    With a learned 30-pixel border prompt added to input images, a frozen mPLUG-Owl2-7B reaches 0.932 SRCC on KADID-10k using about 156K trainable parameters.

Pith tools